Scan quality decides OCR accuracy, and nothing downstream recovers it
Optical character recognition is often blamed for errors that were created before it ran. OCR works by isolating shapes and matching them against learned letterforms, so its accuracy is bounded by how cleanly the shapes can be isolated. If the characters are blurred, skewed, unevenly lit or fighting a patterned background, no amount of post-processing recovers information that the image never contained.
The dominant factor is effective resolution on the text itself, not the megapixels of the camera. Roughly 300 dots per inch across the printed page is where accuracy stops improving much; well below 200, error rates climb sharply because the strokes that distinguish similar glyphs stop being resolved. This is why a 12 megapixel photograph taken at an angle from half a metre away often reads worse than a modest flatbed scan: the pixels exist but few of them are on the text.
The second factor is contrast and evenness. A phone photograph of a page usually has a gradient across it, brighter near the window and darker at the far edge, plus the shadow of the phone itself. A global brightness threshold then either loses the pale text or fills the dark side with noise. The fix is adaptive thresholding, which computes a threshold for each neighbourhood of the image rather than one for the whole page, which is why the cleaning step here handles a lit-from-one-side photograph far better than a contrast slider does.
Skew matters more than it appears to. A page rotated by three degrees still looks fine to a person, but line-finding algorithms segment text into rows, and a slope means a row crosses between lines of text. Deskewing before recognition is often the single largest accuracy gain available on phone-captured documents.
A word on what a searchable PDF actually is, because it is widely misunderstood. It contains the page image exactly as before, plus a layer of invisible text positioned over the words. Your viewer draws the picture and searches the text. This means the recognition errors are still in the file, invisible, and a search for a word the OCR misread will not find it even though you can plainly see the word on screen. Searchable does not mean corrected.