Scan and clean a document

Turn a photo of a page into a clean scan and pull out its text. The image is processed in this browser tab and is not uploaded.

How it works

A phone photo of a page is rarely a clean “scan”: it is greyish, unevenly lit, and not selectable text. This tool fixes both. A cleanup pass converts the photo to a crisp result (greyscale, higher contrast, or true black-and-white) and optical character recognition (OCR) reads the text off the page.

The cleanup runs instantly on your device. The first time you extract text, an English recognition pack downloads (a few megabytes) and is cached; after that OCR works offline too. The image and the recognised text stay in your browser. Nothing is uploaded.

The extracted text is editable, so you can fix the odd character before downloading it.

How to use it

  1. 1 Choose a photo of the page.
  2. 2 Pick a cleanup, Black & white works best for text.
  3. 3 Download the clean image, or press Extract text to run OCR.
  4. 4 Edit the recognised text if needed, then download it as a .txt file.

Good to know

  • The first text extraction downloads an English pack; after that OCR works offline.
  • A flat, well-lit, in-focus photo, cropped to the page, gives the best results.
  • Black & white cleanup usually gives the recogniser the cleanest input.
  • The image and text stay in your browser; nothing is uploaded.

Browser support

  • Chrome 94 and later
  • Firefox 93 and later
  • Safari 16.4 and later
  • Edge 94 and later

These are the oldest versions with everything the tool uses. They are worked out from the tool itself, so they change when it does.

Version numbers checked on 2026-09-04.

Scan quality decides OCR accuracy, and nothing downstream recovers it

Optical character recognition is often blamed for errors that were created before it ran. OCR works by isolating shapes and matching them against learned letterforms, so its accuracy is bounded by how cleanly the shapes can be isolated. If the characters are blurred, skewed, unevenly lit or fighting a patterned background, no amount of post-processing recovers information that the image never contained.

The dominant factor is effective resolution on the text itself, not the megapixels of the camera. Roughly 300 dots per inch across the printed page is where accuracy stops improving much; well below 200, error rates climb sharply because the strokes that distinguish similar glyphs stop being resolved. This is why a 12 megapixel photograph taken at an angle from half a metre away often reads worse than a modest flatbed scan: the pixels exist but few of them are on the text.

The second factor is contrast and evenness. A phone photograph of a page usually has a gradient across it, brighter near the window and darker at the far edge, plus the shadow of the phone itself. A global brightness threshold then either loses the pale text or fills the dark side with noise. The fix is adaptive thresholding, which computes a threshold for each neighbourhood of the image rather than one for the whole page, which is why the cleaning step here handles a lit-from-one-side photograph far better than a contrast slider does.

Skew matters more than it appears to. A page rotated by three degrees still looks fine to a person, but line-finding algorithms segment text into rows, and a slope means a row crosses between lines of text. Deskewing before recognition is often the single largest accuracy gain available on phone-captured documents.

A word on what a searchable PDF actually is, because it is widely misunderstood. It contains the page image exactly as before, plus a layer of invisible text positioned over the words. Your viewer draws the picture and searches the text. This means the recognition errors are still in the file, invisible, and a search for a word the OCR misread will not find it even though you can plainly see the word on screen. Searchable does not mean corrected.

Frequently asked questions

How do I turn a photo of a document into editable text?

Choose the document photo, pick a cleanup (Black & white works best for text), and press Extract text.

Does the OCR work offline?

After the first run downloads the English language pack, yes. Recognition then runs locally without a connection.

Is my document uploaded to a server?

No. Cleanup and text recognition both happen inside your browser tab.

What makes OCR more accurate?

A flat, well-lit, in-focus photo, cropped to the page. The Black & white cleanup usually gives the recogniser the cleanest input.

Can I edit the extracted text?

Yes. The text appears in an editable box so you can correct it, then download it as a .txt file.

Related tools

Your files are processed on your device and are not uploaded. You can verify this in your browser’s developer tools under the Network tab.