Skip to main content
PDF
PDFixel Editorial Team••6 min read

Why You Cannot Edit Text in a Scanned PDF

A scanned page is a photograph, so there is no text to edit. How to tell the difference in seconds, and what OCR can and cannot recover for you.

Direct Answer & Key Takeaway

A scanned PDF contains a photograph of a page, not text, so a PDF editor finds nothing to edit — the characters you can see are pixels, not glyphs. To make the words editable you first need OCR (optical character recognition) to read the image and produce a text layer. You can check which kind of file you have in about five seconds: if you cannot select a word with your cursor, there is no text in the document.

You open a PDF, the text is right there on screen, and your editor refuses to let you change a single character. Nothing is broken and you have not missed a setting. The file simply does not contain what you think it contains, and understanding why takes about two minutes and saves a lot of pointless clicking.

Two Documents That Look Identical

PDF is a container format, and two files that look the same on screen can hold completely different things.

A digital PDF — exported from Word, a browser, a design tool or an accounting system — stores text as text. Each character is a glyph reference with a position and a font. An editor can find those instructions, change them and write them back, which is why editing that kind of file feels ordinary.

A scanned PDF holds a photograph. A scanner or phone camera captured the page, and that image was wrapped in a PDF so it could be shared as a document. What you are reading is an arrangement of dark and light pixels that your eyes resolve into letters. There is no letter "e" anywhere in the file, only pixels that look like one.

This is why an editor appears to do nothing. It is not failing to edit the text; it is correctly reporting that there is no text. The same rule explains why you cannot search a scan, why copying from it produces nothing, and why a screen reader has nothing to read out.

A file being large or slow to open is not what makes it a scan. A one-page photograph can be heavier than a fifty-page text document, because a page of text is a few kilobytes of instructions while a page of pixels is a picture.

The Five-Second Test

Before reaching for any tool, find out which kind of file you have. Open it in any viewer and try to select a word by dragging across it.

If a word highlights, the document has a text layer and an editor will work on it normally. If your cursor draws a selection box over the area but no characters highlight, you have a scan.

Two other checks confirm it. Search for a word you can clearly see on the page: a text PDF finds it, a scan reports nothing. And try to copy a sentence — on a scan, nothing arrives on your clipboard.

One case sits between the two and causes real confusion. A searchable PDF is a scan that has already had OCR applied: the image is still what you see, but an invisible text layer sits behind it. Those files select, search and copy normally even though every visible pixel is still photographic.

  • Drag across a word — does it highlight, or do you just get an empty selection box?
  • Search for a word you can see on the page.
  • Copy a sentence and paste it somewhere.
  • Slightly skewed lines, a grey cast or speckling are usually signs of a camera or scanner original.

What OCR Actually Does

Optical character recognition looks at the image, works out which shapes are characters, and produces text from them. It does not alter the picture — it builds a parallel layer of words with positions, which is what turns a photograph of a page into something you can select, search, copy and edit.

PDFixel runs OCR through a WebAssembly build of Tesseract inside your browser, with support for more than fifty languages. Nothing is uploaded. The first run downloads the recognition engine, which is the largest component on the site, so expect a wait that once; afterwards it is cached.

You have two useful outputs. Exporting plain text hands you the words alone, ready to paste into a document and rework. Exporting a searchable PDF keeps the scan looking exactly as it did while adding the invisible text layer underneath, which is usually the better choice for archiving — the document still looks like the original, and every word in it becomes findable.

Set the recognition language before you run it. Tesseract uses language data to resolve ambiguous shapes, and running English data over a page of Hindi or Arabic produces confident nonsense rather than an error.

For anything you plan to keep rather than rewrite, export a searchable PDF instead of plain text. You keep the original appearance, including signatures and stamps, and gain search over the whole document.

What OCR Cannot Fix

OCR is genuinely useful and it is not magic. Being honest about the limits will save you from trusting a result you should have checked.

Accuracy is set almost entirely by the quality of the original. A clean, straight, evenly lit page reads close to perfectly. A photograph taken at an angle in poor light, a fax, a third-generation photocopy or anything with speckling produces text with errors scattered through it. When output is poor, retaking the photograph will do more for you than any setting will.

Layout is the harder problem. OCR reads characters well and reconstructs structure badly. Multi-column pages, tables, sidebars, footnotes and text wrapped around images often come back in an order that does not match how a human reads the page. Numbers inside tables are the most dangerous case, because a plausible-looking figure in the wrong column reads as correct.

Handwriting is largely out of scope. Tesseract is built for printed type; cursive and casual handwriting are not reliably recognised, and confident wrong answers are the usual outcome rather than a clean failure.

Always proofread OCR output before relying on it, particularly anything with figures in it. Treat the result as a good draft rather than a finished transcription.

Editing a Scan Without OCR

Sometimes you do not need the underlying text at all. If the goal is to sign a scanned contract, fill in a scanned form, black out a name or add a note, you can work straight on top of the image without recognising anything.

An editor can place new text boxes, signatures, shapes and stamps over a scanned page perfectly well. What it cannot do is change the words already photographed into the page, because those are pixels in an image.

Redaction is the one case where this distinction works in your favour. Covering part of a scan with a black box genuinely does hide it, because there is no text layer hiding underneath to be recovered. On a digital PDF the same black box leaves the words fully extractable — the difference catches people out in both directions.

So the decision comes down to what you are trying to achieve: adding to the page needs no OCR, while changing what the page already says does.

  • Signing, form filling, stamping and annotating all work directly on a scan.
  • Page-level work — reordering, rotating, deleting, merging, splitting — never needs OCR.
  • Changing words that were photographed into the page does need OCR first.
  • Making a scan searchable for later needs OCR with a searchable-PDF export.

Summary

A scanned PDF holds a photograph of a page rather than text, so there are no characters for an editor to change and nothing to search or copy. Check which kind of file you have by trying to select a word: an empty selection box means a scan. OCR reads the image and builds a text layer, and exporting a searchable PDF keeps the original appearance while making every word findable — but accuracy depends on the quality of the scan, table and multi-column layouts are reconstructed unreliably, and handwriting is mostly beyond it, so proofread before you trust figures. If you only need to sign, fill, stamp or annotate the page, you can do all of that on the image without OCR at all.

Need to do this now?

Scanned PDF to Text

Run OCR on scanned or image-only PDFs in your browser to turn pictures of words into text you can edit.

Open Tool

Related Articles

View all →