How to Convert a Scanned PDF to Word, and Why Results Vary

Two PDFs can look identical on screen and behave completely differently when you convert them to Word. One was saved from a word processor and contains real text. The other was made by a scanner or a phone camera, and each page is a photograph of text. To turn the second kind into an editable document, a converter has to read the letters off the picture first, then guess how the page was laid out. Both steps can go wrong, and this guide explains where, why, and what you can do about it.

How to Tell Whether Your PDF Is Scanned

The quickest test is to try to select a single word. If your cursor draws a box over the whole page instead of highlighting the word, the page is an image. Searching with Ctrl+F (Cmd+F on a Mac) is the second test: a scanned page finds nothing, even for a word you can see.

Zooming in helps too. Real text stays sharp at any zoom, because it is drawn from a font. Scanned text goes soft and blocky as you zoom, because it is made of pixels.

Some scans are a mixture. Scanner software often runs text recognition when it saves the file and hides the recognized text behind the picture, so you can select and search. That hidden text is only as good as the scanner's guess, though. Copy a paragraph into a plain text editor and you will see exactly what is there, misreadings included.

What OCR Actually Does

OCR, optical character recognition, is the process that turns a picture of text into characters. It works in stages. The image is cleaned up: straightened if it is tilted, and its contrast adjusted. The software then finds the regions that contain text, splits them into lines and words, and matches each shape against what it has learned letters look like, using a dictionary for the language to settle close calls.

The dictionary is important, because it can only help with words. If OCR reads "lnvoice" with a lowercase L, the dictionary nudges it back to "invoice". If it reads an account number as 1O45 instead of 1045, there is nothing to correct it against. The usual confusions are between characters that look alike in many fonts: 0 and O, 1 and l and I, 5 and S, 8 and B, and the pair "rn", which can become "m". That is why numbers, names and codes are where OCR errors hide, and why they deserve the closest check.

Why Results Vary So Much From Scan to Scan

The same converter can produce near-perfect text from one scan and a mess from another. The difference is almost always in the image:

  • Resolution. 300 DPI is the usual recommendation for scanning documents for OCR. Below about 200 DPI, small type loses the detail that separates similar letters. Very small print, such as footnotes and terms on the back of a form, benefits from 400 DPI or more.
  • Angle and curve. A page photographed at an angle, or a book page curving into its spine, gives lines of text that are not straight. Straightening helps, but curved lines often come out with words dropped or merged.
  • Contrast and background. Colored paper, highlighter, stamps across text, and faint copies of copies all make letters harder to separate from what is behind them.
  • Heavy compression. A scan saved at low JPEG quality has blurred edges and smudges around each letter. A fax-quality black-and-white image breaks thin strokes.
  • Handwriting. OCR used for document conversion is built for printed text. Handwritten notes and signatures usually come out wrong, or are left as pictures.
  • Language and layout. Recognition is trained per script and language, so less common scripts and mixed-language pages are harder. Multiple columns, tables without lines and text over pictures make the layout step harder still.

Why the Word File Can Look Different Even When the Text Is Right

Recognizing the characters is only half the job. A converter also has to rebuild the page as something Word understands, and it faces a trade-off. It can write the text as normal flowing paragraphs, which are easy to edit but drift from the original layout. Or it can pin each block of text to its exact position, which looks right but is awkward to edit, because typing in one block does not push the next one down.

Other differences come from information the scan simply does not contain. OCR cannot know which font was used, so Word gets a similar one and line breaks move. A table with ruled lines can be rebuilt as a Word table, but a table without lines may arrive as text separated by tabs or spaces. Headers, footers and page numbers can land in the body of the document, and logos, stamps and signatures stay as images.

Converting a Scan With PDF to Word

Our PDF to Word tool runs OCR automatically on the pages that need it, so a scan comes back as a .docx with editable text, and pages that already contain text are converted directly. The file is sent to a document conversion provider, which states that it deletes the file and the result within three hours. There is no set size limit, but the upload has to finish within about 40 seconds, so a very large scan needs a fast connection. A PDF that needs a password to open has to be unlocked first.

A few things are worth doing before you convert:

  • Turn any sideways or upside-down pages the right way up, and remove blank pages. File Organizer does both without uploading the file.
  • If you are scanning the paper yourself, use 300 DPI in grayscale or color rather than the scanner's black-and-white mode, which can break up thin letters before OCR ever sees them.
  • If you only need a few pages, convert only those. There is less to check afterwards.

What to Check Before You Rely on the Result

OCR makes mistakes on most real scans, and it gives no warning when it does. Before you send or file the converted document, check:

  • Every number that matters. Totals, dates, amounts, account and reference numbers, phone numbers. Compare them against the scan, not against what looks plausible.
  • Names and addresses. They are rarely in a dictionary, so nothing corrects them.
  • Words the spelling checker underlines. Red underlines cluster where OCR guessed wrong, which makes Word's spelling check a rough error finder.
  • Small print and margins. Footnotes, stamps and notes in the margin are the text most often missed altogether.
  • Tables. Make sure each value sits in the right row and column.

What Does Not Work Well

Some documents will disappoint any converter. Handwritten forms, faxes, screenshots of documents taken at screen resolution, and text that is part of a photo or a diagram usually need to be retyped where they matter.

Two of our other tools do not read scans at all. PDF to Excel only reads text that is already in the PDF, so it cannot pull a table out of a scanned page. Translate Document turns scanned PDFs away for the same reason. In both cases, you can convert the scan with PDF to Word first, check the text, and work from the Word file: copy a table into Excel, or translate the .docx.

Conclusion

Converting a scanned PDF to Word is two guesses in a row: what the letters are, and how the page was built. A sharp, straight, 300 DPI scan makes both guesses easier, and a careful check of numbers and names catches what OCR gets wrong. Treat the result as a first draft of the text, not a copy of the original, and you will rarely be caught out.

Have a scan to turn into text? Try the PDF to Word converter.

← Back to the blog