PDF to Excel: How Table Extraction Actually Works
You can see the table in a PDF: rows, columns, a header, a total at the bottom. It seems obvious that a converter should lift it straight into a spreadsheet. The trouble is that the PDF does not contain a table. It contains instructions to draw characters at certain positions and lines at others, and the grid you see is something your eyes put together. Every PDF to Excel converter has to rebuild the table from those scattered pieces, which is why the results range from perfect to unusable. This guide explains how that rebuilding works, where it fails, and how to get clean numbers into Excel either way.
What Is Actually Inside a PDF Table
Take an invoice row with three columns: "Consulting", "12 Mar 2026" and "1,250.00". Inside the PDF, that is a set of separate text fragments, each with an x and y coordinate. The fragments may not match the words: one program writes "1,250.00" as a single piece, another as "1,2" and "50.00", and another places each character on its own. They do not have to appear in reading order either. Some programs draw a whole table one column at a time.
The lines of the grid are separate drawing instructions, unconnected to the text. A thick line may be drawn as many short segments, a shaded row may be a filled rectangle with no lines at all, and many tables have no lines whatsoever. Nothing in the file says "this is a cell", "these two cells are merged" or "this row belongs to the table".
The Two Ways Converters Find a Table
Table extraction tools generally use one of two methods. Open-source tools such as Tabula and Camelot call them "lattice" and "stream".
- Following the ruling lines (lattice). The converter collects the horizontal and vertical lines, works out where they cross, and treats each enclosed area as a cell. Then it drops every piece of text into the cell it falls inside. This works very well on tables with a full grid. It fails when there are no lines, when lines are broken into fragments that do not quite meet, or when shading is used instead of borders.
- Following the whitespace (stream). The converter groups text into rows by height on the page, then looks for vertical gaps that run consistently down the rows and treats them as column boundaries. This handles borderless tables, such as many financial statements. It struggles when columns are close together, when a column is centered rather than aligned, or when a long value in one row bridges the gap that separates two columns in all the others.
Both methods are heuristics. They are reliable on clean, regular tables and unreliable on everything else, which is why no converter gets every table right.
Where Table Extraction Goes Wrong
Even when the columns are found correctly, some problems come up again and again:
- Wrapped cells. A description that wraps onto a second line looks like a new row with only one column filled in.
- Merged and spanning headings. A heading that sits over two columns has to be assigned to one of them, or to neither.
- Page breaks. A table that continues on the next page repeats its header, and a row can be split across the break.
- Numbers that arrive as text. Currency symbols, thousands separators, negatives shown in parentheses and footnote markers attached to a figure can all stop Excel from recognizing a number. A decimal comma, as in "1.250,00", is read correctly only if your Excel uses the same convention.
- Ambiguous dates. 03/04/2026 is the 3rd of April in some countries and the 4th of March in others, and Excel will pick one based on your settings.
- Fonts with no text mapping. Some PDFs embed fonts without the table that maps each glyph back to a character. They display perfectly, but copying or extracting their text produces gibberish.
- Scans. A scanned page has no text at all, only a picture of it, so there is nothing to extract without OCR first.
What Our PDF to Excel Tool Does
Our PDF to Excel converter runs in your browser, so the file is not uploaded. It gives each page of the PDF its own worksheet, named Page 1, Page 2 and so on. Down column A it writes every line of text on the page, from top to bottom, one line per row, and it places each picture at the point in the page where it was drawn.
It does not try to guess where the columns are. That keeps every line, in reading order, with nothing dropped or put in the wrong column by a bad guess. But it also means a table row arrives as one cell, with its values separated by single spaces, and you split it into columns yourself. If the PDF drew its table one column at a time, the values of one row can arrive in consecutive cells instead.
There is no OCR, so a scanned page comes through as a picture rather than text. A PDF that needs a password to open has to be unlocked first.
Splitting One Column Into Real Columns in Excel
Excel has good tools for the second half of the job:
- Text to Columns, on the Data tab, splits a column at a delimiter. Choose Delimited, then Space. This works well when the values themselves contain no spaces, which is common in tables of numbers and codes. Under Advanced, on the last step, you can set the decimal and thousands separators, which fixes "1.250,00" style figures.
- Flash Fill (Ctrl+E) learns from an example. Type the value you want from the first row into the next column, press Ctrl+E, and Excel fills the rest by following the same pattern.
- Formulas, in Excel for Microsoft 365.
=TEXTSPLIT(A2," ")spreads a line across columns, and=TEXTAFTER(A2," ",-1)returns everything after the last space, which is usually the amount at the end of a line. Wrap it inVALUE()to turn it into a number.
Excel for Microsoft 365 on Windows can also import a PDF directly, through Data, Get Data, From File, From PDF. It looks for tables itself and is worth trying on tables with clear borders.
Other Routes Worth Trying
The most reliable way to get a table out of a PDF is not to get it out of the PDF at all. The figures started life in a spreadsheet, an accounting system or a database, and whoever made the PDF can usually export the same data as a spreadsheet or a CSV file. Many banks, for example, let you download a statement as CSV as well as PDF.
For a table with a full grid of lines, PDF to Word is worth a try. When a table comes through it as a real Word table, you can copy it into Excel and the cells stay cells. That is also the route for a scanned table, since PDF to Word runs OCR on scanned pages.
Check the Result Before You Use It
Extraction errors are quiet. A value in the wrong column looks just as valid as one in the right column. Before you rely on extracted data:
- Add up a column with
SUMand compare it with the total printed in the PDF. If they match, the figures almost certainly came through correctly. This is the single best check there is. - Count the rows against the PDF, and look closely at the first and last rows on each page, where page breaks cause trouble.
- Look for numbers stored as text. By default Excel aligns them to the left of the cell and marks them with a small green triangle.
- Check a few dates by hand, especially any where the day is 12 or less.
Conclusion
A PDF table is a picture of a table, made of text and lines that only look connected. Converters rebuild it from ruling lines or from whitespace, and both methods fail on irregular layouts. Knowing that, you can choose the right route: ask for the source data when you can, extract the lines and split them in Excel when you cannot, and always check the totals before the numbers go anywhere important.
Need a PDF's text in a spreadsheet? Try PDF to Excel.
← Back to the blog