How to Extract Tables From PDF to Excel
In short: To move tables from a PDF into Excel, open pdf-to-excel and leave detection on auto. The workbook holds a contents sheet, a text sheet, an image of every page, and a separate sheet of real cells for every table found.
Cluster
how-to guide
Step-by-step instructions for getting a PDF task from input to a reliable result.
Primary tool
PDF to Excel online
Open the tool from this article and complete the operation in the current locale.
Open toolTable of contents
How to Move Tables from a PDF into Excel
A report holds a table, and it has to be recalculated rather than retyped. The tool does more for that than its name suggests: it does not just dump text, it assembles a workbook of several sheets with different jobs. Understanding how that workbook is built pays off before the first run.
What is actually in the downloaded file
The result is not one sheet of data. It is a workbook whose sheets play different roles.
| Sheet | What is on it |
|---|---|
| Contents | An index of pages with links to the other sheets |
| Text | The document's text, one row per page |
| Page N | An image of the whole PDF page |
| Table N.M | A detected table as real cells |
The last row is the point of the exercise. Table sheets hold data you can put formulas on. The page images and the text sheet are insurance in case a table was not detected: the content stays in front of you either way.
The contents sheet works like a table of contents: the page number, a link to that page's sheet, and the first line of its text for orientation. The links jump around inside the workbook.
Three detection modes
The detection parameter changes how the tool looks for cell boundaries, and that changes the result more than anything else does.
| Mode | How it finds a table | What the workbook holds |
|---|---|---|
| Auto | By the ruling lines | Every sheet, the text one included |
| Strict | By lines, with hard requirements on how they intersect | No text sheet at all |
| Loose | By where the text sits, with no reliance on lines | Every sheet, plus the page's full text in a cell comment |
Auto is the default. Strict differs from it in more than fussiness about lines: it drops the text sheet from the workbook, leaving only the tables and the images. That makes sense when the result feeds further processing and spare sheets get in the way.
Loose is the only mode that sees tables with no ruling. That comes at a price, and the price is high.
It has a side effect that is sometimes more useful than its main one: in loose mode the page's entire text also lands in a comment attached to the first cell of the image sheet. Hovering over it shows the page's content without switching to the text sheet.
What the measurement showed
The subject was a document whose first page carries two tables: one with ruling lines, the other with none, just aligned columns.
Auto and strict found only the first, and carried it across perfectly: three rows, three columns, values in place.
Loose found both, but gathered them into one grid and broke words across columns on the way. The word "Region" became two cells, "Re" and "gion". "North" turned into "No" and "rth". The number 900 split into "9" and "00". Empty rows appeared between the data rows.
The practical conclusion: start with auto. If no table is found and the document genuinely has no ruling, switch to loose, but budget time for assembling it by hand. It gives you material, not a finished result.
And the reverse: if the table in the PDF is drawn with ruling and the result is still empty, the mode is not the problem. The document most likely has no text layer.
How to move the table
1. Open pdf-to-excel and upload the document. 2. Leave detection on auto for the first attempt. 3. Run the job and download the workbook. 4. Open the contents sheet and jump to the table sheets. 5. Check the numbers in a couple of rows against the page image on the neighbouring sheet. 6. If no tables turned up, repeat in loose mode.
Step five exists precisely because the page image is in the same workbook. Checking what came across against the original never requires leaving Excel.
Merged cells and sorting
The merge switch is on by default. It concerns cells that span several columns in the source table: switched on, it restores them as genuine merged ranges, reproducing the look of the original.
Switched off, it breaks such cells apart into ordinary ones: the value stays in the first, the rest become empty.
Choose by what you intend to do with the data.
If the table is going to be read and printed, leave merging on: the result matches the original.
If the data is going to be sorted, filtered or pivoted, turn it off. Excel refuses to sort a range containing merged cells and demands they all be made the same size. Unpicking merges by hand afterwards takes longer than flipping the switch beforehand.
Why the file is heavier than expected
Every PDF page enters the workbook as its own sheet carrying a full size image, rendered at nearly double the resolution. Those pictures are almost the entire weight of the result.
On a small test document the gap is especially visible: the source PDF weighed 2 KB and the workbook came out around 45 KB. On real reports the count runs into megabytes, and it grows with the number of pages rather than the number of tables.
If only the data matters and the images are in the way, the spare sheets are deleted inside Excel in a couple of moves, and the file slims down when saved.
A scan, and a table with no text layer
The tool reads text, it does not recognise it. A scanned table is a picture to it: the page image will be in the workbook, the text sheet will be empty, and no mode will find any tables.
The cure comes before the conversion: run the file through ocr-pdf, get a text layer, and convert to a spreadsheet after that. The other order is useless, because there is nothing left for Excel to recognise.
The same trick helps with documents whose text exists but sits crooked: a page shot at an angle recognises worse than a straight one, and straightening it before recognition changes the result noticeably.
Limits, warnings and the neighbouring tools
If table detection fails on some page, the job does not collapse. The workbook arrives whole and the response carries a warning naming those pages: no tables will come from them, the image and the text remain. The same applies to pages whose text could not be read.
A document under an open password cannot be processed; lift the protection with unlock-pdf. A password that only restricts printing or copying does not interfere with table extraction.
When one page holds several tables, each gets its own sheet named after the page number and the table's position on it. Excel caps sheet name length, so long labels are trimmed and collisions are separated with a numeric suffix.
If what you need from the document is text to edit rather than tables, that is pdf-to-word in flowing mode: the same table detector carries them across as genuine Word tables there. And turning the finished workbook back into a PDF is excel-to-pdf.
FAQ
More from this cluster
how-to guide
How to convert PDF to editable Word
The file opened in Word but the cursor will not go into the text? Why the default gives you a picture, how flowing mode differs from exact, and where the running headers went.
how-to guide
How to Extract PDF Pages to a New File
Pulling the sheets you need out of a PDF as a separate file: how the two output modes differ, why the files in the archive keep the original page numbers, and what turning off size preservation does.
how-to guide
How to Crop Margins and Clutter in PDF
Cropping changes only the visible area of a page while the content beyond it stays in the file. What that means in practice, why the margins apply to the whole document at once, and how cropping behaves on rotated pages.
Related tools
PDF to Excel online
Extract table data from a PDF into Excel so it can be calculated, filtered and sorted.
Excel to PDF online
Convert an Excel spreadsheet to PDF to lock its look and print or send it easily.
OCR PDF online
OCR PDF keeps the PDF task in one browser flow: upload the source file, check options, run processing, and download the result.
PDF to Word online
Convert a PDF to a Word document so you can edit the text, extend sections and make changes.
What to do next
All tools
PDF tools catalog: merge, compress, split, convert, rotate, protect and unlock PDF files online, all directly in your browser.
FAQ
Answers to common questions about iHatePDF: whether registration is required, how files are processed, where to check limits, and whether it's safe to upload documents.
Contact
Contact iHatePDF about processing errors, choosing a tool, security, business inquiries, and suggestions for new features.