Skip to content
Article

How to Extract Tables From PDF to Excel

PDF into Excel: which sheets end up in the downloaded workbook, how the three detection modes differ in practice, and why loose mode finds more tables but breaks words apart.

In short: To move tables from a PDF into Excel, open pdf-to-excel and leave detection on auto. The workbook holds a contents sheet, a text sheet, an image of every page, and a separate sheet of real cells for every table found.

Cluster

how-to guide

Step-by-step instructions for getting a PDF task from input to a reliable result.

13 articles

Primary tool

PDF to Excel online

Open the tool from this article and complete the operation in the current locale.

Open tool

Table of contents

How to Move Tables from a PDF into Excel

A report holds a table, and it has to be recalculated rather than retyped. The tool does more for that than its name suggests: it does not just dump text, it assembles a workbook of several sheets with different jobs. Understanding how that workbook is built pays off before the first run.

What is actually in the downloaded file

The result is not one sheet of data. It is a workbook whose sheets play different roles.

SheetWhat is on it
ContentsAn index of pages with links to the other sheets
TextThe document's text, one row per page
Page NAn image of the whole PDF page
Table N.MA detected table as real cells

The last row is the point of the exercise. Table sheets hold data you can put formulas on. The page images and the text sheet are insurance in case a table was not detected: the content stays in front of you either way.

The contents sheet works like a table of contents: the page number, a link to that page's sheet, and the first line of its text for orientation. The links jump around inside the workbook.

Three detection modes

The detection parameter changes how the tool looks for cell boundaries, and that changes the result more than anything else does.

ModeHow it finds a tableWhat the workbook holds
AutoBy the ruling linesEvery sheet, the text one included
StrictBy lines, with hard requirements on how they intersectNo text sheet at all
LooseBy where the text sits, with no reliance on linesEvery sheet, plus the page's full text in a cell comment

Auto is the default. Strict differs from it in more than fussiness about lines: it drops the text sheet from the workbook, leaving only the tables and the images. That makes sense when the result feeds further processing and spare sheets get in the way.

Loose is the only mode that sees tables with no ruling. That comes at a price, and the price is high.

It has a side effect that is sometimes more useful than its main one: in loose mode the page's entire text also lands in a comment attached to the first cell of the image sheet. Hovering over it shows the page's content without switching to the text sheet.

What the measurement showed

The subject was a document whose first page carries two tables: one with ruling lines, the other with none, just aligned columns.

Auto and strict found only the first, and carried it across perfectly: three rows, three columns, values in place.

Loose found both, but gathered them into one grid and broke words across columns on the way. The word "Region" became two cells, "Re" and "gion". "North" turned into "No" and "rth". The number 900 split into "9" and "00". Empty rows appeared between the data rows.

The practical conclusion: start with auto. If no table is found and the document genuinely has no ruling, switch to loose, but budget time for assembling it by hand. It gives you material, not a finished result.

And the reverse: if the table in the PDF is drawn with ruling and the result is still empty, the mode is not the problem. The document most likely has no text layer.

How to move the table

1. Open pdf-to-excel and upload the document. 2. Leave detection on auto for the first attempt. 3. Run the job and download the workbook. 4. Open the contents sheet and jump to the table sheets. 5. Check the numbers in a couple of rows against the page image on the neighbouring sheet. 6. If no tables turned up, repeat in loose mode.

Step five exists precisely because the page image is in the same workbook. Checking what came across against the original never requires leaving Excel.

Merged cells and sorting

The merge switch is on by default. It concerns cells that span several columns in the source table: switched on, it restores them as genuine merged ranges, reproducing the look of the original.

Switched off, it breaks such cells apart into ordinary ones: the value stays in the first, the rest become empty.

Choose by what you intend to do with the data.

If the table is going to be read and printed, leave merging on: the result matches the original.

If the data is going to be sorted, filtered or pivoted, turn it off. Excel refuses to sort a range containing merged cells and demands they all be made the same size. Unpicking merges by hand afterwards takes longer than flipping the switch beforehand.

Why the file is heavier than expected

Every PDF page enters the workbook as its own sheet carrying a full size image, rendered at nearly double the resolution. Those pictures are almost the entire weight of the result.

On a small test document the gap is especially visible: the source PDF weighed 2 KB and the workbook came out around 45 KB. On real reports the count runs into megabytes, and it grows with the number of pages rather than the number of tables.

If only the data matters and the images are in the way, the spare sheets are deleted inside Excel in a couple of moves, and the file slims down when saved.

A scan, and a table with no text layer

The tool reads text, it does not recognise it. A scanned table is a picture to it: the page image will be in the workbook, the text sheet will be empty, and no mode will find any tables.

The cure comes before the conversion: run the file through ocr-pdf, get a text layer, and convert to a spreadsheet after that. The other order is useless, because there is nothing left for Excel to recognise.

The same trick helps with documents whose text exists but sits crooked: a page shot at an angle recognises worse than a straight one, and straightening it before recognition changes the result noticeably.

Limits, warnings and the neighbouring tools

If table detection fails on some page, the job does not collapse. The workbook arrives whole and the response carries a warning naming those pages: no tables will come from them, the image and the text remain. The same applies to pages whose text could not be read.

A document under an open password cannot be processed; lift the protection with unlock-pdf. A password that only restricts printing or copying does not interfere with table extraction.

When one page holds several tables, each gets its own sheet named after the page number and the table's position on it. Excel caps sheet name length, so long labels are trimmed and collisions are separated with a numeric suffix.

If what you need from the document is text to edit rather than tables, that is pdf-to-word in flowing mode: the same table detector carries them across as genuine Word tables there. And turning the finished workbook back into a PDF is excel-to-pdf.

FAQ

First check whether the table has ruling lines: auto and strict look for cell boundaries along the lines. A table with no ruling is visible only to loose mode. If the ruling is there and the result is still empty, the document most likely has no text layer and needs OCR.
It relies on where the text sits rather than on lines, and in the measurement it broke words across columns: Region became Re and gion, and the number 900 split into 9 and 00. It gives you material to assemble by hand, not a finished table.
Because of merged cells: the merge switch is on by default and restores cells that spanned several columns. If the data is going to be sorted or pivoted, turn it off before processing.
Every PDF page enters it as its own sheet carrying a full size image. Those pictures are almost the entire weight. If only the data matters, the spare sheets are deleted inside Excel and the file slims down when saved.
Files are used only for the selected operation and are automatically deleted after processing is finished. We do not use uploaded documents to train AI models.

More from this cluster

Related tools

← All Convert tools

What to do next

If you need a practical next step or service guidance after reading, open these pages.

All tools

PDF tools catalog: merge, compress, split, convert, rotate, protect and unlock PDF files online, all directly in your browser.

FAQ

Answers to common questions about iHatePDF: whether registration is required, how files are processed, where to check limits, and whether it's safe to upload documents.

Contact

Contact iHatePDF about processing errors, choosing a tool, security, business inquiries, and suggestions for new features.