How to convert PDF to editable Word
In short: To get a Word file you can actually edit, switch the result type in pdf-to-word to flowing: the default is exact, which drops every page in as an image with a hidden searchable layer. Flowing mode rebuilds paragraphs, headings and genuine Word tables.
Cluster
how-to guide
Step-by-step instructions for getting a PDF task from input to a reliable result.
Primary tool
PDF to Word online
Open the tool from this article and complete the operation in the current locale.
Open toolTable of contents
Convert a PDF into an Editable Word Document
The request is nearly always the same: a couple of paragraphs in a PDF someone sent need fixing. And the first attempt nearly always disappoints. The file opens in Word, looks perfect, and the cursor will not go into the text. That is not a conversion failure but a consequence of the setting that ships as the default.
By default you get a picture, not text
The tool has two modes, and the switch between them is called the result type. Exact is preselected, and it does exactly what it promises: it preserves the look of the page letter for letter. The method is blunt about it. Each PDF page is rendered to an image and dropped into the Word document whole.
Hence the feeling of a bait and switch. It is a .docx, it opens in Word, the pages match the original in size and orientation, and there is nothing to edit: what sits on the page is pictures.
The second mode, flowing, takes the page apart and rebuilds a real Word document out of the pieces: paragraphs, headings, lists and tables. What it gives up in exchange is the exact placement of everything on the page.
| Mode | What is inside | What you can do |
|---|---|---|
| Exact | An image of every page plus a hidden text layer | Read, search the text, print |
| Flowing | Word paragraphs, headings, lists and tables | Edit the text, restyle it, work with tables |
The choice is not about quality but about the job. A document that has to travel onward in a familiar office format and stay untouched wants exact. A document that has to change wants flowing, and no amount of tuning the exact mode will produce editable text.
What is inside the exact file
The page image is not all of the content. Underneath it goes the text pulled out of the PDF, marked as hidden: it is absent on screen and in print, but a search through the document finds it.
That makes the exact mode useful where it looks pointless. A contract that has to be circulated and later searched for wording works fine this way: the look is preserved, search works, and nobody edits it by accident. The tool does say so itself after processing, warning that the document is not editable text, and that message is worth reading rather than dismissing.
Page size and orientation are taken from the source PDF page by page. A landscape insert in the middle of a portrait document stays landscape.
How to get a document that edits
1. Open pdf-to-word and upload the file. 2. Switch the result type to flowing. This is the one action that matters, and the default does not do it for you. 3. Decide whether you want page previews, covered below, and turn them off if you do not. 4. Run the job and download the .docx. 5. Open the document and walk the navigation pane: it shows which headings were recognised. 6. Check the tables and the first page, where losses show up most often.
Running headers vanish on purpose
Flowing mode runs a separate pass across the whole document looking for repeated lines. Two conditions apply: the line has to sit in the top or bottom tenth of the page, and it has to repeat on at least three pages and on at least half the document. Anything meeting both is dropped from the result.
The comparison is not naive. Runs of digits collapse, so "Page 4" and "Page 5" count as the same line. But the collapsing only kicks in when letters survive in the line: a bare number at the foot of the page is compared exactly, and three different digits on three pages are not equal to each other. In practice a footer reading "Confidential. Page 7" goes away, while a footer that is nothing but a numeral stays behind as a paragraph on every page.
Tested on a four page report: the line `Confidential report Page N` is absent from the flowing result and present in the exact one, where the page is a picture anyway and there is nothing to remove from it.
The mechanism has a downside too. A document title set large at the very top of every page, identical throughout, falls under the same rule and gets deleted along with the running headers. When that top line matters, typing it back is quicker than hunting for the cause.
Headings are decided by font size
Heading styles in the result are not invented. The tool first works out, across the whole document, which font size is the body size. It counts characters rather than lines, so a rare large caption never wins the vote.
Every block is then measured against that size. Roughly 15 percent larger than the body makes a level two heading, half again as large or more makes level one. Everything else stays an ordinary paragraph. The text also has to look like a heading: a long sentence ending in a full stop will not become one, whatever type it is set in.
That produces predictable behaviour on documents laid out without styles. If the whole text is set in one size and the sections are marked only by bold, the Word file will have no headings at all: the size does not differ, so there is nothing to go on. The navigation pane comes out empty and the structure has to be added by hand.
Tables come out as tables
Tables are found by a separate detector, the same one pdf-to-excel is built on. A detected table lands in the document as a genuine Word table with visible gridlines rather than as paragraphs padded with spaces.
Text blocks that fall inside the bounds of a detected table are dropped from the paragraph stream; otherwise the cell contents would appear twice, once as a table and once as prose. The threshold is an overlap of six tenths of the block's area.
Where the table lands is decided by its top edge: it is flushed in before the first paragraph that starts below it. On an ordinary single column page that gives the right order. On a two column page the table can end up somewhere other than where it stood, because the paragraphs of the two columns run sequentially in the file while their heights interleave.
If table detection fails on some page, the job does not collapse: that page comes through as ordinary text and the response carries a warning naming the pages. A silent empty answer is not possible here.
File weight and page previews
Flowing mode adds an image of each page after that page's text by default, so there is something to check the result against. Nothing else affects the weight nearly as much.
Measured on one and the same four page document:
| Result | Size |
|---|---|
| Exact mode | 671 KB |
| Flowing with page previews | 315 KB |
| Flowing without previews | 37 KB |
The source PDF weighed about 4 KB. The gap between the outer rows is eighteenfold, and all of it is pictures. If the document is going into real work rather than into a comparison, turn the previews off straight away: they do not edit, they take no part in the text and they get in the way of proofreading.
Exact mode ignores the previews setting entirely, since its page is an image already.
A scan with no text layer
A scanned document is a picture, and there is nothing in it to extract. In exact mode the result is that same scan repackaged as a .docx. In flowing mode the pages come out empty, with a note in place of each one saying the page was blank.
The cure comes before the conversion, not after: run the file through ocr-pdf, get a PDF with a text layer, and only then convert to Word. The other order does not work, because there will be nothing left for recognition to work on.
A document locked with an open password will not convert either; the job stops and asks for the password. Lift the protection with unlock-pdf. A password that only restricts printing or copying does not interfere.
Once the edits are in, the document usually has to go back to PDF. That is word-to-pdf, which has its own font substitution behaviour worth knowing about in advance.
FAQ
More from this cluster
how-to guide
How to Extract Tables From PDF to Excel
PDF into Excel: which sheets end up in the downloaded workbook, how the three detection modes differ in practice, and why loose mode finds more tables but breaks words apart.
how-to guide
How to Extract PDF Pages to a New File
Pulling the sheets you need out of a PDF as a separate file: how the two output modes differ, why the files in the archive keep the original page numbers, and what turning off size preservation does.
how-to guide
How to Crop Margins and Clutter in PDF
Cropping changes only the visible area of a page while the content beyond it stays in the file. What that means in practice, why the margins apply to the whole document at once, and how cropping behaves on rotated pages.
Related tools
PDF to Word online
Convert a PDF to a Word document so you can edit the text, extend sections and make changes.
Word to PDF online
Convert a Word document to PDF to lock the layout and be sure the file opens the same way everywhere.
PDF to Excel online
Extract table data from a PDF into Excel so it can be calculated, filtered and sorted.
OCR PDF online
OCR PDF keeps the PDF task in one browser flow: upload the source file, check options, run processing, and download the result.
What to do next
All tools
PDF tools catalog: merge, compress, split, convert, rotate, protect and unlock PDF files online, all directly in your browser.
FAQ
Answers to common questions about iHatePDF: whether registration is required, how files are processed, where to check limits, and whether it's safe to upload documents.
Contact
Contact iHatePDF about processing errors, choosing a tool, security, business inquiries, and suggestions for new features.