Skip to content
Article

How to convert PDF to editable Word

The file opened in Word but the cursor will not go into the text? Why the default gives you a picture, how flowing mode differs from exact, and where the running headers went.

In short: To get a Word file you can actually edit, switch the result type in pdf-to-word to flowing: the default is exact, which drops every page in as an image with a hidden searchable layer. Flowing mode rebuilds paragraphs, headings and genuine Word tables.

Cluster

how-to guide

Step-by-step instructions for getting a PDF task from input to a reliable result.

13 articles

Primary tool

PDF to Word online

Open the tool from this article and complete the operation in the current locale.

Open tool

Table of contents

Convert a PDF into an Editable Word Document

The request is nearly always the same: a couple of paragraphs in a PDF someone sent need fixing. And the first attempt nearly always disappoints. The file opens in Word, looks perfect, and the cursor will not go into the text. That is not a conversion failure but a consequence of the setting that ships as the default.

By default you get a picture, not text

The tool has two modes, and the switch between them is called the result type. Exact is preselected, and it does exactly what it promises: it preserves the look of the page letter for letter. The method is blunt about it. Each PDF page is rendered to an image and dropped into the Word document whole.

Hence the feeling of a bait and switch. It is a .docx, it opens in Word, the pages match the original in size and orientation, and there is nothing to edit: what sits on the page is pictures.

The second mode, flowing, takes the page apart and rebuilds a real Word document out of the pieces: paragraphs, headings, lists and tables. What it gives up in exchange is the exact placement of everything on the page.

ModeWhat is insideWhat you can do
ExactAn image of every page plus a hidden text layerRead, search the text, print
FlowingWord paragraphs, headings, lists and tablesEdit the text, restyle it, work with tables

The choice is not about quality but about the job. A document that has to travel onward in a familiar office format and stay untouched wants exact. A document that has to change wants flowing, and no amount of tuning the exact mode will produce editable text.

What is inside the exact file

The page image is not all of the content. Underneath it goes the text pulled out of the PDF, marked as hidden: it is absent on screen and in print, but a search through the document finds it.

That makes the exact mode useful where it looks pointless. A contract that has to be circulated and later searched for wording works fine this way: the look is preserved, search works, and nobody edits it by accident. The tool does say so itself after processing, warning that the document is not editable text, and that message is worth reading rather than dismissing.

Page size and orientation are taken from the source PDF page by page. A landscape insert in the middle of a portrait document stays landscape.

How to get a document that edits

1. Open pdf-to-word and upload the file. 2. Switch the result type to flowing. This is the one action that matters, and the default does not do it for you. 3. Decide whether you want page previews, covered below, and turn them off if you do not. 4. Run the job and download the .docx. 5. Open the document and walk the navigation pane: it shows which headings were recognised. 6. Check the tables and the first page, where losses show up most often.

Running headers vanish on purpose

Flowing mode runs a separate pass across the whole document looking for repeated lines. Two conditions apply: the line has to sit in the top or bottom tenth of the page, and it has to repeat on at least three pages and on at least half the document. Anything meeting both is dropped from the result.

The comparison is not naive. Runs of digits collapse, so "Page 4" and "Page 5" count as the same line. But the collapsing only kicks in when letters survive in the line: a bare number at the foot of the page is compared exactly, and three different digits on three pages are not equal to each other. In practice a footer reading "Confidential. Page 7" goes away, while a footer that is nothing but a numeral stays behind as a paragraph on every page.

Tested on a four page report: the line `Confidential report Page N` is absent from the flowing result and present in the exact one, where the page is a picture anyway and there is nothing to remove from it.

The mechanism has a downside too. A document title set large at the very top of every page, identical throughout, falls under the same rule and gets deleted along with the running headers. When that top line matters, typing it back is quicker than hunting for the cause.

Headings are decided by font size

Heading styles in the result are not invented. The tool first works out, across the whole document, which font size is the body size. It counts characters rather than lines, so a rare large caption never wins the vote.

Every block is then measured against that size. Roughly 15 percent larger than the body makes a level two heading, half again as large or more makes level one. Everything else stays an ordinary paragraph. The text also has to look like a heading: a long sentence ending in a full stop will not become one, whatever type it is set in.

That produces predictable behaviour on documents laid out without styles. If the whole text is set in one size and the sections are marked only by bold, the Word file will have no headings at all: the size does not differ, so there is nothing to go on. The navigation pane comes out empty and the structure has to be added by hand.

Tables come out as tables

Tables are found by a separate detector, the same one pdf-to-excel is built on. A detected table lands in the document as a genuine Word table with visible gridlines rather than as paragraphs padded with spaces.

Text blocks that fall inside the bounds of a detected table are dropped from the paragraph stream; otherwise the cell contents would appear twice, once as a table and once as prose. The threshold is an overlap of six tenths of the block's area.

Where the table lands is decided by its top edge: it is flushed in before the first paragraph that starts below it. On an ordinary single column page that gives the right order. On a two column page the table can end up somewhere other than where it stood, because the paragraphs of the two columns run sequentially in the file while their heights interleave.

If table detection fails on some page, the job does not collapse: that page comes through as ordinary text and the response carries a warning naming the pages. A silent empty answer is not possible here.

File weight and page previews

Flowing mode adds an image of each page after that page's text by default, so there is something to check the result against. Nothing else affects the weight nearly as much.

Measured on one and the same four page document:

ResultSize
Exact mode671 KB
Flowing with page previews315 KB
Flowing without previews37 KB

The source PDF weighed about 4 KB. The gap between the outer rows is eighteenfold, and all of it is pictures. If the document is going into real work rather than into a comparison, turn the previews off straight away: they do not edit, they take no part in the text and they get in the way of proofreading.

Exact mode ignores the previews setting entirely, since its page is an image already.

A scan with no text layer

A scanned document is a picture, and there is nothing in it to extract. In exact mode the result is that same scan repackaged as a .docx. In flowing mode the pages come out empty, with a note in place of each one saying the page was blank.

The cure comes before the conversion, not after: run the file through ocr-pdf, get a PDF with a text layer, and only then convert to Word. The other order does not work, because there will be nothing left for recognition to work on.

A document locked with an open password will not convert either; the job stops and asks for the password. Lift the protection with unlock-pdf. A password that only restricts printing or copying does not interfere.

Once the edits are in, the document usually has to go back to PDF. That is word-to-pdf, which has its own font substitution behaviour worth knowing about in advance.

FAQ

Because the result type defaults to exact: every page is inserted as an image with a hidden text layer underneath it for search. To get an editable document, switch the result type to flowing and convert again.
In flowing mode detected tables come through as genuine Word tables with visible gridlines. If table detection fails on a page, its contents come through as ordinary text and the response carries a warning naming those pages.
Flowing mode removes lines that sit in the top or bottom tenth of the page and repeat on at least three pages and on at least half the document. An identical document title set at the top of every page falls under the same rule.
Not directly. A scan has no text layer, so flowing mode returns empty pages. Run the file through ocr-pdf first, then convert the resulting PDF to Word.
Files are used only for the selected operation and are automatically deleted after processing is finished. We do not use uploaded documents to train AI models.

More from this cluster

Related tools

← All Convert tools

What to do next

If you need a practical next step or service guidance after reading, open these pages.

All tools

PDF tools catalog: merge, compress, split, convert, rotate, protect and unlock PDF files online, all directly in your browser.

FAQ

Answers to common questions about iHatePDF: whether registration is required, how files are processed, where to check limits, and whether it's safe to upload documents.

Contact

Contact iHatePDF about processing errors, choosing a tool, security, business inquiries, and suggestions for new features.