Skip to content
Article

HTML to PDF: Save a Web Page

The tool takes a saved HTML file rather than a page address, and renders it in an isolated browser. Hence the rules: local images are not pulled in, and anything that appears after loading never reaches the PDF.

In short: The page is rendered by a headless browser from the HTML file you upload. Local paths are not fetched at all, and external ones only for the addresses visible in the HTML itself. A reliable archive comes from one self-contained file with everything embedded inside.

Cluster

security

Protection, access control, redaction, and controlled sharing workflows.

4 articles

Primary tool

HTML to PDF online

Open the tool from this article and complete the operation in the current locale.

Open tool

Table of contents

Saving a web page as a PDF: what reaches the archive and what does not

You need a page kept so that a year from now it opens and looks the way it does today: terms of service, a tariff page, an article somebody might edit. PDF suits that better than a screenshot, because the text survives, and better than a bookmark, because it does not depend on somebody else's server.

But between "page in a browser" and "page in a PDF" there are a few steps where content goes missing. Here is where.

The tool takes a file, not a page address

The first thing to understand: you do not type a website address into it. You upload a saved HTML file.

Your browser can save the page. The save dialogue usually offers a choice between HTML only and an option that packs the whole page. The second is noticeably better: it gathers markup, styles and images into one file, and that is what an archive needs.

Why it works that way is explained in the next section, and the same explanation covers most surprising results.

Why images disappear

The page is rendered in an isolated browser, and the content is handed to it directly rather than opened from a file path. Such a page has no location on disk it could read neighbouring files from.

Hence a hard rule: references to local files do not work. If a folder called `page_files` full of images sits next to your `page.html`, not one of them reaches the PDF. The browser simply cannot get at them.

External addresses do work, but not any of them. Before rendering, the HTML is scanned statically, a list of addresses is built from it, and only those are fetched. A resource the markup says nothing about is blocked, and so is a redirect to another address.

There is one practical conclusion: a reliable archive comes from a self-contained file, where images and styles are embedded inside the markup rather than lying beside it.

Anything that appears after loading does not reach the PDF

The second source of loss is dynamics.

A modern page often pulls content in after it opens: images as you scroll, comments on a click, a data table as a separate request. Rendering works from markup rather than from a live page, and whatever loads later is not in it.

For the same reason, collapsed blocks stay collapsed, tabs other than the active one are absent, and everything shown on hover is missing.

You can check in advance whether content will make the archive: switch scripts off in your browser and reload the page. What stays visible will reach the PDF. What vanishes will not.

How PDF archiving differs from a screenshot and a bookmark

Three ways of keeping a page solve different problems, and confusing them is expensive.

MethodWhat is keptWhat is missing
BookmarkNothing but the addressThe page can change or disappear
ScreenshotA picture of the visible areaText cannot be searched or copied, a long page will not fit
PDFText, formatting, every page in fullInteractivity and anything loaded by scripts

For an evidentiary task the third column of the second row is what matters. A screenshot of a tariff page does not let you find a clause by search and does not show what sat below the fold. A PDF does both.

Worth remembering separately: none of the three is legal proof of a page's contents on a date. Notarised certification exists for that, and it works differently.

How to do the saving

1. Open the page in a browser and save it as HTML, choosing the single-file option. 2. Open html-to-pdf and upload the saved file. 3. Choose the page size: A4 or Letter. 4. Leave background graphics on if the design matters. 5. Run the job and compare the result against the original page.

Do step five immediately rather than a month later: if something went missing, the page is still available and can be saved a different way.

Background graphics and why to leave them on

The background graphics switch controls whether fills, background images and coloured panels reach the PDF.

It is on by default, which is right for an archive: switched off, it turns text on a coloured panel into text on white, and part of the meaning of the design is lost. That shows most on warnings marked by colour and on tables with alternating rows.

There is exactly one reason to switch it off: the page is going to a printer and the fills would eat toner. For an archive, leave it on.

Page size and wide tables

The page size defaults to A4. The choice here is only between A4 and Letter; there is no arbitrary format.

One unavoidable problem follows from that. A web page has no width: it adapts to the window. A PDF has a fixed width. A wide table that scrolled sideways in the browser gets clipped at the edge of the sheet.

There is no way around it inside the tool. People work around it outside: reduce the page zoom in the browser before saving, and the table then fits the width.

When HTML becomes plain text

There is an automatic fast path. Very large HTML files holding no images, no tables, no embedded styles and no media are rendered as plain text, with no layout.

That is a deliberate decision: fully rendering a several-megabyte page of pure text is expensive and yields the same thing.

It does not happen silently. The result carries a warning that the document was assembled in a simplified form and the formatting was not preserved. If that result does not suit you, the file was caught by the heuristic by mistake, and the page is worth re-saving in a form that keeps its markup.

What to do with the finished archive

Three steps worth taking straight away.

Check the weight. A page full of images makes a heavy PDF, and compress-pdf will shrink it without touching the text.

Fill in the document properties. In a PDF made from HTML the title field is usually empty or holds a technical string. A meaningful title and date through set-pdf-metadata make an archive far easier to search a year later.

If there are several pages forming one document, assemble them with merge-pdf in reading order rather than in the order you saved them.

And last. For long-term storage it is worth running the result through pdf-to-pdfa: the archival profile requires everything needed for display to live inside the file, which is exactly what an archive wants.

FAQ

No, the tool works from a file. Save the page in your browser as HTML and upload the result. The save option that packs the whole page into a single file works best.
The page is rendered in isolation: requests to local files are blocked, and among external addresses only those visible in the HTML during a static scan are allowed. Images a script pulls in after the page opens never reach the renderer.
That is the automatic fast path: very large HTML files with no images, tables, embedded styles or media are rendered as text. The result carries a warning saying the formatting was not preserved.
No. The basic workflow is available without creating an account.
Files are used only for the selected operation and are automatically deleted after processing is finished. We do not use uploaded documents to train AI models.

More from this cluster

Related tools

← All Convert tools

What to do next

If you need a practical next step or service guidance after reading, open these pages.

All tools

PDF tools catalog: merge, compress, split, convert, rotate, protect and unlock PDF files online, all directly in your browser.

FAQ

Answers to common questions about iHatePDF: whether registration is required, how files are processed, where to check limits, and whether it's safe to upload documents.

Contact

Contact iHatePDF about processing errors, choosing a tool, security, business inquiries, and suggestions for new features.