Skip to content
Article

How to Redact Sensitive Data in a PDF

A black rectangle drawn over text deletes nothing. How the redaction styles differ in substance, why blur does not count as removal, and what happens to the metadata.

In short: The black and white styles genuinely cut the content out from under the rectangle: the covered text disappears from the file while the rest of the page stays text. Blur only masks, and turns the whole document into images as well. Metadata is cleared by default.

Cluster

security

Protection, access control, redaction, and controlled sharing workflows.

4 articles

Primary tool

Redact PDF online

Open the tool from this article and complete the operation in the current locale.

Open tool

Table of contents

Hiding personal data in a PDF so it cannot be pulled back out

The story repeats roughly once a year and makes the news. An organisation publishes a document in which names and addresses are neatly covered by black rectangles. An hour later somebody selects the text with a mouse, pastes it into a notepad, and has everything that was covered.

The reason is simple, and understanding it matters more than remembering a sequence of buttons.

Why a rectangle over text hides nothing

A PDF is built in layers. Text sits in the page content as text, with coordinates and a font. A rectangle drawn over it in an editor is another content element, drawn later and therefore visible on top.

The text goes nowhere. It is still in the file, found by search, selected by mouse and copied. The rectangle covers it visually only, exactly like a sheet of paper laid on a monitor.

Real redaction works differently: the content under the area is first cut out of the page, and only then is a fill laid over the space it left.

Three styles and what each does to the content

StyleWhat happens to the contentDoes the document stay text
BlackThe content under the area is deleted, a black fill goes on topYes, the rest of the page stays text
WhiteThe same, with a white fillYes
BlurThe document is redrawn as images and the area is blurredNo, the text layer disappears entirely

The first two are genuine removal. Verified on a document where a line holding a token number was covered by a black area: after processing that line is found neither by search nor by copying, while the neighbouring lines still select as ordinary text.

Blur masks, it does not remove

The third style stands apart, and it carries two costs.

The first: it does not remove. Blur degrades the image of the area, but the information in it is partly recoverable by processing. For personal data that is unacceptable, and the tool says so outright: choosing blur attaches a warning to the result recommending the black or white style instead.

The second cost is measured in bytes and in capabilities. Blur works from a rasterised copy of the whole document rather than of one area. A one-page test file with twenty lines of text weighed about one kilobyte. After a black redaction it was still that kilobyte. After blur it grew to ninety and lost the text layer completely: such a document cannot be searched.

Blur makes sense in exactly one case: hiding a face in a photograph or a fragment of an image where cutting would leave an incongruous black hole.

How to run a redaction

1. Open redact-pdf and upload the document. 2. Mark the areas on the page itself: the tool shows the document and lets you draw over the spots you need. 3. Choose the style. For data, take black or white. 4. Leave metadata removal switched on. 5. Run the job. 6. Open the result and try to select the covered spot with a mouse, then search the document for the deleted word.

Point six is not a formality. It is the only check that distinguishes genuine removal from a drawn rectangle, and it takes ten seconds.

Areas are mandatory

The tool does not try to guess what should be hidden. With no areas given, the job does not start at all and returns a direct message saying at least one area is required.

That is a deliberate decision, and the right one. Guessing an area in a task where a mistake means a personal-data leak is more dangerous than a refusal: having covered the wrong spot, the tool would report success while the data stayed in plain view.

For the same reason, areas are checked for sense: coordinates outside the page or a zero width produce an error rather than being silently skipped.

Metadata is cleared by default

A separate switch governs clearing the document properties, and it is on by default.

That is not a detail. The properties regularly hold the name of the employee who prepared the document, an internal working title, sometimes a folder path from a disk. A document from which a client's surname was carefully cut, while the author name and a title like "Ivanov case, draft" stayed in the properties, does not solve the problem.

If you switch the clearing off, a separate warning arrives: the original metadata is kept and may contain details identifying people. It does not happen silently.

For finer work with the properties there is set-pdf-metadata, where fields can be filled in with values you choose rather than merely emptied.

A note recording that a redaction happened

The parameters hold a field for a short note written into the finished file's properties, into the subject field.

Its purpose is administrative: the recipient sees that the document was redacted deliberately rather than arriving damaged. Wordings along the lines of "personal data removed" or "withheld under clause such-and-such" save a round of correspondence.

The length is capped at five hundred characters. Longer text is truncated and a warning says so, so the note does not vanish silently, but it does not arrive whole either.

Where data hides besides the body text

A full pass over the document is better done against a list than from memory. The places people forget repeat from document to document.

Headers and footers. They repeat on every page, and if a surname or a case number sits there, every page needs covering, not just the first.

Signatures and stamps. A signature is personal data and shows in full on a scan. A stamp carries an organisation's name and often its address.

Annex tables. The body text gets read carefully, annexes get flipped past, and it is annexes that hold registers of names and phone numbers.

Reverse sides. In a double-sided scan the backs are interleaved with the fronts, and half of them tend to look blank because only the shadow of the text is visible on a thumbnail.

Marginal notes. Handwritten resolutions, registration stamps, stickers with numbers.

A practical trick: before you start, open the document and search it for the surname, the phone number and the address. Search shows every occurrence, including the ones the eye skips.

What the tool does not do

It does not find personal data for you. The tool removes what you drew over and nothing else. Responsibility for covering the whole document stays with a person.

Hence a practical rule: walk the document in full, headers and footers, captions under tables and reverse pages included. That is exactly where data gets forgotten, because attention goes to the body text.

And separately: cropping a page with crop-pdf is not redaction. It changes the visible area of the sheet, while the content beyond it stays in the file and comes back when the crop is reset.

If the document is a scan, there may be no text under the area at all and nothing to cut out; there is simply an image there. In that case the result is reliable for a different reason: a piece of the image is removed. But if the scan went through ocr-pdf beforehand, a text layer appeared under the picture, and it has to be covered by the area too.

FAQ

A rectangle drawn over text is a separate layer, and the text stays underneath, reachable by copying or search. The black and white styles here delete the content from the file and paint the fill over the now-empty space.
No. Blur degrades an image without removing it, and part of the content can be recovered by processing. The tool warns about that explicitly when you choose the style. Real removal needs black or white.
With black and white, yes: the uncovered text stays text and the file barely changes size. With blur, no: the entire document is redrawn as images and the text layer disappears completely.
No. The basic workflow is available without creating an account.
Files are used only for the selected operation and are automatically deleted after processing is finished. We do not use uploaded documents to train AI models.

More from this cluster

Related tools

← All Security tools

What to do next

If you need a practical next step or service guidance after reading, open these pages.

All tools

PDF tools catalog: merge, compress, split, convert, rotate, protect and unlock PDF files online, all directly in your browser.

FAQ

Answers to common questions about iHatePDF: whether registration is required, how files are processed, where to check limits, and whether it's safe to upload documents.

Contact

Contact iHatePDF about processing errors, choosing a tool, security, business inquiries, and suggestions for new features.