How to Redact Sensitive Data in a PDF
In short: The black and white styles genuinely cut the content out from under the rectangle: the covered text disappears from the file while the rest of the page stays text. Blur only masks, and turns the whole document into images as well. Metadata is cleared by default.
Cluster
security
Protection, access control, redaction, and controlled sharing workflows.
Primary tool
Redact PDF online
Open the tool from this article and complete the operation in the current locale.
Open toolTable of contents
Hiding personal data in a PDF so it cannot be pulled back out
The story repeats roughly once a year and makes the news. An organisation publishes a document in which names and addresses are neatly covered by black rectangles. An hour later somebody selects the text with a mouse, pastes it into a notepad, and has everything that was covered.
The reason is simple, and understanding it matters more than remembering a sequence of buttons.
Why a rectangle over text hides nothing
A PDF is built in layers. Text sits in the page content as text, with coordinates and a font. A rectangle drawn over it in an editor is another content element, drawn later and therefore visible on top.
The text goes nowhere. It is still in the file, found by search, selected by mouse and copied. The rectangle covers it visually only, exactly like a sheet of paper laid on a monitor.
Real redaction works differently: the content under the area is first cut out of the page, and only then is a fill laid over the space it left.
Three styles and what each does to the content
| Style | What happens to the content | Does the document stay text |
|---|---|---|
| Black | The content under the area is deleted, a black fill goes on top | Yes, the rest of the page stays text |
| White | The same, with a white fill | Yes |
| Blur | The document is redrawn as images and the area is blurred | No, the text layer disappears entirely |
The first two are genuine removal. Verified on a document where a line holding a token number was covered by a black area: after processing that line is found neither by search nor by copying, while the neighbouring lines still select as ordinary text.
Blur masks, it does not remove
The third style stands apart, and it carries two costs.
The first: it does not remove. Blur degrades the image of the area, but the information in it is partly recoverable by processing. For personal data that is unacceptable, and the tool says so outright: choosing blur attaches a warning to the result recommending the black or white style instead.
The second cost is measured in bytes and in capabilities. Blur works from a rasterised copy of the whole document rather than of one area. A one-page test file with twenty lines of text weighed about one kilobyte. After a black redaction it was still that kilobyte. After blur it grew to ninety and lost the text layer completely: such a document cannot be searched.
Blur makes sense in exactly one case: hiding a face in a photograph or a fragment of an image where cutting would leave an incongruous black hole.
How to run a redaction
1. Open redact-pdf and upload the document. 2. Mark the areas on the page itself: the tool shows the document and lets you draw over the spots you need. 3. Choose the style. For data, take black or white. 4. Leave metadata removal switched on. 5. Run the job. 6. Open the result and try to select the covered spot with a mouse, then search the document for the deleted word.
Point six is not a formality. It is the only check that distinguishes genuine removal from a drawn rectangle, and it takes ten seconds.
Areas are mandatory
The tool does not try to guess what should be hidden. With no areas given, the job does not start at all and returns a direct message saying at least one area is required.
That is a deliberate decision, and the right one. Guessing an area in a task where a mistake means a personal-data leak is more dangerous than a refusal: having covered the wrong spot, the tool would report success while the data stayed in plain view.
For the same reason, areas are checked for sense: coordinates outside the page or a zero width produce an error rather than being silently skipped.
Metadata is cleared by default
A separate switch governs clearing the document properties, and it is on by default.
That is not a detail. The properties regularly hold the name of the employee who prepared the document, an internal working title, sometimes a folder path from a disk. A document from which a client's surname was carefully cut, while the author name and a title like "Ivanov case, draft" stayed in the properties, does not solve the problem.
If you switch the clearing off, a separate warning arrives: the original metadata is kept and may contain details identifying people. It does not happen silently.
For finer work with the properties there is set-pdf-metadata, where fields can be filled in with values you choose rather than merely emptied.
A note recording that a redaction happened
The parameters hold a field for a short note written into the finished file's properties, into the subject field.
Its purpose is administrative: the recipient sees that the document was redacted deliberately rather than arriving damaged. Wordings along the lines of "personal data removed" or "withheld under clause such-and-such" save a round of correspondence.
The length is capped at five hundred characters. Longer text is truncated and a warning says so, so the note does not vanish silently, but it does not arrive whole either.
Where data hides besides the body text
A full pass over the document is better done against a list than from memory. The places people forget repeat from document to document.
Headers and footers. They repeat on every page, and if a surname or a case number sits there, every page needs covering, not just the first.
Signatures and stamps. A signature is personal data and shows in full on a scan. A stamp carries an organisation's name and often its address.
Annex tables. The body text gets read carefully, annexes get flipped past, and it is annexes that hold registers of names and phone numbers.
Reverse sides. In a double-sided scan the backs are interleaved with the fronts, and half of them tend to look blank because only the shadow of the text is visible on a thumbnail.
Marginal notes. Handwritten resolutions, registration stamps, stickers with numbers.
A practical trick: before you start, open the document and search it for the surname, the phone number and the address. Search shows every occurrence, including the ones the eye skips.
What the tool does not do
It does not find personal data for you. The tool removes what you drew over and nothing else. Responsibility for covering the whole document stays with a person.
Hence a practical rule: walk the document in full, headers and footers, captions under tables and reverse pages included. That is exactly where data gets forgotten, because attention goes to the body text.
And separately: cropping a page with crop-pdf is not redaction. It changes the visible area of the sheet, while the content beyond it stays in the file and comes back when the crop is reset.
If the document is a scan, there may be no text under the area at all and nothing to cut out; there is simply an image there. In that case the result is reliable for a different reason: a piece of the image is removed. But if the scan went through ocr-pdf beforehand, a text layer appeared under the picture, and it has to be covered by the area too.
FAQ
More from this cluster
security
Protect a PDF with a Password
How to password-protect a PDF before emailing it, what it actually does, and why the password must travel separately from the file.
security
How to Remove a PDF Password
A PDF carries two protection mechanisms and only one of them is real. What encryption gives you, why a copy restriction is a request, and why an empty password field sometimes settles the job.
security
HTML to PDF: Save a Web Page
The tool takes a saved HTML file rather than a page address, and renders it in an isolated browser. Hence the rules: local images are not pulled in, and anything that appears after loading never reaches the PDF.
Related tools
Redact PDF online
Cover personal details and confidential content in a PDF, apply the redaction, and inspect the downloaded copy before sharing it. A practical use case is to prepare an ID scan, statement, or contract without exposing unnecessary personal data.
Crop PDF online
Crop the visible area of PDF pages to remove wide margins, dark scan edges or extra space.
Set PDF metadata
Set or change PDF metadata - title, author, subject and keywords - to make documents easier to organize.
OCR PDF online
OCR PDF keeps the PDF task in one browser flow: upload the source file, check options, run processing, and download the result.
What to do next
All tools
PDF tools catalog: merge, compress, split, convert, rotate, protect and unlock PDF files online, all directly in your browser.
FAQ
Answers to common questions about iHatePDF: whether registration is required, how files are processed, where to check limits, and whether it's safe to upload documents.
Contact
Contact iHatePDF about processing errors, choosing a tool, security, business inquiries, and suggestions for new features.