Redact PDF metadata online
A PDF carries a title, an author, the software that produced it, creation and modification timestamps, and often an XMP block repeating all of it. It can also carry annotations, form data and whole attached files.
And it carries the single most misunderstood thing in document privacy: the black box that does not redact anything.
Not switched on yet. PDF scanning and cleaning are built and tested in our engine but are not available in the website today — the checker currently accepts photos and Word documents. This page explains the problem; join the waitlist to hear when PDFs are live.
The black rectangle problem
Drawing a filled rectangle over text in a PDF adds a rectangle. It does not remove the text. The characters are still in the content stream underneath, with their positions, and selecting the area and copying gets them back. So does any text extraction tool, in one line of code.
This has caused real disclosures — court filings, government documents and redacted contracts whose covered text was recoverable by selecting it. It keeps happening because the document looks correct on screen.
Real redaction removes the characters from the content stream and then draws the box. Our engine detects the failed kind: text that sits underneath an opaque filled rectangle is reported as a redaction that does not redact.
What PDF metadata holds
Two places, and tools that clean one often miss the other.
- The Info dictionary — title, author, subject, keywords, creator, producer, dates
- An XMP packet — frequently a duplicate of the above, plus editing history
- Annotations — sticky notes and comments, each with an author
- Embedded attachments — entire files carried inside the PDF
- Form field values, including ones no longer displayed
- The producing software and version, which narrows down the machine it came from
Why "online" is the wrong instinct here
Most online PDF tools work by uploading your document to a server. For a document you are redacting — which by definition contains something you do not want seen — that is an odd first step.
When PDF support arrives here it will work the way the rest of the site does: in your browser, with the file never leaving your device. We would rather ship it late than ship it as an upload.
What to do today
Until PDF support is live, the honest advice is: do not rely on a black box. If text must be removed, remove it in the source document and export a fresh PDF, or use a tool that states it removes content rather than covering it. Then check the result by selecting the redacted area and trying to copy.
For photos and Word documents, the checker on this site works today.
Hear when this is switched on
One message when it launches. Nothing else, ever.
Questions
No. PDF support is implemented in the engine and covered by tests, but it is not switched on in the website. We would rather say that than have you upload a sensitive document and find out.
Because it is an additional drawing instruction, not a deletion. The text remains in the content stream with its coordinates, and copy-paste or any extraction library retrieves it. Redaction has to remove the characters themselves.
Printing to PDF or exporting as an image does drop the underlying text, which is why the advice circulates. It also drops the real text layer, making the document unsearchable and inaccessible to screen readers. It is a workaround, not a fix.
There is no date we are willing to promise. Join the waitlist and you will get one message when it launches.
Related
- File privacy scannerScan a file for personal data, hidden metadata, author names and edit history before you share it. Runs in you…
- Hidden data in Excel filesHidden worksheets, hidden rows and columns, cell comments and revealing headings — what a spreadsheet carries …
- Remove the author from a Word documentStrip the author, last-edited-by, company and comments from a .docx file, and accept tracked changes so delete…
We report no issues found, never “safe”. Absence of detections is not proof of absence, and detection is best-effort.