PDF Privacy Scanner
See what a PDF says about you before you send it — metadata, paths and personal data.
The PDF Privacy Scanner reads your document on your device and masks every example it reports. Nothing about the file is uploaded or logged.
Redact what is printed on the page
About PDF Privacy Scanner
Every PDF carries more than the page shows. The properties dictionary usually names an author, records which application produced the file and when, and sometimes preserves a local file path complete with a username and a folder structure. The text layer can hold plenty more: an email address in a footer, an account number in a table, a card number in an appendix nobody re-read. This scanner reports both. Metadata findings are graded by how identifying they are, and the text is matched against patterns for addresses, card numbers, national identifiers and account codes — every example masked, so a scan never republishes what it found.
Features
- Reads author, title, subject, keywords, producer, creator and timestamps
- Detects local file paths, which usually expose a username and folder tree
- Matches emails, phone numbers, IBANs, card numbers, IP addresses and national IDs
- Card numbers checked against the Luhn algorithm to cut false positives
- Every example masked in the report rather than shown in full
- A single risk score with a plain-language reading of what it means
- One-click strip of every metadata field, saved as a new document
How to use the PDF Privacy Scanner
- Drop the PDF onto the page
- Read the metadata findings and the personal data matched in the text
- Note the risk score and decide what has to change
- Strip the metadata, and redact the page content if the text findings matter
Example
Input
tender-response.pdf
Output
High risk: author "J. Whitfield", path C:\Users\jwhitfield\Bids\…, 4 email addresses in the text.
The path is the finding people are most surprised by — it names the machine's user account.
Common errors & troubleshooting
- The scan reports nothing at all. — Some producers write no metadata, which is a good result. Check the text findings separately — a clean properties dictionary says nothing about what is printed on the page.
- A phone number match is actually a reference number. — Digit patterns overlap. The scanner errs towards reporting, so read the masked example and dismiss the ones that are not what they look like.
- Stripping metadata did not remove a name from the page. — Metadata and page content are different things. A name printed in the document has to be redacted from the page itself.
- The text scan found nothing in an obviously data-heavy document. — The PDF is a scan with no text layer, so there is nothing to match. Recognise the text first, then scan again.
Frequently asked questions
- What is the most common thing a PDF leaks?
- The author field, followed by the producing application. Both are set automatically by whatever made the file, and neither is visible when you read the document.
- Is my document uploaded to be scanned?
- No. Reading the metadata and matching the text both happen in this tab — uploading a document to check its privacy would defeat the exercise.
- Why are the examples masked?
- So the report itself is safe to screenshot or share. You can see that an address was found and roughly which one, without the full value on screen.
- Does stripping metadata change the pages?
- No. Only the document properties are rewritten; every page, its text and its fonts stay exactly as they were.
- What score should I be comfortable sending?
- Anything the scanner rates low. Medium usually means a name and a timestamp, which is fine internally and worth clearing before a document goes outside.
Related tools
- PDF Metadata Editor — View, edit or strip a PDF’s title, author, subject, keywords and producer.
- Redact PDF — Black out words and areas so they are gone from the file.
- Flatten PDF — Lock filled form fields into the page so a PDF looks the same everywhere.
- EXIF Viewer & Remover — View and strip EXIF metadata (including GPS) from photos — nothing is uploaded.
- Encrypt PDF — Password-protect a PDF with AES-256 or RC4 encryption.
- PDF Form Data Extractor — Pull the answers out of a filled PDF form as JSON or CSV, field by field.
- PDF to Text — Extract selectable text from a PDF as plain text or Markdown.
All ArrayKit tools