PII detection in images, with OCR
Personally identifiable information in an image is rarely in a tidy field. It is a card photographed on a desk, a passport page sent to a letting agent, a screenshot of a banking app pasted into a support chat.
Finding it means reading the picture, then testing what was read. This page explains how that works here, and is straight about which half runs where.
Partly available. Metadata, barcode and face detection work in the browser today. OCR — reading text out of the picture itself — currently runs through the API only.
What counts as PII in an image
Two sources, and a complete answer needs both: the metadata block, and the pixels.
- Payment card numbers — validated with the Luhn checksum, so a random sixteen digits does not fire
- IBANs — validated mod-97, the check built into the standard
- Passport MRZ lines — validated with the check digits the format specifies
- National identifiers, such as US Social Security numbers, with the ranges that are never issued excluded
- Dates of birth, email addresses, phone numbers and street addresses
- GPS coordinates from the metadata, which identify a person by where they were
- Faces, located so they can be covered
How the OCR path works
The detection library has no network access by design, so it cannot fetch a model. OCR needs a trained model, which means it cannot live inside that boundary. Instead the library defines an interface and the caller supplies an engine, pointed at assets they host themselves.
The practical consequence is that asking for OCR without supplying an engine raises an error rather than quietly reaching out to a CDN. We would rather fail loudly than have a privacy tool make a request nobody expected.
On the API, OCR is a query parameter: add ?ocr=true to a scan. In the browser it is not switched on yet, and the toggle on the checker says so.
Why checksums matter for PII
Regular expressions over OCR output are a false-positive machine. OCR misreads characters, documents contain reference numbers that look like card numbers, and a detector that cries wolf gets switched off.
Validating against the checksum that is built into the format — Luhn, mod-97, MRZ check digits — removes most of that noise. Every detector has tests for things it must not match as well as things it must.
Masking: results never show the value
A finding carries a masked preview and nothing else. A card number shows its last four digits, an email shows its first character and domain, a token shows its first four characters. There is no field in the result type that can hold a raw value, which is enforced by the type system rather than by remembering.
Questions
Through the API, yes — pass ?ocr=true on a scan and supply an OCR engine. In the browser, not yet. We list it on the checker as not switched on rather than leaving you to find out.
No. The engine runs where your code runs, against assets you host. The core library is prevented from making network calls by a lint rule that bans fetch, XMLHttpRequest, WebSocket and every node module, so this is enforced rather than promised.
Images and Word documents in the browser today. PDFs, Excel and CSV are implemented in the engine and tested, but are not switched on in the website yet.
Nothing is stored. There is no column anywhere in the database that could hold a file, extracted text or a finding preview, and nothing sensitive is written to logs.
Related
- Detect sensitive information in imagesFind GPS coordinates, faces, ID numbers, card numbers and boarding-pass barcodes in an image before you share …
- File privacy scannerScan a file for personal data, hidden metadata, author names and edit history before you share it. Runs in you…
- Check a document before you upload it to ChatGPTSee what a file really carries before you paste it into an AI chatbot: author names, comments, edit history an…
We report no issues found, never “safe”. Absence of detections is not proof of absence, and detection is best-effort.