Skip to content
Partly available

PII detection in images, with OCR

Personally identifiable information in an image is rarely in a tidy field. It is a card photographed on a desk, a passport page sent to a letting agent, a screenshot of a banking app pasted into a support chat.

Finding it means reading the picture, then testing what was read. This page explains how that works here, and is straight about which half runs where.

Partly available. Metadata, barcode and face detection work in the browser today. OCR — reading text out of the picture itself — currently runs through the API only.

What counts as PII in an image

Two sources, and a complete answer needs both: the metadata block, and the pixels.

  • Payment card numbers — validated with the Luhn checksum, so a random sixteen digits does not fire
  • IBANs — validated mod-97, the check built into the standard
  • Passport MRZ lines — validated with the check digits the format specifies
  • National identifiers, such as US Social Security numbers, with the ranges that are never issued excluded
  • Dates of birth, email addresses, phone numbers and street addresses
  • GPS coordinates from the metadata, which identify a person by where they were
  • Faces, located so they can be covered

How the OCR path works

The detection library has no network access by design, so it cannot fetch a model. OCR needs a trained model, which means it cannot live inside that boundary. Instead the library defines an interface and the caller supplies an engine, pointed at assets they host themselves.

The practical consequence is that asking for OCR without supplying an engine raises an error rather than quietly reaching out to a CDN. We would rather fail loudly than have a privacy tool make a request nobody expected.

On the API, OCR is a query parameter: add ?ocr=true to a scan. In the browser it is not switched on yet, and the toggle on the checker says so.

Why checksums matter for PII

Regular expressions over OCR output are a false-positive machine. OCR misreads characters, documents contain reference numbers that look like card numbers, and a detector that cries wolf gets switched off.

Validating against the checksum that is built into the format — Luhn, mod-97, MRZ check digits — removes most of that noise. Every detector has tests for things it must not match as well as things it must.

Masking: results never show the value

A finding carries a masked preview and nothing else. A card number shows its last four digits, an email shows its first character and domain, a token shows its first four characters. There is no field in the result type that can hold a raw value, which is enforced by the type system rather than by remembering.

Questions

  • Through the API, yes — pass ?ocr=true on a scan and supply an OCR engine. In the browser, not yet. We list it on the checker as not switched on rather than leaving you to find out.

Related

We report no issues found, never “safe”. Absence of detections is not proof of absence, and detection is best-effort.