A passport photo is the most sensitive file most apps ever receive. It carries a face, a full name, a date of birth and a document number, and it is enough to open an account somewhere else. Yet in many upload flows the picture is copied five or six times before anyone asks whether it needs to exist at all.
This post looks at the question from the engineering side. It quotes what the GDPR says about keeping data, walks through the places where a passport image quietly piles up, and describes a pattern we call read, return, forget: the image is read, the result is returned, and nothing of the picture is kept. It ends with the part that stays your job, because some businesses are required to keep a copy, and from July 2027 the EU spells that out in a new law.
This is the blog of doc.cheap, a Passport and ID OCR API that reads identity documents and returns the result as JSON. Read the product parts with that in mind.
What the GDPR actually asks
The GDPR does not say "never store a passport image". It says something more useful: keep what you need, for as long as you need it, and no longer. Art. 5(1) of Reg. (EU) 2016/679 sets the principles. Two of them decide most of the design.
| Principle | Text of Art. 5(1) | What it means for an image |
|---|---|---|
| Data minimisation, point (c) | "adequate, relevant and limited to what is necessary in relation to the purposes for which they are processed" | If your process needs the name, date of birth and document number, the picture itself may not be necessary once they are read. |
| Storage limitation, point (e) | "kept in a form which permits identification of data subjects for no longer than is necessary for the purposes for which the personal data are processed" | Every copy needs an end date, and "we never got round to deleting it" is not one. |
Art. 5(2) adds accountability: you must be able to show that you follow these principles. And Art. 25(1), data protection by design, asks for "appropriate technical and organisational measures, such as pseudonymisation, which are designed to implement data-protection principles, such as data minimisation, in an effective manner".
Read together, they turn a legal question into an engineering one. The fewer copies of an image exist, the fewer places you have to describe, secure, back up and eventually empty. A copy that was never written is the only one that needs none of that work.
Where passport images end up
Most teams store the image on purpose in one place. The trouble is the places nobody chose. Here is a list to walk through against your own flow.
| Place | How the image gets there |
|---|---|
| Upload bucket | The client uploads to object storage first, and the backend reads it from there. The object outlives the request. |
| Request logs | A logging middleware writes request bodies, and a base64 image is a request body. |
| Error reports | An exception tracker attaches the payload of the request that failed. |
| Queues and retries | A job message carries the image, and a dead-letter queue keeps the failed ones for weeks. |
| Backups and snapshots | A database or disk snapshot taken that day holds every image written before it, long after the row is deleted. |
| Support tickets | A user mails the photo again "because the upload did not work". |
| Analytics and session replay | A tool records the page, including the preview of the selected file. |
| The OCR provider | The service that reads the document keeps its own copy, under its own retention rules. |
The last row is the one you control least. Your own logs you can fix. A copy held by a vendor is governed by that vendor's settings, and you have to ask what they are.
The pattern: read, return, forget
The pattern is simple to state. The image exists only in memory, for the length of one request. What leaves the request is the reading: the extracted values. The picture does not leave at all.
- Send the image straight to recognition. No upload bucket in between. If a bucket is unavoidable for large files, give the object a lifetime of minutes and delete it when the call returns.
- Use what you need from the response at once. Crops such as the holder's photo exist only in that response. If your flow compares a selfie to the portrait, do it now.
- Keep the reading, not the picture. Store the fields your process needs, under your own retention rule. A date of birth and a document number are still personal data, so they get an end date too.
- Keep the image out of logs and error reports. Strip bodies at the boundary, in one place, rather than trusting every caller to remember.
- Write down the result. For each place in the table above, note whether the image can reach it and why not. That note is the accountability Art. 5(2) asks for.
What our API does with the image
Here is how doc.cheap handles the same question, as its data retention and privacy page describes it.
- The image is never stored. It lives in memory for the request, is handed to the recognition engine, and is gone when the response is written. No disk, object store or log receives it.
- The crops are not stored either. The document crop, the holder's photo and the signature come back in the response of the call that made them. A scan read back later through
GET /v1/scans/{id}has every image slot set tonull. - What can be kept is the reading, and only for the window you ask for. The
retain_hoursoption sets it per request, from 0 up to 8760 hours (one year). An explicit value always wins over the account setting. retain_hours: 0writes nothing. Not a row that expires at once: no row at all. There is nothing to sweep, nothing in a backup and nothing to export. The scan still counts as a scan.- The account default covers the rest. When a request names no window, the account's own history setting applies: 24 hours, 7 days, 1 month or 1 year. New accounts start on 1 year, so the dashboard shows a history. Shortening the setting applies to rows already stored, each measured from its own creation time.
- A retained row keeps one small picture: a thumbnail of at most 96 px on its longest side and at most 16 KiB, shown in the dashboard's operations log so a row can be recognised. It is not readable through the API. The thumbnail goes when the row goes.
- One scan can be deleted early. A live key sends
DELETE /v1/scans/{id}, which removes the result, the history row and the thumbnail. It is final.
The control history retention guide covers the settings step by step, and our page on how we process data gives the summary.
Here is a zero-retention call in Python with requests. The public sandbox key sk_sandbox_public is printed in the docs and needs no signup: it gives 10 free recognised documents per address in all, and at most 10 requests an hour. Zero retention is a setting of your own account, so it needs your live key. The public sandbox is not an account: it keeps a record of every scan, with its small picture, for the service's own log, so send it a test image, never a real document.
import base64
import uuid
import requests
API = "https://api.doc.cheap/v1/scans"
KEY = "sk_sandbox_public" # your own live key in production
def read_and_forget(path):
with open(path, "rb") as f:
image = base64.b64encode(f.read()).decode("ascii")
response = requests.post(
API,
headers={
"Authorization": f"Bearer {KEY}",
"Idempotency-Key": str(uuid.uuid4()),
},
json={
"image": image,
# 0: with a live key, nothing about this scan is written down on the API side.
# False: no portrait crop, because this flow does not use one.
"options": {"retain_hours": 0, "return_portrait": False},
},
timeout=30,
)
response.raise_for_status()
scan = response.json()
del image # the local copy goes as soon as the call returns
if scan["meta"]["status"] != "recognized":
return None
# Keep the reading your process needs, under your own retention rule.
return {
"scan_id": scan["meta"]["id"],
"document_number": scan["document"]["number"],
"expiry_date": scan["document"]["expiry_date"],
"birth_date": scan["holder"]["birth_date"],
"mrz_status": scan["mrz"]["status"],
}
Zero retention has one cost worth knowing before it surprises you. An Idempotency-Key normally lets a retry return the first result. With retain_hours: 0 there is no stored result to return, so for 24 hours a retry under the same key is refused with HTTP 409 and the code idempotency_replay_unavailable, rather than answered twice. Treat that answer as "the first call went through" and use the result you already have.
Reading the scan back shows the other side of the design. A sandbox key reads nothing back at all, whatever the id. We sent this request with sk_sandbox_public on 5 October 2026:
curl https://api.doc.cheap/v1/scans/<SCAN_ID> \
-H "Authorization: Bearer sk_sandbox_public"
It came back with HTTP 404 (the message is trimmed):
{
"error": {
"code": "not_found",
"message": "No scan with id …",
"docs_url": "https://doc.cheap/docs/errors/not_found"
}
}
A live key gets the same 404 for a scan made with retain_hours: 0, and for any scan once its window has passed. If you are comparing services on this point, the passport OCR API comparison is a place to start; ask each one where the image goes, not only what it returns.
What stays your job
Read, return, forget removes the copies on the API side. It does not decide what your business has to keep. For some businesses the answer is "a copy", and the law says so.
The EU's new anti-money-laundering law, Reg. (EU) 2024/1624, applies from 10 July 2027. Art. 90 says: "It shall apply from 10 July 2027, except in relation to obliged entities referred to in Article 3, points (3)(n) and (o), to which it shall apply from 10 July 2029." Its Art. 77 on record retention makes obliged entities, such as banks and other financial firms, keep:
"a copy of the documents and information obtained in the performance of customer due diligence pursuant to Chapter III, including information obtained through electronic identification means;"
Art. 77(3) sets the length: the records are "retained for a period of 5 years commencing on the date of the termination of the business relationship", and then "obliged entities shall delete personal data upon expiry of the five-year period". Art. 77(2) allows, under conditions, "a retention of the references to such information" in place of copies.
So if you are an obliged entity, zero retention at the API does not remove your duty to keep a record. It changes where the record lives. Your own store becomes the one copy, and the storage limitation principle above still applies to it: five years after the relationship ends, it goes. The design work is to make that store deliberate, with one place, one owner, encryption, access control and a deletion job, rather than the accidental pile from the table above.
If you are not an obliged entity, ask the plain question first: does anything in your process need the picture after the fields are read? Often the honest answer is no.
A checklist
- Every place in the "where images end up" table is checked against your flow.
- The image goes straight to recognition, or through a bucket with a lifetime of minutes.
- Crops are used inside the response handler and not written anywhere.
- The OCR call sets its retention on purpose,
retain_hours: 0when nothing needs reading back. - Retries handle the 409
idempotency_replay_unavailableanswer. - Logs and error reports strip request bodies at one boundary.
- The fields you keep have an end date, and something deletes them.
- If a law requires a copy, it lives in one deliberate store with its own deletion date.
This is an engineering summary, not legal advice. If you find a place where an image can leak that this post misses, write to admin@doc.cheap.
A question for you: where did you last find a copy of an identity document that nobody meant to keep? Tell us in the comments below.