A passport reader in Python is about twenty lines the first time and about a hundred the time it has to survive production. The extra eighty are not about OCR. They are about three questions every paid API raises: what happens when the photo is bad, what happens when a request times out and you send it again, and how your own records know which calls cost money.

This tutorial builds that script with requests and nothing else. It uses doc.cheap, a Passport and ID OCR API, and this is doc.cheap's own blog, so weigh the product choices accordingly. The patterns (one idempotency key per image, branching on a stable error code, storing a per-call cost flag) carry over to any paid API you call from Python.

One request with the sandbox key

The docs print a public sandbox key, sk_sandbox_public. It runs the same recognition as a paid key and needs no account: 10 free recognised documents per IP address in total, and at most 10 requests an hour whatever the answer. The free passport OCR API page has the same first call as a single curl. An account later adds 100 free documents every month, with no card.

import base64
import requests

with open("specimen.jpg", "rb") as f:
    image = base64.b64encode(f.read()).decode("ascii")

response = requests.post(
    "https://api.doc.cheap/v1/scans",
    headers={"Authorization": "Bearer sk_sandbox_public"},
    json={"image": image},
    timeout=60,
)
scan = response.json()
print(response.status_code, scan["meta"]["status"], scan["meta"]["billed"])

json= sets Content-Type: application/json for you. The image travels as base64 inside the body, JPEG or PNG. The call is synchronous: the fields come back in this response, with no job id to poll and no webhook to host.

Test with a synthetic specimen, never your own passport. ICAO's fictitious "Utopia" documents and the specimen pages many issuers publish exist for exactly this.

Shrink the photo first. A phone picture can be several megabytes before base64 adds a third. The passport guide recommends about 1600 px on the long edge at JPEG quality 85; a body over the ceiling is refused with payload_too_large before anything runs.

Reading the answer

Every key of the response is always present, and a value that is not known is None after json(), never a missing key. Four parts decide what your code does next:

  • meta.status says whether the document was read. It is one of five strings: recognized, no_document_found, unreadable, unsupported_document, rejected. Only the first carries data.
  • document and holder hold the curated values: document.kind (passport, an ID card and so on), the country as ISO 3166-1 alpha-3, the number, the dates as ISO YYYY-MM-DD; holder.surname, holder.given_names, holder.birth_date. Each group is None as a whole when the scan produced nothing for it.
  • mrz.status is passed, failed or absent: whether the machine-readable zone was found and its check digits agree. mrz.text is the zone as read, so you can run the check digits yourself.
  • meta.billed says whether this call was charged to the balance.

The point that shapes the whole client: a photo that could not be read is HTTP 200, not an error. So the code branches twice, on the HTTP error code for refusals and on meta.status for outcomes. Raise on no_document_found and your retry loop resends a photo that will never read. Treat it as success and you store a document with no fields.

The billed flag

On a live key, billed is True only when a document was actually recognised. Nothing found, an unreadable image, an unsupported type, a failure on the service side: none of those is charged. A billed document costs $0.01, flat, at any volume; the passport OCR API comparison puts that next to the prices other services publish.

The sandbox charges nothing at all. On sk_sandbox_public, billed still tells you whether the same scan would have been charged on a live key, which is what makes it worth testing against. Store the flag next to each result: summing your own billed rows for a month then needs no reconciliation with an invoice.

Retries that cannot charge twice

The risky retry is the one after a timeout. You do not know whether the first request reached the server, and on a paid API a blind retry can pay twice for one image.

The answer is an Idempotency-Key header on POST /v1/scans, 1 to 255 characters. Generate it once per image and send the same value on every attempt. On a live key, a retry with the same key and the same body gets the first result back instead of a second recognition, and a replay costs nothing. Three rules from the reference shape the code:

  • The key is bound to the body. The same key with a different image or different options is 409 idempotency_conflict. Build the body once, outside the loop.
  • A second attempt can arrive while the first is still running. That is 409 idempotency_in_progress: wait and retry with the same key.
  • No stored result, no replay. With retain_hours: 0 nothing is stored, so a retry under the same key within 24 hours gets 409 idempotency_replay_unavailable rather than a second run. Keeping nothing and replaying a result do not go together; choose per use case.

The sandbox keys accept the header but it decides nothing there, since nothing is billed. The code path is still worth testing.

Which errors are worth a retry. Every error body has the same shape (code, message, docs_url, request_id, event_id), and the code is what to branch on, because one HTTP status can carry codes that need opposite handling. The handle errors guide sorts them into buckets:

Codes What to do
rate_limited, document_repeated, internal_error, engine_unavailable, service_unavailable, maintenance Wait (honour Retry-After), then retry
idempotency_in_progress Wait, retry with the same key
validation_failed, invalid_request, payload_too_large, unsupported_media_type Fix the request; a retry fails again
unauthorized, registration_required, insufficient_credits Fix the key or the account; waiting changes nothing

Retry-After is whole seconds. The sandbox's hourly limit can ask for most of an hour, and no caller wants one function call to sleep that long, so the client below gives up when the wait is over a minute.

The whole client

# scan.py
import base64
import os
import time
import uuid

import requests

API = "https://api.doc.cheap/v1/scans"
KEY = os.environ.get("DOC_CHEAP_API_KEY", "sk_sandbox_public")
RETRY = {
    "rate_limited", "document_repeated", "internal_error", "engine_unavailable",
    "service_unavailable", "maintenance", "idempotency_in_progress",
}
MAX_WAIT = 60  # seconds; a longer Retry-After is reported, not slept through


class ScanError(Exception):
    def __init__(self, status, error):
        super().__init__(f"{error['code']} ({status}): {error['message']}")
        self.code = error["code"]
        self.docs_url = error["docs_url"]
        self.request_id = error["request_id"]


def retry_after(response, attempt):
    try:
        seconds = int(response.headers.get("Retry-After", ""))
    except ValueError:
        seconds = 0
    return seconds if seconds > 0 else 2 ** attempt


def scan(path, reference=None, attempts=4, session=None):
    session = session or requests.Session()
    with open(path, "rb") as f:
        image = base64.b64encode(f.read()).decode("ascii")
    body = {"image": image, "reference": reference, "options": {"return_portrait": False}}
    headers = {
        "Authorization": f"Bearer {KEY}",
        "Idempotency-Key": str(uuid.uuid4()),  # one key per image, reused on every attempt
    }

    for attempt in range(1, attempts + 1):
        last = attempt == attempts
        try:
            response = session.post(API, headers=headers, json=body, timeout=60)
            payload = response.json()
        except (requests.RequestException, ValueError):
            # A timeout, a dropped connection, or a proxy answering with HTML.
            if last:
                raise
            time.sleep(2 ** attempt)
            continue

        if response.ok:
            return payload
        error = payload["error"]
        if error["code"] not in RETRY or last:
            raise ScanError(response.status_code, error)
        wait = retry_after(response, attempt)
        if wait > MAX_WAIT:
            raise ScanError(response.status_code, error)
        time.sleep(wait)


def summarise(scan):
    meta, document, holder, mrz = scan["meta"], scan["document"], scan["holder"], scan["mrz"]
    if meta["status"] != "recognized":
        return {"ok": False, "status": meta["status"], "billed": meta["billed"]}
    document, holder = document or {}, holder or {}
    return {
        "ok": True,
        "billed": meta["billed"],
        "kind": document.get("kind"),
        "country": document.get("country"),
        "expiry_date": document.get("expiry_date"),
        "surname": holder.get("surname"),
        "birth_date": holder.get("birth_date"),
        "mrz": mrz["status"],
    }


if __name__ == "__main__":
    import sys

    try:
        print(summarise(scan(sys.argv[1], reference="demo-1")))
    except ScanError as err:
        print(err, err.docs_url, err.request_id, file=sys.stderr)
        sys.exit(1)

Run it as python scan.py specimen.jpg. A few choices worth explaining:

  • requests does not raise on a 4xx or 5xx unless you call raise_for_status(). The client reads the JSON body either way, because the error body carries the code; raise_for_status() would throw that away.
  • The body and the key are built once, before the loop. That is what makes every attempt the same request in the server's eyes.
  • A Session reuses the connection across retries and across a batch of images. Pass one in when you scan many files.
  • return_portrait: False leaves out the crop of the holder's photo. You get the page crop and the fields, and one fewer face sits in your logs and storage.
  • reference comes back as meta.reference (up to 128 characters): the easy way to join a scan to your own order or user. The two sides of an ID card are two calls; give them the same reference.
  • Never log the request body. It is an identity document. Log code, request_id and docs_url; that is everything support needs.

Keeping less data

The uploaded image is held in memory for the request and never written to durable storage. The result is kept so you can fetch it again with GET /v1/scans/{id}, for a window you choose: per request, options.retain_hours takes 0 to 8760, and 0 stores nothing at all. If you need the JSON only once, send retain_hours: 0 and accept the replay trade-off above. Processing happens in the EU.

Going live

Set DOC_CHEAP_API_KEY to your own key and nothing else changes: same endpoint, same response shape, same client. A registered key allows 60 requests a minute, so a batch job that respects Retry-After will not need its own rate limiter. Credits are bought with cryptocurrency (BTC, ETH, TRX, or USDT on Ethereum or Tron) from $1; there is no card checkout today.

If something in the response is awkward to handle from Python, write to admin@doc.cheap.