Local-first OCR pipeline for passports and Indian KYC documents. It preprocesses scans, classifies the document, runs targeted OCR with RapidOCR (PP-OCRv5), and extracts structured fields—including passport MRZ data, Indian passport back-page fields, and identifier/holder fields for PAN, Aadhaar, driving licences, voter IDs, MGNREGA/NREGA job cards, and National Population Register (NPR) name-and-address letters.
Ships as a Python package with a FastAPI server, plus an npm wrapper at packages/passport-ocr that auto-spawns the Python server for Node.js consumers.
Important
This project is beta software for data extraction. It does not establish document authenticity, verify identity against an issuing authority, detect fraud, or by itself satisfy KYC obligations. Evaluate it on your own legally obtained dataset before production use.
| Document | Key fields extracted | Identifier validation |
|---|---|---|
| Passport (biodata) | name, passport no., nationality, DOB, sex, expiry/issue dates, place of birth | TD3 MRZ ICAO check digits |
| Passport (back page) | father/mother/spouse, address, file no., old passport details | — |
| PAN card | PAN, name, father's name, DOB | PAN format + holder-type char |
| Aadhaar (front/back) | Aadhaar no. (+ VID, masked-card support), name, DOB/YOB, gender, address, pincode | Verhoeff checksum |
| Driving licence | DL no., name, DOB, issue/validity dates (NT + TR), address, blood group, vehicle class | DL format |
| Voter ID (EPIC) | EPIC no., name, relation name + type, gender, DOB/age | EPIC format |
| MGNREGA/NREGA job card | job-card no., household head, registration/validity, location, category/BPL, adult members | Conservative hierarchical format |
| NPR name/address letter | reference no., resident name, address, pincode, issue date | Not offline-verifiable |
Together with the existing passport path, this covers the officially valid document categories listed in the RBI KYC Master Direction; PAN is supported as a separate tax identifier.
The document type is detected automatically; /scan returns the matching
field block (fields/backPageFields for passports, or panFields,
aadhaarFields, drivingLicenceFields, voterIdFields,
nregaJobCardFields, or nprLetterFields) keyed by documentType.
NREGA and NPR support is experimental until measured against a representative private image dataset. The extraction layer handles multiple label/layout variants, but deterministic text-region tests are not evidence of real-image accuracy.
To keep passport OCR behavior unchanged, a positive passport-page probe is never overridden by the KYC router. A driving licence or voter card whose crop contains multiple passport-like labels can therefore remain ambiguous; an explicit KYC-only entry point is tracked as future work.
Requires Python 3.12+ and uv.
make install
make dev # uvicorn on :8000 with reloadScan a document:
curl -F image=@/path/to/document.jpg \
http://localhost:8000/scanmake docker-build
docker run --rm -p 8000:8000 passport-ocrThe image pre-downloads PP-OCRv5 models at build time so the first request is fast.
npm install document-ocrimport { DocumentOCR } from 'document-ocr';
const ocr = new DocumentOCR();
const result = await ocr.scan(imageBuffer);
if (result.status === 'success' && result.documentType === 'passport') {
console.log(result.fields.passportNumber, result.mrzValid);
}
await ocr.stop();The package auto-creates a .venv, installs the Python deps, and manages the local server lifecycle. See packages/passport-ocr/README.md for full options including HTTP mode.
| Method | Path | Description |
|---|---|---|
| GET | /health |
Liveness |
| GET | /ready |
503 until OCR models finish loading |
| POST | /scan |
Multipart image upload, returns scan result JSON |
/scan accepts images and PDFs up to 10 MB. Concurrency is serialized
internally. Set DOCUMENT_OCR_API_TOKEN to require
Authorization: Bearer <token> on /scan. For internet-facing deployments,
use platform IAM or an API gateway in addition to application-level controls.
Incomplete or semantically invalid non-passport extractions return HTTP 422
with the full structured failure result.
preprocess— orientation, document boundary detection, perspective correction, quality checksclassify_passport_page— biodata vs non-biodata vs not-a-passport (cheap bottom-crop probe)- passport path:
run_ocr(RapidOCR PP-OCRv5, full-page fallback when MRZ is missing) →parse_mrz(TD3 MRZ with per-field + overall checksum validation) →extract_back_page(bilingual label-aware extraction) →validate(cross-checks MRZ vs visual fields, computes confidence) - non-passport path:
classify_documentroutes full-page OCR to the matching extractor (pan/aadhaar/driving_licence/voter_id/nrega_job_card/npr_letter), checks minimum required fields, validates identifiers where possible, and fails closed on partial or semantically implausible records
For non-passport documents, status: "success" means the extractor returned
the document-specific minimum field set and passed a conservative OCR-region
confidence gate. Alternate model readings for the same detected geometry count
once. identifierValid means only an offline format/checksum check passed. NPR
references have no universal public checksum, so this value is null; it is
never evidence of authenticity.
The default recognition behavior is unchanged (en, with the existing
automatic fallback). If the expected KYC population uses a known script, run
up to four recognition passes only on the non-passport path:
DOCUMENT_OCR_KYC_LANGS=en,devanagari make devAvailable values are en, latin, devanagari, ka (Kannada), ta
(Tamil), and te (Telugu). Extra models increase latency and memory. The
bundled RapidOCR version has no Bengali recognition model: Bengali-labelled
extractor fixtures prove parsing behavior only, not Bengali image OCR. Select
languages and release thresholds from measured benchmark slices. Server
readiness initializes every configured model so model-download or startup
failures are reported before the first scan. Unsupported names or more than
four configured models fail readiness instead of silently falling back.
Single entry point: core.pipeline.scan(image_input).
The same pipeline runs as a container, on Cloud Run, or on AWS Lambda. Copy
.env.deploy.example to .env.deploy.<env> and fill in the relevant values
first (Docker Hub and/or GCP).
Cloud Run and Lambda deployments are authenticated by default. Do not expose an identity-document endpoint publicly without authentication, authorization, rate limiting, encrypted transport, and an explicit data-retention policy.
# Build deploy/docker/Dockerfile, push to Artifact Registry, deploy the service.
bash deploy/cloudrun/deploy.sh production
# or via Cloud Build:
gcloud builds submit --config deploy/cloudrun/cloudbuild.yaml \
--substitutions=_REGION=asia-south1,_SERVICE=document-ocrService config lives in deploy/cloudrun/service.yaml
(2 GiB / 2 vCPU, containerConcurrency: 1 since OCR is serialized, scale-to-zero).
The SDK's mode: 'lambda' invokes a function that takes { "image_base64": "..." }.
cd deploy/lambda
sam build && sam deploy --guided # uses Dockerfile.lambda + template.yamlThe handler (deploy/lambda/handler.py) returns the
raw scan result on success and an {statusCode, body} envelope for errors.
The included SAM template supports IAM-authenticated direct invocation and does
not create a public API Gateway endpoint.
make test
make buildThe per-document extractors are covered by deterministic TextRegion fixtures
under tests/python/test_*_extractor.py. These tests verify parsing and
validation behavior; they are not a claim of real-world OCR accuracy.
For the legacy passport image benchmark, place its private dataset and
manifest.json under benchmark-data/ and run:
make benchmarkFor a non-passport KYC dataset, use the versioned manifest and release gates:
make benchmark-kyc \
KYC_MANIFEST=/secure/kyc-eval/manifest.json \
KYC_DATASET_ROOT=/secure/kyc-eval \
KYC_REPORT=/secure/kyc-eval/reports/main.jsonThe KYC evaluator reports classification, acceptance/false-success,
exact/normalized field, complete-record, runtime, per-document, and
design/year/issuer/language/capture-quality slice metrics without copying
ground-truth values into its report. Required per-document slices fail manifest
validation when a declared variant is absent. See
benchmarks/KYC_DATASET.md for the variant matrix
and annotation workflow.
benchmark-data/ is ignored by Git. Never commit identity documents or
personal data. See CONTRIBUTING.md for fixture rules.
Local Python and npm-local modes process documents on the same machine and do not include telemetry or document storage. HTTP and Lambda modes transmit documents to infrastructure controlled by the operator.
Read PRIVACY.md before deploying and report vulnerabilities using SECURITY.md.
MIT. Dependencies retain their own licenses; see THIRD_PARTY_NOTICES.md.
{ "status": "success", // success | failure | unsupported_page "documentType": "passport", // passport | pan | aadhaar | driving_licence | voter_id | nrega_job_card | npr_letter | unknown "pageType": "passport_biodata", // same document values, plus passport_biodata | passport_non_biodata | unknown "confidence": 0.91, "fields": { "surname": "...", "givenNames": "...", "fullName": "...", "passportNumber": "...", "nationality": "IND", "dateOfBirth": "1990-05-21", "sex": "M", "expiryDate": "2030-04-12", "issueDate": "2020-04-13", "placeOfBirth": "...", "countryCode": "IND" }, "backPageFields": { "fatherName": "...", "motherName": "...", "spouseName": "...", "address": "...", "pincode": "...", "city": "...", "state": "...", "fileNumber": "...", "oldPassportNumber": "...", "oldPassportDateOfIssue": "...", "oldPassportPlaceOfIssue": "..." }, // Exactly one document block is populated, keyed by documentType; the rest // are null. (passport uses fields/backPageFields above.) "panFields": { "panNumber": "...", "name": "...", "fatherName": "...", "dateOfBirth": "..." }, "aadhaarFields": { "aadhaarNumber": "...", "name": "...", "dateOfBirth": "...", "yearOfBirth": null, "gender": "...", "address": "...", "pincode": "...", "checksumValid": true, "aadhaarMasked": false, "aadhaarLast4": "...", "vid": null }, "drivingLicenceFields": { "dlNumber": "...", "name": "...", "dateOfBirth": "...", "issueDate": "...", "validityDate": "...", "address": "...", "relationName": "...", "bloodGroup": "...", "classOfVehicle": "...", "validityDateTransport": null }, "voterIdFields": { "epicNumber": "...", "name": "...", "relationName": "...", "relationType": "father", "gender": "...", "dateOfBirth": "...", "age": null }, "nregaJobCardFields": { "jobCardNumber": "...", "headOfHousehold": "...", "category": "SC", "registrationDate": "...", "validityFrom": "...", "validityTo": "...", "address": "...", "village": "...", "gramPanchayat": "...", "block": "...", "district": "...", "state": "...", "bplStatus": true, "familyId": "...", "members": [{ "serialNumber": "1", "name": "...", "fatherOrHusbandName": "...", "gender": "FEMALE", "age": 39 }] }, "nprLetterFields": { "referenceNumber": "...", "name": "...", "address": "...", "pincode": "...", "issueDate": "..." }, "mrzRaw": ["P<IND...", "..."], "mrzValid": true, "lowConfidence": false, "identifierValid": null, // format/checksum only; never an authenticity verdict "missingRequiredFields": [], "errors": [], "warnings": [], "processingMs": 412 }