Skip to content

Documents ​

Applicants and operators upload files (identity documents, proofs of address, contracts) that land in object storage, never on the API. The API only issues short-lived presigned URLs, verifies what was uploaded and records the result; the worker scans and neutralises the file afterwards. A collection-flow file field holds a reference to a document, and a submission cannot be submitted until every referenced document is clean.

Packages:

  • @aletheia-dev/storage — the ObjectStorage interface, the S3 implementation (createS3Storage, storageFromEnv) and MemoryStorage for tests.
  • @aletheia-dev/documents — checkUpload (sniffing and hashing), processDocument (re-encoding, PDF inspection), the DocumentScanner interface and NoopScanner.
  • @aletheia-dev/db — repos.documents (create, get, getMany, update, guarded transition, list) over the documents table (migration 0008, row-level security).
  • @aletheia-dev/workflow-engine — startDocumentScan and the scanDocument workflow + activity.
  • apps/api — /documents routes; apps/worker — runs the scan activity.

Upload flow ​

  1. Ticket. POST /documents with { fileName, contentType, sizeBytes, subjectId? | submissionId?, fieldKey?, caseId? } creates a pending row and returns a DocumentUploadTicket:

    json
    {
      "documentId": "…",
      "upload": {
        "method": "PUT",
        "url": "http://localhost:3900/aletheia-documents/<tenant>/<subject>/<document>?X-Amz-…",
        "headers": { "Content-Type": "image/png" },
        "expiresAt": "…"
      }
    }

    The URL is a presigned PUT for 15 minutes, signed for the object's key, the declared Content-Type and the declared sizeBytes as Content-Length: storage refuses a body of another type or size, or under another key. The size is the one the client declared, not the global maximum.

  2. Direct upload. The browser sends the file as the body of a PUT to upload.url with upload.headers (uploadToStorage in @aletheia-dev/api-client); the browser sets the Content-Length itself. With curl: curl -X PUT -H 'Content-Type: image/png' --data-binary @file.png '<upload.url>'. Nothing passes through the API.

  3. Finalize. POST /documents/:id/finalize reads the object back and verifies it (below). On success the document becomes uploaded and the scan workflow is started; on a mismatch the object is deleted and the document becomes rejected with a rejectionReason and a rejectionCode (Rejections). Finalize is idempotent: a document already past pending is returned as-is (rejected and infected answer 409).

  4. Scan. The worker runs scanDocument (Temporal, on the aletheia-apps task queue, TEMPORAL_APP_TASK_QUEUE): scanning, then clean or infected. Clean images are replaced by a re-encoded copy. The verdict is stored in scan: { engine, signature?, scannedAt }.

  5. Reference. The collection flow's file field stores { fileId: documentId, name? }, or an array of them when the field takes several files (multiple, at most maxFiles, 10 unless set lower). A field's capture (camera, upload or both) and documentKind (passport, id_card, proof_of_address, selfie) shape how the collection terminal asks for the file: the camera alone or the device's files alone, the front camera for a selfie, and a line saying what to take.

Status model ​

StatusMeaning
pendingRow exists, nothing verified yet. The ticket may or may not have been used.
uploadedFinalize verified size, type and hash; the scan is queued.
scanningThe scan workflow owns the document.
cleanScanned and neutralised; downloadable; accepted by submit.
infectedThe scanner found malware; the object is deleted, the row keeps the verdict.
rejectedFinalize or processing refused the file (type mismatch, size, undecodable image, PDF with active content). The object is deleted.

Transitions out of pending and uploaded/scanning go through repos.documents.transition, an UPDATE … WHERE status = ANY(from), so a finalize retried concurrently or racing the scan cannot overwrite a later state.

Rejections ​

A rejected document carries rejectionReason, written for operators, and rejectionCode, which clients explain in their own words (the collection terminal never shows the reason):

CodeSet byMeaning
emptyfinalizethe file is empty
too_largefinalize, the scanover the 20 MiB limit
size_mismatchfinalizenot the size the client declared
type_mismatchfinalizean accepted type, but not the one declared (a renamed file)
unsupported_typefinalizenot an image or a PDF that is accepted
undecodable_imagethe scanan image that cannot be read (or over 50 megapixels)
pdf_active_contentthe scana PDF with scripts, actions or embedded files
object_missingthe scannothing in storage when the scan read it

What finalize checks ​

  • head of the object: 409 conflict ("nothing uploaded yet") when it does not exist.
  • The object is read with a hard cap of DOCUMENT_MAX_BYTES (20 MiB); larger objects are rejected.
  • checkUpload (@aletheia-dev/documents): the file is not empty, its size equals the declared sizeBytes, the type sniffed from the magic bytes (file-type) is one of the allowed types and equals the declared one. The declared type is never trusted; the sniffed one is what gets stored. A text file declared as image/png is rejected with declared image/png but the file is not a recognised image or PDF.
  • sha256 of the bytes is recorded.

What the scan step does ​

processDocument(bytes, { contentType, fileName }, scanner) in @aletheia-dev/documents:

  1. Runs the configured DocumentScanner (scan(bytes, meta) -> { verdict, signature? }). infected ends processing; the object is deleted.
  2. Images (image/jpeg, image/png, image/webp) are re-encoded with sharp (limitInputPixels 50 MP): this strips metadata, neutralises polyglot files and rejects anything that does not decode. The re-encoded bytes replace the object.
  3. PDFs are inspected for active content: /JavaScript, /JS, /OpenAction, /Launch (see PDF_DANGEROUS_TOKENS) cause a rejection. PDFs are otherwise stored unchanged.

A clean image or PDF then gets a thumbnail (renderThumbnail): a grayscale WebP that fits in 320 by 480 pixels, of the image or of the PDF's first page. PDFium (WebAssembly) renders the page in a worker thread that is stopped after 15 seconds, so a hostile file can hold up only that thread. The thumbnail is stored beside the original (<key>.thumbnail.webp) and its key recorded as thumbnailKey; a file that does not render is still clean, without a thumbnail.

The scanner interface ships with NoopScanner (engine: 'none', everything clean); it is what DOCUMENT_SCANNER=none selects and the only engine today. A ClamAV implementation is deferred (docs/plans/deferred.md). The status model and the scanning state exist so a real scanner is a drop-in.

Storage layout ​

Keys are server-chosen: ${tenantId}/${subjectId}/${documentId}, and the thumbnail's is the same with .thumbnail.webp appended. Clients never pick paths; the policy is bound to the exact key. One bucket (STORAGE_BUCKET) holds every tenant; isolation is by prefix plus the API's tenant scoping. Rejected and infected objects are deleted; the row keeps the reason or verdict.

Environment ​

VariablePurpose
STORAGE_ENDPOINTS3 endpoint the API and worker talk to. Unset = documents disabled (/documents answers 503 storage_unavailable, the worker registers no scan activity).
STORAGE_PUBLIC_ENDPOINTEndpoint browsers reach for the presigned URLs (defaults to STORAGE_ENDPOINT).
STORAGE_REGIONSigning region (garage for the dev stack).
STORAGE_BUCKETBucket name (aletheia-documents).
STORAGE_ACCESS_KEY, STORAGE_SECRET_KEYCredentials; required when the endpoint is set (startup fails otherwise).
STORAGE_CORS_ORIGINSComma-separated browser origins. The API applies them to the bucket at startup (ensureCors) and warns when empty.
DOCUMENT_SCANNERScanner engine for the worker; none (default) is the only value today.

The API logs documents disabled: STORAGE_ENDPOINT not set when storage is off, and the configured origins when it is on. The API start-up log reports documents: true|false; readiness (GET /health/ready) checks the database and Temporal only, while the console's health summary (GET /health/summary) also probes the bucket and reports storage off when it is not configured.

Dev stack: Garage ​

docker compose up -d starts garage (dxflrs/garage:v2.4.1, single node, S3 on port 3900) with the bucket and key created from GARAGE_DEFAULT_* on first start. Values for a local API and worker:

STORAGE_ENDPOINT=http://localhost:3900
STORAGE_PUBLIC_ENDPOINT=http://localhost:3900
STORAGE_REGION=garage
STORAGE_BUCKET=aletheia-documents
STORAGE_ACCESS_KEY=GKaletheiadev0000000000000
STORAGE_SECRET_KEY=aletheiadevsecret0000000000000000000000000000000000000000000000ab
STORAGE_CORS_ORIGINS=http://localhost:5176,http://localhost:5177

MinIO is not used: its community images were discontinued in 2025.

Production: S3-compatible storage ​

The S3 implementation uses @aws-sdk/client-s3 with path-style addressing, and asks a store for nothing beyond the S3 API's common core: presigned PUTs and GETs, head, get, put and delete, and the bucket's CORS rules (sent with a Content-MD5 checksum, which every store accepts). AWS S3, Cloudflare R2, OVH Object Storage, MinIO and Garage all serve it. For Cloudflare R2:

bash
STORAGE_ENDPOINT=https://<account id>.r2.cloudflarestorage.com
STORAGE_REGION=auto
STORAGE_BUCKET=<bucket>
STORAGE_ACCESS_KEY=<R2 API token access key id>
STORAGE_SECRET_KEY=<R2 API token secret>

The token needs Object Read & Write on the bucket and the right to change its CORS rules (an Admin Read & Write token, or set the rules in the dashboard and leave STORAGE_CORS_ORIGINS empty).

Use a private bucket with server-side encryption. Bucket versioning and Object Lock (evidence retention) are deferred: the interface does not depend on them and nothing in the API would change.

Permissions ​

PermissionRoles
documents:readadmin, analyst, applicants (own submission only)
documents:writeadmin, integration, applicants (own submission only)

Applicants carry a submission token. The auth hook authorizes them only when the route's resource selector names their submission, so applicant calls must say which submission they act on:

  • POST /documents — body.submissionId must be the applicant's own submission, and fieldKey must be a file field of that submission's flow (400 otherwise). The subject is the submission's subject. Operators and integrations send subjectId (or submissionId, from which the subject is derived). A submissionId must name an in_progress submission (409 otherwise).
  • POST /documents/:id/finalize, GET /documents/:id, GET /documents/:id/download, GET /documents/:id/thumbnail — applicants add ?submissionId=<own>; the handler also checks document.submissionId matches (403 otherwise).
  • GET /documents — applicants pass ?submissionId=<own>; everyone else must filter by at least one of subjectId, submissionId, caseId (400 otherwise). status, limit, offset are optional.
  • DELETE /documents/:id?submissionId=<own> — applicants only (403 for everyone else), while their submission is in_progress (409 otherwise): a file replaced or removed in the form does not stay attached for operators. The row, the object and the thumbnail are deleted; a document that is uploaded or scanning answers 409 until the scan settles, and an infected one stays as the scanner's evidence (409). An answer that still names the file fails the submit check.

Downloads are 60-second presigned GETs with Content-Disposition: attachment, only for clean documents (409 otherwise). GET /documents/:id/thumbnail answers the same kind of URL for the thumbnail ({ url, expiresAt }), or 404 when the document is not clean or has none; unlike a download it is not audited. Files are never served from the API origin. The collection terminal shows the thumbnail of a PDF and of a file saved on another device (its own picture of an image picked in the tab otherwise), so its Content-Security-Policy allows images from object storage.

Audit actions ​

document.created, document.uploaded ({ sha256, contentType, sizeBytes }), document.rejected ({ reason, code }), document.scanned (verdict, from the worker), document.downloaded, document.deleted ({ fileName, submissionId, fieldKey, status }, by the applicant). Resource type document.

Submit rule ​

POST /collection-submissions/:id/submit runs the schema validation first, then fileFieldValues(flow, data) (@aletheia-dev/collection-flow) collects the { fileId } values of the visible file fields on the data's path through the flow, each file of a field that takes several on its own. Every referenced document must exist, belong to this submission and be clean; otherwise the submit is a 400 validation_error with one issue per offending file, named by its field, and by its place in the field for one of several (statements.1):

json
{ "issues": [{ "path": "proof", "message": "document <id> is rejected" }] }

Messages: document <id> not found, document belongs to another submission, document <id> is <status>.

Verification ​

Document verification is an app capability, document.verify: a call_app step hands a clean document to a verification app and gets back a verdict. Two platform apps provide it: doc-verify-mock (deterministic, no account needed; what the seed and the smoke use) and sumsub (the real vendor, see App catalogue). Both expose an asynchronous action verify with input { documentId, applicant: { fullName? | firstName?, lastName?, dateOfBirth? } } and output

json
{
  "outcome": "approved | declined | review",
  "checks": { "nameMatch": true, "expired": false, "mrzValid": true },
  "extracted": { "fullName": "Jane Doe", "documentNumber": "…", "expiryDate": "2031-01-01" },
  "engine": "mock"
}

Apps never receive a presigned URL or a storage credential. An app that declares needs: ['documents'] reads a document through the host: as a request body the host streams to the vendor, or into the module with document_read. The host fetches it from the API's internal listener (GET /internal/app-calls/:callId/documents/:id) with a token that names the tenant and the call and expires after five minutes, and the API serves only the tenant's own documents with status clean, at most ten per call: a pending, rejected or infected document, or another tenant's id, is refused. The bytes are the processed copy described above, never the raw upload (Apps: host interface).

How the seeded KYB flow uses it ​

The kyb-basic flow's review step carries an optional identityDocument file field ("Identity document of the owner"). The kyb-onboarding workflow is

collect -> screen -> has_document -> idv -> rules -> route
                         \___________________/

has_document is a branch on submission.identityDocument.fileId (exists): with a document the run calls verify of the doc-verify-mock app with { documentId: submission.identityDocument.fileId, applicant: { fullName: submission.uboName } } and stores the output under idv; without one it goes straight to the rules. The seeded rule identity_not_verified (CEL, warn, weight 40) fails when idv.outcome exists and is not approved, so a submission without a document passes it and a declined or inconclusive one adds 40 to the risk score. It is part of the kyb-onboarding rule set.

The step is asynchronous: the run sits in waiting_callback (step idv) until the vendor's webhook arrives at POST /webhooks/apps/doc-verify-mock, which completes the app's session and signals the run (Apps: reference). The mock calls no vendor, so on the development stack nothing posts that webhook by itself: pnpm mock:callback <callback-url> <external-id> --idv plays the vendor, with the external id from idvExternalId in the run context (App catalogue); the smoke posts it itself.

What the applicant sees ​

After submitting, the collection terminal polls GET /collection-submissions/:id/status every two seconds, for up to ten minutes (submissions:read, resource = the submission, so applicant tokens work on their own submission; operators and integrations on any):

json
{
  "submission": { "id": "…", "status": "submitted", "submittedAt": "…" },
  "run": { "id": "…", "status": "waiting_callback", "currentStepId": "idv", "decision": null },
  "next": null
}

run is null for a submission started without a run; decision is { outcome } once the run has one. next names the form the same run now waits for from the applicant (a request for more information, or a later collection step) as { submissionId, link, message }, where only the applicant's own token gets a link; otherwise it is null. Keys are always present (null, never omitted). Nothing else about the run (the context, rule results, the case) is exposed to applicants. The terminal shows "Checking your details" while run.status is waiting_callback and "We're reviewing your application" for waiting_manual, unless the flow shows a thank-you only (outcome.visibility: neutral). Once the run completes it shows a neutral "Thank you", or the decision itself when the flow's outcome.visibility is decision (unset: when the deployment sets COLLECTION_TERMINAL_SHOW_DECISION=true).

The mock vendor ​

doc-verify-mock reads the document through the host and decides from its first bytes and its file name, without OCR: a file that is not a PNG, JPEG or PDF, is shorter than 64 bytes or is named blank reads as nothing (review, every check null); the words expired, invalid and fail in the file name give an expired document, a bad machine-readable zone and a vendor failure; other words that do not describe the file are the name on the document, else the applicant's. Three checks follow: nameMatch (Dice similarity between the applicant's name and the name on the document, threshold nameMatchThreshold, default 0.8), expired and mrzValid. Any failed check is declined; otherwise the verdict is approved when a check ran and review when none could. So passport-valid.png is approved for its applicant and passport-expired.png declined.

It then answers pending with the verdict inside the external id, and the signed webhook { id, externalId } (x-mock-idv-signature: the hex HMAC-SHA256 of the raw body under the install's webhookSecret, default mock-idv-webhook-secret) completes the call with it. The details are in the App catalogue.

Sumsub ​

sumsub is the real vendor: verify creates an applicant with the document (the host streams it into the upload), and applicantReviewed webhooks (signed with X-Payload-Digest) are mapped to the same output shape. While Sumsub still needs something from the applicant (a liveness check), the collection terminal shows its WebSDK after the submit, with tokens from the app's hand-off. Sumsub takes one webhook URL per account in its dashboard, so the app declares where the body carries the applicant's id (webhookExternalId) instead of relying on ?externalId=. Secrets, the webhook setup and sandbox simulation are described in the App catalogue. Switching the seeded workflow to it is a matter of installing sumsub and changing the idv step's app.

Limits ​

  • Size: 20 MiB (DOCUMENT_MAX_BYTES), checked when the document is created, enforced by the presigned URL (signed for the declared size) and again on finalize.
  • Types: image/jpeg, image/png, image/webp, application/pdf (DOCUMENT_CONTENT_TYPES). The declared type must equal the sniffed type.
  • HEIC is excluded: images are re-encoded on ingest and the libvips bundled with sharp cannot decode HEVC, so HEIC uploads would always fail processing. Clients convert to JPEG before uploading.
  • Images: at most 50 megapixels (IMAGE_MAX_PIXELS); larger images are rejected as undecodable.

Released under the Apache-2.0 License.