Appearance
Run a pilot
A checklist for a first real deployment: one tenant, a handful of reviewers, real traffic at pilot volume. It sequences the operate pages (Deployment, Operations, Releasing, Performance) and links to their sections rather than repeating them. Work through it top to bottom; each item names what "done" looks like.
1. Choose how to run it
- [ ] One VM with Docker Compose or Kubernetes with the Helm chart. Both run the same four images (
api,worker,migrate,web) published to GHCR with each platform release (Images). Compose is the shorter path for a pilot (Compose runbook); the chart (Helm chart,deploy/helm/README.mdin the repository) suits a team that already runs Kubernetes and brings its own Postgres, Temporal, Zitadel and object storage. - [ ] Pick the platform version (
vX.Y.Ztags;-rcmarks a prerelease) and read its release notes, in particular the migration and workflow-compatibility sections (Versioning). - [ ] Size it from the load-test numbers: for about 3 runs/s sustained, one API replica, one worker replica, a Temporal server with 2 CPUs and a dedicated Postgres with 2 CPUs and 4 GB is the starting point (Sizing).
2. DNS and TLS
- [ ] Five host names under your domain point at the deployment:
api.,console.(admin console),collect.(collection terminal),auth.(Zitadel) ands3.(object storage, reached directly by browsers for uploads) (DNS and the host). - [ ] Certificates: the compose stack's Caddy obtains and renews them from Let's Encrypt and proxies Zitadel with HTTP/2 (TLS); on Kubernetes your ingress does the same.
- [ ] The browser origins are configured on the bucket (
STORAGE_CORS_ORIGINS: theconsole.andcollect.origins) and, for merchants embedding the collection flow, in the tenant's embed origins and the web image'sEMBED_ORIGINS(Origin allow-list). The console and the terminal reach the API through their own hosts, soAPI_CORS_ORIGINSneeds neither.
3. Zitadel and the tenant
- [ ] Zitadel runs with its own masterkey and database; the bootstrap PAT it writes on first start is the only credential the seed needs (First start).
- [ ]
pnpm --filter @aletheia-dev/auth seedagainst the publicauth.host created the project, the roles and the two applications (the console's sign-in and the API's token introspection); the printed ids and the API client secret are in the environment file (Seed Zitadel). - [ ] The organisation is linked to a tenant (
pnpm tenant:add --org <org id> --name <name>); the Zitadel admin has changed the initial password and your reviewers have accounts with theadminoranalystrole, your backend a service user with theintegrationrole and a PAT. - [ ] Tenant settings are set, in the console's Settings › SLA and Settings › Tenant: the default case SLA, whether publishing needs approval and four-eyes (Approval settings), the embed origins.
4. Secrets
- [ ] Every secret is generated (
openssl rand -hex 32), stored outside the repository and never in an image: database passwords, the Zitadel masterkey and API client secret, the directory PAT,SUBMISSION_TOKEN_SECRET(32 characters or more), the object-storage keys,SECRET_STORE_KEY,APP_RUNNER_TOKENandAPP_CALL_TOKEN_SECRET(Environment file, Environment reference). - [ ] The environment file (
chmod 600) or the Kubernetes Secret is backed up with the data: without the Zitadel masterkey the Zitadel database cannot be read, and withoutSECRET_STORE_KEYthe stored app secrets and webhook signing secrets cannot be opened. - [ ] The vendor apps the pilot needs are installed for the tenant, their secrets stored in the console (Connect › Apps) and a test call made from each app's page; the ones it does not need stay uninstalled (App catalogue).
5. Migrate, start, verify
- [ ] The
migratejob ran first and exited 0; the API and worker started after it (Migrate and start).SCHEMA_CHECK=strictis set on the API and worker (the production compose does; the chart documents it) so a process refuses to start when the schema is behind its migrations (Environment reference). - [ ]
GET /health/readyon the public API reportsdbandtemporalok, and the API log at start-up names the configured storage origins rather thandocuments disabled. - [ ] The app runner answers
GET /health/readyinside the network, the API loggedcatalogue reconciledat start-up, and Connect › Apps lists the platform apps (the health indicator reports appsoffwhen the runner or object storage is missing). - [ ] A reviewer can sign in to
console.; your backend can call the API with its PAT; a collection link opens oncollect.. Run the tutorial against the pilot API with the pilot's own tokens: it is the acceptance test.
6. Backups
- [ ] Postgres (Aletheia, Temporal and Zitadel share it in the compose stack) and the object storage volume are dumped on a schedule, consistently, and a restore has been rehearsed once (Backups).
- [ ] Retention is agreed: the audit trail is append-only and the export with its digest is how you hand it to an auditor (Exports).
7. Observability
- [ ]
OTEL_EXPORTER_OTLP_ENDPOINTpoints at a collector; traces, metrics and logs carry the request id and the tenant (Telemetry environment, Logs and correlation). - [ ] Prometheus scrapes the API and worker
/metricsinside the network; the provisioned Grafana dashboards (Operations, Decisions) and alert rules fromdeploy/grafana/are loaded (Metrics). - [ ] Someone receives the alerts, and each alert's runbook is a click away (Runbooks).
8. Before every upgrade
- [ ] Read the release notes for
migration:destructiveandworkflow:incompatible; a destructive migration is a major release with an operator note (Migration gates, Workflow compatibility). - [ ] Replay a sample of in-flight runs' histories against the new worker image before rolling it; a
FAILmeans drain the running workflows first, whatever the label says (Replaying production histories before a deploy). - [ ] Roll in order: migrate, worker, API, app runner, UIs (Upgrades); to roll back, redeploy the previous version without the migrate step (Rolling back).
9. The first week
Watch the five alerts and work their runbooks:
- [ ] App failures: vendor outages, rate limits, rotated secrets, the app runner (runbook).
- [ ] Workflow task failures: almost always a non-deterministic change after a deploy (runbook).
- [ ] HTTP errors: 5xx by route, with the request id to find the log line (runbook).
- [ ] SLA breaches: a staffing signal; the
Breached SLAqueue is where reviewers look (runbook, SLAs). - [ ] Stale approvals: usually four-eyes with a single admin (runbook).
Also look at the decision mix (aletheia.decision.count by outcome and source) and the case queue depth: a policy that sends everything to manual review is the most common pilot finding, and tuning it is a rule set change with a diff and a backtest, not a deploy (Rule sets and aggregation, Testing a rule before it decides anything).