Skip to content

Run a pilot ​

A checklist for a first real deployment: one tenant, a handful of reviewers, real traffic at pilot volume. It sequences the operate pages (Deployment, Operations, Releasing, Performance) and links to their sections rather than repeating them. Work through it top to bottom; each item names what "done" looks like.

1. Choose how to run it ​

  • [ ] One VM with Docker Compose or Kubernetes with the Helm chart. Both run the same four images (api, worker, migrate, web) published to GHCR with each platform release (Images). Compose is the shorter path for a pilot (Compose runbook); the chart (Helm chart, deploy/helm/README.md in the repository) suits a team that already runs Kubernetes and brings its own Postgres, Temporal, Zitadel and object storage.
  • [ ] Pick the platform version (vX.Y.Z tags; -rc marks a prerelease) and read its release notes, in particular the migration and workflow-compatibility sections (Versioning).
  • [ ] Size it from the load-test numbers: for about 3 runs/s sustained, one API replica, one worker replica, a Temporal server with 2 CPUs and a dedicated Postgres with 2 CPUs and 4 GB is the starting point (Sizing).

2. DNS and TLS ​

  • [ ] Five host names under your domain point at the deployment: api., console. (admin console), collect. (collection terminal), auth. (Zitadel) and s3. (object storage, reached directly by browsers for uploads) (DNS and the host).
  • [ ] Certificates: the compose stack's Caddy obtains and renews them from Let's Encrypt and proxies Zitadel with HTTP/2 (TLS); on Kubernetes your ingress does the same.
  • [ ] The browser origins are configured on the bucket (STORAGE_CORS_ORIGINS: the console. and collect. origins) and, for merchants embedding the collection flow, in the tenant's embed origins and the web image's EMBED_ORIGINS (Origin allow-list). The console and the terminal reach the API through their own hosts, so API_CORS_ORIGINS needs neither.

3. Zitadel and the tenant ​

  • [ ] Zitadel runs with its own masterkey and database; the bootstrap PAT it writes on first start is the only credential the seed needs (First start).
  • [ ] pnpm --filter @aletheia-dev/auth seed against the public auth. host created the project, the roles and the two applications (the console's sign-in and the API's token introspection); the printed ids and the API client secret are in the environment file (Seed Zitadel).
  • [ ] The organisation is linked to a tenant (pnpm tenant:add --org <org id> --name <name>); the Zitadel admin has changed the initial password and your reviewers have accounts with the admin or analyst role, your backend a service user with the integration role and a PAT.
  • [ ] Tenant settings are set, in the console's Settings › SLA and Settings › Tenant: the default case SLA, whether publishing needs approval and four-eyes (Approval settings), the embed origins.

4. Secrets ​

  • [ ] Every secret is generated (openssl rand -hex 32), stored outside the repository and never in an image: database passwords, the Zitadel masterkey and API client secret, the directory PAT, SUBMISSION_TOKEN_SECRET (32 characters or more), the object-storage keys, SECRET_STORE_KEY, APP_RUNNER_TOKEN and APP_CALL_TOKEN_SECRET (Environment file, Environment reference).
  • [ ] The environment file (chmod 600) or the Kubernetes Secret is backed up with the data: without the Zitadel masterkey the Zitadel database cannot be read, and without SECRET_STORE_KEY the stored app secrets and webhook signing secrets cannot be opened.
  • [ ] The vendor apps the pilot needs are installed for the tenant, their secrets stored in the console (Connect › Apps) and a test call made from each app's page; the ones it does not need stay uninstalled (App catalogue).

5. Migrate, start, verify ​

  • [ ] The migrate job ran first and exited 0; the API and worker started after it (Migrate and start). SCHEMA_CHECK=strict is set on the API and worker (the production compose does; the chart documents it) so a process refuses to start when the schema is behind its migrations (Environment reference).
  • [ ] GET /health/ready on the public API reports db and temporal ok, and the API log at start-up names the configured storage origins rather than documents disabled.
  • [ ] The app runner answers GET /health/ready inside the network, the API logged catalogue reconciled at start-up, and Connect › Apps lists the platform apps (the health indicator reports apps off when the runner or object storage is missing).
  • [ ] A reviewer can sign in to console.; your backend can call the API with its PAT; a collection link opens on collect.. Run the tutorial against the pilot API with the pilot's own tokens: it is the acceptance test.

6. Backups ​

  • [ ] Postgres (Aletheia, Temporal and Zitadel share it in the compose stack) and the object storage volume are dumped on a schedule, consistently, and a restore has been rehearsed once (Backups).
  • [ ] Retention is agreed: the audit trail is append-only and the export with its digest is how you hand it to an auditor (Exports).

7. Observability ​

  • [ ] OTEL_EXPORTER_OTLP_ENDPOINT points at a collector; traces, metrics and logs carry the request id and the tenant (Telemetry environment, Logs and correlation).
  • [ ] Prometheus scrapes the API and worker /metrics inside the network; the provisioned Grafana dashboards (Operations, Decisions) and alert rules from deploy/grafana/ are loaded (Metrics).
  • [ ] Someone receives the alerts, and each alert's runbook is a click away (Runbooks).

8. Before every upgrade ​

  • [ ] Read the release notes for migration:destructive and workflow:incompatible; a destructive migration is a major release with an operator note (Migration gates, Workflow compatibility).
  • [ ] Replay a sample of in-flight runs' histories against the new worker image before rolling it; a FAIL means drain the running workflows first, whatever the label says (Replaying production histories before a deploy).
  • [ ] Roll in order: migrate, worker, API, app runner, UIs (Upgrades); to roll back, redeploy the previous version without the migrate step (Rolling back).

9. The first week ​

Watch the five alerts and work their runbooks:

  • [ ] App failures: vendor outages, rate limits, rotated secrets, the app runner (runbook).
  • [ ] Workflow task failures: almost always a non-deterministic change after a deploy (runbook).
  • [ ] HTTP errors: 5xx by route, with the request id to find the log line (runbook).
  • [ ] SLA breaches: a staffing signal; the Breached SLA queue is where reviewers look (runbook, SLAs).
  • [ ] Stale approvals: usually four-eyes with a single admin (runbook).

Also look at the decision mix (aletheia.decision.count by outcome and source) and the case queue depth: a policy that sends everything to manual review is the most common pilot finding, and tuning it is a rule set change with a diff and a backtest, not a deploy (Rule sets and aggregation, Testing a rule before it decides anything).

Getting help ​

  • Questions and bugs: by email, as Support describes, with what a bug report needs and what response to expect. A security problem goes through Security.
  • What is deliberately not there yet, and why: the roadmap.

Released under the Apache-2.0 License.