Skip to content

Deployment ​

Aletheia ships as five container images built from one Dockerfile, a production docker compose stack for a single VM, and a Helm chart for Kubernetes. This page covers the images and the compose deployment; the chart has its own section at the end.

Images ​

ImageTargetRunsPort(s)
ghcr.io/akhiljames/aletheia-apiapinode dist/main.js (apps/api)4000, 4001 (API_INTERNAL_PORT, the app runner's; never published)
ghcr.io/akhiljames/aletheia-workerworkernode dist/main.js (apps/worker)WORKER_HEALTH_PORT, METRICS_PORT (optional)
ghcr.io/akhiljames/aletheia-app-runnerapp-runnernode dist/main.js (apps/app-runner): the WebAssembly apps4100, METRICS_PORT (optional)
ghcr.io/akhiljames/aletheia-migratemigrate@aletheia-dev/db dist/migrate.js, exits 0 when donenone
ghcr.io/akhiljames/aletheia-webwebnginx serving the admin console, the collection terminal and the landing page5176, 5177, 5178

Tags, pushed by pnpm ship (scripts/ci-local/main/images.sh: one docker buildx bake run over docker-bake.hcl, so the shared build stages run once for all five images):

  • sha-<short commit> for linux/amd64 when a merge into main changes something an image is built from. Commits that only touch tests, docs, workflows or deployment files build nothing, so not every commit on main has a tag: deploy the newest sha- tag that exists (gh api /users/akhiljames/packages/container/aletheia-api/versions --jq '.[].metadata.container.tags[]').
  • <version>, latest and sha-<short commit> for linux/amd64 and linux/arm64 when a release is shipped (v1.2.3 gives 1.2.3). One of the two architectures is emulated and several times slower, which is why only releases build both.

Pin a version or a sha- tag in production; latest moves with every release. The images carry org.opencontainers.image.source, .version and .revision labels.

Build locally for the host architecture:

bash
docker build --target api     -t aletheia-api:dev .
docker build --target worker  -t aletheia-worker:dev .
docker build --target app-runner -t aletheia-app-runner:dev .
docker build --target migrate -t aletheia-migrate:dev .
docker build --target web     -t aletheia-web:dev .

Dockerfile layout ​

Multi-stage, one workspace install shared by every target:

  1. base: node:22-bookworm-slim, tini, the pnpm version pinned in package.json (packageManager) through corepack.
  2. deps: copies only pnpm-lock.yaml, pnpm-workspace.yaml and the root package.json, then pnpm fetch fills the store. This layer survives source changes.
  3. compile: on node:22-trixie-slim for the machine running the build (extism-js needs glibc 2.39, newer than bookworm's), installs the pinned extism-js and binaryen (pnpm tools:extism, in a layer only the pin file invalidates), then runs turbo run build scoped to the three Node apps, the three web apps and the catalogue's apps, and pnpm catalog:index, which writes catalog/dist (App catalogue). Its output is the dist/ folders, which do not depend on the platform.
  4. build: copies the source, pnpm install --offline --frozen-lockfile and unpacks the compile stage's dist/ folders. It ends with pnpm deploy --prod --legacy for @aletheia-dev/api, @aletheia-dev/worker and @aletheia-dev/app-runner, which writes a self-contained tree per app under /out/<app>: the app's dist/ and package.json, every @aletheia-dev/* workspace package copied as a published package would be (its files, i.e. dist/ and, for @aletheia-dev/db, drizzle/), and the third-party production dependencies from the store. The worker therefore resolves @aletheia-dev/workflow-engine/workflows on disk for Temporal's workflow bundler, and sharp and @temporalio/core-bridge use the prebuilt binary pnpm selected for the image platform.
  5. api, worker: base plus /out/<app> at /app, user node, tini as PID 1, NODE_ENV=production. The API image also carries the catalogue, catalog/dist, at /app/catalog (APP_CATALOG_DIR). The worker target loads sharp and @temporalio/worker and finds PDFium's WebAssembly file (document thumbnails) during the build, so a broken dependency fails the image build rather than the first activity. Health checks: GET /health on the API; GET /healthz on the worker when WORKER_HEALTH_PORT is set (the check passes when it is not).
  6. app-runner: base plus /out/app-runner at /app, the same user, tini and NODE_ENV. It runs the WebAssembly apps (packages/app-host on @extism/extism) for the API and the worker on port 4100, and health-checks GET /health. The image holds no database driver credentials, storage keys or Zitadel settings: the process needs APP_RUNNER_TOKEN (the bearer token shared with the API and the worker; in production it refuses to start without one), APP_RUNNER_API_URL (the API, where it fetches modules by hash and documents by call token), and reaches the internet for the apps' HTTP requests. APP_RUNNER_CONCURRENCY (8) and APP_RUNNER_TENANT_CONCURRENCY (4) cap the calls in flight per process and per tenant; the compiled modules go to APP_RUNNER_CACHE_DIR (the system temp dir by default). The API and the worker find it through APP_RUNNER_URL.
  7. migrate: the api image with a different command (@aletheia-dev/db's compiled dist/migrate.js, which reads DATABASE_URL and applies drizzle/). No health check.
  8. web: nginx:1.27-alpine with the three Vite builds under /usr/share/nginx/html/<app>, deploy/web/nginx.conf.template and deploy/web/entrypoint.sh. Runs as the nginx user on unprivileged ports.

Nothing in the file is architecture specific; CI builds both platforms from the same stages. Sizes for one architecture: api and migrate about 475 MB, the worker more (Temporal core and sharp), app-runner about 300 MB, web about 55 MB.

pnpm deploy keeps node_modules strict: sharp is reachable from @aletheia-dev/documents, not from /app itself. To poke at a dependency inside a container, resolve it from the package that owns it (the build-time check in the worker target shows how) rather than node -e "require('sharp')" at the top level.

The web image ​

One image serves every environment. At start-up entrypoint.sh writes the admin console's and the landing page's /config.json from the environment and renders the nginx configuration:

VariableUsed for
API_URLthe API's base URL: where the apps' /api requests go unless an _API_UPSTREAM below names another address; never sent to the browser
ZITADEL_ISSUER, ZITADEL_PROJECT_ID, ZITADEL_ORG_DOMAINthe admin console's zitadelIssuer, zitadelProjectId, zitadelOrgDomain
COLLECTION_FLOW_URLthe collection terminal's public address, the admin console's collectionFlowUrl (embed snippets)
EMBED_ORIGINScomma-separated origins every tenant's form links may be framed by; each link's own tenant's embedOrigins join them, from the API (Embedding)
ZITADEL_CLIENT_ID_ADMIN_CONSOLE (optional)zitadelClientId of the admin console; without it the console cannot sign in
ADMIN_CONSOLE_API_UPSTREAM (optional)where nginx forwards the admin console's /api requests; defaults to API_URL
ADMIN_CONSOLE_STORAGE_ORIGINS (optional)object-storage origins allowed by the admin console's Content-Security-Policy; any https: origin when unset
COLLECTION_TERMINAL_API_UPSTREAM (optional)where nginx forwards the collection terminal's /api requests; defaults to API_URL
COLLECTION_TERMINAL_STORAGE_ORIGINS (optional)object-storage origins the terminal's Content-Security-Policy allows (uploads and thumbnails); the console's, then any https:
COLLECTION_TERMINAL_IMAGE_ORIGINS (optional)image origins the terminal's Content-Security-Policy allows for logos, tenants' and the API's default; any https: when empty
COLLECTION_TERMINAL_HANDOFF_SDKS (optional)vendor SDKs applicants may finish a check in after the submit, comma-separated: sumsub (Sumsub's WebSDK, for the sumsub app). Each adds the vendor's hosts to the terminal's Content-Security-Policy and lets their frames use the camera and microphone (Embedding); empty: none
LANDING_API_UPSTREAM (optional)where nginx forwards the landing page's POST /api/signup; defaults to API_URL
LANDING_CONSOLE_URL, LANDING_DOCS_URL (optional)the landing page's consoleUrl and docsUrl: its "Sign in" link and its documentation links (the published docs site by default)
LANDING_TURNSTILE_SITE_KEY (optional)turnstileSiteKey: the Cloudflare Turnstile widget on the signup form. Set it with the API's TURNSTILE_SECRET_KEY or not at all

No app calls the API from the browser directly. The admin console (5176) requests /api/* on its own origin and nginx forwards those to ADMIN_CONSOLE_API_UPSTREAM, normally the API's internal address (http://api:4000 in the compose file, the API Service in the chart). Its config.json has no apiUrl, the API needs no CORS entry for its host, and headers that describe the upstream are removed from the responses. Its pages are sent with a Content-Security-Policy (scripts from its own origin only, with eval allowed for the schema-form validator; requests only to itself, the identity provider and object storage), frame-ancestors 'none', HSTS and a Permissions-Policy. It uploads case attachments straight to object storage, so its origin belongs in STORAGE_CORS_ORIGINS.

The collection terminal (5177), the applicant app (apps/collection-terminal), reaches the API the same way: /api/* on its own origin, forwarded to COLLECTION_TERMINAL_API_UPSTREAM. It has no config.json: each form's look comes from the API, the tenant's branding over the deployment's defaults (the API's COLLECTION_TERMINAL_* variables, Environment), and the page shows only a loader until it arrives. Its pages are sent with a stricter Content-Security-Policy (scripts and styles from its own origin only, no eval; requests only to itself and object storage), frame-ancestors 'self' <EMBED_ORIGINS> followed, on a form link, by its tenant's embed origins (nginx asks the API for them), HSTS and a Permissions-Policy that allows the camera for "Take a photo". COLLECTION_TERMINAL_HANDOFF_SDKS adds a vendor's SDK to both: for sumsub, Sumsub's script, frames, requests and images, and the camera and microphone for its frames only. Its access log records paths without query strings, so the token in a submission link never reaches it, and an address that is not a form answers 404. Applicants upload documents straight to object storage, so its origin belongs in STORAGE_CORS_ORIGINS. COLLECTION_FLOW_URL is its public address: the API and the worker build applicant links (/flow/<id>#token=…) on it, and the API the admin console's flow preview links (/preview#token=…: a draft opened in preview mode, nothing saved). It also serves the embed loader (/embed/aletheia-collect.js, /embed/aletheia-collect.esm.js) to pages on any origin.

The landing page (5178, apps/landing) is the public site: the product with interactive examples (a rule set, a workflow run, a collection form), a page per industry at /solutions/<industry> with worked examples and that industry's policy to try, and the signup form. nginx serves its index.html for every path, and the page picks the view from the path. It reaches the API for one route only: nginx forwards POST /api/signup to LANDING_API_UPSTREAM, limits each client address (CF-Connecting-IP behind Cloudflare, the peer otherwise) to six attempts a minute with a burst of four (429 rate_limited), refuses other methods and answers 404 for every other /api path. Its Content-Security-Policy allows scripts from its own origin and Cloudflare Turnstile's, frames from Turnstile only, requests to itself, and frame-ancestors 'none'. A signup makes a workspace: a Zitadel organisation named after the company with the person as its admin, the Aletheia project granted to it, an empty tenant linked to it, and an email (sent with Resend) with the console link, the organisation domain and a temporary password that Zitadel makes them change at the first sign-in. The API does this when ZITADEL_SIGNUP_PAT, RESEND_API_KEY, SIGNUP_EMAIL_FROM and SIGNUP_CONSOLE_URL are set (Environment); otherwise the form answers 503 signup_unavailable. The API also takes one workspace per email, at most SIGNUP_MAX_PER_HOUR an hour, and with TURNSTILE_SECRET_KEY only forms whose Turnstile token passes. The Zitadel seed creates the aletheia-signup machine user (instance role IAM_ORG_MANAGER) and writes its PAT to signup.pat in its state directory. The compose stack below does not expose the landing page; the Helm chart does with ingress.hosts.landing.

All three servers cache /assets/* (hashed by Vite) as immutable for a year, answer a missing asset with 404 rather than index.html, and send Cache-Control: no-cache for everything else, including index.html, the config.json files and the latest embed loader; a pinned loader (/embed/v1/) is immutable for a year.

Environment reference ​

Every setting the API, worker and migrate images read is listed in docs/env.md, generated from EnvSchema in packages/core/src/env.ts (defaults, descriptions and which processes read them); the app runner has its own schema in apps/app-runner. The compose file below sets the deploy-specific ones from the env file and derives the URLs from DOMAIN.

Two settings matter for how a deployment behaves:

  • SCHEMA_CHECK is strict in production: the API and worker compare the database with the migrations bundled in their image at start-up and refuse to serve when the database is behind. Run the migrate image first, always.
  • WORKER_HEALTH_PORT enables GET /healthz on the worker; the compose file and the chart set it so the health check and readiness probes have something to call.
  • SECRET_STORE_KEY (32 random bytes, base64: openssl rand -base64 32) turns on the secret store, where admins store apps' secrets from the console and where the signing secrets of outbound webhooks live. The API and every worker need the same value, and the stored secrets cannot be opened without it: keep it with the database backups. Without it, no app that needs a secret can be enabled and outbound webhooks are off.
  • Scheduled audit exports run on the workers (backtest queue) and write their files to the object storage documents use, under audit-exports/; without STORAGE_ENDPOINT the API refuses to create schedules. Keep the bucket's retention in line with your audit policy.
  • Outbound webhooks leave from the workers: they need egress to the tenants' https endpoints. Every worker sends deliveries, one at a time reads the audit trail for new events, and both run on an extra owner connection pool of two per worker (DATABASE_URL). WEBHOOK_ALLOW_PRIVATE_NETWORKS lets endpoints use http and private addresses; leave it unset in production.
  • APP_RUNNER_TOKEN (openssl rand -base64 32) authenticates the API and the worker to the app runner and the runner to the API's internal routes; all three need the same value, and the runner refuses to start in production without one. APP_RUNNER_URL on the API and the worker and APP_RUNNER_API_URL on the runner are internal addresses; the compose file and the chart set them.
  • The API listens twice: the public port (API_PORT, 4000, behind Caddy or the Ingress) and API_INTERNAL_PORT (4001), which serves only the runner's routes (/internal/app-modules/… under APP_RUNNER_TOKEN, /internal/app-calls/… under a call token) and must never be published or routed: APP_RUNNER_API_URL points at it (http://api:4001 in the compose file, the API Service's internal port in the chart), and the public port answers 404 for those paths. APP_CALL_TOKEN_SECRET (openssl rand -base64 32) signs the call tokens: the API and the worker share it, the runner never holds it, and without it apps run but cannot read documents. Apps are on only with object storage (STORAGE_ENDPOINT), where the modules live; the API image carries the platform catalogue (APP_CATALOG_DIR=/app/catalog) and publishes it at start. App secrets are stored through the API, so SECRET_STORE_KEY is required wherever apps run.
  • ZITADEL_MANAGEMENT_PAT lets the API manage service users and their keys (Connect › API keys) and invite people, change their roles and deactivate them (Settings › Team); without it both pages send admins to the Zitadel console.
  • A tenant may set its own collection flow URL (Settings › Tenant): links then point at that host, which must route to the web image's collection terminal (a DNS name and a reverse-proxy rule per tenant domain; in the chart, ingress.hosts.collect.extraHosts with the certificates under ingress.tls). The terminal works on any host; the token names the tenant.

Compose runbook (one VM) ​

Files: deploy/compose/docker-compose.prod.yml, deploy/compose/Caddyfile, deploy/compose/.env.example and deploy/compose/postgres-init/. The stack runs Caddy, the five Aletheia images, Postgres 16, Temporal (auto-setup), Zitadel and Garage on one Docker host with resource limits, health checks, unless-stopped restarts and log rotation (json-file, 50 MB x 5 per container).

1. DNS and the host ​

Point these records at the VM: api., console. (admin console), collect. (collection terminal), auth. (Zitadel) and s3. (object storage, reached directly by browsers for uploads), all under DOMAIN. Open 80 and 443 (TCP, plus 443/UDP for HTTP/3). Install Docker Engine with the compose plugin and clone the repository (the compose file mounts docker/temporal and docker/garage from it).

2. Environment file ​

bash
cp deploy/compose/.env.example deploy/compose/.env
chmod 600 deploy/compose/.env

Fill in every value. Generate secrets with openssl rand -hex 32; the Zitadel masterkey must be exactly 32 characters (openssl rand -hex 16). The Zitadel client ids stay empty until step 4. The real .env is never committed (.gitignore covers .env and the Dockerfile context ignores it too).

3. First start ​

bash
cd deploy/compose
docker compose --env-file .env -f docker-compose.prod.yml pull
docker compose --env-file .env -f docker-compose.prod.yml up -d postgres temporal zitadel garage
docker compose --env-file .env -f docker-compose.prod.yml ps   # wait for "healthy"

On the first start of the empty data directory postgres-init/10-app-role.sh creates the aletheia_app role with DATABASE_APP_PASSWORD; Temporal creates its databases; Zitadel initialises its instance, the admin user and the bootstrap machine user whose PAT is written to the zitadelbootstrap volume. Caddy requests certificates as soon as it is up.

4. Seed Zitadel ​

The API and the admin console need a Zitadel project, roles, an API application and the console's OIDC application. The seed creates them on any instance, and gives Zitadel's login page the console's look (the instance's default branding, which every organisation inherits). Run it from a checkout that can reach the public auth. host, with the bootstrap PAT copied out of the volume:

bash
docker compose --env-file .env -f docker-compose.prod.yml cp zitadel:/zitadel/bootstrap/bootstrap.pat /tmp/bootstrap.pat
ZITADEL_ISSUER=https://auth.$DOMAIN \
ZITADEL_BOOTSTRAP_PAT_FILE=/tmp/bootstrap.pat \
ZITADEL_ADMIN_CONSOLE_ORIGIN=https://console.$DOMAIN \
ZITADEL_ADMIN_EMAIL=$ZITADEL_ADMIN_EMAIL \
ZITADEL_SEED_STATE_DIR=$HOME/.aletheia/seed-$DOMAIN \
pnpm --filter @aletheia-dev/auth seed

Against anything but the development stack the seed refuses to guess, so the three extra variables are required: the console origin becomes the OIDC redirect URI (the development default is http://localhost:5176), ZITADEL_ADMIN_EMAIL is the admin from the env file, and ZITADEL_SEED_STATE_DIR is where seed.json, directory.pat and management.pat are written. That directory holds the API client secret and the two PATs, which Zitadel only returns once: keep it (a rerun reuses it instead of rotating them) and protect it like the env file. The directory user (aletheia-directory) holds ORG_OWNER_VIEWER and the management user (aletheia-management) ORG_USER_MANAGER in the seeded organisation; neither holds a project role. For tenants in other organisations, give the management user ORG_USER_MANAGER there too, or IAM_USER_MANAGER on the instance. The development users (an analyst with a published password and two smoke-test machine users, one with the admin role) are not created on a real instance; ZITADEL_SEED_DEV_USERS=true asks for them, for a staging stack that runs the smoke test.

It prints the project id, the client ids, the API client secret and both PATs; copy them into ZITADEL_PROJECT_ID, ZITADEL_CLIENT_ID_ADMIN_CONSOLE, ZITADEL_API_CLIENT_ID, ZITADEL_API_CLIENT_SECRET, ZITADEL_DIRECTORY_PAT, ZITADEL_MANAGEMENT_PAT and ZITADEL_ORG_DOMAIN in .env, then delete /tmp/bootstrap.pat. Link the organisation to a tenant with the command the seed prints (pnpm tenant:add --org <id> --name <tenant name>, which needs DATABASE_URL).

5. Migrate and start the applications ​

bash
docker compose --env-file .env -f docker-compose.prod.yml up -d
docker compose --env-file .env -f docker-compose.prod.yml logs migrate   # "migrations applied"
curl -fsS https://api.$DOMAIN/health/ready

migrate runs once (restart: "no"); api and worker start only after it exited 0 (depends_on: service_completed_successfully), and after Temporal, Zitadel and Garage are healthy; app-runner starts once the API is healthy and publishes no port (the API and the worker reach it as http://app-runner:4100 inside the compose network, and it reaches the API's internal listener as http://api:4001, which Caddy never fronts). Open https://console.$DOMAIN and sign in with the Zitadel admin (password change is required on first login).

Upgrades ​

Change ALETHEIA_VERSION in .env, then roll in this order so the schema is never ahead of or behind a running process:

bash
docker compose --env-file .env -f docker-compose.prod.yml pull
docker compose --env-file .env -f docker-compose.prod.yml up -d migrate       # 1. schema
docker compose --env-file .env -f docker-compose.prod.yml up -d worker        # 2. worker
docker compose --env-file .env -f docker-compose.prod.yml up -d api           # 3. API
docker compose --env-file .env -f docker-compose.prod.yml up -d app-runner    # 4. app runner
docker compose --env-file .env -f docker-compose.prod.yml up -d web           # 5. UIs

Migrations are additive, so the previous worker and API keep working while the new schema is applied, a new worker drains in-flight activities (the compose file gives it a 60 s stop grace period) and a new app runner finishes the calls it holds (70 s). To roll back an application version, set the previous tag and run the same sequence without the migrate step; a schema rollback is a restore (below).

Backups ​

Postgres is the state of Aletheia, Temporal and Zitadel (databases aletheia, temporal, temporal_visibility and zitadel in the pgdata volume); Garage holds the documents in the garagedata volume. Back up both, consistently:

bash
docker compose --env-file .env -f docker-compose.prod.yml exec -T postgres \
  pg_dumpall -U "$POSTGRES_USER" | gzip > aletheia-$(date +%F).sql.gz
docker run --rm -v aletheia_garagedata:/data -v "$PWD":/backup alpine \
  tar czf /backup/garage-$(date +%F).tgz -C /data .

Keep the .env file (it holds the Zitadel masterkey without which the Zitadel database cannot be read) and the Caddy caddydata volume (certificates; re-issuable) with the backups. Restore by stopping the stack, recreating the volumes from the archives and starting it again; the migrate image is idempotent and will find nothing to apply.

TLS ​

Caddy obtains certificates from Let's Encrypt for the five host names (ACME_EMAIL receives expiry notices), renews them automatically and redirects HTTP to HTTPS. Zitadel runs with --tlsMode external, so it trusts Caddy for the external URL, and the auth. route proxies with h2c because the Zitadel gRPC APIs need HTTP/2. To use your own certificates, replace the email global option with a tls <cert> <key> directive per site and mount the files into the caddy container.

Observability ​

Set OTEL_EXPORTER_OTLP_ENDPOINT in .env to export traces and metrics to a collector reachable from the containers (docs/operations.md). API_METRICS_PORT, WORKER_METRICS_PORT and APP_RUNNER_METRICS_PORT expose Prometheus endpoints inside the compose network only; scrape them from a sidecar or an in-network Prometheus, not through Caddy.

Helm chart ​

The chart in deploy/helm/aletheia deploys the API, worker, app runner and web images against external Postgres, Temporal, Zitadel and object storage; its README (deploy/helm/README.md) covers installation, the migration job and SCHEMA_CHECK, external dependencies, every value, worker autoscaling on task slots, verification and the kind-based tests. The app runner pod (appRunner.*) gets APP_RUNNER_TOKEN from the Secret and nothing else secret (the API and the worker get the whole Secret, APP_CALL_TOKEN_SECRET included), keeps compiled modules on an emptyDir at /tmp (1Gi), asks for 250m CPU and 512Mi with a 1Gi limit, never mounts the ServiceAccount token, and its NetworkPolicy (on by default) admits ingress from the API and worker pods only and lets it reach the API pods' internal port (api.internalPort, the internal port of the API Service and never an Ingress backend), kube-dns and public addresses only, every private range excluded: an app cannot reach Postgres, Temporal, Zitadel or object storage through the runner. Released charts are pushed to oci://ghcr.io/akhiljames/charts when a release is shipped (see releasing.md), and the local CI's helm-lint and helm-kind jobs lint, snapshot and install the chart on kind on every change under deploy/helm.

Sizing ​

The numbers come from the load test (docs/performance.md, "Load test"; scripts/load/README.md). On the reference machine (Apple M2 Pro, Docker Desktop with 5 CPUs and 8 GB for Postgres, Temporal, Zitadel and Garage), one API process and one worker process with the default task slots (40 workflow and 100 activity per queue) served 10 concurrent integrations (k6 virtual users, each starting a run and waiting for it) at 3.5 completed runs/s with API p95 under 200 ms and run completion p95 of 6 s. At 50 concurrent integrations the API still answered at p95 150 ms with no failed requests, but throughput fell to 2.5 runs/s, run completion p95 reached 27 s and 14 % of runs outlived their 30 s budget, while the worker's slots peaked at 13 of 40 workflow and 15 of 100 activity. What saturated was Postgres (which Temporal's persistence shares with the application in the compose stacks, about 95 % of a CPU) and the Temporal server (about 85 %); the API and worker processes stayed around 40 % and 30 % of a core.

Rules of thumb that follow: the API is the cheapest tier (one replica handled 160 requests/s at these latencies; add replicas on CPU above 70 %). Run throughput is bounded by Temporal and its database before the worker, so give Temporal's database its own instance with fast storage (the production compose and the chart's temporal values) and watch its CPU and the Temporal schedule-to-start latencies first. Add worker replicas, each bringing another 40 workflow and 100 activity slots per queue, when temporal_worker_task_slots_available approaches 0 or when run completion time climbs while Temporal's own latencies stay flat; a worker replica is also the unit for rule CPU. App calls run on the app runner, which takes APP_RUNNER_CONCURRENCY calls at once per replica (APP_RUNNER_TENANT_CONCURRENCY per tenant): add runner replicas when aletheia.app.calls.in_flight stays near that limit or calls fail with busy. Per 50 concurrent integrations sustaining about 3 runs/s, start from one API replica (1 CPU, 512 MB), one worker replica (1 CPU, 1 GB), a Temporal server with 2 CPUs and a dedicated Postgres with 2 CPUs and 4 GB, then measure with bash scripts/load/run.sh against the real deployment: the reference machine shares its five CPUs between every container, so these are starting points, not guarantees.

Released under the Apache-2.0 License.