Skip to content

Operations ​

How to run Aletheia with traces, metrics and correlated logs, what the API, the worker and the app runner emit, and how to follow one workflow run end to end. Containers and deployment are covered in Deployment.

Telemetry environment ​

The API, the worker and the app runner read the standard OpenTelemetry variables (packages/core/src/env.ts):

VariableDefaultEffect
OTEL_EXPORTER_OTLP_ENDPOINTunsetOTLP/HTTP base URL (/v1/traces and /v1/metrics are appended). Unset: telemetry is off entirely.
OTEL_SDK_DISABLEDfalsetrue keeps telemetry off even when an endpoint is set.
OTEL_SERVICE_NAMEaletheia-api / aletheia-worker / aletheia-app-runnerOverrides service.name.
OTEL_TRACES_SAMPLERparentbased_always_onOne of always_on, always_off, traceidratio, parentbased_always_on, parentbased_always_off, parentbased_traceidratio.
OTEL_TRACES_SAMPLER_ARG1Ratio in [0, 1] for the traceidratio samplers.
OTEL_PG_STATEMENTSfalsePuts SQL text on database spans. Off by default: statements can carry personal data.
TRACE_URL_TEMPLATEunsetAPI only: a link to one trace in your backend with {traceId} in it. A run stores the id of the sampled trace that started it, and its page links there for admins (ops:read).
METRICS_PORTunsetServes Prometheus /metrics on 0.0.0.0:<port>; the worker also binds Temporal core metrics on port + 1. See Metrics.
WORKER_HEALTH_PORTunsetWorker only: GET /healthz on this port. See Worker health.
NODE_ENVdevelopmentBecomes deployment.environment.name on the resource and env on log lines.
LOG_LEVELinfopino level; the worker forwards Temporal core logs at the same level.

service.version is the app's package.json version. Nothing is exported until an endpoint or METRICS_PORT is set, so unit tests and the default dev stack are unaffected.

Startup order matters: apps/api/src/main.ts and apps/worker/src/main.ts call bootstrap(), which loads the environment, creates the logger, calls startTelemetry and only then imports server.ts dynamically. The instrumentations patch http, fastify and undici as those modules load; a static import of the server would hoist them above the SDK and produce no spans.

ESM note: the apps are ES modules, and the package instrumentation (fastify) patches through require-in-the-middle, which an ESM import never passes through; only the http and undici built-ins are traced without help. So bootstrap(), when telemetry is on, also calls module.register('@opentelemetry/instrumentation/hook.mjs') (OpenTelemetry's import-in-the-middle loader hook) before importing server.ts; the hook rewrites every module loaded from then on so the instrumentations see them. Because the registration lives in bootstrap.ts rather than in a node --import flag, pnpm start (node dist/main.js), CI and pnpm dev (tsx watch) all get the same spans; @opentelemetry/instrumentation is a direct dependency of both apps so the hook resolves from their dist/. Without an OTLP endpoint nothing is registered and modules load exactly as before.

What is traced ​

SourceSpans
API requests (http, fastify)One server span per request; /health* and /metrics are ignored. The auth hook adds the tenant and principal.
Temporal client calls (workflow-engine)temporal.<operation> client spans (startRun, startBacktest, startDocumentScan, signal*, queryRunState) with the tenant, run and workflow ids. createTemporalClient adds OpenTelemetryWorkflowClientInterceptor, which nests StartWorkflow:<type> (and SignalWorkflow:*, QueryWorkflow:*) under them and injects the span into the workflow headers, so the worker's RunWorkflow:<type> continues the API request's trace.
Workflows and activities (worker)@temporalio/interceptors-opentelemetry's OpenTelemetryPlugin: RunWorkflow:<type> in the sandbox, StartActivity:<name>/RunActivity:<name> on the worker, linked through the Temporal headers the client interceptor sets. Each activity body runs in its own activity.<name> span with aletheia.run_id, aletheia.step_id, the tenant, the workflow id and the definition key.
App calls (worker)app.invoke per attempt with the app's name, the action, the attempt, the call id, the idempotency key and the outcome; the request to the runner nests below it.
The app runnerIts server span per call continues the caller's trace (service aletheia-app-runner), with a client span for each vendor request the host sends.
Rules (rule-engine)One span per rule set evaluation and one per rule with key, type and outcome.
Postgres (@aletheia-dev/db)db.query spans from the postgres.js wrapper in @aletheia-dev/db (operation, table, optional statement text with OTEL_PG_STATEMENTS=1); db.transaction around each scoped repository call.

The worker passes the interceptors' Temporal plugin to every Worker.create only when telemetry is enabled, so a worker without an endpoint runs exactly as before. Workflow spans leave the sandbox through a sink built in @aletheia-dev/telemetry (workflowSpanSink); apps/worker/src/temporal-runtime.ts bridges the Temporal package's SDK 1.x span shape to the 2.x exporter.

Attribute glossary (aletheia.*) ​

AttributeSet byMeaning
aletheia.request_idAPI request-id hookThe x-request-id of the request (generated or caller-supplied).
aletheia.tenant_idAPI auth hook, activitiesTenant the span belongs to. On spans only, never on the resource.
aletheia.principal.kindAPI auth hookhuman, machine or applicant.
aletheia.principal.idAPI auth hookSubject of humans and machines; applicants carry a sha256 prefix of the submission id.
aletheia.run_idactivitiesAletheia workflow run id (search key for one run).
aletheia.step_idactivitiesStep of the definition the activity executes.
aletheia.workflow_idactivitiesTemporal workflow id.
aletheia.definition.keyactivitiesWorkflow definition key.
aletheia.app.name, .action, .attempt, .call_id, .outcome, .idempotency_keyapp invoker (worker)One app.invoke span per attempt; the call id is the invocation record's id.
aletheia.rule_set.key, .band, aletheia.rule.count, aletheia.rules.outcomerule enginePer rule set evaluation.
aletheia.rule.key, .type, .outcomerule enginePer rule.

Temporal's own interceptors add run_id, temporalWorkflowId and temporalActivityId.

Logs and correlation ​

Output is pino JSON (pino-pretty outside production). Every line carries service, version, env, pid and hostname; when a span is recording, the mixin adds trace_id and span_id, which is what links a log line to its trace in Grafana.

FieldWhereNotes
requestIdevery API line of a requestSame value as the x-request-id response header and error.requestId in error bodies.
tenantId, principalKindAPI lines after authenticationBound on request.log by the auth hook; the request completed line always has them.
trace_id, span_idany line inside a spanHex ids of the active span.
taskQueueworker start-up linesOne Temporal worker per queue.
tenantId, runId, stepId, workflowIdactivity linesBound on the activity logger from its input.
tenantId, app, action, attemptworker, app invokerOne line per attempt, app invocation <status>: warn for an error or a timeout.
tenantId, appName, callIdapp runnerEvery line a module logs through log, and the host's app http_request lines.
component: temporal-sdkworkerTemporal TypeScript and (forwarded) Rust core logs.

Request ids: the API accepts a caller's x-request-id when it is 1 to 128 characters of [A-Za-z0-9._-] and generates a UUID otherwise; the id is echoed on every response, 4xx and 5xx included, and repeated in every error body ({ "error": { "code", "message", "details"?, "requestId" } }), so a client can quote it in a support request. ApiError in @aletheia-dev/api-client exposes it as requestId.

Redaction ​

createLogger censors these keys at the top level and one level down ([redacted]): authorization (also under req.headers and headers), token, secrets, apiKey, secretKey, webhookSecret, password, submissionUrl, link. Deeper paths need an explicit entry (redact option); log secrets as a top-level key or not at all.

Sampling ​

The default is parentbased_always_on: everything is recorded, and a span joins its parent's decision. For production set

bash
OTEL_TRACES_SAMPLER=parentbased_traceidratio
OTEL_TRACES_SAMPLER_ARG=0.1

so 10 % of API requests start a trace and the workflow, activity, app and rule spans of those requests follow the parent decision, keeping each sampled trace complete. Errors are not force-sampled yet: a failed request outside the 10 % leaves logs (with requestId) but no trace. Tail-based sampling for errors belongs in the collector (an item for docs/plans/deferred.md).

Local stack: Grafana LGTM ​

bash
pnpm infra:observability          # grafana/otel-lgtm behind the `observability` compose profile
export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318
pnpm dev                          # or start the API and worker however you normally do
open http://localhost:3000        # Grafana (no login)

See docker/otel/README.md for ports and volumes. To follow one run end to end:

  1. Take the run id from POST /workflow-runs (or the runId on a worker log line).
  2. In Grafana choose Explore, data source Tempo, query type Search, and filter on the span attribute aletheia.run_id = the id. Service aletheia-api shows the request that started the run; aletheia-worker the RunWorkflow, RunActivity, app.invoke and rule spans, and aletheia-app-runner the calls it ran for them.
  3. Open the API trace: the root server span carries aletheia.request_id, the tenant and the principal, and the worker's RunWorkflow span continues the same trace under StartWorkflow (see What is traced).
  4. A vendor callback is its own trace (it is a new request to POST /webhooks/apps/…); the session it completed, with its external id and status, is in GET /app-callbacks?workflowRunId=<run>.
  5. For logs, copy trace_id from any span into Loki ({service="aletheia-worker"} |= "<trace_id>") or search by requestId for the API request.

CI: the OTLP sink ​

The e2e job does not run the LGTM image. It starts scripts/otlp-sink.mjs (an HTTP server that answers 200 {} to every POST and appends <time> POST <path> to a log file), points the API and the worker at it with OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4319, and sets SMOKE_OTLP_SINK_LOG=otlp.log. Smoke step 21 (scripts/smoke-operations.sh) checks the x-request-id behaviour and then polls that file for up to 20 s for a /v1/traces line (the batch processor exports every 5 s). Without SMOKE_OTLP_SINK_LOG the export assertion is skipped with a warning, so pnpm smoke against a stack without telemetry still passes.

To reproduce locally:

bash
node scripts/otlp-sink.mjs 4319 /tmp/otlp.log &
OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4319 pnpm --filter @aletheia-dev/api start &
SMOKE_OTLP_SINK_LOG=/tmp/otlp.log pnpm smoke

Metrics ​

Every instrument is defined once in @aletheia-dev/telemetry (packages/telemetry/src/meters.ts: names in METRIC, instruments in meters) and recorded by the code path that owns it. The rule engine and the workflow engine do not depend on that package (it carries the SDK), so each mirrors the few instruments it records with the same names, units and bucket advice (metrics.ts next to the recording site). Instruments are created lazily on whichever meter provider is global, so without startTelemetry every record is a no-op.

InstrumentTypeUnitLabelsRecorded by
aletheia.rule.evaluation.durationhistogrammstenant, rule_type, outcomeevaluateRule (rule engine), one point per rule; outcome pass/fail/error/skipped
aletheia.rule_set.evaluation.durationhistogrammstenant, rule_set_key, bandthe evaluateRules activity, rules plus aggregation; band min..max or none
aletheia.decision.countcounter{decision}tenant, outcome, source, definition_keyemitDecision (automated), POST /cases/:id/decide (manual)
aletheia.workflow.run.durationhistogramstenant, definition_key, statusrecordRunState when the run finishes (completed/failed/cancelled), from startedAt
aletheia.workflow.run.activegauge{run}tenant, statusgauge collector (worker), non-terminal statuses
aletheia.app.invocation.durationhistogrammstenant, app, action, statusthe app invoker (worker), one point per attempt, the runner's round trip included; status success/error/timeout/pending
aletheia.app.callback.waithistogramstenant, app, action, resultapp webhook routes (completed/failed), timeoutAppCallback (timed_out); first completion wins
aletheia.app.call.durationhistogrammsapp, export, outcomethe app runner, one point per call of a module's export; outcome ok/pending/error/timeout/trapped/bad_output
aletheia.app.calls.in_flightup-down counter{call}nonethe app runner, calls running at the moment
aletheia.case.opengauge{case}tenant, priority, sla_stategauge collector (worker), open and in_review cases
aletheia.case.time_to_decisionhistogramstenant, typePOST /cases/:id/decide, decidedAt - createdAt
aletheia.approval.pendinggauge{approval}tenant, kindgauge collector (worker), publish requests without a decision; kind app_version for an app's
aletheia.document.pendinggauge{document}tenant, statusgauge collector (worker), pending/uploaded/scanning
aletheia.document.processed.countcounter{document}tenant, statusprocessDocument, final status clean/infected/rejected
aletheia.audit.export.rowscounter{row}tenant, formatGET /audit-events/export after the last row is written
http.server.request.durationhistogramsHTTP semantic conventions@opentelemetry/instrumentation-http (the API; /health* and /metrics excluded)

Buckets: millisecond histograms advise 1, 2, 5, 10, 25, 50, 100, 250, 500, 1000; second histograms 1, 5, 15, 60, 300, 900, 3600, 21600, 86400. Keys and ids are never labels: a rule key is unbounded per tenant, so the per-rule histogram carries the rule type, and the dashboards drill into traces (aletheia.rule.key on rule.evaluate spans) for a specific rule.

Tenant label. tenant is on every instrument: the product is multi-tenant by design and cardinality is bounded by the tenant count, which an operator controls. A deployment that runs one tenant per installation drops it in the dashboard with a recording rule or a sum without (tenant) view; the shipped Grafana dashboards aggregate across tenants by default and offer tenant as a variable.

OTLP and /metrics. With OTEL_EXPORTER_OTLP_ENDPOINT set, metrics are pushed over OTLP/HTTP every 15 s next to the traces. With METRICS_PORT set, the same meter provider also serves Prometheus text exposition on 0.0.0.0:<port>/metrics, with or without an endpoint, so a deployment without a collector scrapes directly. The two paths render names differently: the direct exporter turns dots into underscores, suffixes counters with _total and puts the unit on a # UNIT line only (aletheia_rule_evaluation_duration_bucket, aletheia_decision_count_total, aletheia_workflow_run_active, scope as otel_scope_name="aletheia"); the OTLP receivers of Prometheus, Mimir and the LGTM image add the unit to the name (aletheia_rule_evaluation_duration_milliseconds_bucket, aletheia_workflow_run_duration_seconds_bucket) and drop brace units. The dashboards target the OTLP spelling; smoke step 22 asserts the direct one. A metrics-only process (METRICS_PORT without an endpoint) registers no tracer provider: spans stay non-recording and nothing is sent.

The worker's two ports. The API and the worker are separate processes with their own environment, so each reads METRICS_PORT itself and the deployment sets different values (CI: API 9464, worker 9466). The worker binds a second server on METRICS_PORT + 1 for Temporal core's own metrics (temporal_*: task slots, poll latency, workflow task failures) through Runtime.install({ telemetryOptions: { metrics: { prometheus } } }); core's OTLP option is gRPC only, so it is not wired to the OTLP/HTTP endpoint. Both ports are logged at startup.

Gauge collector. The four observable gauges are grouped counts from MetricsRepo (packages/db/src/repos/metrics.ts) over the worker's single-connection owner handle (the app role sees no rows across tenants). registerGaugeCollector runs the four queries on collection, never per request, and caches the result for 30 s, so a process exporting over OTLP every 15 s and being scraped at the same time still issues one query set per 30 s. A failing query logs a warning and leaves the gauges out of that export rather than reporting stale counts. The gauges are only registered when telemetry is on.

Worker health. With WORKER_HEALTH_PORT set, the worker answers GET /healthz with { status, workers: [{ taskQueue, state }], temporal }: 200 when every Temporal worker is RUNNING, 503 otherwise (status and temporal are then unavailable; a running worker holds a live connection and polls, which is the connectivity signal the SDK exposes). The server closes with the workers. The worker image's health check and the chart's probes call it.

To reproduce smoke step 22 locally:

bash
METRICS_PORT=9466 WORKER_HEALTH_PORT=9470 pnpm --filter @aletheia-dev/worker start &
METRICS_PORT=9464 pnpm --filter @aletheia-dev/api start &
SMOKE_METRICS_API=http://localhost:9464 SMOKE_METRICS_WORKER=http://localhost:9466 \
SMOKE_METRICS_WORKER_CORE=http://localhost:9467 SMOKE_WORKER_HEALTH=http://localhost:9470 pnpm smoke
curl -s localhost:9464/metrics | grep '^aletheia_'

Without the SMOKE_* variables the step skips its assertions with a warning.

Runbooks ​

The provisioned alert rules (deploy/grafana/provisioning/alerting/aletheia.yaml) link here. Each entry says what the alert measures, what to look at first, and the usual causes.

App failures ​

Fires when more than 5 % of app invocations for one tenant and app end in error or timeout over ten minutes. Open the Apps row of the Decisions dashboard for that tenant, then the traces for app.invoke spans with aletheia.app.outcome in error|timeout. From the API, with a token of that tenant, GET /apps/health?windowMinutes=60 gives each app's calls, failures, timeouts, p95 duration and last failure over the window (5 to 1440 minutes), with the open circuits and the installs on a blocked module; GET /app-invocations?appName=<name>&status=error&status=timeout lists the failed calls newest first, each with the version it ran on and what it used at the runner, and with the error text for a caller holding runs:context (an admin). Connect › Apps shows the same per app. Usual causes: a vendor outage or rate limit, a rotated secret (store the new value with POST /apps/:name/secrets/:secret), an upgrade to a version that calls the vendor differently (install the previous version again), or the runner itself: a runner that is unreachable or out of call slots fails every app at once, which the runner's own aletheia.app.call.duration and aletheia.app.calls.in_flight show. A module the operator listed in APP_BLOCKED_SHA256 fails every call with app_blocked. Disabling the install (PUT /apps/:name/install with enabled: false) stops new invocations; runs already waiting on a callback keep waiting until its timeout, unless a forced uninstall (DELETE /apps/:name/install?force=true) fails them at once.

A worker stops calling an app by itself after five failed calls in a row for one tenant: its circuit opens, calls fail at once with app_circuit_open for 30 seconds, and then one trial call goes out (each worker process keeps its own circuits). Only calls that ran the module count: a busy or unreachable runner does not open an app's circuit. GET /apps/health lists the open and half-open circuits as circuits, the audit trail has app.circuit.opened and app.circuit.closed, and Connect › Apps and Home name the app. Check the vendor with the app page's test call (POST /apps/:name/test), which bypasses the circuit.

Webhook deliveries ​

There is no alert rule: a delivery given up after five attempts raises a notification for the tenant's admins and shows on Home under Needs your attention, and the worker logs it at warn as webhook delivery given up with the tenant, endpoint and delivery ids. Connect › Webhooks, or GET /webhook-deliveries?status=failed&status=exhausted, lists what failed with the response code or error; Send test ping checks the endpoint now. Usual causes: the receiver is down or answers non-2xx (its own logs say why), a name that resolves to a private address, a receiver that takes longer than 10 seconds, or a receiver still on the secret before a rotation more than 24 hours ago (it answers 401 or 403). Deliveries wait while every worker lacks SECRET_STORE_KEY (webhook deliveries paused at start-up) and stay queued for a disabled endpoint. The worker reads new events every two seconds and logs outbound event dispatch failed when it cannot; events are not lost while it fails, it carries on from where it stopped.

Scheduled audit exports ​

There is no alert rule either: a failed run is recorded with its error, Home lists it for a day and the worker logs scheduled audit export failed at warn. Audit › Exports › Runs shows the error. Usual causes: object storage refusing the write (credentials, bucket, quota), a filter so wide its file passes 256 MiB (narrow it or schedule it more often), or the worker without STORAGE_ENDPOINT (the run never starts: the Temporal UI shows the auditExportSchedule workflow failing on an unregistered activity). The next run covers the failed run's window again.

Workflow task failures ​

Fires when the worker reports any workflow task failure in five minutes. These are not business failures: a workflow task fails when the interpreter throws outside an activity, which after a deploy almost always means a non-deterministic change replayed against an in-flight run. Check the worker log for DeterminismViolationError and the workflow:incompatible label on the released change; roll the worker back or let the affected runs be reset from the Temporal UI. Activity failures are a separate metric and retry on their own.

HTTP errors ​

Fires when more than 1 % of responses on one route are 5xx over five minutes. Find the route in the Operations dashboard, then search traces by http.route and status 500; the server span's aletheia.request_id matches the requestId in the API log line, which carries the error. Usual causes: the database or Temporal being unreachable (/health/ready reports which), or an object-storage error on document routes.

Rate limits ​

Each API instance gives every caller API_RATE_LIMIT_PER_MINUTE requests a minute (600 by default, 0 turns it off) and then answers 429 rate_limited with a Retry-After header and details.retryAfterSeconds; responses carry x-ratelimit-limit, x-ratelimit-remaining and x-ratelimit-reset. A caller is a tenant's user or service user, an applicant's submission (every tab and device on one link share it), or on the public routes (vendor webhooks, flow previews) the client address; health checks are not counted. The collection terminal waits as long as Retry-After says before it saves or polls again. The count is kept per instance, so with several replicas a caller's effective budget is the limit times the replicas it reaches. A caller that keeps hitting it is usually a polling loop: raise the limit only after checking the caller's request pattern in the API logs (principalKind, url).

Every API response carries Cache-Control: no-store unless the route sets its own: responses hold personal data and links that no browser or proxy cache should keep.

SLA breaches ​

Fires when the number of open cases in breached state for a tenant rises over fifteen minutes. This is a staffing signal, not a system fault: open the admin console's Operate › Cases on the "Breached SLA" queue, assign or close. If every tenant breaches at once, check that the worker is running (the SLA timer lives in the run) and that the aletheia.case.open gauge is fresh.

Stale approvals ​

Informational: a definition has waited more than a day for approval. Open the admin console's Operate › Approvals, "Stale · over 1 day" view; the requester is in the alert labels. A request nobody can approve usually means the tenant enabled four-eyes with a single admin; either add an approver or have the admin reject and re-request with approvals turned off.

Released under the Apache-2.0 License.