Appearance
Operations
How to run Aletheia with traces, metrics and correlated logs, what the API, the worker and the app runner emit, and how to follow one workflow run end to end. Containers and deployment are covered in Deployment.
Telemetry environment
The API, the worker and the app runner read the standard OpenTelemetry variables (packages/core/src/env.ts):
| Variable | Default | Effect |
|---|---|---|
OTEL_EXPORTER_OTLP_ENDPOINT | unset | OTLP/HTTP base URL (/v1/traces and /v1/metrics are appended). Unset: telemetry is off entirely. |
OTEL_SDK_DISABLED | false | true keeps telemetry off even when an endpoint is set. |
OTEL_SERVICE_NAME | aletheia-api / aletheia-worker / aletheia-app-runner | Overrides service.name. |
OTEL_TRACES_SAMPLER | parentbased_always_on | One of always_on, always_off, traceidratio, parentbased_always_on, parentbased_always_off, parentbased_traceidratio. |
OTEL_TRACES_SAMPLER_ARG | 1 | Ratio in [0, 1] for the traceidratio samplers. |
OTEL_PG_STATEMENTS | false | Puts SQL text on database spans. Off by default: statements can carry personal data. |
TRACE_URL_TEMPLATE | unset | API only: a link to one trace in your backend with {traceId} in it. A run stores the id of the sampled trace that started it, and its page links there for admins (ops:read). |
METRICS_PORT | unset | Serves Prometheus /metrics on 0.0.0.0:<port>; the worker also binds Temporal core metrics on port + 1. See Metrics. |
WORKER_HEALTH_PORT | unset | Worker only: GET /healthz on this port. See Worker health. |
NODE_ENV | development | Becomes deployment.environment.name on the resource and env on log lines. |
LOG_LEVEL | info | pino level; the worker forwards Temporal core logs at the same level. |
service.version is the app's package.json version. Nothing is exported until an endpoint or METRICS_PORT is set, so unit tests and the default dev stack are unaffected.
Startup order matters: apps/api/src/main.ts and apps/worker/src/main.ts call bootstrap(), which loads the environment, creates the logger, calls startTelemetry and only then imports server.ts dynamically. The instrumentations patch http, fastify and undici as those modules load; a static import of the server would hoist them above the SDK and produce no spans.
ESM note: the apps are ES modules, and the package instrumentation (fastify) patches through require-in-the-middle, which an ESM import never passes through; only the http and undici built-ins are traced without help. So bootstrap(), when telemetry is on, also calls module.register('@opentelemetry/instrumentation/hook.mjs') (OpenTelemetry's import-in-the-middle loader hook) before importing server.ts; the hook rewrites every module loaded from then on so the instrumentations see them. Because the registration lives in bootstrap.ts rather than in a node --import flag, pnpm start (node dist/main.js), CI and pnpm dev (tsx watch) all get the same spans; @opentelemetry/instrumentation is a direct dependency of both apps so the hook resolves from their dist/. Without an OTLP endpoint nothing is registered and modules load exactly as before.
What is traced
| Source | Spans |
|---|---|
API requests (http, fastify) | One server span per request; /health* and /metrics are ignored. The auth hook adds the tenant and principal. |
Temporal client calls (workflow-engine) | temporal.<operation> client spans (startRun, startBacktest, startDocumentScan, signal*, queryRunState) with the tenant, run and workflow ids. createTemporalClient adds OpenTelemetryWorkflowClientInterceptor, which nests StartWorkflow:<type> (and SignalWorkflow:*, QueryWorkflow:*) under them and injects the span into the workflow headers, so the worker's RunWorkflow:<type> continues the API request's trace. |
| Workflows and activities (worker) | @temporalio/interceptors-opentelemetry's OpenTelemetryPlugin: RunWorkflow:<type> in the sandbox, StartActivity:<name>/RunActivity:<name> on the worker, linked through the Temporal headers the client interceptor sets. Each activity body runs in its own activity.<name> span with aletheia.run_id, aletheia.step_id, the tenant, the workflow id and the definition key. |
| App calls (worker) | app.invoke per attempt with the app's name, the action, the attempt, the call id, the idempotency key and the outcome; the request to the runner nests below it. |
| The app runner | Its server span per call continues the caller's trace (service aletheia-app-runner), with a client span for each vendor request the host sends. |
Rules (rule-engine) | One span per rule set evaluation and one per rule with key, type and outcome. |
Postgres (@aletheia-dev/db) | db.query spans from the postgres.js wrapper in @aletheia-dev/db (operation, table, optional statement text with OTEL_PG_STATEMENTS=1); db.transaction around each scoped repository call. |
The worker passes the interceptors' Temporal plugin to every Worker.create only when telemetry is enabled, so a worker without an endpoint runs exactly as before. Workflow spans leave the sandbox through a sink built in @aletheia-dev/telemetry (workflowSpanSink); apps/worker/src/temporal-runtime.ts bridges the Temporal package's SDK 1.x span shape to the 2.x exporter.
Attribute glossary (aletheia.*)
| Attribute | Set by | Meaning |
|---|---|---|
aletheia.request_id | API request-id hook | The x-request-id of the request (generated or caller-supplied). |
aletheia.tenant_id | API auth hook, activities | Tenant the span belongs to. On spans only, never on the resource. |
aletheia.principal.kind | API auth hook | human, machine or applicant. |
aletheia.principal.id | API auth hook | Subject of humans and machines; applicants carry a sha256 prefix of the submission id. |
aletheia.run_id | activities | Aletheia workflow run id (search key for one run). |
aletheia.step_id | activities | Step of the definition the activity executes. |
aletheia.workflow_id | activities | Temporal workflow id. |
aletheia.definition.key | activities | Workflow definition key. |
aletheia.app.name, .action, .attempt, .call_id, .outcome, .idempotency_key | app invoker (worker) | One app.invoke span per attempt; the call id is the invocation record's id. |
aletheia.rule_set.key, .band, aletheia.rule.count, aletheia.rules.outcome | rule engine | Per rule set evaluation. |
aletheia.rule.key, .type, .outcome | rule engine | Per rule. |
Temporal's own interceptors add run_id, temporalWorkflowId and temporalActivityId.
Logs and correlation
Output is pino JSON (pino-pretty outside production). Every line carries service, version, env, pid and hostname; when a span is recording, the mixin adds trace_id and span_id, which is what links a log line to its trace in Grafana.
| Field | Where | Notes |
|---|---|---|
requestId | every API line of a request | Same value as the x-request-id response header and error.requestId in error bodies. |
tenantId, principalKind | API lines after authentication | Bound on request.log by the auth hook; the request completed line always has them. |
trace_id, span_id | any line inside a span | Hex ids of the active span. |
taskQueue | worker start-up lines | One Temporal worker per queue. |
tenantId, runId, stepId, workflowId | activity lines | Bound on the activity logger from its input. |
tenantId, app, action, attempt | worker, app invoker | One line per attempt, app invocation <status>: warn for an error or a timeout. |
tenantId, appName, callId | app runner | Every line a module logs through log, and the host's app http_request lines. |
component: temporal-sdk | worker | Temporal TypeScript and (forwarded) Rust core logs. |
Request ids: the API accepts a caller's x-request-id when it is 1 to 128 characters of [A-Za-z0-9._-] and generates a UUID otherwise; the id is echoed on every response, 4xx and 5xx included, and repeated in every error body ({ "error": { "code", "message", "details"?, "requestId" } }), so a client can quote it in a support request. ApiError in @aletheia-dev/api-client exposes it as requestId.
Redaction
createLogger censors these keys at the top level and one level down ([redacted]): authorization (also under req.headers and headers), token, secrets, apiKey, secretKey, webhookSecret, password, submissionUrl, link. Deeper paths need an explicit entry (redact option); log secrets as a top-level key or not at all.
Sampling
The default is parentbased_always_on: everything is recorded, and a span joins its parent's decision. For production set
bash
OTEL_TRACES_SAMPLER=parentbased_traceidratio
OTEL_TRACES_SAMPLER_ARG=0.1so 10 % of API requests start a trace and the workflow, activity, app and rule spans of those requests follow the parent decision, keeping each sampled trace complete. Errors are not force-sampled yet: a failed request outside the 10 % leaves logs (with requestId) but no trace. Tail-based sampling for errors belongs in the collector (an item for docs/plans/deferred.md).
Local stack: Grafana LGTM
bash
pnpm infra:observability # grafana/otel-lgtm behind the `observability` compose profile
export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318
pnpm dev # or start the API and worker however you normally do
open http://localhost:3000 # Grafana (no login)See docker/otel/README.md for ports and volumes. To follow one run end to end:
- Take the run id from
POST /workflow-runs(or therunIdon a worker log line). - In Grafana choose Explore, data source Tempo, query type Search, and filter on the span attribute
aletheia.run_id= the id. Servicealetheia-apishows the request that started the run;aletheia-workertheRunWorkflow,RunActivity,app.invokeand rule spans, andaletheia-app-runnerthe calls it ran for them. - Open the API trace: the root server span carries
aletheia.request_id, the tenant and the principal, and the worker'sRunWorkflowspan continues the same trace underStartWorkflow(see What is traced). - A vendor callback is its own trace (it is a new request to
POST /webhooks/apps/…); the session it completed, with its external id and status, is inGET /app-callbacks?workflowRunId=<run>. - For logs, copy
trace_idfrom any span into Loki ({service="aletheia-worker"} |= "<trace_id>") or search byrequestIdfor the API request.
CI: the OTLP sink
The e2e job does not run the LGTM image. It starts scripts/otlp-sink.mjs (an HTTP server that answers 200 {} to every POST and appends <time> POST <path> to a log file), points the API and the worker at it with OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4319, and sets SMOKE_OTLP_SINK_LOG=otlp.log. Smoke step 21 (scripts/smoke-operations.sh) checks the x-request-id behaviour and then polls that file for up to 20 s for a /v1/traces line (the batch processor exports every 5 s). Without SMOKE_OTLP_SINK_LOG the export assertion is skipped with a warning, so pnpm smoke against a stack without telemetry still passes.
To reproduce locally:
bash
node scripts/otlp-sink.mjs 4319 /tmp/otlp.log &
OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4319 pnpm --filter @aletheia-dev/api start &
SMOKE_OTLP_SINK_LOG=/tmp/otlp.log pnpm smokeMetrics
Every instrument is defined once in @aletheia-dev/telemetry (packages/telemetry/src/meters.ts: names in METRIC, instruments in meters) and recorded by the code path that owns it. The rule engine and the workflow engine do not depend on that package (it carries the SDK), so each mirrors the few instruments it records with the same names, units and bucket advice (metrics.ts next to the recording site). Instruments are created lazily on whichever meter provider is global, so without startTelemetry every record is a no-op.
| Instrument | Type | Unit | Labels | Recorded by |
|---|---|---|---|---|
aletheia.rule.evaluation.duration | histogram | ms | tenant, rule_type, outcome | evaluateRule (rule engine), one point per rule; outcome pass/fail/error/skipped |
aletheia.rule_set.evaluation.duration | histogram | ms | tenant, rule_set_key, band | the evaluateRules activity, rules plus aggregation; band min..max or none |
aletheia.decision.count | counter | {decision} | tenant, outcome, source, definition_key | emitDecision (automated), POST /cases/:id/decide (manual) |
aletheia.workflow.run.duration | histogram | s | tenant, definition_key, status | recordRunState when the run finishes (completed/failed/cancelled), from startedAt |
aletheia.workflow.run.active | gauge | {run} | tenant, status | gauge collector (worker), non-terminal statuses |
aletheia.app.invocation.duration | histogram | ms | tenant, app, action, status | the app invoker (worker), one point per attempt, the runner's round trip included; status success/error/timeout/pending |
aletheia.app.callback.wait | histogram | s | tenant, app, action, result | app webhook routes (completed/failed), timeoutAppCallback (timed_out); first completion wins |
aletheia.app.call.duration | histogram | ms | app, export, outcome | the app runner, one point per call of a module's export; outcome ok/pending/error/timeout/trapped/bad_output |
aletheia.app.calls.in_flight | up-down counter | {call} | none | the app runner, calls running at the moment |
aletheia.case.open | gauge | {case} | tenant, priority, sla_state | gauge collector (worker), open and in_review cases |
aletheia.case.time_to_decision | histogram | s | tenant, type | POST /cases/:id/decide, decidedAt - createdAt |
aletheia.approval.pending | gauge | {approval} | tenant, kind | gauge collector (worker), publish requests without a decision; kind app_version for an app's |
aletheia.document.pending | gauge | {document} | tenant, status | gauge collector (worker), pending/uploaded/scanning |
aletheia.document.processed.count | counter | {document} | tenant, status | processDocument, final status clean/infected/rejected |
aletheia.audit.export.rows | counter | {row} | tenant, format | GET /audit-events/export after the last row is written |
http.server.request.duration | histogram | s | HTTP semantic conventions | @opentelemetry/instrumentation-http (the API; /health* and /metrics excluded) |
Buckets: millisecond histograms advise 1, 2, 5, 10, 25, 50, 100, 250, 500, 1000; second histograms 1, 5, 15, 60, 300, 900, 3600, 21600, 86400. Keys and ids are never labels: a rule key is unbounded per tenant, so the per-rule histogram carries the rule type, and the dashboards drill into traces (aletheia.rule.key on rule.evaluate spans) for a specific rule.
Tenant label. tenant is on every instrument: the product is multi-tenant by design and cardinality is bounded by the tenant count, which an operator controls. A deployment that runs one tenant per installation drops it in the dashboard with a recording rule or a sum without (tenant) view; the shipped Grafana dashboards aggregate across tenants by default and offer tenant as a variable.
OTLP and /metrics. With OTEL_EXPORTER_OTLP_ENDPOINT set, metrics are pushed over OTLP/HTTP every 15 s next to the traces. With METRICS_PORT set, the same meter provider also serves Prometheus text exposition on 0.0.0.0:<port>/metrics, with or without an endpoint, so a deployment without a collector scrapes directly. The two paths render names differently: the direct exporter turns dots into underscores, suffixes counters with _total and puts the unit on a # UNIT line only (aletheia_rule_evaluation_duration_bucket, aletheia_decision_count_total, aletheia_workflow_run_active, scope as otel_scope_name="aletheia"); the OTLP receivers of Prometheus, Mimir and the LGTM image add the unit to the name (aletheia_rule_evaluation_duration_milliseconds_bucket, aletheia_workflow_run_duration_seconds_bucket) and drop brace units. The dashboards target the OTLP spelling; smoke step 22 asserts the direct one. A metrics-only process (METRICS_PORT without an endpoint) registers no tracer provider: spans stay non-recording and nothing is sent.
The worker's two ports. The API and the worker are separate processes with their own environment, so each reads METRICS_PORT itself and the deployment sets different values (CI: API 9464, worker 9466). The worker binds a second server on METRICS_PORT + 1 for Temporal core's own metrics (temporal_*: task slots, poll latency, workflow task failures) through Runtime.install({ telemetryOptions: { metrics: { prometheus } } }); core's OTLP option is gRPC only, so it is not wired to the OTLP/HTTP endpoint. Both ports are logged at startup.
Gauge collector. The four observable gauges are grouped counts from MetricsRepo (packages/db/src/repos/metrics.ts) over the worker's single-connection owner handle (the app role sees no rows across tenants). registerGaugeCollector runs the four queries on collection, never per request, and caches the result for 30 s, so a process exporting over OTLP every 15 s and being scraped at the same time still issues one query set per 30 s. A failing query logs a warning and leaves the gauges out of that export rather than reporting stale counts. The gauges are only registered when telemetry is on.
Worker health. With WORKER_HEALTH_PORT set, the worker answers GET /healthz with { status, workers: [{ taskQueue, state }], temporal }: 200 when every Temporal worker is RUNNING, 503 otherwise (status and temporal are then unavailable; a running worker holds a live connection and polls, which is the connectivity signal the SDK exposes). The server closes with the workers. The worker image's health check and the chart's probes call it.
To reproduce smoke step 22 locally:
bash
METRICS_PORT=9466 WORKER_HEALTH_PORT=9470 pnpm --filter @aletheia-dev/worker start &
METRICS_PORT=9464 pnpm --filter @aletheia-dev/api start &
SMOKE_METRICS_API=http://localhost:9464 SMOKE_METRICS_WORKER=http://localhost:9466 \
SMOKE_METRICS_WORKER_CORE=http://localhost:9467 SMOKE_WORKER_HEALTH=http://localhost:9470 pnpm smoke
curl -s localhost:9464/metrics | grep '^aletheia_'Without the SMOKE_* variables the step skips its assertions with a warning.
Runbooks
The provisioned alert rules (deploy/grafana/provisioning/alerting/aletheia.yaml) link here. Each entry says what the alert measures, what to look at first, and the usual causes.
App failures
Fires when more than 5 % of app invocations for one tenant and app end in error or timeout over ten minutes. Open the Apps row of the Decisions dashboard for that tenant, then the traces for app.invoke spans with aletheia.app.outcome in error|timeout. From the API, with a token of that tenant, GET /apps/health?windowMinutes=60 gives each app's calls, failures, timeouts, p95 duration and last failure over the window (5 to 1440 minutes), with the open circuits and the installs on a blocked module; GET /app-invocations?appName=<name>&status=error&status=timeout lists the failed calls newest first, each with the version it ran on and what it used at the runner, and with the error text for a caller holding runs:context (an admin). Connect › Apps shows the same per app. Usual causes: a vendor outage or rate limit, a rotated secret (store the new value with POST /apps/:name/secrets/:secret), an upgrade to a version that calls the vendor differently (install the previous version again), or the runner itself: a runner that is unreachable or out of call slots fails every app at once, which the runner's own aletheia.app.call.duration and aletheia.app.calls.in_flight show. A module the operator listed in APP_BLOCKED_SHA256 fails every call with app_blocked. Disabling the install (PUT /apps/:name/install with enabled: false) stops new invocations; runs already waiting on a callback keep waiting until its timeout, unless a forced uninstall (DELETE /apps/:name/install?force=true) fails them at once.
A worker stops calling an app by itself after five failed calls in a row for one tenant: its circuit opens, calls fail at once with app_circuit_open for 30 seconds, and then one trial call goes out (each worker process keeps its own circuits). Only calls that ran the module count: a busy or unreachable runner does not open an app's circuit. GET /apps/health lists the open and half-open circuits as circuits, the audit trail has app.circuit.opened and app.circuit.closed, and Connect › Apps and Home name the app. Check the vendor with the app page's test call (POST /apps/:name/test), which bypasses the circuit.
Webhook deliveries
There is no alert rule: a delivery given up after five attempts raises a notification for the tenant's admins and shows on Home under Needs your attention, and the worker logs it at warn as webhook delivery given up with the tenant, endpoint and delivery ids. Connect › Webhooks, or GET /webhook-deliveries?status=failed&status=exhausted, lists what failed with the response code or error; Send test ping checks the endpoint now. Usual causes: the receiver is down or answers non-2xx (its own logs say why), a name that resolves to a private address, a receiver that takes longer than 10 seconds, or a receiver still on the secret before a rotation more than 24 hours ago (it answers 401 or 403). Deliveries wait while every worker lacks SECRET_STORE_KEY (webhook deliveries paused at start-up) and stay queued for a disabled endpoint. The worker reads new events every two seconds and logs outbound event dispatch failed when it cannot; events are not lost while it fails, it carries on from where it stopped.
Scheduled audit exports
There is no alert rule either: a failed run is recorded with its error, Home lists it for a day and the worker logs scheduled audit export failed at warn. Audit › Exports › Runs shows the error. Usual causes: object storage refusing the write (credentials, bucket, quota), a filter so wide its file passes 256 MiB (narrow it or schedule it more often), or the worker without STORAGE_ENDPOINT (the run never starts: the Temporal UI shows the auditExportSchedule workflow failing on an unregistered activity). The next run covers the failed run's window again.
Workflow task failures
Fires when the worker reports any workflow task failure in five minutes. These are not business failures: a workflow task fails when the interpreter throws outside an activity, which after a deploy almost always means a non-deterministic change replayed against an in-flight run. Check the worker log for DeterminismViolationError and the workflow:incompatible label on the released change; roll the worker back or let the affected runs be reset from the Temporal UI. Activity failures are a separate metric and retry on their own.
HTTP errors
Fires when more than 1 % of responses on one route are 5xx over five minutes. Find the route in the Operations dashboard, then search traces by http.route and status 500; the server span's aletheia.request_id matches the requestId in the API log line, which carries the error. Usual causes: the database or Temporal being unreachable (/health/ready reports which), or an object-storage error on document routes.
Rate limits
Each API instance gives every caller API_RATE_LIMIT_PER_MINUTE requests a minute (600 by default, 0 turns it off) and then answers 429 rate_limited with a Retry-After header and details.retryAfterSeconds; responses carry x-ratelimit-limit, x-ratelimit-remaining and x-ratelimit-reset. A caller is a tenant's user or service user, an applicant's submission (every tab and device on one link share it), or on the public routes (vendor webhooks, flow previews) the client address; health checks are not counted. The collection terminal waits as long as Retry-After says before it saves or polls again. The count is kept per instance, so with several replicas a caller's effective budget is the limit times the replicas it reaches. A caller that keeps hitting it is usually a polling loop: raise the limit only after checking the caller's request pattern in the API logs (principalKind, url).
Every API response carries Cache-Control: no-store unless the route sets its own: responses hold personal data and links that no browser or proxy cache should keep.
SLA breaches
Fires when the number of open cases in breached state for a tenant rises over fifteen minutes. This is a staffing signal, not a system fault: open the admin console's Operate › Cases on the "Breached SLA" queue, assign or close. If every tenant breaches at once, check that the worker is running (the SLA timer lives in the run) and that the aletheia.case.open gauge is fresh.
Stale approvals
Informational: a definition has waited more than a day for approval. Open the admin console's Operate › Approvals, "Stale · over 1 day" view; the requester is in the alert labels. A request nobody can approve usually means the tenant enabled four-eyes with a single admin; either add an approver or have the admin reject and re-request with approvals turned off.