Appearance
Performance
Replay tests, rule-evaluation benchmarks and the load test for the API and worker.
Replay
Temporal re-executes workflow code against the recorded event history whenever a run resumes (after a signal, a timer, a worker restart or a deploy). The interpreter in @aletheia-dev/workflow-engine must therefore issue the same commands, in the same order, for the same inputs as the version that recorded the history, or runs in flight fail with a non-determinism error after the worker is upgraded. The replay test is the guard that every interpreter change must pass before merge.
What the histories cover
packages/workflow-engine/histories/*.json are event histories recorded on Temporal's time-skipping test server with mocked activities, one per path through the seeded definitions (packages/db/src/seed-data.ts) and the other workflows (histories/README.md lists them):
| History | Path |
|---|---|
hello-world.approve / .reject | evaluate_rules -> automated decision |
kyb.manual-review.approve | collection signal -> sanctions app -> rule set -> create_case -> manual decision |
kyb.manual-review.sla-breach | as above, SLA timer fires and escalates the case before the (late) decision |
kyb.async-app.completed | asynchronous call_app resumed by the vendor callback signal (appCallback) |
kyb.async-app.failed | failed callback -> run fails (AppCallbackFailed) |
kyb.async-app.timeout | callback timer fires -> timeoutAppCallback -> run fails (CallbackTimeout) |
kyb.collection-timeout | collection deadline fires -> run fails (CollectionTimeout) |
kyb.branch-no-document | has_document false path -> rule set -> automated approve |
kyb.document-verified.approve | has_document true path -> document verification app -> automated approve |
transaction-monitoring.approve | velocity rule -> branch default -> automated decision |
backtest.two-batches | backtestRule: 250 ids in two batches, chained summary |
scan-document.clean | scanDocument on the app queue |
src/__tests__/replay.test.ts bundles the current workflow code once and streams every history through one replay worker (Worker.runReplayHistories); it needs no Temporal server and runs in the default pnpm test, in about 8 s. A history that fails replay is reported by name with the SDK's DeterminismViolationError (Core's TMPRL1100 message names the mismatching command). The suite also replays a deliberately tampered copy of hello-world.approve (an activity type renamed in a scheduled event) and asserts the violation, so a replay worker that silently accepted everything would fail the suite.
Re-recording and the compatibility policy
bash
pnpm --filter @aletheia-dev/workflow-engine record-historiesruns src/__tests__/record-histories.ts on the time-skipping server (RUN_INTEGRATION=1, first run downloads the server binary), rewrites histories/*.json and formats them. The scenario ids are fixed, so a re-recording differs only in timestamps, run ids and the host identity. Re-record when:
- a history scenario changes on purpose (a new step kind, a new seeded path) or a new path needs cover: add a scenario to the recorder and record it; commit the JSON with the change;
- the replay test fails and the change is meant to be incompatible.
A change that fails replay is incompatible: runs that are mid-flight when the worker is upgraded (waiting for a collection, a callback, a manual decision or a timer) will fail their next workflow task. Such a pull request re-records the histories and carries the workflow:incompatible label (or that string in the description or commit message; docs/releasing.md, "Workflow compatibility"). The release notes then open with the instruction to drain running workflows before upgrading the worker, and the release is a major version. A change that fails replay without the label does not merge.
Prefer compatible changes. Reordering activities, adding an activity or timer before an existing one, changing an activity's name or the arguments to condition() deadlines are all incompatible; adding a step kind the recorded definitions do not use, changing activity implementations, retry policies that only affect new attempts and query handlers are compatible.
patched() / deprecatePatch() from @temporalio/workflow are the alternative for a hot fix while long-running runs are in flight and draining is not an option: the new code branches on patched('fix-name') so old histories replay the old path and new runs take the new one, and the patch is removed (deprecatePatch) once every run started before it has finished. They are reserved for that case; routine changes go through the label and a drained upgrade, because every patch is a permanent branch in the interpreter until it is retired.
Replaying production histories before a deploy
The worker image ships dist/replay.js, which replays histories against the worker's production bundle (the same workflowsPath the workers serve, so it also proves the built code, not only the test bundle, is replay-safe). GET /workflow-runs/:id/history (admin only; event payloads carry the unredacted run context; 409 when the API has no Temporal client) returns a run's history as proto JSON, and temporal workflow show --output json produces the same format:
bash
# Histories exported to files
node apps/worker/dist/replay.js history-a.json history-b.json
# Straight from the API (admin PAT), one or more runs
API_URL=https://api.example.com API_TOKEN=... \
node apps/worker/dist/replay.js --from-api <runId> --from-api <runId>It prints ok <name> or FAIL <name>: <error> per history and exits 1 on any failure (2 on a usage error). Before upgrading a worker with runs in flight, replay a sample of their histories against the new image (docker run --rm -v $PWD/histories:/h <worker-image> node dist/replay.js /h/*.json); a FAIL means the release needs the drain, whatever the label says. The smoke test (scripts/smoke-operations.sh, step 24) does this for a run that went through collection, screening and a manual decision.
Benchmarks
packages/rule-engine/bench/evaluate.bench.ts measures rule evaluation on synthetic data with vitest bench (tinybench underneath). Every scenario runs the real evaluator path: config validation, the handler, the rule.evaluate span and the duration histogram (both inert without a provider). Fixtures live in bench/fixtures.ts: a merchant with the fields the seeded rules read plus 20 extra scalars, a kyb-basic submission, Set-backed lists of 10 and 10 000 entries, 10 000 events scanned linearly for velocity and an app service that answers on the current tick. The seeded kyb-onboarding rules are copied there (rule-engine must not depend on db); keep the copy in step with packages/db/src/seed-data.ts.
Run it with pnpm --filter @aletheia-dev/rule-engine bench (or pnpm bench at the root, which runs every package's bench). Each scenario warms up for 200 ms and then samples for 1 s; the table shows ops/s, mean, p75, p99 and the relative margin of error. The JSON report lands in packages/rule-engine/bench/.results/results.json (git-ignored) or wherever BENCH_JSON points; node packages/rule-engine/bench/summary.mjs <results.json> renders it as a Markdown table and appends it to $GITHUB_STEP_SUMMARY when that is set.
Scenarios (names are stable; CI and this page refer to them):
| Scenario | What it measures |
|---|---|
rule:comparison | subject.country in [US, GB, DE]: path lookup and condition evaluation |
rule:list_lookup_small | submission.country against the 10-entry blocked_countries list |
rule:list_lookup_large | subject.legalName against the 10 000-entry blocked_entities list |
rule:score_threshold | sanctions.score <= 70, contributing the value to the risk score |
rule:cel_simple | CEL data.subject.country != "US": environment build, parse, type-check and evaluation |
rule:cel_haspath_arith | CEL with hasPath, arithmetic and three field reads (the seeded UBO rule's shape) |
rule:cel_inlist_large | CEL with two inList lookups against the large list (async function calls inside the expression) |
rule:velocity_count_24h | count of the subject's payment events in 24 h over the 10 000-event store |
rule:app | input mapping, an immediate app call and the resultPath check |
set:kyb-onboarding (6 rules, sum_weights) | the seeded set end to end: evaluateRules (concurrent) plus aggregateDecision with the seeded clamp and bands |
set:synthetic-50 (50 rules, sum_weights) | 50 rules cycling through every type above (a third of them CEL), the same policy |
Numbers from an Apple Silicon dev machine (M2 Pro, Node 22), one run of pnpm bench with other work idle. Treat them as the shape of the cost, not as a reference: a second run on the same machine while other processes were busy came out 30 to 40 % slower across the board, and the bench runs the TypeScript source through Vitest's module runner, whose export getters add a small fixed cost per call that the production bundle does not pay.
| Scenario | ops/s | mean (ms) | p75 (ms) | p99 (ms) |
|---|---|---|---|---|
rule:comparison | 1 196 000 | 0.001 | 0.001 | 0.003 |
rule:score_threshold | 1 163 000 | 0.001 | 0.001 | 0.003 |
rule:list_lookup_large | 1 016 000 | 0.002 | 0.001 | 0.007 |
rule:list_lookup_small | 962 000 | 0.002 | 0.002 | 0.007 |
rule:app | 542 000 | 0.003 | 0.002 | 0.014 |
rule:cel_simple | 34 300 | 0.037 | 0.027 | 0.196 |
rule:cel_haspath_arith | 29 400 | 0.041 | 0.033 | 0.212 |
rule:cel_inlist_large | 27 700 | 0.052 | 0.034 | 0.303 |
rule:velocity_count_24h | 15 700 | 0.087 | 0.060 | 0.425 |
set:kyb-onboarding (6 rules, sum_weights) | 14 400 | 0.082 | 0.067 | 0.355 |
set:synthetic-50 (50 rules, sum_weights) | 995 | 1.166 | 1.307 | 3.649 |
Reading it: the non-CEL handlers cost about a microsecond; a CEL rule costs about 40 µs because the environment is built and the expression parsed and type-checked on every evaluation (the registry validates the config, which checks the expression, and the handler parses it again), so a set's cost is dominated by its CEL rule count; list size does not matter for a Set (it is one Postgres round trip per lookup in production, which the bench does not model); velocity is a linear scan here and an indexed query in production.
Budget guards
packages/rule-engine/src/__tests__/evaluate.budget.test.ts runs with pnpm test (about half a second). After 50 warm-up iterations it times 200 iterations of every single-rule scenario and of the seeded set and asserts p95 under 50 ms per rule and under 250 ms for the set; the failure message carries the measured p95. These are regression guards, not SLOs: the bounds sit 15 to 1 000 times above the p99 measured above so that a shared CI runner passes and a tenfold slowdown (a handler that starts awaiting something per call, a CEL environment that got expensive, a span that allocates) fails. Tighten them only after a quieter runner exists; comparing against the previous run is deferred (docs/plans/deferred.md).
The third guard asserts that the CEL budget (CEL_LIMITS.timeoutMs, 50 ms) cuts a pathological expression off with outcome error and the reason expression timed out after 50ms, within 50 to 1 000 ms. The expression nests two comprehensions over a 500-entry list and calls inList for every pair (up to 250 000 lookups) through a list service that resolves on the next macrotask, as the Postgres-backed one does. That detail matters: the budget is a Promise.race against a timer, so it interrupts evaluation only where the expression awaits (a list lookup). A synchronous comprehension cannot be interrupted at all (a 1 000 by 1 000 nested map ran to completion in about 140 ms on the dev machine) and is bounded by the parse limits instead (500 AST nodes, depth 32, 1 000 list elements), and a lookup service that resolves in a microtask starves the timer the same way (the same nested lookup ran 500 ms uninterrupted with the Set-backed service). In production the list service does I/O, so the budget holds there; an in-memory list implementation must yield (setImmediate) to keep it.
After a merge
pnpm ship runs the benchmark after every merge into main (scripts/ci-local/main/bench.sh; never on pull requests): it builds @aletheia-dev/rule-engine and its dependencies, runs the bench with BENCH_JSON set, prints the table with bench/summary.mjs and keeps the JSON as .git/local-ci/bench/<commit>.json (the --reporter=json output: testResults[].assertionResults[].benchmarks[].tasks[], each task with latency and throughput statistics in tinybench's shape, period and rank). Compare two results only when both ran on the same machine and the rme columns are a few percent; the job is informational and does not fail on slower numbers.
Load test
scripts/load/k6-onboarding.js (k6; scripts/load/README.md) drives the API and worker like integrations do: each virtual user (VU) creates a subject, starts a run and waits for it to complete, in a loop. hello runs hello-world to an automated approve; kyb runs kyb-onboarding to waiting_collection, fills and submits the form as the applicant (token from context.submissionUrl) and waits for the automated approve; mixed (the default) is 70 % hello and 30 % kyb. Every status code is checked; custom metrics are run_completion_time (start to completed), run_failed (failed, cancelled or over the 30 s per-run budget) and submit_latency. Thresholds in the script: http_req_failed < 1 %, API p95 (2xx responses) < 300 ms, run completion p95 < 10 s. Every VU calls as the same service user, so the API under test runs with API_RATE_LIMIT_PER_MINUTE=0; run.sh stops when it sees the limit. bash scripts/load/run.sh adds the admin PAT, the worker gauges before and after (aletheia_workflow_run_active, temporal_worker_task_slots_available / _used) and their peak during the run. It is informational and run by hand, for example against a release candidate.
Reference runs, SCENARIO=mixed, on an Apple M2 Pro development machine (10 cores, 16 GB; Docker Desktop with 5 CPUs and 8 GB running Postgres, Temporal, Zitadel and Garage), one API process and one worker process with the default task slots (40 workflow, 100 activity per queue), both as node dist/main.js on the host:
| VUs | Duration | API p50 / p95 / p99 (ms) | Requests/s | Runs/s | Completion p50 / p95 (s) | Runs failed | Peak slots used (aletheia-workflows) | Thresholds |
|---|---|---|---|---|---|---|---|---|
| 10 | 1 min | 29 / 191 / 311 | 43 | 3.5 | 2.6 / 6.1 | 0 of 213 | 4 of 40 workflow, 5 of 100 activity | all pass |
| 50 | 2 min | 42 / 152 / 797 | 160 | 2.5 | 14.5 / 26.7 | 52 of 365 | 13 of 40 workflow, 15 of 100 activity | completion |
Both runs had zero failed HTTP requests and all status checks passed; the submit call (the heaviest write) was p95 327 ms at 10 VU and 1.4 s at 50 VU. The 50 VU run fails only the completion threshold: run throughput on this machine tops out at about 3 to 3.5 runs/s, so extra VUs queue rather than add throughput, and 14 % of runs outlived the 30 s budget (19 hello, 30 kyb after submitting, 3 kyb before reaching collection). The worker was not the limit: its task slots were at most a third used and the process sat near 30 % of a core, the API near 40 %. Postgres (shared between Temporal's persistence and the application in the dev compose) ran at about 95 % of a CPU and the Temporal server at about 85 %, so Temporal scheduling on a CPU-starved shared database is what saturates first. Treat the 10 VU row as this machine's reference; at 50 VU the API threshold (p95 under 300 ms) holds here and the completion threshold does not, and both should be re-measured on the reference VM (docs/deployment.md, "Sizing") before being quoted.