Development Workflow
This document covers running tests, integration tests, e2e tests, and the test-infrastructure contracts that govern cross-runtime parity verification. For top-level build/typecheck/lint commands see the top-level CLAUDE.md.
⚠️ You CAN run every durable runtime locally — do NOT skip it
Before claiming a runtime (DBOS, Cloudflare, Temporal, Redis) "can't be tested locally," spin up the infra below and actually run the test. Postgres comes from Docker; Cloudflare (Miniflare/workerd) needs no Docker at all. A change that touches a runtime is not "robustly tested" until its real integration/e2e tests have run green — unit tests with mocked stores do not prove cross-runtime behavior.
| Runtime / store | How to run locally | Docker? |
|---|---|---|
| DBOS (Postgres-backed) | RUNTIME_DBOS_INTEG=1 POSTGRES_URL=postgresql://test:test@localhost:5433/helix_agents_test pnpm exec vitest run --config vitest.integ.config.ts <file> (from packages/runtime-dbos). The DBOS integ suite is paused unless RUNTIME_DBOS_INTEG=1 is set. | Yes (Postgres) |
| Cloudflare workers/DO | pnpm run test:cloudflare (.cf.test.ts, real workerd) | No |
| Cloudflare Workflows | pnpm run test:workflows (.workflow.test.ts + each .wf-noiso.test.ts in its own isolate) | No |
| Cloudflare Workflows (real engine, chat hosts) | pnpm --filter @helix-agents/e2e run test:cfw-engine (packages/e2e/src/__tests__/*.wf.test.ts in workerd against the real Miniflare Workflows engine, D1 and the DO stream; its own CI job, not a turbo task; build first) | No |
| store-cloudflare (D1/DO) | pnpm --filter @helix-agents/store-cloudflare run test:integration (Miniflare) | No |
| Redis store | REDIS_URL=redis://localhost:6379 pnpm --filter @helix-agents/store-redis run test:integration | Yes (Redis) |
| Temporal e2e | mostly in-process via @temporalio/testing (no Docker); a few Redis-backed Temporal e2e need docker compose ... temporal | Sometimes |
Bring up Postgres + Redis: docker compose -f docker-compose.test.yml up -d postgres-store redis.
Sandbox / environment gotchas (learned the hard way):
- Port 5433 already taken? Another project's Postgres may hold it. Run a dedicated container on a free port and point
POSTGRES_URLat it:docker run -d --name helix-test-pg -e POSTGRES_USER=test -e POSTGRES_PASSWORD=test -e POSTGRES_DB=helix_agents_test -p 5544:5432 postgres:16-alpinethenPOSTGRES_URL=postgresql://test:test@localhost:5544/helix_agents_test. - Temporal "Failed to connect before the deadline" with no server running is the in-process
@temporalio/testingenv failing to download its test-server binary — a network-restricted-sandbox symptom, NOT "Temporal can't be tested." It runs fine in CI / with network. It does not mean Postgres or Miniflare are unavailable — those have no such dependency. - Miniflare
mkdirat: File name too longon a.wf-noiso.test.tsis a temp-dir path-length limit triggered by a deeply-nested checkout path (Miniflare encodes the absolute path into the temp dir name). Environmental, not a code failure; run from a shorter path if it blocks you — e.g. a short-path worktree of your branch:git worktree add /tmp/hx-wf <branch>, thenpnpm install,pnpm run buildandpnpm run test:workflowsfrom there. - Conformance cf-do persist dirs. The conformance suite's cf-do backend (
packages/e2e/src/harness/cf-do/miniflare-backend.ts) keeps each Miniflare's DO-SQLite under a short literal/tmp/hxc-XXXXXXdirectory, not underos.tmpdir(): macOS's per-user temp path is long enough to overflow workerd's persistence paths. Each backend removes its dir on dispose; a killed run can leave/tmp/hxc-*dirs behind, and they are safe to delete once no conformance run is active.
Node-side Miniflare (the *.integ.test.ts suites) uses Miniflare 5. The integration suites in store-cloudflare, runtime-cloudflare, memory-cloudflare and e2e construct new Miniflare(...) in Node and reach bindings through Miniflare's Node proxy (getD1Database, getDurableObjectNamespace). They pin miniflare@5.20260925.0-alpha (the version wrangler@4.141.0 ships), which contains the SynchronousFetcher fix for FU-FLAKE-MINIFLARE-SYNC-PROXY-DESYNC (cloudflare/workers-sdk#15552). Miniflare 5 replaced the flat v3/v4 option shape (modules, script, durableObjects, d1Databases, ...) with workers: [{ config, dev }]. Keep writing the familiar v4 shape and wrap it in convertV4MiniflareOptions(...) (exported by miniflare), as every existing suite does. The converter drops the v4 per-plugin persistence options (durableObjectsPersist, d1Persist, kvPersist) without an error: use the top-level resourcePersistencePath instead, or storage silently stops persisting across a dispose and re-create. @cloudflare/vitest-pool-workers (the .cf.test.ts / .workflow.test.ts suites) stays on ^0.9.x with its own bundled Miniflare 4. Tests there run inside workerd, so they never use the Node proxy. Newer pool-workers releases require vitest 4, which this repo has not adopted.
Package manager (pnpm)
The repo uses pnpm; the exact version is pinned in package.json (packageManager), and any pnpm 10+ on your PATH switches to it automatically. Workspace settings (overrides, allowed install scripts, hoisting) live in pnpm-workspace.yaml.
- Declare every import in the importing package's own
package.json. pnpm does not hoist dependencies, so a package can only resolve what it (or the workspace root) declares. That includes test-only imports,@types/*packages needed bytsc, and binaries a script runs. An import that works in one package because a sibling declares it will fail in the other. - Internal dependencies use
"@helix-agents/<pkg>": "workspace:^".pnpm publish(whichchangeset publishuses) rewrites these to^<current version>. - Pass arguments without
--:pnpm --filter @helix-agents/core run test src/foo.test.ts. pnpm forwards a literal--to the script, and vitest then ignores the file filter and runs the whole suite. - Dependencies with install scripts must be listed under
allowBuildsinpnpm-workspace.yaml;pnpm installfails on an unreviewed one. Preferfalsefor optional native accelerators: atrueentry that compiles with node-gyp fetches the Node headers from the network on every fresh install, and CI has no cache for them (FU-CI-NATIVE-BUILD-CPU-FEATURES-HANG). @helix-agents/e2eis linked into the rootnode_modules(publicHoistPattern) because runtime-cloudflare's workflows tests import its parity harness, and declaring that dependency would create a package-graph cycle.
Running a single test
# Run tests in a specific package
pnpm --filter @helix-agents/core run test
# Run a specific test file
pnpm --filter @helix-agents/core run test src/__tests__/orchestration/step-processor.test.ts
# Run tests matching a pattern
pnpm --filter @helix-agents/core run test --testNamePattern="should"
# Watch mode for a package
pnpm --filter @helix-agents/core run test --watchIntegration tests
Integration tests require external services. The docker-compose.test.yml provides Redis, PostgreSQL, and Temporal.
Start test infrastructure:
docker compose -f docker-compose.test.yml up -dRun integration tests:
# All integration tests (requires all services running). turbo passes the
# test env vars through (see "Environment variables under turbo" below), so
# RUNTIME_DBOS_INTEG=1 / HELIX_TEMPORAL_BACKEND / TEMPORAL_ADDRESS set here
# reach the suites.
RUNTIME_DBOS_INTEG=1 pnpm run test:integration
# Redis store tests only
REDIS_URL=redis://localhost:6379 pnpm --filter @helix-agents/store-redis run test:integration
# Postgres store tests only
POSTGRES_URL=postgresql://test:test@localhost:5433/helix_agents_test pnpm --filter @helix-agents/store-postgres run test:integration
# Cloudflare store tests (uses Miniflare, no external services needed)
pnpm --filter @helix-agents/store-cloudflare run test:integration
# E2E tests (requires Redis and optionally Temporal)
pnpm run test:e2eStop test infrastructure:
docker compose -f docker-compose.test.yml downEnvironment variables under turbo, and loud test gates
turbo 2 runs every task in strict env mode: a task only sees the variables turbo.json names, plus a small built-in set (PATH, HOME, CI, NODE_OPTIONS, ...). Everything else is dropped silently. Until the CI-health MR, pnpm run test:integration at the root dropped RUNTIME_DBOS_INTEG (the runtime-dbos suite collected zero files and exited 0) and TEMPORAL_ADDRESS (Temporal tests fell back to localhost:7233, which worked in CI only because GitLab's Kubernetes executor exposes service containers on localhost).
globalPassThroughEnvinturbo.jsonlists every variable the test suites read: the gates (RUNTIME_DBOS_INTEG,HELIX_DOCKER_TESTS,REMOTE_MATRIX_FULL,SKIP_*_E2E, ...), backend selection (HELIX_TEMPORAL_BACKEND,TEMPORAL_ADDRESS,TEMPORAL_URL,HELIX_TEST_BACKEND_FILTER), connection strings, API keys and the conformance variables. The list applies to every turbo task, sopnpm run test:integration,test:e2e,test:cloudflare, ... all pass them through. Running a suite package-direct (pnpm --filter <pkg> run test:integration) never went through turbo and still works as before.pnpm run check:turbo-env(scripts/check-turbo-env.mjs, run by the CIlintjob) scans the test surface forprocess.env.Xreads and fails when a name is neither in turbo.json nor in the script'sIGNOREDlist (variables a test sets for itself). When you add a new env var to a test, add it toglobalPassThroughEnvor the check fails.- The one whole-suite gate is going away.
RUNTIME_DBOS_INTEGis the only gate that can make a suite collect zero files and exit 0 (the runtime-dbos wrapper prints aPAUSEDwarning when it is unset). MR !259 (dbos-integ-ci-unpause, GL #108) removes the gate and runs the suite by default, locally and in CI, so the CI-health MR does not change the DBOS wrapper or thetest:integ:dbosjob. Until !259 lands, passRUNTIME_DBOS_INTEG=1to run the suite; with this passthrough it reaches the suite through turbo too. - Temporal: in CI,
createDockerTestEnvironmentwarns whenTEMPORAL_ADDRESSis unset (it is a warning, not an error, because the helper is published as@helix-agents/runtime-temporal/testing).
Per-test describe.skipIf(...) gates (for example HELIX_DOCKER_TESTS, REMOTE_MATRIX_FULL) still show up in vitest's summary as skipped tests. Only a whole-suite gate can produce a silent "0 files, exit 0", which is why a dropped variable is caught by check:turbo-env rather than left to a skip.
Services provided by docker-compose.test.yml:
- Redis Stack on port 6379 - for store-redis and memory-redis integration tests (includes RediSearch module for vector search)
- PostgreSQL (store) on port 5433 - for store-postgres integration tests (user: test, password: test, database: helix_agents_test)
- Temporal on port 7233 - lightweight
temporal server start-dev(Go-based, in-memory, no Postgres dep). Required for all Temporal integration tests — the framework defaults to this Docker backend. SetHELIX_TEMPORAL_BACKEND=test-serverto opt back into the in-process Java@temporalio/testingserver (no docker needed, but loses C-1 stress coverage at N=20). See "Temporal backend selection" below.
Notes:
- The store-cloudflare tests use Miniflare to simulate Cloudflare Workers locally, so they don't need Docker
- Some tests in store-redis are skipped due to pub/sub timing flakiness (documented in test comments)
E2E test helpers
The packages/e2e/src/helpers/ directory provides reusable infrastructure for writing cross-runtime e2e tests:
temporal-setup.ts—createTemporalTestContext()creates an in-process Temporal environment using@temporalio/testing. No Docker/external Temporal server needed. Returns{ runner, registry, stateStore, streamManager, llmAdapter, cleanup }.cloudflare-executor-setup.ts/cloudflare-pool-setup.ts— Miniflare-based Cloudflare Workers test environment (no Docker needed).redis-setup.ts— Redis connection helpers for integration tests that need Redis.stream-collector.ts—collectChunks()for collecting stream output into an array.assertions.ts— Common test assertions:assertNoErrors(),assertHasOutput().test-workflows.ts— Temporal workflow wrapper (testAgentWorkflow) used by the test runner.
When writing new e2e tests, use existing helpers rather than building from scratch. Cross-runtime tests follow established patterns in agent-hooks-parity.integ.test.ts, ai-sdk-temporal.integ.test.ts, tracing-temporal.integ.test.ts, etc.
Temporal tests use a Docker-hosted real Temporal server by default — a temporal server start-dev container provided by docker-compose.test.yml (the temporal service on port 7233). Bring it up before running any Temporal integration test:
docker compose -f docker-compose.test.yml up -d temporalTemporal backend selection (HELIX_TEMPORAL_BACKEND)
The Temporal test harness can run against either backend, selected at test-process start via env var:
- Default (no env var, or
HELIX_TEMPORAL_BACKEND=docker) — connects to a real Temporal server (defaults tolocalhost:7233; override viaTEMPORAL_ADDRESS). Required for stress tests that exceed the embedded Java test server's hard limits — specifically the grpc-java 16384-byteSETTINGS_MAX_HEADER_LIST_SIZE(grpc/grpc-java#4284 / #8919 / #11430) which theconcurrency-invariants-stress > temporal-* > C-1test hits at N=20. Also avoids thetemporalio/sdk-typescript#876teardown race. HELIX_TEMPORAL_BACKEND=test-server— opt back into the in-process JavaTestWorkflowEnvironmentfrom@temporalio/testing. Zero docker deps, supports time-skipping. Useful for offline / no-docker work, but cannot run the C-1 stress invariant at N=20 (will silently skip via the harness check or hit the grpc-java header limit). CI does not exercise this path.
Both backends expose the same structural TemporalTestEnvironment interface to consumers (TemporalTestRunner, the e2e harness adapter) so test code is backend-agnostic. See packages/runtime-temporal/src/__tests__/testing/test-environment.ts.
Test file suffixes
*.test.ts— Unit tests. Mocks/stubs only, no external services. Run withpnpm run test.*.integ.test.ts— Integration tests. May need Redis, Temporal, or Miniflare. Run withpnpm run test:integration.*.cf.test.ts— Cloudflare worker tests. Run inside the actualworkerdruntime via@cloudflare/vitest-pool-workers. Run withpnpm run test:cloudflare.*.workflow.test.ts— Cloudflare Workflows infrastructure tests. Also run insideworkerd. Run withpnpm run test:workflows.*.wf.test.ts(packages/e2e only) — chat-host suites against the REAL Cloudflare Workflows engine, run in the worker's own isolate (vitest.cfw-engine.config.ts). Run withpnpm --filter @helix-agents/e2e run test:cfw-engine(build first; it importsdist/). The e2e package's default vitest config excludes them.
Build before direct
.cf.test.ts/.workflow.test.tsruns. These tests import@helix-agents/runtime-cloudflare(and other workspace packages) from their builtdist/, NOT from source. The turbo-orchestrated scripts (turbo run test:cloudflare/test:e2e) are safe — theirdependsOn: ["build", "^build"]rebuilds dependencies first. But invoking vitest directly (e.g.pnpm --filter @helix-agents/e2e run test:cloudflare <file>orvitest run --config vitest.cloudflare.config.ts) bypasses turbo and runs against whateverdist/exists — STALE dist will silently exercise old code. Runpnpm run build(orpnpm --filter @helix-agents/runtime-cloudflare run build) after editing source and before a direct cf/workflow test run.
Integration tests include runtime parity tests — many features are tested across multiple runtime/store combinations to ensure consistent behavior. Common patterns in the e2e package:
interrupt-resume-{js,redis,cloudflare,temporal}— same feature across runtimesai-sdk-{js,redis-js,temporal,redis-temporal,cloudflare}— AI SDK with different backendsusage-{tracking,redis-js,redis-temporal,temporal,cloudflare}— usage tracking across stacks
Integration tests live both in the e2e package (cross-package flows) and within individual packages (store-redis, store-cloudflare, ai-sdk, agent-server, etc.).
When making changes (features, bug fixes, refactors), validate across the full test surface area. Changes to core or shared interfaces can break any runtime or store implementation. Run unit tests first for fast feedback, then integration and e2e tests to catch cross-runtime issues. Don't assume a change that passes JS runtime tests will work on Temporal or Cloudflare — the parity tests exist because runtimes have different execution models and failure modes.
Test-infrastructure contracts
A few runtime-level test infrastructure invariants worth knowing about before debugging cross-file flake:
- DBOS per-incarnation
executorIDisolation (commit3a2f6e453) —createDBOSE2EContext()andcreateDBOSTestContext()mint a freshexecutorIDper call. Vitest's default forks pool runs e2e files in parallel, but every fork shares the SAME Postgresdbos.*system tables. Without per-fork executorIDs, file B'sDBOS.launch()would recover PENDING workflows leaked by file A and re-execute them under file B's globally-boundCallLLMStep.llmAdapter. The mint-per-context fix pins each invocation to its own recovery bucket;restartProcess()reuses the captured executorID so recovery-after-crash tests still work.listWorkflows()andcancelWorkflow()cleanup is also scoped to the local executorID so cleanup never sabotages a sibling fork's in-flight workflows. MockLLMAdaptergracefulExhaustmode (commit51cf75b60) — opt-in vianew MockLLMAdapter([...], { gracefulExhaust: true }). When set, exhaust returns an errorStepResult(withshouldStop: true) instead of throwing — Temporal-runtime tests that fire-and-forgetresultPromiseno longer accumulate uncaughtWorkflowFailedErrorrejections after vitest teardown.createTemporalTestContext()opts in by default. The JS / Cloudflare / DBOS test helpers don't need this — those runtimes don't have an activity-level retry layer that would amplify exhaust into a retry storm. Default remains throw-on-exhaust to preserve existing semantics.- StreamChunk schema cross-package parity (commit
68ca082a5) —packages/e2e/src/__tests__/stream-chunk-roundtrip-parity.integ.test.tsenumerates everyStreamChunkSchemavariant (33 fixtures across 31 chunk type literals + 3suspension_marker.payloadvariants) and asserts each round-trips identically through bothInMemoryStreamManagerandRedisStreamManager(which validates viaStreamChunkSchema.parseon read). Coverage is asserted at runtime — adding a new schema variant without a fixture fails the guardian test rather than silently regressing the consumer-side validation.
Runtime conformance suite
Outcome-level scenarios run across every runtime × store × tier with a strict expected-failure ledger. Design: docs/superpowers/specs/2026-09-23-runtime-conformance-suite-design.md. Catalogue: docs/conformance/findings.md. Burn-down: docs/conformance/status.md. Writing a scenario or editing your own finding record: docs/dev/conformance-scenario-authoring.md.
- Unit tests (no infra, but needs built dist — see
pnpm run buildbelow):pnpm --filter @helix-agents/e2e run test:conformance-unit— coverssrc/conformance/__tests__and the CI scripts' own tests inpackages/e2e/scripts/__tests__(benign-teardown, soak, verify) - cf-do harness self-tests (Miniflare; no Docker,
pnpm run buildfirst):pnpm --filter @helix-agents/e2e run test:conformance-harness. They guard the cell contract, so CI runs them only in the non-retryingconformance:verifygate; the retrying e2e config excludessrc/harness/cf-do/**, and the cf-do unit tests run withtest:conformance-unit. - Floating-promise lint (type-aware; an un-awaited
check()/settle()is a silent pass) oversrc/conformance,src/harnessand theconformance-*.integ.test.tsentrypoints:pnpm --filter @helix-agents/e2e run lint:conformance - Matrix (infra up;
pnpm run buildfirst):pnpm --filter @helix-agents/e2e run test:conformanceandrun test:conformance-dbos.test:conformanceruns four files: the JS, Temporal and canary executor-tier entrypoints, plusconformance-cf.integ.test.ts(the host and client tiers on cf-do and js-chat-host). - Host and client tiers alone, for iteration (cf-do under Miniflare + the JS contrast host; no Docker or services,
pnpm run buildfirst — the cf-do worker bundle resolves@helix-agents/*fromdist/):pnpm --filter @helix-agents/e2e run test:conformance-cf. It runs that 4th file through the samevitest.conformance.config.ts, so do not point it at a results dir that a fulltest:conformancealso writes to:conformance:verifyfails a cell that ran more than once. - Verify merged results:
CONFORMANCE_RESULTS_DIR=… pnpm --filter @helix-agents/e2e run conformance:verify - Regenerate docs after editing a record under
catalogue/:pnpm --filter @helix-agents/e2e run conformance:docs
Rules: a ledgered cell must fail on exactly its listed assertion ids; an unexpected pass or a different failure fails CI. Never edit a finding's ledger fragments to turn CI green without understanding the change — see the header comment in catalogue/index.ts.
Declared gaps (DECLARED_GAPS in gaps.ts) are distinct from a product capability being rejected. When a scenario's needs (driver primitives, e.g. resume/attach) exceed what the tier's driver can express, the cell touches nothing and records declared-gap — this is a limitation of the observation tier, not the runtime under test, and is decided BEFORE the backend health gate below (a gap cell reports declared-gap even on a poisoned/setup-failed backend). It counts as passing only if listed in DECLARED_GAPS with a reason: CI fails an unlisted gap ("UNDECLARED GAP") and a listed gap the driver can now express ("STALE DECLARED GAP"). A missing PRODUCT capability (Scenario.requires, e.g. workspaces) is a different thing — the runner substitutes the scenario's rejection.attempt, which must observe a typed framework_not_supported error thrown by the driver call or reported as outcome.thrown (a run that starts and only fails late does not count).
Backend rebuild on crash / unsettled cell. A cell that crashes (the scenario threw outside check(), a harness/scenario defect) or is left UNSETTLED (a named ctx.settle(assertionId, promise, boundMs) timed out — itself a ledgerable failure under assertionId — or the cell's own outer timeout fired; either way the run may still be executing) makes the suite dispose the current backend instance and build a FRESH one (model, env, driver) before the next cell — so one hung run doesn't contaminate every later cell's verdict on that backend. Only a FAILED dispose or rebuild poisons the backend (GATED (<reason>) for every later cell on it). See the design spec §9 for the full rationale (this supersedes the earlier "poison on any crash" rule).
test:conformance / test:conformance-dbos run through scripts/run-vitest-benign.mjs --config <cfg> (not the looser run-both-vitest.mjs the rest of the integration suite uses): it streams the child vitest process live and forgives an exit ONLY when every test that ran passed AND the only unhandled error(s) are a known SDK teardown race, matched strictly against each error block's own headline (benign-teardown.mjs's isStrictBenignTeardownExit — a race string appearing elsewhere in an unrelated block, or after/inside the wrong place in the output, does not count). A real test failure is never forgiven regardless of what else is in the output. conformance:verify also runs a ci-guard.ts check first: the number of literal files in vitest.conformance.config.ts's include array must equal test:conformance's parallel: count in .gitlab-ci.yml (GitLab's --shard splits by file) — update both together.
The matrix jobs (test:conformance, test:conformance:dbos) run against the same docker-compose.test.yml infrastructure as the rest of the integration suite, with CONFORMANCE_RESULTS_DIR set so each shard's .jsonl results land in a shared, artifact-uploaded directory. Cells record their result as vitest task meta; the host-side ConformanceReporter (packages/e2e/src/conformance/reporter.ts, registered in both vitest.conformance*.config.ts) writes results-<job>-<pid>.jsonl and fails the run if any cell recorded no result. A CLI --reporter flag replaces the config's reporters, so keep the ConformanceReporter when you override them. A -t filter excuses only the cells it filtered out. Watch mode is refused (the reporter throws and writes nothing), because reruns would append duplicate results. Running the whole flow locally against the standard docker-compose.test.yml services (see "Integration tests" above — Redis on its default port 6379, Postgres on 5433):
docker compose -f docker-compose.test.yml up -d
export REDIS_URL=redis://localhost:6379
export POSTGRES_URL=postgresql://test:test@localhost:5433/helix_agents_test
export CONFORMANCE_RESULTS_DIR=$PWD/packages/e2e/.conformance-results
rm -rf "$CONFORMANCE_RESULTS_DIR"
# Unit tests also need a built dist of store-memory / workspace-memory / core
# (not just core) — build the whole workspace, not a single package.
pnpm run build
# All four files, including the host/client-tier cf entrypoint.
pnpm --filter @helix-agents/e2e run test:conformance
pnpm --filter @helix-agents/e2e run test:conformance-dbos
pnpm --filter @helix-agents/e2e run conformance:verifyAlternatively, against a standalone Redis/Postgres pair on non-default ports and the in-process Temporal test server (no Docker Temporal container needed) — e.g. for a sandbox that already has other services bound to 6379/5432/7233:
export REDIS_URL=redis://localhost:6390
export POSTGRES_URL=postgresql://test:test@localhost:5433/helix_agents_test
export HELIX_TEMPORAL_BACKEND=test-server
export CONFORMANCE_RESULTS_DIR=$PWD/packages/e2e/.conformance-results
rm -rf "$CONFORMANCE_RESULTS_DIR"
# Unit tests also need a built dist of store-memory / workspace-memory / core
# (not just core) — build the whole workspace, not a single package.
pnpm run build
pnpm --filter @helix-agents/e2e run test:conformance
pnpm --filter @helix-agents/e2e run test:conformance-dbos
pnpm --filter @helix-agents/e2e run conformance:verifyCI runs the same matrix sharded across test:conformance (4-way: one file per Node runtime family, the canaries, and the host/client-tier conformance-cf.integ.test.ts — see vitest.conformance.config.ts) and test:conformance:dbos. The cf shard needs no service (Miniflare runs workerd in-process) — only the build job's packages/*/dist/ artifacts, which every conformance job downloads through needs: [build]. Then conformance:verify merges every shard's results, checks the FULL expected-cell set (every scenario × tier × backend with a running lane in lanes.ts — executor-tier scenarios × every Node backend, host/client-tier scenarios × cf-do and js-chat-host — not just the ledgered ones: a cell that silently produced no result line at all is a failure, e.g. an empty shard or a dropped backend) in CI, and additionally runs test:conformance-unit, test:conformance-harness, lint:conformance and conformance-docs.ts --check. Every step runs even when an earlier one failed. The job collects their exit codes and fails at the end, so the report is always written. Its summary is printed to the job log and published as the Conformance report artifact exposed on the merge request (packages/e2e/.conformance-report/summary.md). It holds each gate step's result, the verdict (or the ci-guard / results problem that stopped verification), and the findings burn-down, the same table status.md renders. conformance:verify runs (needs: [build, test:conformance, test:conformance:dbos], rules: ... when: always) even when a matrix shard failed, so its report is never silently skipped — the pipeline still blocks overall via the red shard. None of the conformance jobs retry on script_failure (like every integration job, see "CI retry and flake policy" below), because a rerun must never paper over a ledger mismatch (they do still retry once on runner_system_failure/unknown_failure, genuine infra flake). A nightly conformance:soak job (schedule-triggered on the default branch only — needs: [build] requires build to have run in the same pipeline, which its own rules only allow for merge_request_event or main) runs the full matrix three times (scripts/conformance-soak-run.sh) and fails if any cell's verdict is inconsistent across runs, is skipped in any run, or is missing a result line in any run — including a cell that is missing from EVERY run, via an expected-cell set — and also if any run's job exited non-zero (each job's post-wrapper exit code is recorded in the run's exit-codes.txt; canary failures and teardown leaks show up only there) or any single run fails conformance:verify's CI-mode checks (a cell failing the same way in every run is not flaky, but still fails the soak) (scripts/conformance-soak.ts / src/conformance/soak.ts). To run a local soak (e.g. 2 iterations) with the same env as the local matrix above:
rm -rf packages/e2e/.soak # the script refuses a non-empty soak dir (a rerun would double-count)
bash packages/e2e/scripts/conformance-soak-run.sh 2 $PWD/packages/e2e/.soakThe repo's Claude Code Stop hook (.claude/hooks/check-docs-stop.sh) has its own test suite, run in CI by the cheap test:claude-hooks job (bash + git only): bash .claude/hooks/__tests__/check-docs-stop.test.sh.
CI retry and flake policy
CI retries infrastructure failures only, in every job. The pipeline-wide default (top of .gitlab-ci.yml) retries runner_system_failure, unknown_failure and stuck_or_timeout_failure. The integration shards (.integration_base: test:integ:e2e, test:integ:dbos, test:integ:rest) and the conformance jobs (.conformance_base) retry once, on runner_system_failure and unknown_failure only. No job retries on script_failure, so a failing test is never turned into a green by a second attempt.
Two MRs removed the last genuine-failure retries:
- CI health (!272) stopped
.integration_baseretryingscript_failure. A sweep of the 45 pipelines before it found 20 integration shards that passed only on that retry; each is listed with its cause and owner indocs/dev/follow-ups.mdunder "CI health (MR !272)". - CI health 2 stopped
test:cloudflareretryingscript_failure. In the 164 pipelines it swept, that job's first attempt was red in 36 of the 78 after !272, and the retry turned every one green. It also removed every in-test vitest retry (below). The sweep, the causes and their fixes are under "CI health 2" indocs/dev/follow-ups.md.
No in-test retries either. Vitest retry (in a config or as a per-test { retry: N } option) re-runs a failing test inside a green job, and the CI reporter does not even print it. No vitest config or test sets one; don't add one. The example apps' Playwright configs still set retries (Playwright does print those, as "N flaky"). Those flakes and their removal are tracked as FU-FLAKE-APPROVAL-GATES-PLAYWRIGHT and FU-FLAKE-OPENNEXT-REFRESH-PLAYWRIGHT (owner: CI-health-3).
Write tests that load cannot break. A CI runner can be 100× slower than a laptop: a single Durable Object proxy read has taken seconds, and a test that takes 10ms locally took 13s. So:
- No wall-clock windows on positive outcomes. Never assert that something happens within N ms (
expect(elapsed).toBeLessThan(…), aPromise.raceagainst a timer, a loop that gives up after a deadline). Await the event itself; if it never happens, the test hangs until vitest's test timeout, which fails it. "Still waiting after N ms" checks are fine: load can only make them hold. - No fixed sleeps before acting on a run. Don't
sleep(150)and then interrupt; wait for the chunk, commit or state that proves the run reached the point the test needs (waitForStreamChunk, astep_committedmarker, a polled commit). Then the run's progress, not the runner's speed, decides what the test sees. - Prove a wake, not a latency. To show a reader is woken by an event rather than by a fallback poll, disable the poll (the StreamManager contract runs
DOStreamManager/RedisStreamManagerwithsafetyNetPollMs: Infinity) and await the read. - Prove a fix under load. Run the changed test at least 15 times in a row with
--retry=0under CI-like load (see "CI test jobs" below for how to get it).
When a job is red:
- Read the failure. If it is new, or plausibly caused by your diff, fix it. Never retry it away.
- A retry is acceptable only for
runner_system_failure, or, for a genuine test failure, when ALL of these hold (the reliability programme's rule):- the failure is a known bug already on
mainand unrelated to your diff; - you disclose the retry in the MR with links to both the failed and the retried job;
- the bug has a filed entry (a
follow-ups.mdFU or a conformance finding) with a named owner, and the MR links it; - every test your MR adds or changes passed on its first attempt.
- the failure is a known bug already on
- Otherwise the MR is blocked. A flaky test is a bug: root-cause it, or file it and hand it to its owner.
Quarantines. Only a user decision can quarantine a test (test.fixme, it.skip, or a harness skip gate). Each one is listed, with its bug, owner and exit condition, in the single register at Quarantined tests. That register covers packages/e2e tests as well as example apps. Add every new quarantine there, along with its follow-ups.md entry.
CI test jobs
A test job never rebuilds. turbo's test tasks (test, test:integration, test:cloudflare, test:workflows, typecheck, …) depend on build, and turbo has no cache in CI. So a plain turbo run test:cloudflare in a test job rebuilt every workspace package (every example, five next builds, the docs) concurrently with its tests, which made the tests 10–100× slower and was behind most of the test:cloudflare flakes. Instead:
- the
buildjob builds once and uploadspackages/*/dist/andexamples/*/dist/as artifacts; - every test job has
needs: [build]and runs turbo with--only(pnpm exec turbo run test:cloudflare --only), so it runs the tests and nothing else; - a job that needs an output beyond
dist/builds only that package: the Next.js example jobs runturbo run build --only --filter=<their app>.
Locally, pnpm run test:cloudflare (no --only) still builds first, as before.
A red job lists every failing package. The multi-package turbo jobs (test:unit, test:integ:rest, test:cloudflare, test:workflows, typecheck, lint) run with --continue=always, so one package's failure no longer cancels the rest: the job shows every failing package, not just the first, and still exits non-zero. (turbo's default, --continue=never, once stopped test:unit at 6 of 36 tasks, so an MR's own tests never ran; FU-CI-UNIT-TURBO-STOPS-AT-FIRST-FAILURE.)
Where jobs run. arm64 jobs (build, test:unit, test:cloudflare, test:workflows, test:cfw-engine, typecheck, lint, …) run on the jmif-dev-srv runner, which is a developer machine's Docker VM shared with other projects' pipelines. Heavy local runs on that machine (CPU stress, parallel matrices) slow those jobs down. amd64 jobs (the integration shards, conformance, the example apps) run on a separate Kubernetes runner.
Reproducing CI-like load. On that machine, run the job's command in a node:24 container (linux/arm64, like the runner) on a copy of the repo: git ls-files -z | COPYFILE_DISABLE=1 tar --null -T - -cf - | docker run -i -v <vol>:/w node:24 tar -xf - -C /w/agents (without COPYFILE_DISABLE=1, macOS tar adds ._* files that vitest then tries to load). Then install and build as the job does (.pnpm_install, pnpm run build), and loop the job's test command with --retry=0 while a forced rebuild (pnpm exec turbo run build --force) runs beside it.
Example app E2E tests (Playwright + real workerd)
The example apps under examples/opennext-cloudflare-do/ and examples/research-assistant-cloudflare-do/ have Playwright E2E tests that run against a real wrangler dev (real workerd, real DOs, real D1). The only thing mocked is the OpenAI HTTP API, via a record/replay proxy at examples/*/test-utils/openai-mock-server.ts.
Read ./example-app-testing.md before modifying example tests. It documents the discipline these tests must follow: fixtures must be real OpenAI recordings (no hand-crafting), prompts must be natural language (no strict "EXACTLY ONCE" directives), assertions must test SDK invariants (no LLM-compliance checks), and the mock must stay structural (semantic hashing, not selective request matching). Several "obvious" shortcuts are explicitly forbidden — the doc explains why.
CI runs these in OPENAI_MOCK_MODE=replay against checked-in fixtures. If a test changes prompts, agent system prompts, or tool definitions, fixtures need to be re-recorded — see the doc for the procedure.