Test infrastructure roadmap
This doc tracks deferred test-infrastructure and framework work surfaced during sub-project D (test infrastructure overhaul, branch omnara/stateless-suspension-redesign). Each item is captured with scope, estimated effort, dependencies, and priority. Items shipping in v7.x get tracked here until they land.
Atomicity stress (Option B — chaos / fuzz / property-based)
Sub-project D landed Option A scope (per-store concurrent-writer bumps to 20-50). The remaining Option B items below are all correctness-focused (chaos injection, fuzz, property-based). The earlier A1/A2/A3 scale-stress items (100-writer Postgres CAS, 1k-session deadline fan-out, 10k-message × 10 concurrent suspends) were removed as part of the perf-test purge — they were primarily about scaling existing tests' volume rather than surfacing new correctness invariants, and there's no appetite for that flavor of work right now. If a real race surfaces in production that needs higher-volume reproduction, re-add the specific case with intent.
A4. Postgres connection-kill mid-transaction
- Scope: Network partition simulation. Mid-transaction, kill the Postgres connection. Verify the writer sees a
StaleStateErroror equivalent typed failure (not silent corruption). - Effort: ~3 days. Postgres test infra needs a network-layer chaos helper (analogous to existing
redis-chaos-helpers.ts). - Priority: v7.x.
A5. D1 batch-failure injection
- Scope: Force the D1
batch()call to fail mid-batch (e.g., simulated rate-limit response). Assert the recovery path resolves cleanly per the fix in commit7509872e3. - Effort: ~2 days.
- Priority: v7.x.
A6. Long-running fuzz
- Scope: Random schedule of submits, registers, expires, cleanups for N hours (parameterizable). Property-based invariants checked continuously. Best run nightly with separate CI lane.
- Effort: ~5 days.
- Priority: v7.x.
A7. Property-based tests via fast-check
- Scope: Define
saveStateAndPromoteStaginginvariants formally (atomicity, version monotonicity, message append idempotency, staging cleanup) and verify viafast-checkshrinking on each store. - Effort: ~5 days.
- Priority: v7.x.
Framework gaps surfaced during D (sub-project A or B)
Note: F1-F5 + F7 are runtime-specific gaps. F6 (atomic-counter merge across parallel tools) is feature work, grouped in its own section below since it's not a runtime-specific gap.
F1. Temporal submitToolResult routing for top-level entries
- Status: ✅ DONE (sub-project A.2 Task 2.1, commit
e50fee6be). - Fix:
commitSuspendedStepnow writes top-levelclientToolCallOwnershipinline as part of the atomic save batch for self-rooted entries (rootSessionId === ownerSessionIdORrootSessionId === undefined). Sub-agent entries continue to use the existingwriteRootOwnershippost-save path.
F2. Temporal INTERRUPT_POLL_CADENCE_MS = 30_000
- Status: ✅ DONE (sub-project A.2 Task 3.2, commit
b8294adb7). - Fix: Workflow body checks the durable interrupt flag at every step boundary via the new
consumeInterruptFlagactivity (sub-second observable). TheINTERRUPT_POLL_CADENCE_MSconstant + the 30swf.conditiontimeout are deleted.
F3. Temporal stateStore.listRuns() returns empty
- Status: ✅ DONE (sub-project A.2 Tasks 1.1/1.2/2.1, commits
a05173d2e,7ad2c9d80,e50fee6be). - Fix:
executor.execute()andexecutor.resume()write run records viastateStore.createRun()(withturn,startSequence,startMessageCount,startUIMessageCountmatching cross-runtime metadata).commitSuspendedStepwrites a run-record status update on suspension.applyResultsAndReloadwrites a fresh run record on resume entry.listRuns(sessionId)now returns observable runs on Temporal-backed sessions.
F4. Temporal handle.result() doesn't resolve on suspension
- Status: ✅ DONE (sub-project A.2 Tasks 1.2/3.2 + commit
6b5f4b309). - Fix: Workflow exits cleanly on every HITL boundary returning
AgentWorkflowResult { status: 'suspended_*' }. The executor's handle factory projects the workflow result ontoAgentResult.status(commit6b5f4b309was the missing projection in the active-handle factory).handle.result()now resolves with the'suspended_*'family for HITL agents that paused mid-run, matching the cross-runtime contract.
F5. Temporal teardown gRPC cleanup race
- Scope: When multiple Temporal backends run sequentially in the same vitest file (e.g., temporal-memory + temporal-redis + temporal-postgres), the second/third backends fail with
Channel has been shut down/Hook timed out in 30000mserrors duringafterAll. Per-backend filter mode (HELIX_TEST_BACKEND_FILTER=temporal-memory) works fine. - Discovered in: Cluster 7 + Cluster 8 verification runs.
- Sub-project owner: B or test-infra follow-up. Likely fix: ensure
TestWorkflowEnvironmentis fully torn down (not ref-count-deferred) between Temporal backend invocations in the same file. Or split temporal backends into separate files. - Still reproduces in test-server mode (2026-09-26, conformance Phase B Task 9): with
HELIX_TEMPORAL_BACKEND=test-server, the e2eremote-agent-matrix.integ.test.ts(a vitest worker dies:Channel has been shut down/Worker exited unexpectedly, 7 of 38 tests never finish) andconcurrency-invariants-stress.integ.test.ts(C-1 on atemporal-*backend:waitFor timeout after 90000ms, or the same worker crash) fail on origin/main and on the branch alike (3/3 on main). Against a real Temporal server (TEMPORAL_ADDRESS, the default mode and CI's) both pass 3/3. Use a real server when running the full e2e suite locally.
F7. DBOS resume() doesn't enforce contract for terminal/active sessions
- Status: ✅ DONE (sub-project A.3, commits
8ec773936,16e4acb99,7d9498b8e). - Root cause: truthy-object pitfall on
packages/runtime-dbos/src/lifecycle/resume.ts:79— the dual-CAS fallback never rejected becausecasFromInterrupted(the discriminated union object) was always truthy. The unit tests mockedcompareAndSetStatusreturning booleans, masking the bug. - Fix: single multi-status CAS (
['interrupted', 'failed'] → 'active') with explicit.okcheck, mirroring runtime-temporal's post-A.2 pattern.
Atomic-counter merge semantics (Cluster 14 deferred)
F6. Atomic counter merge across parallel tools
- Scope: Today's framework: parallel scalar customState writes are last-write-wins on every runtime + every state store (verified in
packages/runtime-dbos/src/__tests__/integration/staging-atomicity.integ.test.tsandpackages/e2e/src/__tests__/staged-state.integ.test.ts:335-338). Arrays witharrayDeltaMode: truecorrectly merge via deltas. Counters / numeric scalars don't. - Feature request: Add new
MergeChangesopcodes for increment/decrement on numeric keys, OR a per-key server-side merge function. Either path requires:- Extending the
MergeChangesschema in@helix-agents/core - Updating per-store
applyMergeToCustomStateto interpret new opcodes - Tests across all 5 stores
- Extending the
- Effort: ~1 week.
- Sub-project owner: Likely a separate sub-spec (this is real feature work, not test infra).
- Workaround: Use array-append semantics (push entries, assert array length) or serialize via a single tool.
Other deferred items
O1. LLM-based fuzzing
- Scope: Generate random tool-execution scenarios via small LLM prompts to find edge cases humans wouldn't think of.
- Effort: Unknown; experimental.
- Priority: v7.x; nice-to-have.
O2. Consolidation of per-backend test files
- Scope: Migrate
interrupt-resume-{js,redis,cloudflare,temporal}family,ai-sdk-*family, and similar same-feature-different-backend files into harness-driven parity tests. Sub-project D deliberately did NOT migrate these (would have risked breaking working tests). - Dependencies: Sub-project B shipped first, so feature gaps are filled and consolidation doesn't accidentally hide a missing-backend signal.
- Effort: ~5 days.
- Priority: v7.x; post-B.
O3. Capability metadata drift detection
- Scope: Lint that asserts
BackendDescriptor.capabilitiesmatches actual runtime fail-fast guards. E.g., ifTEMPORAL_CAPSincludes'workspaces'but Temporal still has the workspace fail-fast guard, lint fails. - Effort: ~2 days.
- Priority: v7.x.
O4. Real-LLM gated CI job
- Scope: New CI job
e2e:real-llmthat runs against real OpenAI / Anthropic, gated on env keys. Covers happy-path client-tool flows, approval gates, persistent sub-agents. - Dependencies: API keys provisioned in CI variables.
- Effort: ~3 days.
- Priority: v7.x; nice-to-have.
O5. Workerd-context CF DO + CFW Workflows harness setup helpers
- Status — Phase 1 ✅ DONE.
setupCfDoD1(stub.fetchAgentExecutoradapter over a generic injectableHarnessAgentServerDO) andsetupCfwWorkflowsD1(env.AGENT_WORKFLOW.create()+ poll adapter) are now REAL workerd-context implementations. The reusable machinery — context-split registries (node-backends/cf-backends),selectViable(registry, opts), therecordedHooks+prepareAgentBackendEnvbridge, the Node-import-free shared scenario module (parity/scenarios/lifecycle-hooks.ts), and the genericHarnessAgentServerDO — all landed. TheIMPL_PENDING_O5marker is dropped from both descriptors;cf-do-d1+cfw-workflows-d1are workerd parity-matrix participants verified byharness-smoke.cf.test.ts. - Phase 1 proof: the
lifecycle-hooks-paritysuite is now parity-complete across ALL runtimes — Node (7 backends), CF-DO (cf-do-d1, e2e workerd pool, Scenarios 1+2), and CFW (cfw-workflows-d1, runtime-cloudflare workflows pool, Scenarios 1+3). Two divergences are gated (NOT weakened): FU-O5-CF-DO-APPROVAL- STREAM-READ (CF-DO Scenario 3) + FU-O5-CFW-TRACING-CONTEXT-PERSISTENCE (CFW Scenario 2). Seedocs/dev/follow-ups.md. - Phases 2..N (open): apply the established recipe to the OTHER harness parity suites (usage-subagent, expired-session, concurrency- invariants-stress, approval-gate-hook-parity, atomic-suspend-write). Each is a thin entrypoint over the now-shared machinery.
- Pattern reference:
packages/runtime-cloudflare/src/__tests__/lifecycle-hooks-parity.wf-noiso.test.ts(CFW) +packages/e2e/src/__tests__/lifecycle-hooks-parity.cf.test.ts(CF-DO). - Sub-project owner: B (feature parity matrix).
- Update (conformance Phase B increment 1, 2026-09-26): the
cf-do-d1REGISTRY row was replaced bycf-do(packages/e2e/src/harness/registries/cf-backends.ts): a realDurableObjectAgentBasesubclass on DO-SQLite under Miniflare, driven from Node, with real evict (dispose + recreate over the same persist dir) and wake. It backs the conformance suite's host and client tiers (docs/dev/conformance-scenario-authoring.md), not the parity suites. The workerdsetupCfDoD1adapter is still used bylifecycle-hooks-parity.cf.test.ts, which builds its own descriptor for it. Phases 2..N above are unchanged.
Done
The following items shipped across the v7 stateless-suspension work train. Each item's detailed closure write-up lives inline in its original section above (the ✅ DONE status lines); this section is a flat index for quick scanning.
- F1 — Temporal
submitToolResultrouting for top-level entries (sub-project A.2 Task 2.1, commite50fee6be). Full write-up at the### F1section above. - F2 — Temporal
INTERRUPT_POLL_CADENCE_MSdeletion (sub-project A.2 Task 3.2, commitb8294adb7). Workflow body now checks the durable interrupt flag at every step boundary; 30s poll cadence +wf.conditiontimeout deleted. - F3 — Temporal
stateStore.listRuns()returns observable rows (sub-project A.2 Tasks 1.1/1.2/2.1, commitsa05173d2e,7ad2c9d80,e50fee6be).executor.execute()/resume()/commitSuspendedStep/applyResultsAndReloadall write run records viastateStore.createRun(). - F4 — Temporal
handle.result()resolves on suspension (sub-project A.2 Tasks 1.2/3.2 + commit6b5f4b309). Workflow exits cleanly withAgentWorkflowResult.status: 'suspended_*'; executor handle projects the result ontoAgentResult.status. - F7 — DBOS resume() contract for terminal/active sessions (sub-project A.3, commits
8ec773936,16e4acb99,7d9498b8e). Truthy-object pitfall on the discriminated-union CAS result fixed; unified to a single multi-status CAS mirroring the post-A.2 Temporal pattern.
Still open: F5 (Temporal teardown gRPC race), F6 (atomic counter merge — feature work, not a runtime gap), and the O1-O4 "other deferred" cluster. O5 Phase 1 is DONE (lifecycle-hooks parity-complete on cf-do-d1 + cfw-workflows-d1); O5 Phases 2..N (the remaining harness parity suites) remain via the established recipe.