Skip to content

Test infrastructure roadmap ​

This doc tracks deferred test-infrastructure and framework work surfaced during sub-project D (test infrastructure overhaul, branch omnara/stateless-suspension-redesign). Each item is captured with scope, estimated effort, dependencies, and priority. Items shipping in v7.x get tracked here until they land.

Atomicity stress (Option B — chaos / fuzz / property-based) ​

Sub-project D landed Option A scope (per-store concurrent-writer bumps to 20-50). The remaining Option B items below are all correctness-focused (chaos injection, fuzz, property-based). The earlier A1/A2/A3 scale-stress items (100-writer Postgres CAS, 1k-session deadline fan-out, 10k-message × 10 concurrent suspends) were removed as part of the perf-test purge — they were primarily about scaling existing tests' volume rather than surfacing new correctness invariants, and there's no appetite for that flavor of work right now. If a real race surfaces in production that needs higher-volume reproduction, re-add the specific case with intent.

A4. Postgres connection-kill mid-transaction ​

  • Scope: Network partition simulation. Mid-transaction, kill the Postgres connection. Verify the writer sees a StaleStateError or equivalent typed failure (not silent corruption).
  • Effort: ~3 days. Postgres test infra needs a network-layer chaos helper (analogous to existing redis-chaos-helpers.ts).
  • Priority: v7.x.

A5. D1 batch-failure injection ​

  • Scope: Force the D1 batch() call to fail mid-batch (e.g., simulated rate-limit response). Assert the recovery path resolves cleanly per the fix in commit 7509872e3.
  • Effort: ~2 days.
  • Priority: v7.x.

A6. Long-running fuzz ​

  • Scope: Random schedule of submits, registers, expires, cleanups for N hours (parameterizable). Property-based invariants checked continuously. Best run nightly with separate CI lane.
  • Effort: ~5 days.
  • Priority: v7.x.

A7. Property-based tests via fast-check ​

  • Scope: Define saveStateAndPromoteStaging invariants formally (atomicity, version monotonicity, message append idempotency, staging cleanup) and verify via fast-check shrinking on each store.
  • Effort: ~5 days.
  • Priority: v7.x.

Framework gaps surfaced during D (sub-project A or B) ​

Note: F1-F5 + F7 are runtime-specific gaps. F6 (atomic-counter merge across parallel tools) is feature work, grouped in its own section below since it's not a runtime-specific gap.

F1. Temporal submitToolResult routing for top-level entries ​

  • Status: ✅ DONE (sub-project A.2 Task 2.1, commit e50fee6be).
  • Fix: commitSuspendedStep now writes top-level clientToolCallOwnership inline as part of the atomic save batch for self-rooted entries (rootSessionId === ownerSessionId OR rootSessionId === undefined). Sub-agent entries continue to use the existing writeRootOwnership post-save path.

F2. Temporal INTERRUPT_POLL_CADENCE_MS = 30_000 ​

  • Status: ✅ DONE (sub-project A.2 Task 3.2, commit b8294adb7).
  • Fix: Workflow body checks the durable interrupt flag at every step boundary via the new consumeInterruptFlag activity (sub-second observable). The INTERRUPT_POLL_CADENCE_MS constant + the 30s wf.condition timeout are deleted.

F3. Temporal stateStore.listRuns() returns empty ​

  • Status: ✅ DONE (sub-project A.2 Tasks 1.1/1.2/2.1, commits a05173d2e, 7ad2c9d80, e50fee6be).
  • Fix: executor.execute() and executor.resume() write run records via stateStore.createRun() (with turn, startSequence, startMessageCount, startUIMessageCount matching cross-runtime metadata). commitSuspendedStep writes a run-record status update on suspension. applyResultsAndReload writes a fresh run record on resume entry. listRuns(sessionId) now returns observable runs on Temporal-backed sessions.

F4. Temporal handle.result() doesn't resolve on suspension ​

  • Status: ✅ DONE (sub-project A.2 Tasks 1.2/3.2 + commit 6b5f4b309).
  • Fix: Workflow exits cleanly on every HITL boundary returning AgentWorkflowResult { status: 'suspended_*' }. The executor's handle factory projects the workflow result onto AgentResult.status (commit 6b5f4b309 was the missing projection in the active-handle factory). handle.result() now resolves with the 'suspended_*' family for HITL agents that paused mid-run, matching the cross-runtime contract.

F5. Temporal teardown gRPC cleanup race ​

  • Scope: When multiple Temporal backends run sequentially in the same vitest file (e.g., temporal-memory + temporal-redis + temporal-postgres), the second/third backends fail with Channel has been shut down / Hook timed out in 30000ms errors during afterAll. Per-backend filter mode (HELIX_TEST_BACKEND_FILTER=temporal-memory) works fine.
  • Discovered in: Cluster 7 + Cluster 8 verification runs.
  • Sub-project owner: B or test-infra follow-up. Likely fix: ensure TestWorkflowEnvironment is fully torn down (not ref-count-deferred) between Temporal backend invocations in the same file. Or split temporal backends into separate files.
  • Still reproduces in test-server mode (2026-09-26, conformance Phase B Task 9): with HELIX_TEMPORAL_BACKEND=test-server, the e2e remote-agent-matrix.integ.test.ts (a vitest worker dies: Channel has been shut down / Worker exited unexpectedly, 7 of 38 tests never finish) and concurrency-invariants-stress.integ.test.ts (C-1 on a temporal-* backend: waitFor timeout after 90000ms, or the same worker crash) fail on origin/main and on the branch alike (3/3 on main). Against a real Temporal server (TEMPORAL_ADDRESS, the default mode and CI's) both pass 3/3. Use a real server when running the full e2e suite locally.

F7. DBOS resume() doesn't enforce contract for terminal/active sessions ​

  • Status: ✅ DONE (sub-project A.3, commits 8ec773936, 16e4acb99, 7d9498b8e).
  • Root cause: truthy-object pitfall on packages/runtime-dbos/src/lifecycle/resume.ts:79 — the dual-CAS fallback never rejected because casFromInterrupted (the discriminated union object) was always truthy. The unit tests mocked compareAndSetStatus returning booleans, masking the bug.
  • Fix: single multi-status CAS (['interrupted', 'failed'] → 'active') with explicit .ok check, mirroring runtime-temporal's post-A.2 pattern.

Atomic-counter merge semantics (Cluster 14 deferred) ​

F6. Atomic counter merge across parallel tools ​

  • Scope: Today's framework: parallel scalar customState writes are last-write-wins on every runtime + every state store (verified in packages/runtime-dbos/src/__tests__/integration/staging-atomicity.integ.test.ts and packages/e2e/src/__tests__/staged-state.integ.test.ts:335-338). Arrays with arrayDeltaMode: true correctly merge via deltas. Counters / numeric scalars don't.
  • Feature request: Add new MergeChanges opcodes for increment/decrement on numeric keys, OR a per-key server-side merge function. Either path requires:
    • Extending the MergeChanges schema in @helix-agents/core
    • Updating per-store applyMergeToCustomState to interpret new opcodes
    • Tests across all 5 stores
  • Effort: ~1 week.
  • Sub-project owner: Likely a separate sub-spec (this is real feature work, not test infra).
  • Workaround: Use array-append semantics (push entries, assert array length) or serialize via a single tool.

Other deferred items ​

O1. LLM-based fuzzing ​

  • Scope: Generate random tool-execution scenarios via small LLM prompts to find edge cases humans wouldn't think of.
  • Effort: Unknown; experimental.
  • Priority: v7.x; nice-to-have.

O2. Consolidation of per-backend test files ​

  • Scope: Migrate interrupt-resume-{js,redis,cloudflare,temporal} family, ai-sdk-* family, and similar same-feature-different-backend files into harness-driven parity tests. Sub-project D deliberately did NOT migrate these (would have risked breaking working tests).
  • Dependencies: Sub-project B shipped first, so feature gaps are filled and consolidation doesn't accidentally hide a missing-backend signal.
  • Effort: ~5 days.
  • Priority: v7.x; post-B.

O3. Capability metadata drift detection ​

  • Scope: Lint that asserts BackendDescriptor.capabilities matches actual runtime fail-fast guards. E.g., if TEMPORAL_CAPS includes 'workspaces' but Temporal still has the workspace fail-fast guard, lint fails.
  • Effort: ~2 days.
  • Priority: v7.x.

O4. Real-LLM gated CI job ​

  • Scope: New CI job e2e:real-llm that runs against real OpenAI / Anthropic, gated on env keys. Covers happy-path client-tool flows, approval gates, persistent sub-agents.
  • Dependencies: API keys provisioned in CI variables.
  • Effort: ~3 days.
  • Priority: v7.x; nice-to-have.

O5. Workerd-context CF DO + CFW Workflows harness setup helpers ​

  • Status — Phase 1 ✅ DONE. setupCfDoD1 (stub.fetch AgentExecutor adapter over a generic injectable HarnessAgentServer DO) and setupCfwWorkflowsD1 (env.AGENT_WORKFLOW.create() + poll adapter) are now REAL workerd-context implementations. The reusable machinery — context-split registries (node-backends / cf-backends), selectViable(registry, opts), the recordedHooks + prepareAgentBackendEnv bridge, the Node-import-free shared scenario module (parity/scenarios/lifecycle-hooks.ts), and the generic HarnessAgentServer DO — all landed. The IMPL_PENDING_O5 marker is dropped from both descriptors; cf-do-d1 + cfw-workflows-d1 are workerd parity-matrix participants verified by harness-smoke.cf.test.ts.
  • Phase 1 proof: the lifecycle-hooks-parity suite is now parity-complete across ALL runtimes — Node (7 backends), CF-DO (cf-do-d1, e2e workerd pool, Scenarios 1+2), and CFW (cfw-workflows-d1, runtime-cloudflare workflows pool, Scenarios 1+3). Two divergences are gated (NOT weakened): FU-O5-CF-DO-APPROVAL- STREAM-READ (CF-DO Scenario 3) + FU-O5-CFW-TRACING-CONTEXT-PERSISTENCE (CFW Scenario 2). See docs/dev/follow-ups.md.
  • Phases 2..N (open): apply the established recipe to the OTHER harness parity suites (usage-subagent, expired-session, concurrency- invariants-stress, approval-gate-hook-parity, atomic-suspend-write). Each is a thin entrypoint over the now-shared machinery.
  • Pattern reference:packages/runtime-cloudflare/src/__tests__/lifecycle-hooks-parity.wf-noiso.test.ts (CFW) + packages/e2e/src/__tests__/lifecycle-hooks-parity.cf.test.ts (CF-DO).
  • Sub-project owner: B (feature parity matrix).
  • Update (conformance Phase B increment 1, 2026-09-26): the cf-do-d1 REGISTRY row was replaced by cf-do (packages/e2e/src/harness/registries/cf-backends.ts): a real DurableObjectAgentBase subclass on DO-SQLite under Miniflare, driven from Node, with real evict (dispose + recreate over the same persist dir) and wake. It backs the conformance suite's host and client tiers (docs/dev/conformance-scenario-authoring.md), not the parity suites. The workerd setupCfDoD1 adapter is still used by lifecycle-hooks-parity.cf.test.ts, which builds its own descriptor for it. Phases 2..N above are unchanged.

Done ​

The following items shipped across the v7 stateless-suspension work train. Each item's detailed closure write-up lives inline in its original section above (the ✅ DONE status lines); this section is a flat index for quick scanning.

  • F1 — Temporal submitToolResult routing for top-level entries (sub-project A.2 Task 2.1, commit e50fee6be). Full write-up at the ### F1 section above.
  • F2 — Temporal INTERRUPT_POLL_CADENCE_MS deletion (sub-project A.2 Task 3.2, commit b8294adb7). Workflow body now checks the durable interrupt flag at every step boundary; 30s poll cadence + wf.condition timeout deleted.
  • F3 — Temporal stateStore.listRuns() returns observable rows (sub-project A.2 Tasks 1.1/1.2/2.1, commits a05173d2e, 7ad2c9d80, e50fee6be). executor.execute() / resume() / commitSuspendedStep / applyResultsAndReload all write run records via stateStore.createRun().
  • F4 — Temporal handle.result() resolves on suspension (sub-project A.2 Tasks 1.2/3.2 + commit 6b5f4b309). Workflow exits cleanly with AgentWorkflowResult.status: 'suspended_*'; executor handle projects the result onto AgentResult.status.
  • F7 — DBOS resume() contract for terminal/active sessions (sub-project A.3, commits 8ec773936, 16e4acb99, 7d9498b8e). Truthy-object pitfall on the discriminated-union CAS result fixed; unified to a single multi-status CAS mirroring the post-A.2 Temporal pattern.

Still open: F5 (Temporal teardown gRPC race), F6 (atomic counter merge — feature work, not a runtime gap), and the O1-O4 "other deferred" cluster. O5 Phase 1 is DONE (lifecycle-hooks parity-complete on cf-do-d1 + cfw-workflows-d1); O5 Phases 2..N (the remaining harness parity suites) remain via the established recipe.

Released under the MIT License.