Skip to content

Follow-up items and deferred work ​

This doc tracks deferred work and known incomplete items surfaced during the sub-project sequence on branch omnara/stateless-suspension-redesign. Items here are NOT blocking the current sub-project's acceptance gate — they are intentionally scoped out and tracked for follow-up.

For test-infrastructure-specific deferred items (atomicity stress, harness backlog, etc.), see test-infrastructure-roadmap.md.

How this doc works ​

  • Each item has a stable ID (e.g. FU-A2-01) so commit messages can reference it.
  • Items are organized by surfacing sub-project (A.1, A.2, etc.).
  • Status: open, in-progress, done.
  • Once closed, items move to the bottom of the doc with a closing-PR reference.

Known CI flakes ​

A running log of intermittent CI failures with root cause + mitigation, so a red pipeline on one of these is recognized rather than re-investigated from scratch. Recognizing one is not permission to retry it: since MR !272 no CI job retries on script_failure, and a manual retry of a genuine failure follows the disclosed-retry rule in development-workflow.md.

FU-FLAKE-RUNTIME-JS-K5-COMPLETED-STATUS-WRITE: runtime-js commit-checkpoint-truth-fixes-2.test.ts K5 completed-status-write case resolves instead of rejecting — done ​

Closed by: this MR (CI-health-3 CI-flakes bundle). Owner: CI-health-3 (routed by the orchestrator). Surfaced by: MR !297 pipeline 2912057160, job 16928257935. Seen once before on !288's branch, job 16900882183 (2026-10-02).

  • Test: runtime-js src/__tests__/commit-checkpoint-truth-fixes-2.test.ts › "Task 11 fix round 2 — runtime-js › a failed completed-status write after K5 reports completed, and the next entry finishes (never supersedes) that run".
  • Signature: AssertionError: promise resolved "{ sessionId: 'k5-status', …(11) }" instead of rejecting. The other ELIFECYCLE lines in the job are turbo cascade kills, not failures.
  • Not load: nothing else was stressing the host at the time (00:4x, 2026-10-05), so this is treated as a genuine intermittent test or product race, not CI starvation. !297 changes no runtime-js code.
  • Cause (test design, not a product bug): the test gave the run a real 300 ms lease, judged on the store's Date.now(). writeRunStatus stops renewing it right before the terminal status write, which the test fails 3× (25 + 50 ms backoff). The 'early' entry then had to reach its run-start commit inside what was left of 300 ms of wall time to be refused. When it arrived after the lapse it legitimately ended the run and won. That is the documented JS lease contract (the FU-FLAKE-JS-LEASE-BLOCKED-RENEWAL family).
  • Red-first: a 350 ms stall injected before the 'early' entry reproduces the CI assertion on the old test.
  • Fix: Date (the store's lease clock and the lease's local clock) is frozen for the test and advanced past the ttl instead of sleeping. The same stall passes.
  • Siblings (same latent race): commit-checkpoint-truth-fixes.test.ts I3 and I4 held a real 300 ms lease through 500 + 200 ms hook sleeps and asserted the intruder is refused. A 400 ms busy-wait in the hook failed both (expected 'superseded' to be 'completed'). They now hold 5 + 2 periods of lease time on the frozen clock, advancing by ttl/3 only after a renewal has landed, so a run whose renewals stopped times out instead of passing on luck.

FU-FLAKE-AGENT-SERVER-STATUS-LIVENESS-SLEEP: agent-server tests slept a fixed window for a run to finish — done ​

Closed by: this MR (CI-health-3 CI-flakes bundle). Surfaced by: !299 pipeline 2914044694, test:unit job 16942478423 (retried once under the disclosed-retry rule as 16944739774).

  • Signature: agent-server-status-liveness.test.ts › "reports isExecuting:false after a run completes and the handle is released": expected 'running' to be 'completed'.
  • Cause (test design): it waited a bare setTimeout(r, 250) for the run to complete. On a loaded runner the run had not, so getStatus honestly reported running.
  • Red-first: a 400 ms model step (a loaded host) fails the old test with the CI signature, and the fixed one passes.
  • Fix: an event wait, the existing waitForStatus helper, plus a new waitForNotExecuting for a released handle. Both bounds are failure signals, never performance budgets.
  • The family (every bare sleep in the agent-server tests, classified):
    • Fixed, sleep before a positive assumption: agent-server.test.ts ×2 and chat-routes.integ.test.ts ×2 (interrupt and abort of a "terminal" session), http-handler.test.ts (replica A's released handle), agent-server-end-frame.test.ts and integration/lifecycle.integ.test.ts (interrupt, abort and isExecuting "while running"), and integration/workspace-route.integ.test.ts ×2 (the registry cleared after completion). The two slow tools now signal when they start and run until aborted (each test ends by aborting), instead of completing on their own after 2 s of wall time.
    • Kept, with a comment: express-adapter.test.ts ×3 (each delay precedes a negative "not called" check, which a slow host can only pass vacuously, never fail falsely) and runtime-parity.integ.test.ts cleanup (a teardown pause with no assertion after it).
    • Fine as is: poll intervals inside polling loops (agent-server-stateless, runtime-parity, helpers/test-serverwaitForStatus, workspace-route polling), and the workspace-route mock provider's deliberate window delay.

FU-CI-UNIT-TURBO-STOPS-AT-FIRST-FAILURE: a turbo CI job stopped at its first failing package — done ​

Closed by: this MR (CI-health-3 CI-flakes bundle). Surfaced by: !299 pipeline 2914044694, test:unit job 16942478423: Tasks: 6 successful, 36 total, so runtime-cloudflare's tests never ran.

  • turbo's default --continue=never cancels every other task once one fails, so a red job hid whether the other packages passed (and the MR's own tests could have no CI result at all).
  • Fix: --continue=always on the multi-package turbo jobs (test:unit, test:integ:rest, test:cloudflare, test:workflows, typecheck, lint). Build and single-package jobs are unchanged.
  • Demo (6 packages, one with a deliberately failing test, --concurrency=1, ×2): with never, 0 and then 4 of 6 tasks ran to success; with always, 5 of 6 both times. Both modes exit 1.

FU-FLAKE-CFW-STOP-CONDITIONS-RUN-STATUS-READ: CFW stop-conditions tests read the run record before its terminal status — done ​

Closed by: !299 (CI-health-3). Surfaced by: !298 pipeline 2913874706, test:workflows job 16941132802 (retried once under the disclosed-retry rule as job 16941808210, green).

  • Signature: runtime-cloudflare cfw-stop-conditions.wf-noiso.test.ts › "retry.max-steps-continues-without-message (C5, no step boundary)": expected undefined to be 'max_steps' at the run record's completionReason, while the result's errorDetail.code was already framework_max_steps_exhausted.
  • Cause (test, not product): the file's waitForTurn returned once the session was terminal, which the K5 state commit writes. The terminal order on every runtime writes the run's terminal status, carrying completionReason, after K5. An assertion on the run record could read that gap.
  • Red-first: a store wrapper (testWorkflowInjection.setStateStoreWrapper) that delays the workflow's terminal updateRunStatus by 2 s fails the CI case with its exact signature, and a sibling (.fresh-budget, expected +0 to be 2). With the fix, both pass under the same delay. The demo wrapper is not committed.
  • Fix: waitForTurn also waits until the current run has left the live statuses (isLiveRunStatus), so every run-record assertion in the file reads after the run's own terminal write.

FU-FLAKE-CLIENT-TOOL-RECOVERY-REDIS-JS-SUSPEND-RACE: client-tool recovery tests resumed before the original run finished suspending — done ​

Closed by: !306.

Surfaced by: !305 job 16956379754 (a text-only MR).

  • Signature: e2e client-tool-recovery-redis-js.integ.test.ts › "orphan drains with runtime_restarted on fresh-executor resume": AgentAlreadyRunningError … (status: suspended_client_tool) from runtime-js run-start.ts (the CP-96 expected_mismatch loser), called from executorB.resume().
  • Cause (test): the test waits only for pendingClientToolCalls['orphan-tc'] to persist, then flips the session to interrupted and resumes from a fresh executor B. On runtime-js the suspension order is frames → suspension commit (the pending entry persists here) → the run's suspended_client_tool status → pauseStream, and executor A is still alive finishing it. B planned against A's running run; A's status write landed before B's commit, so B correctly lost the run-start CAS.
  • Not a product bug: in every variant A is live and still writing, so refusing B is the CP-96 gate doing its job. A genuinely crashed A holds its run only until its JS lease lapses on the store clock (documented).
  • Fix: wait for the stream's paused status (it comes last, and it is the documented "suspension committed" signal) as well as the pending entry, in this test and its two siblings with the same wait (client-tool-recovery-js, client-tool-error-code-invariant-js).
  • Red-first: a store wrapper on executor A that delays its suspended_client_tool run-status write by 300 ms fails the old test 1/1 (AgentAlreadyRunningError) and passes the new one. The wrapper is not committed.

FU-FLAKE-SKILL-FS-PROVIDER-HOOK-TIMEOUT: skill-fs filesystem-provider.test.ts setup hook timed out at 10 s — open (diagnostics in) ​

Owner: CI-health-3. Surfaced by: !304 pipeline 2915625900, test:unit job 16953822082 (2026-10-05).

  • Signature: the file fails at file level with Hook timed out in 10000ms, then TypeError: The "path" argument must be of type string … Received undefined at afterAll (rm(root, …)).
  • What the trace shows: root was still undefined in afterAll, so the beforeAll's first await, mkdtemp(join(tmpdir(), 'helix-skills-')), had not resolved after 10 s. The hook timer fired on time (10 014 ms), so the event loop was running; the filesystem call was what stalled.
  • Context:
    • The job ran on the shared local runner ("Joe Dev Server", arm64, Docker executor): this Mac's Docker VM. Its disk was at 83% (339/410 GB), its executor prepare took 63 s to pull node:24, and in the same job store-memory's collect took 42 s.
    • At about the same time, CI-health-3 was running single-file store-postgres integ and unit runs against its private Postgres container on that same VM disk (within the shared-host rules, but extra I/O).
    • vitest's default forks pool gives each file its own process, so nothing else in that process was using its libuv threadpool.
    • The working hypothesis is an I/O stall on the VM's disk, not this test's work. Unproven.
  • Frequency: 1 of the 19 failed test:unit jobs since 2026-09-26 (the first sighting).
  • Done (!306): the setup is self-diagnosing and the teardown null-safe, with no timeout raise. Each fs step is timed; if the hook never finishes, afterAll fails with the step still pending and how long (fixture setup did not finish: step "mkdtemp" still pending after 10005ms …), not with a TypeError that masks it. Demo with a never-resolving mkdtemp: the old file reproduces the CI signature exactly; the new one names the step.
  • Open: the root cause, from the next sighting's named step. On a recurrence, an unrelated-flake retry with disclosure (§A3).

FU-FLAKE-CF-MULTIRUN-150MS: fixed sleep before interrupt() lets mock responses leak into the resumed run — open (partly fixed by !306) ​

Partly fixed by: !306 (the two files below); the siblings are open.

Origin: !264 (cf-multirun-flake-fix, closed as superseded; carried forward).

  • Cause: the multi-run tests queued run 1's mock responses (text, then a slow tool call), called execute(), slept a fixed 50–150 ms and interrupted. On a slow runner run 1 had not reached the tool call yet, so the scripted responses stayed queued and the resumed run 2 consumed them. Setting the sleep to 0 reproduces it every time.
  • Fixed:
    • multi-run-streaming.integ.test.ts (4 sites): a local waitForToolStart waits for the slow tool's tool_start chunk and asserts getRemainingResponseCount() === 0, so a leak cannot pass silently.
    • ai-sdk-d1-with-js.cf.test.ts "interrupt and resume with state preservation via DO stream": it waited only for the increment commit, which lands before step 2's LLM call, so tc-abort could leak into the resume. It now also waits for tc-abort's tool_start and asserts no response remains.
  • Open: the same pattern elsewhere. A heuristic grep (a setTimeout(r|resolve, N) within 4 lines before interrupt(/stop(/abort() in packages/e2e/src/__tests__: interrupt-resume-js (15), interrupt-resume-d1-with-js (13), interrupt-resume-redis (12), redis-cas-races (11), interrupt-resume-dbos (11), interrupt-resume-subagent (9), deep-subagent-abort (7), concurrent-coordination (7), stream-ordering (6), error-recovery (6), and fewer elsewhere (!264's counts). Some are poll loops or deliberate "interrupt at an arbitrary point" races, so audit each before converting it.

FU-FLAKE-WAITFOR-NO-FINAL-RECHECK: poll helpers that give up without checking once more after the deadline — open (partly fixed by !306) ​

Partly fixed by: !306 (the shared helpers); the local copies are open.

Origin: !266 (unit-timing-flakes-fix, closed as superseded; carried forward).

  • Defect: a while (Date.now() - start < timeout) { check; sleep } loop followed by throw. On a loaded host the last sleep can overrun the deadline, and a condition that became true during it fails the wait anyway.
  • Fixed:
    • the shared helpers, one change each covering many tests: runtime-dbos integration/helpers/poll.ts (waitForCondition, also behind waitForSessionStatus), agent-server helpers/test-server.ts (waitForStatus) and helpers/wait-for-status.ts (waitForStatus, waitForNotExecuting), e2e helpers/stream-collector.ts (waitFor), helpers/timing-helpers.ts (waitForCondition) and helpers/do-subagent-bridge.ts (waitForCompletion). Each now checks, and only then breaks on the deadline, so its last check lands at or after it.
    • red-first: agent-server/src/__tests__/wait-for-status-final-check.test.ts (fake clock; the condition holds only after the deadline): 2 of 3 fail on the old helpers, 3 of 3 pass now.
    • !266's local fixes in client-tool-real-crash-recovery, companion-usage-scoping and the research-assistant-cloudflare-do smoke test.
    • Where a real causal signal exists, no poll at all: completed-marker-reaper.integ.test.ts awaits handle.result() (the suspension commit precedes it), and research-assistant concurrency.test.ts awaits execute() / retry() (their run-start commit has landed; the mock LLM answers only after a setTimeout(10) macrotask, so the concurrent call always meets the live run).
  • Open: local copies of the same loop. A structural scan (each waitFor* function's body; the statement after its while loop is a bare throw) finds, after this fix (b57ac80864), 86 helpers in 80 files: e2e 57, runtime-cloudflare 15, runtime-temporal 1, runtime-js 1, agent-server 1, and 11 in examples (opennext-cloudflare-do 5, research-assistant-cloudflare-do 2, workspaces-showcase, research-assistant-dbos, remote-agents-cloudflare-do, nextjs-redis 1 each). Outside test:unit, the dominant flake driver there is service latency, not this one-interval window, so they are left for a per-job diagnosis rather than an 80-file sweep; switching a file to a shared helper fixes it for free.

FU-FLAKE-FIXED-SLEEP-THEN-ASSERT: a fixed sleep, then an assertion on detached background work — open (partly fixed by !306) ​

Partly fixed by: !306 (run-loop-persistent); the js-agent-executor* sites are open.

Origin: !266 (carried forward).

  • Defect: a test starts fire-and-forget work (a non-blocking persistent child, a background resume), sleeps a fixed 50–80 ms and asserts on its effect. Under load the work has not landed; unlike a poll there is no timeout error, the assertion just fails (or passes on a stale value).
  • Fixed, runtime-js run-loop-persistent.test.ts: every positive site waits for the real signal (a ref registered or terminal, the child's LLM called, its abort observed, the follow-up message appended).
    • Two of !266's conversions were vacuous: the 'recur' and 'r1' re-spawn tests waited for a ref that is seeded before the run, so the wait was true at once and the completionDelivered assertion still raced the reset. Both now wait for the child's new user message, which the continue / re-spawn path appends only after the ref reset.
    • Two sites stay as documented negative windows (a regression would start a second loop or call the child's LLM later; absence needs a window): the failed-history-read resume (the failure itself is already in, inside the awaited resumeChildAgent) and the CAS-loser continuation.
  • Open: js-agent-executor-interrupt.test.ts (138–329, 377, 414, 440, 469, 503, 525, 594, 626), js-agent-executor.test.ts (973, 1976, 2071), js-agent-executor-resume.test.ts (327, 1078, 1143): a fixed 30–500 ms sleep after starting a slow tool, then abort() / interrupt() / canResume(). A tool's staged updateState is not visible in the store until its step commits, so a store poll does not work; the tool's tool_start stream chunk does (as in FU-FLAKE-CF-MULTIRUN-150MS), or an in-test gate the tool signals.

FU-FLAKE-DBOS-CLIENT-TOOL-PARITY-TIMING: DBOS client-tool submit / timeout races in the e2e parity tests — open ​

Owner: the DBOS bucket (was B8). Origin: !266 (carried forward; referenced in the CI-health table above with no entry until now).

  • client-tool-happy-parity.integ.test.ts › dbos-postgres: "exactly one submit wins" fails with AgentAlreadyRunningError … (status: active); "emits tool_end with client_tool_timeout error and continues" fails with expected 'completed' to be 'suspended_client_tool'. CI job 16758483177 (!268, passed on its then auto-retry).
  • Same family: client-tool-timeout-dbos.integ.test.ts returned suspended_client_tool instead of completed (1 of 2 isolated runs on 59d47f1e0f).
  • FU-A2-43 marked the parity test fixed, so this is a recurrence. Suspected: fixed waits around the DBOS client-tool timeout and submit. Not root-caused; it belongs with the DBOS integ work (!259's revival).

FU-FLAKE-LOCAL-LOAD-TEMPORAL-DURABLE-STREAM: local-only flakes under heavy host load — open (no owner) ​

Origin: !266 (carried forward). Seen locally only, never in CI, at load average 90–138 on 14 cores; both passed standalone:

  • Temporal gRPC Failed to connect before the deadline, only inside a full concurrent turbo run test:integration.
  • A durable-stream-contract concurrency-timing miss.

Recorded so a future local repro is not re-diagnosed from scratch.

FU-SUBAGENT-RESULT-RECORDING-RELOADS-PARENT: recording each sub-agent result re-reads the parent's whole state and history about three times — open ​

Owner: unassigned (perf; CFW / Temporal; the orchestrator routes it). Surfaced by: CI-health-3's profile of FU-FLAKE-CF-CONCURRENT-SUBAGENT-FOLD-TIMEOUT. Related: FU-CFW-PER-STEP-COST.

  • Observed cost: for every child that completes, the parent's full state and history are read about three times (recordSubSessionMessage → loadAgentState, core recordSubSessionResult, the D1 append commit), so O(children × history) D1 work per wave in production. The measurements and the profile harness are in FU-FLAKE-CF-CONCURRENT-SUBAGENT-FOLD-TIMEOUT.
  • Fix direction: pass the already-loaded parent state into recordSubSessionResult, and look up the toolCallId's existing result without scanning the whole history (an indexed lookup in each store). Spans core, runtime-cloudflare steps.ts and runtime-temporal (activities.ts, workflow.ts, workflow-dispatch.ts, its other callers).

FU-FLAKE-DO-SUPERSEDED-ALSO-COMPLETE: DO hooks.superseded-no-terminal-hook saw ['complete', 'superseded'] — done (test race fixed; product window split out) ​

Closed by: CI-health-3's flakes-5 MR. Surfaced by: main pipeline 2918408193 (6b9f649c0d), test:integ:rest, job 16973198710 (the first sighting).

  • Signature: runtime-cloudflare do-run-guards.integ.test.ts › "hooks.superseded-no-terminal-hook (DO): a superseded run fires onAgentSuperseded only": expected [ 'complete', 'superseded' ] to deeply equal [ 'superseded' ].
  • Test race (fixed): the test worker's hookLog (and toolGate) were module-level, shared by every session's DO in the workerd process, and /test/inspect returned the whole log. A background execution of an earlier test (a wake alarm, an R26 reset, or any run waiting on the shared gate) that completes during this test is counted as the superseded run's own. Forced interleave (scratch): a foreign session held at its model call, released into its gate tool after this test's reset, completes on this test's release-gate; the unscoped log then reads ["superseded","complete"]. Fix: each hook entry records { event, sessionId, runId }; /test/inspect returns only the inspected session's events plus hookRuns, and the test asserts its single entry belongs to the superseded run (never the intruder). With the fix, the same interleave gives ["superseded"].
  • Product window (split out): a supersede landing between onAgentComplete and the K5 commit makes one run fire complete then superseded (runtime-js repro: the agent's own onAgentComplete makes an intruder run-start commit). The user ruled "allow both, last wins", like R40's ['complete', 'fail']; documented and contract-tested on every runtime by FU-HOOKS-COMPLETE-THEN-SUPERSEDED-WINDOW.
  • Which one CI hit cannot be told from its log (no run ids). In this test the supersede lands before the gate is released, so the fenced step commit after the gate tool should stop the run before completion; with run ids in the log now, a recurrence names the run.

FU-HOOKS-COMPLETE-THEN-SUPERSEDED-WINDOW: a supersede between onAgentComplete and K5 fires both terminal hooks — done ​

Closed by: CI-health-3's hooks-window MR. Ruling (user): allow both, last wins.

  • Terminal hooks precede K5 by design (their customState writes land in it). A supersede in that window ends the run superseded after onAgentComplete already fired: complete then superseded, the same shape as R40's ['complete', 'fail']. The last terminal hook is authoritative.
  • Verified on every runtime with a contract test that supersedes the run from inside onAgentComplete (exactly [complete, superseded], both hooks carrying the losing run's id; that run's status superseded; the intruder current; and, where a hook can write state, a marker onAgentComplete writes is not persisted. DBOS terminal hook contexts have no updateState, so that check does not apply there): runtime-js k5-failure-hooks.test.ts; the CF DO do-run-guards.integ.test.ts (a per-session armed side effect in the test worker); CF Workflows commit-checkpoint-truth.wf-noiso.test.ts (a beforeStoreCall seam before the run's own K5); DBOS and Temporal commit-checkpoint-truth.integ.test.ts. None produced a bare superseded or the reverse order.
  • Tracing-langfuse fix: the second closing hook re-emitted the root span as unknown-agent with no input (the run's pending data was dropped by the first emission), overwriting the trace's name; this also affected R40's complete → fail. Emitted runs now keep their identity in a bounded map for a later terminal hook. Best effort, same process only: when the second hook runs on another worker (Temporal, DBOS, CF Workflows) it still emits unknown-agent.
  • Docs: the hook-order table and onAgentSuperseded in the hooks guide and reference, concepts.md §Superseded runs, execution-flow, the tracing guide, the runtime guides and references, hook-type TSDoc and CLAUDE.md (user-approved).

FU-FLAKE-CF-CONCURRENT-SUBAGENT-FOLD-TIMEOUT: CF D1 concurrent sub-agent fold test timed out at 60 s — open ​

Owner: CI-health-3 (root cause after the DBOS measurement). Surfaced by: main pipeline 2913857752 (17ebc280ce), test:integ:rest job 16940995440 (amd64 k8s), not retried.

  • Signature: runtime-cloudflare concurrent-subagent-fold.integ.test.ts › "concurrent sub-agent customState fold (D1 integration) › composes all 16 parallel afterSubAgent folds (append)": Test timed out in 60000ms.
  • The first such failure in about 12 failed test:integ:rest jobs. A 16-way parallel fold on D1 running out of 60 s may be a real serialisation or contention problem in the fold path, not just a slow runner: root-cause before calling it load.
  • CI-health-3 analysis (2026-10-06), from job history: not a hang. In the failing job (amd64 k8s, whose trace carries "Insufficient memory" scheduling notes) the file took 334 s; (modify) passed at 56 979 ms and (append) timed out at 60 316 ms. In 8 passing test:integ:rest jobs (Intel DigitalOcean runners, 2026-10-05/06) the 6-test file took 106–257 s: (modify) 17.8–46.6 s, (append) 15.8–42.5 s, the half-failing wave 14–38 s. The tests are intrinsically slow, vary about 3×, and run with a thin margin.
  • A second defect: the timed-out test's workflow kept running. Its mock step.do retries threw D1 state getCurrentRun failed … TypeError - fetch failed while the next describe ("removal …", with its own runtime context) ran, so a timed-out run keeps driving D1 into later tests.
  • Profile (local, idle host; a counting proxy on the D1 binding, scratch, not committed): each test makes 405–420 D1 calls whose latencies sum to 57–83 s inside a 6–8 s wall clock: the 16-way wave's calls queue behind one another through Miniflare's proxy, about 150 ms each. Under CI load every round trip slows and the wave stretches linearly. Per child, the PARENT's full state and history are read about three times: steps.ts recordSubSessionMessage → loadAgentState (state, every message, refs, interrupt flag); core recordSubSessionResult (state, refs, and every message again, to look for one toolCallId); and the D1 store's append commit (state and messages a third time). That is O(children × history) per wave in production too.
  • Leak fixed (!306): every test bumps a module-level counter, and a mock step (do, waitForEvent) created for an earlier test throws a NonRetryableError, so a timed-out run stops at its next step. Demo (scratch): test A starts a slow 4-child wave and returns without awaiting it; while test B runs, D1 calls bound to A's session went from 106 to 7. The 7 that remain are unexplained (they are not through the mock's do or waitForEvent; likely the stale run's own failure path, not verified).
  • Open (product): the per-child re-reads, filed as FU-SUBAGENT-RESULT-RECORDING-RELOADS-PARENT. The timeout itself stays at 60 s.

FU-FLAKE-CFW-ENGINE-REWIND-SNAPSHOT-UNSTABLE: rewind-resync-cfw-engine.wf.test.ts "retry same-turn" refresh read is snapshot_unstable — open (T1 path) ​

Owner: CI-health-3 (T1 priority, orchestrator). Surfaced by: MR !291 pipeline 2910439346, job 16919362444.

  • Test: rewind stream resync (RM-60, R34) — cfw-engine-d1-do › retry same-turn: a page refreshed mid the rewinding run renders the rewound transcript (ai 6.0.281).
  • Signature: SnapshotUnavailableError: … no stable snapshot read in 3 passes — the fenced snapshot read (readSnapshotView) saw the session's status, current run or latest truncate change in every pass, or a pass timed out.
  • A/B (!291, same machine, sequential, --retry=0, no pipeline running): head d73f5e177d 5/5 and main 5bb1eaaf2b 5/5 — not reproduced on either side; no product change between the 26/26-green 876627b81c and the failing head. Likely CI-load timing in the 3-pass fence window; RM-60 rewind-refresh is a T1 path, so it is worth a root cause rather than a wider window.
  • CI timing (CI-health-3): the failing case took 33.4 s.
    • fetchSnapshotWithRetry makes 4 fetches with 1.75 s of total backoff, each over 3 passes capped at 2 s.
    • So every pass hit its 2 s cap. A torn fence fails a pass in milliseconds, and the run was parked at the model gate, so its fence could not move.
    • That leaves two candidates: (a) slow D1 / DO reads under load, or (b) a read that hangs.
  • Not reproduced on macOS (CI-health-3):
    • A lane-level wrapper timed every snapshot read (alert over 250 ms or still pending at 1.5 s). It is uncommitted; the patch is in the reliability dossier.
    • Runs: the 12 "refreshed mid the rewinding run" cases × 5, and the full test:cfw-engine suite × 6 (79/79 each). Both ran at --retry=0 beside forced example rebuilds, at load average up to 111.
    • The wrapper fired zero times and nothing failed. That argues against (a), at least on macOS.
  • Diagnostics (!293): snapshot_unstable now records each pass's outcome on the error (snapshotPasses): a timeout with the read in flight, or a fence_changed with the field and its before/after values. The message carries a short summary, so the next CI occurrence names its cause.
  • Next:
    • Reproduce in a linux/arm64 node:24 container (closer to the runner) with the same instrumentation, ≥10× under capped load.
    • A product change to the pass budget is on hold until evidence shows slow-but-consistent reads; that change would be the user's decision.

FU-FLAKE-CONF-HARNESS-BRIDGE-NETWORK-LOST: conformance harness host-driver.integ.test.ts client-tool cell loses the Miniflare bridge — ✅ fixed (plan 0c MR-1b) ​

Fixed (plan 0c MR-1b): CI-health-3 found the race: workers reach the bridge through an external service binding, so workerd pools keep-alive sockets to it and expires idle ones at about 5 s, while Node's bridge server closed them at its keepAliveTimeout (5 s) plus a ~1 s buffer. The failing job ran 41.7 s on a starved runner, where workerd's timer fired late, Node closed first, and the non-idempotent hook.fired POST went out on the dead socket. The bridge now sets server.keepAliveTimeout = 0 (packages/e2e/src/harness/cf-do/bridge-server.ts), so it never closes an idle socket first; there are no client retries. Same mechanism as cloudflare/workers-sdk#14850. Regression test: packages/e2e/src/harness/cf-do/__tests__/bridge-keepalive.test.ts (one raw keep-alive socket, a POST, a 6.5 s idle, a second POST on the same socket); it fails with keepAliveTimeout = 5_000 and passes with 0. Pending: CI-health-3's loaded before/after proof (at least 15 runs); its numbers are not recorded yet.

Owner: CI-health-3. Surfaced by: MR !291 pipeline 2910439346, conformance:verify job 16919362440.

  • Test: host driver live on 'cf-do' › a client tool is answered by the responder (the product auto-send POST); the run completes; the client rejection is only recorded.
  • Signature: bridge hook.fired: fetch failed: Network connection lost. — the stream failed (['failed'], expected ['paused','ended']) and endCell recorded a harness failure.
  • A/B (!291, vitest.conformance-harness.config.ts): head d73f5e177d 5/5 and main 5bb1eaaf2b 5/5 — not reproduced; a harness-side Miniflare bridge connection drop under CI load.

FU-FLAKE-CF-E2E-TEST-WORKER-INVALIDATED: e2e test:cloudflare Durable Object invalidated by "test-worker.ts changed" mid-run — ✅ closed (CI health 2: benign log, never a failure) ​

Closed (CI health 2): the message is vitest-pool-workers' own: after each test file it calls vi.resetModules(), so the next file re-imports test-worker.ts with new class identities, and a Durable Object instance still alive from the earlier file (with isolatedStorage: false they outlive the file) fails the wrapper's constructor-identity check when anything touches it again. Nothing writes the file. In the 164-pipeline sweep every job that logged it failed on another, now-fixed cause (CP-52, the fixed-sleep interrupts, the D1 parity timeout). The 22 per-test { retry: 2 } options that cited it are removed, and the e2e test:cloudflare suite ran 15 of 15 green at --retry=0 under CI-like load while logging it 376 times. If a test ever fails on it, give that test DOs that no other file touches.

Owner: B8 (CI health). Not caused by any one MR.

@helix-agents/e2e's test:cloudflare sometimes fails with workerd reporting packages/e2e/src/test-worker.ts changed, invalidating this Durable Object. Please retry the DurableObjectStub#fetch() call. (broken.inputGateBroken) partway through the run. Nothing in the failing MRs edits that file, so something in the job changes a file in the worker's module graph while it runs. The first suspect is a sibling test or build step writing under packages/e2e/src/: compare FU-E2E-CROSS-SERVICE-BUNDLES-IN-SRC, where esbuild bundles are left under src/. Seen on unrelated MRs, including catalogue/docs-only !273 (16758514869), !266 (16759542448), !268 (16757428811), !272 (16760323616) and !275 (16764206430). To close: find the writer (for example, by watching packages/e2e/src for writes during a CI-equivalent run), keep generated output out of the worker's module graph, and pin that with a check.

FU-FLAKE-CF-D1-DO-SEQUENCE-PRESERVE: ai-sdk-d1-with-js.cf.test.ts interrupt/resume stream never reaches ended in 5s — ✅ fixed (CI health 2) ​

Fixed (CI health 2): a wall-clock bound, not a stream that never ends. The test interrupted Run 1 after a fixed 150ms sleep, so on a loaded runner Run 1 was interrupted before it consumed its scripted text and Run 2 ran the slow 2s tool instead; waitForStreamStatus then gave up after 5s. The test now interrupts once the tc-abort tool_start chunk is on the stream (waitForStreamChunk), and waitForStreamStatus has no deadline (the test timeout catches a stream that never ends).

Owner: B8 (CI health). Not caused by any one MR.

"should preserve chunk sequence across interrupt/resume with DO" fails with Timeout waiting for stream 'do-sequence-preserve-…' to reach status 'ended' after 5000ms. Seen on unrelated MRs: !268 (16757428811), !269 (16758270746), !272 (16760323616) and !275 (16764371786). It is either a wall-clock bound too tight for a loaded runner (the same family as FU-FLAKE-CF-DO-STREAM-WALLCLOCK-BOUND) or a real resumed stream that is never ended. That must be decided before the bound is widened. If the stream truly never ends, it is a product finding for the conformance owner.

FU-FLAKE-CF-D1-GETCHUNKS-AFTER-RESUME: ai-sdk-d1-with-js.cf.test.ts "should filter chunks correctly with getChunksFromSequence after resume" — ✅ fixed (CI health 2) ​

Fixed (CI health 2): confirmed cause. With the sleep at 0ms (a loaded runner) the original test fails every time with the CI signature ('Filter Run 1 contentFilter Run 2 cont…' not to contain 'Run 1'): Run 2 consumed Run 1's scripted text. All four sleep-then-interrupt sites in the file now wait for the slow tool's tool_start chunk. 6 failures in the sweep (main 16850833009, !269, !279, !285, …).

Owner: orchestrator (re-brief). Not caused by any one MR.

The e2e test (packages/e2e/src/__tests__/ai-sdk-d1-with-js.cf.test.ts:544, assertion near :613) failed once in CI on the B2 MR (16764777038) and passed on retry (16764870705). Locally it passed 25 of 25 on both B2 and main, including under CPU load. Suspected cause: the test's real-clock setTimeout(150) before the interrupt (around :561-563) — under CI load the run can reach a different point by the time the interrupt lands, so the chunks after the resume sequence differ. To close: replace the sleep with a condition wait on the observed stream state (for example, wait until the first tool_start chunk is visible) before interrupting, and pin it with a repeated run.

FU-FLAKE-WAIT-STATUS-TRANSITION: wait-for-status-transition.test.ts real-clock race — ✅ fixed ​

Status: done · Closed by: branch chore/followups-flakes-and-cf-docs-demos

The suite drove status transitions via real setTimeout and polled against the real clock with a 500ms deadline. On a saturated CI worker the 100ms status-change timer can be delayed past the deadline, so a must-succeed scenario observed false ("expected false to be true") — hit on !223 test:unit. Fixed by converting to vi.useFakeTimers() + advanceTimersByTimeAsync so the outcome is a pure function of the controlled clock (the production helper's deadline behavior is correct and unchanged). Reproduced deterministically by blocking the event loop past the deadline; verified 0/30 failures under heavy event-loop load.

FU-FLAKE-REMOTE-DUP-START-RACE: duplicate-/start tests raced the child run to completion — ✅ fixed ​

Status: done · Closed by: branch fix-remote-dup-start-race (!271) · Jobs:test:integ:e2e 2/4

Symptom: remote-agent-flow.integ.test.ts › "duplicate start rejection › duplicate start throws ALREADY_RUNNING for active session" failed at its third /start with Remote agent start failed (409): {"error":"Session idem-session already completed","code":"ALREADY_COMPLETED"} instead of RemoteAgentAlreadyRunningError. 10 of the last 40 pipelines hit it, on main (jobs 16755525951, 16747461115) and on 7 MRs (16756192309, 16756128346, 16756105472, 16756403531, 16756029835, 16756279865, 16755705432, 16755700151, 16750110303). The earliest occurrence is job 16721832769 (!254, 2026-09-24), before any change on the path.

Root cause (test defect, not a product regression): the test assumed the child session was still active when the duplicate /start calls arrived, but its un-gated MockLLMAdapter queue (text, then __finish__, commented "Slow") completes in roughly 30 ms. The third /start, two HTTP round trips later, sometimes arrived after completion, and AgentServer.startAgent then correctly answered ALREADY_COMPLETED. Between the last green main (8a6114093c) and the first failures, nothing changed in agent-server, runtime-js, store-memory, MockLLMAdapter, the HTTP transport, the remote-agent helper or the test file, and vitest's shard-2/4 file set stayed the same. The race dates from f6cf57aa91 (2026-08-03), which added the third /start. The onset coincides with CI load changes (the conformance jobs added by !254 run on the same runner pool), not with a code change.

Fix: the test now holds the child's first model call on a ScriptedModel hold gate, waits for whenHeld, asserts that the session is active, and makes both duplicate calls. It then releases the gate, waits for completed with nothing left scripted, and asserts that a later duplicate gets ALREADY_COMPLETED. Two sibling tests made the same assumption and got the same kind of fix:

  • remote-agent-matrix.integ.test.ts › "ALREADY_RUNNING → attach" holds the JS, Temporal and DBOS producers' next generateStep.
  • agent-server http-flow.integ.test.ts › "POST /start returns 409 for already-running session" uses ScriptedModel hold.

A 300 ms (flow and agent-server) or 5 s (matrix) delay before the duplicate call reproduced the failure on the old tests (flow 10/10, agent-server 10/10, matrix 8/9 cells). With the same delay, the fixed tests passed every time. After the fix, all three tests passed 30/30 on a normal machine and 10/10 under CPU load (one yes worker per core).

FU-FLAKE-WORKERD-INTROSPECT-TEARDOWN: vitest-pool-workers workflow-introspection teardown race — open (benign log; no retry) ​

Status: open (no safe in-code fix) · Jobs: test:workflows (full-execution.workflow.test.ts), test:cloudflare / test:workflows (lifecycle-hooks-parity.wf-noiso.test.ts)

Symptom: intermittent job failure citing jsg.Error: Instance dispose (workerd/api/actor-state.c++: broken.outputGateBroken) or Network connection lost, emitted when the workflow Durable Object actor is disposed at test/isolate teardown.

Root cause (common): these tests start a REAL Workflow instance and read its terminal status via cloudflare:test's introspectWorkflowInstance(...). Disposing the workflow Durable Object actor races with the actor's output gate, emitting Instance dispose / broken.outputGateBroken from workerd C++ (io-context) — NOT a JS-catchable rejection, so it cannot be suppressed from test code. Locally the tests pass and the log is benign; on a loaded CI runner the race nondeterministically escalates to a job failure (Network connection lost). The two affected files dispose the actor on different schedules:

  • full-execution.workflow.test.ts runs under isolatedStorage: true (vitest.workflows.config.ts), where the introspection handle MUST be disposed per-test (await using) for the storage-isolation teardown assertion to pass — so the actor is disposed at the END OF EACH TEST. Removing the await using is NOT an option: it breaks that assertion (verified — 2 tests drop to failing).
  • lifecycle-hooks-parity.wf-noiso.test.ts runs under isolatedStorage: false (vitest.workflows-subagent.config.ts; the .wf-noiso suffix exists specifically to sidestep the isolatedStorage teardown assertion) and does NOT await using the handle — the actor is instead disposed at ISOLATE TEARDOWN.

Either way the disposal-vs-output-gate race is the same, and no test-code change closes it.

Investigated (branch chore/followups-flakes-and-cf-docs-demos): reproduced the teardown exception locally for full-execution.workflow.test.ts (fires on every run as benign stderr — tests still pass 7/7 at the vitest level); confirmed removing its await using breaks the isolatedStorage teardown assertion.

Update (CI health 2): CI no longer retries test:workflows or test:cloudflare on a test failure, and no failure in the 164-pipeline sweep came from this race: the Instance dispose lines are logged on every run and are harmless, and the Network connection lost failures in these jobs were the loopback keep-alive race (FU-FLAKE-POOL-WORKERS-STACKED-STORAGE-LOST, fixed). If this race ever fails a test again, it is a red job to investigate, not a retry. The original mitigation was:

Mitigation (historical): retry the failed job. Consistent with the documented precedent for pool-workers / CI-capacity flakes ("the root cause was CI infrastructure capacity, not an in-code defect — there is no code change that would close it", see the post-!199 merge-train note below). A real fix would require an upstream @cloudflare/vitest-pool-workers improvement to introspection-handle disposal (currently pinned ^0.9.0); revisit on the next pool-workers bump.

FU-FLAKE-DBOS-PER-CALL-HOOKS-TURN1: per-call-hooks.integ.test.ts persistent SEND test reads the hook log before turn 1's onAgentStart is recorded — fixed by MR !259 (not yet on main) ​

Status: fixed by MR !259 (B8, dbos-integ-ci-unpause, GL #108), which is open and not yet on main. !259 waits on the turn-1 assistant reply (waitForAssistantText(…, 'turn-1-reply')) instead of the persisted user message, which removes the race this entry names. Move this entry to Done items when !259 merges. · Class: test-side race, reproduced on origin/main @ 3560d6c72a (3 failures in 11 runs of the file, private Postgres, a heavily loaded laptop) and on the conformance drain branch (2 in 4), which does not touch runtime-dbos. · Test: packages/runtime-dbos/src/__tests__/integration/per-call-hooks.integ.test.ts, "persistent SEND path: only the first execute()-time hooks fire on subsequent turns".

Signature: AssertionError: turn-1 onAgentStart fires the first execute()-time hooks: expected [] to include 'first' (line ~517).

Cause (suspected): the test's waitForCondition gate waits only for the session to be active with the turn-1 user message persisted. Nothing orders the onAgentStart hook step's recording before that state becomes visible, so under load the snapshot of the hook log can be taken before the hook runs.

Fix direction: make the wait condition also require the firstonAgentStart record (or poll the log with a bound), rather than widening a timeout. If the hook genuinely can fire after the turn is persisted, confirm that ordering is intended before changing the test.

Owner: B8 (MR !259).

FU-FLAKE-MINIFLARE-SYNC-PROXY-DESYNC: Miniflare Node binding proxy desynchronises under CPU contention — ✅ fixed (MR !272) ​

Status: fixed by MR !272 (the Miniflare 5 bump; see "Resolution" at the end of this entry) · Class: test-harness race in a third-party dependency (Miniflare's Node-side binding proxy), triggered by CI CPU contention. Not a D1 store defect. · Jobs: test:integ:e2e 3/4 on MR !267 (failed 16755572339, passed on automatic retry 16755916593); same signature earlier in test:integ:rest jobs 16746282265 (MR !257, runtime-cloudflare steps.integ.test.ts), 14622985680 (MR !222) and 14429769579 (MR !172).

Symptom: partway through a Node-pool Miniflare suite (here packages/e2e/src/__tests__/ai-sdk-d1-with-js.integ.test.ts), every later D1 call fails with D1 state loadState/createSession/getMessages failed ... AssertionError ... (message?.id === id), thrown from miniflare/src/plugins/core/proxy/fetch-sync.ts:138 under D1StateStore → db.prepare(...). The first test that trips it can instead show sessionId: undefined: its POST's execute() threw inside handleFreshTurnPath, which returns a 200 rejectedOnlyResponse (ai-sdk/src/handler/handle-chat-stream.ts:1030, :2155-2163) with no X-Session-Id header. Async calls (d1.exec in clearD1Tables) keep working, so each beforeEach still passes.

Root cause: D1Database#prepare is synchronous, so Miniflare's Node proxy serves it through SynchronousFetcher (miniflare@3.20250718.3/dist/src/index.js:13047-13114): the caller posts {id} to a helper worker thread, Atomics.waits on a shared 0/1 flag, then takes one message with receiveMessageOnPort and asserts its id. The helper posts the reply, then Atomics.store(flag, 1), then Atomics.notify (:13040-13041) as two separate steps. If the caller reaches Atomics.wait only after the helper has stored 1 (the caller was paused, e.g. GC or descheduled, after posting), the wait returns at once. If the helper is then preempted between the store and the notify, the caller has time to post the next request and wait again, and the late notify for the old request wakes it. The port is empty, the assertion fires, and the next reply is left queued. From then on every call reads the previous call's reply, so the proxy is off by one for the rest of that Miniflare instance's life, which is why one flake fails the rest of the file. The same code is in every Miniflare we install (3.20250718.3, 4.20251011.0 and 4.20260205.0). Upstream fixed it in cloudflare/workers-sdk#15552 (merged 2026-09-16, which uses per-request generations instead of the 0/1 flag). It ships in miniflare 5.20260916.0-alpha and wrangler 4.134.0; we pin miniflare@^3 in e2e, runtime-cloudflare, store-cloudflare and memory-cloudflare. Production is unaffected: in a Worker, D1StateStore gets a real workerd binding, and SynchronousFetcher only exists in Miniflare's Node proxy (getD1Database / getBindings / getPlatformProxy).

Reproduced (MR !267 DoD pass): a --require preload that pauses the main thread for 300 ms between posting and Atomics.wait, and pauses the helper 600 ms between Atomics.store and Atomics.notify, once, against a real Miniflare 3 D1 binding. Result: call 1 ok, then calls 2-10 all fail with (message?.id === id). Without the preload, 10/10 pass. The suite itself passed 3/3 locally unmodified. Not a regression: it occurs in May 2026 MR pipelines (before the pnpm migration), and no recent main commit touches the proxy path.

Mitigation: the .integration_base automatic retry (retry: max 1, when: script_failure, .gitlab-ci.yml:165-171) absorbs it. The retry started 2 s after the failure. Fix options: move the Node-pool Miniflare suites to a Miniflare with the upstream fix once a stable release carries it (or a pnpm patch of SynchronousFetcher backporting the generation counter); or in the harness, rebuild the Miniflare instance when a call throws this assertion instead of letting the rest of the file cascade.

Resolution (MR !272): the Node-pool suites (e2e, runtime-cloudflare, store-cloudflare, memory-cloudflare) moved from miniflare@^3 to miniflare@5.20260925.0-alpha, the version wrangler@4.141.0 ships. Its SynchronousFetcher carries the upstream generation fix (Atomics.store publishes (id + 1) | 0 and the caller waits for its own generation), so a late notify can no longer wake the wrong call. The .integration_base script_failure retry that used to absorb this is gone (same MR). Evidence: ai-sdk-d1-with-js.integ.test.ts 10/10 green on the bump (7 on an idle machine, 3 under a 10-process yes CPU stress), no (message?.id === id) in any run. The deterministic --require preload repro was not re-run (it targets the Miniflare 3 code path, which is no longer installed).


FU-FLAKE-DO-STREAM-CONTRACT-TRIMMED-AFTER-ATTACH: CU.trimmed-after-attach reads [] on a loaded runner — ✅ fixed (CI health 2) ​

Fixed (CI health 2): the case reads the three retained chunks, then reads until the reader's own caughtUp resolves (no 300ms window). The whole contract lost its wall-clock windows; see FU-FLAKE-CF-DO-STREAM-WALLCLOCK-BOUND.

Owner: orchestrator (routing). Not caused by any one MR.

runtime-cloudflare's test:cloudflare fails do-stream-manager-contract.cf.test.ts > "CU.trimmed-after-attach: backlog chunks cleaned up after attach are not waited for" with AssertionError: expected [] to deeply equal [ 1, 2, 3 ] (the case is in core/src/testing/stream-manager-contract.ts, shared by every stream manager). The case gives each it.next() a 300 ms window (settledWithin(..., 300)) and stops reading at the first one that does not settle, so [] means the FIRST backlog read did not arrive in 300 ms. On the DO, every read is a runInDurableObject proxy call. Seen on !287 at 650491a381: job 16913998801 failed, and the job's own retry: when: script_failure passed it as 16914066439. That runner was loaded: the same run's G1.concurrent-writers took 4.6 s. It did not reproduce locally on main bd26dc9b57 or on !287's head: the file alone 30/30 on each, and the whole runtime-cloudflare test:cloudflare suite 10/10 on each (alternating). main's test:cloudflare also fails on its first attempt in several recent pipelines on other DOStreamManager contract cases (e.g. 16882767276). To close: replace the wall-clock window with a deterministic end of the backlog (the reader's catchUpHead / caughtUp, which the next assertion already uses), keeping the case's point: chunks cleaned up after attach are never waited for.

FU-E2E-DBOS-LOAD-SENSITIVITY: DBOS e2e integration cells fail under parallel turbo load — open ​

Update (CI health 2): reproducing it (the three named files' DBOS cells on a fresh Postgres, under CPU load) failed 3 of 4 runs, but on a different cause: concurrent first DBOS.launch()s racing DBOS's schema migration (FU-E2E-DBOS-LAUNCH-RACE, fixed: 15 of 15 fresh-database runs green after). The "Workflow … has been cancelled" cells did not reproduce in those runs; they were only ever seen under a full local turbo run test:integration (every package at once). Left open until reproduced that way.

Owner: orchestrator (routing). Not caused by any one MR. Seen locally, not in CI.

Under turbo run test:integration (all packages at once, --concurrency=2), a few DBOS cells in @helix-agents/e2e fail, a different set on each run, and pass when the same files run alone. On !287's head (650491a381) the run failed five, all passing package-direct (the e2e DBOS files 272/272) and in CI:

  • cross-runtime-resume-matrix.integ.test.ts (dbos + postgres + redis): "B18/A2 C7 (terminal route)" and "B2 C7 (interrupted route, completion)" (expected 'failed' to be 'completed'), and "two concurrent resume(from_checkpoint) on a suspended session — exactly one wins" (Workflow … has been cancelled);
  • session-model-branching.integ.test.ts ('dbos-postgres'): "C11: branching with a checkpointId from ANOTHER session is rejected typed" (Workflow … has been cancelled);
  • remote-agent-matrix.integ.test.ts (producer agent-server/dbos): "late-join snapshot + fromSequence patches == terminal loadState" (Matcher did not succeed in 30000ms).

Two failures report a DBOS workflow cancelled mid-test, which suggests one cell's cleanup (or the shared DBOS executor's shutdown) cancelling another cell's workflow when the machine is slow. To close: reproduce under load (for example the e2e DBOS files with a CPU hog alongside), find who cancels the workflow, and fix the isolation or the bound.

Conformance catalogue leads (not yet catalogued) ​

Leads the conformance owner has not verified yet; each becomes a finding (or is withdrawn) in a later catalogue batch.

(None open. FU-CONF-STOPWHEN-PARITY was catalogued as DI-29 in batch 6 and moved to Done items.)

Conformance follow-ups from !288 (owner-approved designs) ​

FU-CONF-RM60-HOST-CLIENT-CELLS: host / client rewind-resync cells for RM-60 — open ​

!288 fixed RM-60 at the executor tier (one rollback stream_resync per from_checkpoint rewind on every runtime) and dropped its unwired host / client drafts. The owner-approved design:

  • retry and resume-from-checkpoint are separate host capabilities;
  • the outcome is a bounded poll of the snapshot's terminal status;
  • the js-chat-host retry is a declared gap;
  • a client cell may run the host op and then reload.

RM-60 stays open until these cells exist.

FU-CONF-CP74-DBOS-HOOKS: DBOS hook seams for the CP-74 dual-window cell — open ​

The owner approved the direction, on conditions:

  • the seams are awaited only inside recorded step bodies, so replay never re-runs them;
  • they are @internal and not exported from the package index;
  • they must pass the W1 / W2 window proofs;
  • they ship in their own MR with the full DoD.

CP-74 stays open until then.

Conformance follow-ups from MR-1a (agent bundling) ​

FU-CONF-PARTIAL-BATCH-T0-WINDOWS: submit.partial-batch-resume|dbos-postgres fails under load because its windows are anchored on t0 — open ​

The cell's settle windows are measured from the cell's start (t0), not from an observed event. On a loaded machine the first DBOS step can take longer than about 7 s, so the batch is not yet pending when a window closes and the cell fails with a timing assertion, not a product verdict. It surfaced as a load flake around the MR-1a work.

  • Fix: anchor each window on "batch pending observed" (the pending client-tool batch seen through the observation channel), the way the other partial-batch cells wait for their gate, instead of on t0. Keep the bounds as ceilings, not as the sequencing.
  • Owner: the conformance owner. It lands after MR-1a (do not fold it into the pure-refactor MR).

Conformance follow-ups from plan 0c MR-1b (the CF Workflows harness) ​

FU-CONF-CFW-PRODUCT-WAITS-PRE-FLIP: decide the CF Workflows product waits before the lane runs — open ​

The cfw-workflows-d1 executor lane is built and declared not-yet-run. Before MR-2 flips it to running, the plan 0c spec's §2 rule ("product waits versus named settles") must be applied to every product wait the lane measured (entry gate in spec §9, MR-2). The MR-1b self-tests (packages/e2e/src/harness/cfw/__tests__/lane.integ.test.ts) measured these on Miniflare:

  • stop.mid-tool-honors-abort costs 8.4-8.8 s per cell, against 0.7-2.3 s for the other three live-smoke cells. Review attributed it to handle.result()'s poll backoff, a product wait. That is not confirmed: it needs a long-run measurement of where the time goes.
  • No settle wait occurred (0 of 10 two-turn sessions in each of 3 runs), so the prior-instance settle-wait has no duration sample. Stop→resume and retry settle waits were not measured.

The decision is the owner's, taken with the user, and is never a per-cell bound or a per-backend tweak (spec §11, "Product waits against bounded settles"). Record the figures and the decision in the spec's MR-2 as-built section.

Owner: the conformance owner (plan 0c MR-2, before T7).

FU-CONF-CFW-BODY-TERMINATE-UNLEDGERED: body-side terminates are not ledgered on the CF Workflows lane — open (MR-2 entry gate) ​

The worker's body-side binding wrapper (observedBinding, packages/e2e/src/harness/cfw/worker/cfw-deps.ts) reports each successful create (wf.created) but returns the inner get handles unwrapped. So the product's unknown-owner terminate (workflowBindingOwnerProbe's terminate, which does binding.get(id) then terminate(), called by the escape hatch in finishCfwEntryStream; both runtime-cloudflare/src/run-entry.ts) is never ledgered as terminatedBy: 'product' and its session is never fenced. That deviates from plan 0c spec §3.3 / §3.6 (body terminates "reported through the wrapper"). The terminated instance's late traffic is still loud, but it fails the NEXT cell (as closed-session late traffic), so the failure is blamed on the wrong cell.

To close (before MR-2 flips the lane): wrap the body get handles in observedBinding and report a successful terminate() to the bridge (a new session-scoped, epoch-fenced kind), which the Node side folds into the ledger as terminatedBy: 'product' so endCell fences it.

Owner: the conformance owner (plan 0c MR-2 entry gate).

FU-CONF-CFW-POOL-ROW-RETIREMENT: retire the cfw-workflows-d1-pool row and the legacy workerd helper — open ​

Plan 0c MR-1b repointed cfw-workflows-d1 to the Node-driven conformance lane and kept the legacy workerd adapter (packages/e2e/src/harness/setup-helpers/cfw-workflows-d1.ts, setupCfwWorkflowsD1) as a separate cfw-workflows-d1-pool row with CFW_POOL_CAPS (no branching), because live consumers still use it: runtime-cloudflare's lifecycle-hooks-parity.wf-noiso.test.ts, checkpoint-truth-matrix.wf-noiso.test.ts and cfw-d1-backend-smoke.wf-noiso.test.ts, and e2e's harness-smoke.cf.test.ts.

To close (MR-2 T9): for each consumer, check whether a cfw lane cell now covers the same contract; retire what is covered and keep the rest with a stated reason here. Retiring the last consumer removes the helper, its e2e/package.json export and the cfw-workflows-d1-pool descriptor. The pool-workers test:cfw-engine lane and runtime-cloudflare's own *.wf-noiso suites that do not use the helper are out of scope.

Owner: the conformance owner (plan 0c MR-2, T9).

CF Workflows per-step cost ​

FU-CFW-PER-STEP-COST: !288 makes a CF Workflows LLM step about 1.25× slower than before — open ​

Measured on the T9 workload (cfw-terminal-truth.wf-noiso, 50 steps), on the same machine and harness:

main 5bd30e20da!288
wall time, idle117 ms/step144 ms/step (1.22×)
wall time, under load216 ms/step276 ms/step (1.28×)
step.do per LLM step67
D1 round-trips per LLM step~55~44
stream-DO calls per LLM step~6~9

D1 got cheaper. The extra time is the committed-count step.do, plus three stream-DO calls: an /info head read for the commit cursor, an extra /init, and the step_committed marker /write. Real CF Workflows users pay this.

Consolidation to do:

  • fold committed-count into the step commit's step.do;
  • piggyback the head read and the step_committed write on calls the step already makes;
  • drop the redundant /init.

Guard it with a step / round-trip count test.

UI-chunk schema conformance (DI-20) follow-ups ​

Surfaced by the DI-20 fix (docs/superpowers/specs/2026-09-26-ui-chunk-schema-conformance-design.md). Finding IDs are catalogued by the conformance owner.

FU-DI20-01: @helix-agents/llm-vercel never reports LLM step boundaries (RM-48) — open ​

vercel-adapter.ts maps chunks inside streamText({ onChunk }), and ai only calls onChunk for text/reasoning/source/tool/raw parts, never start-step / finish-step (the adapter iterates fullStream but ignores those parts). So the adapter never fires onStepStart / onStepEnd: no core step_start / step_end chunks, no per-step usage / cachedTokens / finishReason on the Helix stream, and includeStepEvents emits nothing with the production adapter. The opennext-cloudflare-do example sets includeStepEvents: true but gets no step events (its comment now says so).

FU-DI20-02: DBOS emits step_start but never step_end (RM-49) — open ​

runtime-dbos/src/steps/call-llm.ts writes step_start (without stepId) directly; there is no step_end anywhere in runtime-dbos. packages/e2e/src/__tests__/ai-sdk-dbos.integ.test.ts pins the gap with an it.fails finish-step non-vacuity test that turns red once DBOS emits step_end.

FU-DI20-03: live vs reload shape of a denied approval differs — open (pre-existing) ​

A live denial settles the tool part via a synthetic tool-output-available (deny string), while the replay path (replay-events.ts, output-denied state) emits tool-output-denied, so the part state differs between the live stream and a reload. The replay output-denied state has no production producer today.

FU-DI20-04: snapshot converter emits a top-level filename on file parts — open (pre-existing) ​

helix-to-aisdk-converter.ts puts filename on reloaded file parts, while the live file chunk now carries it in providerMetadata.helix.filename (DI-20). No core file chunk producer exists today.


LLM attempt retries and the error chunk (E5) follow-ups ​

E5 made every runtime write an LLM step's error chunk only for a failure that ends the run (core holdLLMErrorReport). The Temporal attempt fence and the DBOS exhausted-step settle came with it. These two items are what is left.

FU-E5-LLM-ERROR-STEP-NOT-RETRIED: llm-vercel error steps are never retried on Temporal / DBOS — open ​

@helix-agents/llm-vercel never throws: it returns every provider failure as an error step (shouldStop: true), including a transient, retryable: true one (a 429, a 5xx, a dropped connection). Temporal's runLLMStep and DBOS's runCallLLM retry only a THROWN generateStep (the activity retry / the step retry), so on those runtimes a transient provider error fails the run at once. CF Workflows: done (CF-D1 MR 1b, !287). Its LLM step rethrows a retryable or unclassified error step so step.do retries it under llmStepRetry, and its onError goes through holdLLMErrorReport. The item stays open for Temporal and DBOS (owners: CF-D1 MR 2 / MR 3). A fix there rethrows a retryable error step inside the activity / step, so the durable retry runs it again with no error chunk (E5 already holds the report). It must also keep the exhausted retry's last classified detail (DI-24).

FU-E5-EXHAUSTED-DRAFT-NOT-DISCARDED: a failed attempt's draft is not marked discarded when the retry runs out — open ​

When every attempt of an LLM step fails (Temporal's activity retry, DBOS's step retry), the run fails with the last attempt's error. The attempts' streamed text is never marked discarded (step_discarded, as discardInterruptedStep does for an interrupt). A connected client keeps the dead draft on screen next to the error, and a refresh does not show it. The terminal failure path should emit step_discarded for the in-flight step before the error chunk on Temporal (finalizeTerminalRun after a runLLMStep failure) and DBOS (the turn settle in runAgentLoopOneTurn).

CF-D1 MR 1b (!287) follow-ups ​

!287 made runtime-js, the CF DO and CF Workflows end a failed run with its real, classified error. These items are the Temporal and DBOS halves, which the CF-D1 MR 2 (Temporal) and MR 3 (DBOS) rounds own.

FU-CFD1-DI03-TEMPORAL-DBOS: a hard abort on Temporal / DBOS carries no framework_cancelled — open ​

Owners: CF-D1 MR 2 (Temporal), CF-D1 MR 3 (DBOS).

DI-03's record names runtime-js, the CF DO and CF Workflows, which now record framework_cancelled (category framework, retryable: false) for a hard abort. Neither runtime-temporal nor runtime-dbos writes that code anywhere (grep framework_cancelled finds nothing outside tests), so an abort there is indistinguishable from any other uncoded failure. To close: record the typed cancellation at each runtime's abort seam (the handle's abort() and the workflow side that observes it), and prove it with an executor-tier cell. That cell needs a hard-abort driver primitive, which the conformance harness does not have yet (it has only the soft stop()); ask the conformance owner for one.

FU-CFD1-DI22-TEMPORAL-DBOS: a classified cause on Temporal / DBOS terminal seams — open ​

Owners: CF-D1 MR 2 (Temporal), CF-D1 MR 3 (DBOS).

resolveErrorDetail() now walks the cause chain (DI-22) and both runtimes call it in their handle's catch, but their workflow-side terminal seams persist only the detail that reaches them. Temporal's top-level workflow catch falls back to the wrapping ActivityFailure (DI-04), and an ordinary LLM error reaches both as { message } (DI-24). To close: carry the classified detail across the activity / step boundary so the walk has something to find, then prove it with the discipline.error-step-detail cells (ledgered under DI-24 today) plus a wrapped-cause case.

FU-CFD1-CP71-DURABLE-RUNTIMES: a throw inside the agent loop on DBOS / CF Workflows — open ​

Owners: CF-D1 MR 3 (DBOS); CF Workflows when its executor lane lands (plan 0c).

CP-71 (a throw inside the loop, for example a throwing systemPrompt function, strands the session) is fixed on runtime-js and the CF DO and proven by lifecycle.loop-throw-fails-run on js-*. The record never checked the durable runtimes; the cell now does:

  • Temporal passes it on temporal-*, on main as well: it never stranded the session.
  • DBOS fails it on dbos-postgres: execute() / result() throws the prompt's error instead of resolving a failed result, the persisted session never turns terminal, and a follow-up execute() throws the same way. Catalogued as CP-94 (the prompt is resolved in execute(), before any run).
  • CF Workflows has an executor lane since plan 0c MR-1b (cfw-workflows-d1), declared not-yet-run; the cell runs there once MR-2 flips the lane to running.

CF-D1 MR 1c (!291) follow-ups ​

!291 gave runtime-js, the CF DO and CF Workflows one stop decision (core resolveStopDecision), a typed framework_max_steps_exhausted failure for a maxSteps cut-off, retry() that continues such a run, and RunMetadata.completionReason for stop_when / max_steps stops. Verbatim string tool results (RM-32) reached all five runtimes. These items are the Temporal and DBOS halves, a store finding, a JS / CFW divergence and an open product question.

FU-CFD1C-DI29-DI25-TEMPORAL: Temporal ignores stopWhen and completes a maxSteps cut-off — open ​

Owner: CF-D1 MR 2 (Temporal).

The Temporal workflow still calls shouldStopExecution with maxSteps alone (DI-29) and ends a cut-off run of an agent without outputSchema completed with no output (DI-25). To close: decide stop / continue through core resolveStopDecision with the agent's stopWhen (C1, C2: a throwing predicate fails the run once, never retried; the decision must be replay-deterministic, so evaluate it in an activity or record it); fail a max_steps stop with createMaxStepsExhaustedError through the terminal path so the detail reaches the result, the session, the stream and onAgentFail (C3); write completionReason: 'stop_when' | 'max_steps' on the run record (C4) so planRetry's C5 rule continues a cut-off run with a fresh budget (the retry entry must accept an empty plan.input). Note that Temporal child workflows are also uncapped (FU-TEMPORAL-CHILD-MAX-STEPS). Proof: flip the pinned divergence in e2e/src/__tests__/stop-conditions-parity.integ.test.ts (assertADivergent / assertBDivergent → assertA / assertB) and clear the ledgered discipline.stop-when-honoured, discipline.max-steps-exhausted and lifecycle.retry-after-max-steps cells on temporal-*.

FU-CFD1C-DI29-DI25-DBOS: DBOS ignores stopWhen and completes a maxSteps cut-off — open ​

Owner: CF-D1 MR 3 (DBOS).

The same two gaps on DBOS (workflows/shared.ts passes maxSteps alone). To close: wire resolveStopDecision (C1, C2) inside a recorded step, fail a max_steps stop typed (C3), and write completionReason (C4) so planRetry's C5 rule applies to DBOS retry(). DBOS's own completionReason mapping (DI-27, 'failed' for every text answer) must not write 'max_steps' on a completed run any more. Proof: flip the pinned divergence in e2e/src/__tests__/stop-conditions-parity-dbos.integ.test.ts and clear the ledgered cells on dbos-postgres.

FU-DELETE-SESSION-ORPHANS-CHILDREN: deleteSession(parent) leaves its child sessions behind — open ​

Owner: the store-integrity bucket (CF-D2), or a later core MR. Finding: RM-64 (reserved; catalogued in owner batch 10). Surfaced by: CF-D1 MR 1c, Task 8 (C7).

deleteSession(parentSessionId) deletes only the parent's own rows on memory, Redis, Postgres and D1 (store-cloudflare/src/d1-state.ts: WHERE session_id = ?); DO-SQLite does not support deleting sessions. A child session id is derived from the parent session id and the tool-call id (core generateSubSessionId, <parent>-sub-<callId>, on runtime-js and the DO; <parent>__sub__<callId> on CF Workflows), so a reused parent session id (execute → deleteSession → execute) whose model repeats a tool-call id makes the new child reuse the old child's session, history included, and the child's model sees the deleted conversation's turns. Seen on the CF Workflows real engine (the C7 case: the second child ran a new turn on the preserved <sid>__sub__dup-call session); in principle on every runtime, since each derives the child session id from the parent session id and the call id. Real providers mint unique tool-call ids, so it needs a deterministic model or a provider that numbers calls per conversation. To close: cascade deleteSession to the session's sub-sessions (its sub_session_refs, recursively) on every store, or key child session ids on something a delete retires. Prove it with a store contract case and a cross-runtime e2e case.

FU-RESUME-AT-CAP-JS-CFW-DIVERGENCE: a resume at or past the step cap runs one step on JS, none on CFW — open ​

Owner: CF-D1 MR 2 / the reliability orchestrator. Surfaced by: the MR 1c final review (C5). Pre-existing.

A run that suspends on its cap step (for example a client tool called on step maxSteps) records the cap as the resume's entryStepCount. On runtime-js (and the CF DO, which runs its executor) the resume runs ONE more step: the step iterator decides stop / continue only after a step (resolveStopDecision(stepResult, stepCount, …)), so the model is called once and the run is cut off after it. On CF Workflows the loop guard stepCount < maxSteps (workflow.ts) is false at entry, so the resume calls the model zero times and fails max_steps with no step commit. Both end failed with framework_max_steps_exhausted, and retry() continues either with the full budget (C5), but the number of model calls a resume at the cap gets differs. Seen in runtime-js/src/__tests__/stop-conditions.test.ts (the resumed run's extra step) and runtime-cloudflare/src/__tests__/cfw-stop-conditions.wf-noiso.test.ts ("no step boundary": zero model calls). To close: pick one rule (a resume is a continuation of the same turn's budget, so most likely "no step past the cap", or "a resume always gets at least one step") and apply it on every runtime, Temporal and DBOS included, with a conformance cell.

FU-MAX-STEPS-COMPANION-CONTINUATION: should a max_steps-failed companion be continuable like a completed one? — open ​

Owner: user / reliability orchestrator decision. Surfaced by: the MR 1c final review.

MR 1c (U1) made a maxSteps cut-off of an agent without outputSchema end failed. For a persistent companion that means a re-consult no longer continues it: companion__sendMessage throws not active (status: failed), and companion__spawnAgent re-spawns it fresh, losing its memory (core companion-tool-dispatch.ts continues only a completed child; the CF DO's persistent-companion-tools.ts mirrors it). Before MR 1c the companion completed and was continuable. The behaviour is documented (upgrade guide §1, the sub-agents critic-loop section, the JS / DO / CFW runtime pages: use stopWhen or a larger maxSteps). The open question: should a companion whose run record says completionReason: 'max_steps' be continued (a new turn with a fresh budget, as retry() does for a root run) rather than re-spawned? If yes, change both continuation sites (core dispatcher and the DO's tools) and every runtime's continuation primitive, with a cross-runtime e2e case.

FU-RUNSTART-FINISH-COMPLETION-REASON: a run start that ends a stranded max_steps run writes no completionReason — open ​

Owner: CF-D1 MR 2 / the reliability orchestrator. Surfaced by: the MR 1c final review (C5). Narrow crash window.

When a run's K5 terminal commit landed but its run-status write did not (a crash between the two), the next entry ends that stranded run inside its own run-start commit (StartRun.finishCurrent, core RunStartFinish in core/src/store/run-start.ts, R27/R35). RunStartFinish carries status and error only, so the run record gets no completionReason. finishCommittedRun derives 'max_steps' from the session's errorDetail code on recovery (!291); this path does not. So if a run failed framework_max_steps_exhausted and crashed in that window, a later retry() reads run.completionReason as unset and C5 does not apply: with a step boundary the retry needs a message and restores the boundary's step count; with no step boundary a resume's retry throws validation_error. To close: either let planRetry also read the session's errorDetail.code when the failed run's record has no reason, or add a completionReason field to RunStartFinish / finishCurrent and write it on every store (memory, Redis, Postgres, D1, DO-SQLite) with a contract case.

Related, low risk: a max_steps retry (C5) of a resumed run with no trigger and no drain continues with input: [], so the retried run's first model call sees exactly the kept prefix. CF Workflows (run-entry.ts, buildTrailingHealMessages over the kept prefix) and DBOS heal trailing unpaired tool calls in that prefix before the call; runtime-js and Temporal do not heal it on this path. The kept prefix ends at a step or drain commit, so it is normally paired (RM-29); only a prefix that somehow ends in an unpaired call would reach the model unhealed. To close: run the C5 continuation heal on the empty-input retry path on runtime-js and Temporal too, with a unit case each.

T1 refresh path: stream readers lose a chunk after a resume cleanup (T1-A) ​

FU-T1A-REDIS-REBUILD-DROPS-LIVE-CHUNK: a Redis resumable reader drops a live chunk during its truncation rebuild — done ​

Closed by: this MR (CI-health-3 T1-A). Found by CP-96 on !295's CI.

  • Symptom: e2e refresh-c1 js-redis › "refresh while the client tool is pending" rendered "Thanks, it is" for the snapshot's "Thanks, teal it is" (pipeline 2911894663, job 16927336679): client B lost one continuation text_delta. It is about 1 in 10 to 20 under load on main.
  • Cause (RedisStreamManager.createResumableReader):
    1. The resume of runtime-js (and the CF DO, which runs it), DBOS and CF Workflows calls a run-scoped cleanupToStep (Temporal never calls it), which sets __truncated_at_step even when it removes nothing (G4). The continuation run then streams with the marker set.
    2. With the marker set, each drained next() rebuilds its buffer from LRANGE on the main connection.
    3. It then dropped every pending pub/sub chunk missing from that snapshot as "cleaned up". A chunk written just after the LRANGE reaches pendingChunks over the subscriber connection, so it was dropped. The next chunk went straight to the waiting consumer, and the dropped one was never re-read.
  • Fix: the rebuild reads the writer's counter and the chunk list in one MULTI. A pending chunk above that counter was written after the snapshot, so it is kept.
  • Proof: new contract case G4.rebuild-keeps-live-chunks, with an optional holdNextRebuildRead force.
    • Redis holds the rebuild read's reply until a newer chunk reaches the subscriber.
    • On main it fails with the exact signature: expected 'c5' to be 'c4'.
    • It passes on every manager after the fix.
  • Audit of the other managers:
    • Memory and the in-process DO read one authoritative source with no separate pending queue, so they are not affected.
    • runtime-cloudflare's DOStreamManagerClient (the reader behind createCloudflareChatHandler, e.g. the opennext example) reads the DO agent's own /sse, served by the in-process DO manager. It has no truncation handling and no client-side prune, so it is not affected either.
    • The binding-side DO reader had a different loss in the same family: FU-T1A-BINDING-PRUNE-IGNORES-FLOOR.

FU-T1A-BINDING-PRUNE-IGNORES-FLOOR: a binding-side DO reader drops buffered chunks a floored cleanup kept — done ​

Closed by: this MR (CI-health-3 T1-A audit).

  • Cause: on a truncated event, DurableObjectStreamManager's readers pruned every unread buffered chunk with step > truncatedAtStep. The DO's cleanupToStep only removes such chunks above its minSequence floor. The resume of runtime-js (and the CF DO, which runs it), DBOS and CF Workflows passes the previous run's startSequence as that floor, and their interrupt the interrupted run's when its record reads. Unread chunks of earlier turns with a higher per-turn step number were dropped, although the DO kept them.
  • Fix: StreamTruncatedEvent carries minSequence. The DO broadcasts it on WebSocket and SSE, and the reader prunes only chunks above it.
  • Proof: new contract case G4.floor-kept-backlog-survives. On main it fails as expected [ 't2s1' ] to deeply equal [ 't1s2', 't1s3', 't2s1' ], and it passes on every manager after the fix.

FU-REDIS-REBUILD-LTRIM-PENDING: a Redis rebuild silently drops a pending chunk that maxChunks trimmed — open ​

Owner: unassigned (reliability orchestrator to route). Pre-existing; out of T1-A's scope.

  • The G4 rebuild drops a pending chunk with a sequence at or below the snapshot's counter that the snapshot does not hold, as "cleaned up". A chunk that WRITE_CHUNK_SCRIPT's LTRIM removed (retention, maxChunks) meets the same test, so a slow reader loses it silently instead of getting the retention-floor signal (TRUNCATED, the floor in firstSequence).
  • Fix direction: in the rebuild's MULTI, also read firstSequence. A pending chunk at or below it was trimmed, not cleaned up, and should surface the retention gap the way the attach path does.

FU-G4-NOOP-CLEANUP-MARKER: a resume cleanup that removes nothing still sets the truncation marker — open ​

Owner: unassigned (G4 contract; reliability orchestrator to route).

  • Every resume's cleanup sets the marker, including no-op ones. On Redis, each drained reader next() for the rest of the continuation run then pays an O(n) MULTI/LRANGE rebuild. The in-process DO pays a SELECT, and the DO binding broadcasts a truncated event.
  • That is a cost, not a correctness issue, after T1-A.
  • Options:
    • Record the marker only when the cleanup removed a chunk. CLEANUP_TO_STEP_SCRIPT knows that atomically.
    • Clear the marker when the resumed run claims the stream.
  • Either one changes the G4 contract ("a no-op cleanup is still observable"), so it needs a contract decision and matching changes in every manager and in the contract suite.

T1 refresh path: a CFW SEND answered before its run claimed the stream (T1-B) ​

FU-T1B-CFW-ENTRY-BEFORE-CLAIM: a CFW execute() returned before its run claimed its stream — done ​

Closed by: this MR (CI-health-3 T1-B). Found by CP-96 on !295's CI.

  • Symptom: cfw-engine refresh-attempt-retry › "live attempt-switch re-anchor" (ai 6.0.49): no resync trigger re-anchored within 30000ms of the hold (data parts: ["data-resume-rejected"]) (pipeline 2911894663, job 16927336694). A load flake on main since S5.
  • Cause: CloudflareAgentExecutor.entryStarted returned on the run row alone for a new or active session, right after the run-start commit. The instance claims and makes the stream live after that commit. With a claim lag of about 1 s on a loaded runner, handleChatStream's streamHandle got handle.stream() → null on all four attempts. It answered the SEND with data-resume-rejected / stream_unavailable, and the run went on with no listener.
  • Fix: for a run that owns its stream, the entry has started once the stream holds a chunk stamped with this run at or after its start sequence (its run_started; Temporal's rule), or the run left running. The stream's owner and active status are not enough: the binding-side DO's /info reads a never-created stream as active, a claim does not create it (only the entry's writer /init before run_started does), and a manager may implement no claimStream. A child writing into its parent's stream (isForeignStream, R41) keeps the run-record check. The existing fail-safe still bounds the wait.
  • Proof:
    • Unit (runtime-cloudflare cfw-entry-stream-live.test.ts):
      • red on main for new and active sessions;
      • a DO-like stream manager whose claim does not create the stream: red on a claim-and-active rule;
      • a manager with no claimStream: that rule waited for the fail-safe;
      • failed-before-claim and foreign-stream controls.
    • e2e (send-claim-lag-cfw-engine.wf.test.ts): the harness's beforeClaim / beforeCreateWriter hooks hold the entry at its claim, or after it and before the writer's /init, until the SEND has answered.
      • On main both holds failed for all three client profiles, with the CI's data-resume-rejected.
      • With a claim-and-active rule, the claim hold passed but the writer hold still failed.
      • With the run_started rule, all six pass.

FU-DBOS-ENTRY-FALLBACK-BEFORE-CLAIM: a DBOS entry wait that times out can return before the claim — open ​

Owner: DBOS bucket (CF-D1 MR 3). Found by the T1-B cross-runtime audit.

  • On its normal path DBOS's execute() / resume() / retry() wait for the workflow's RUN_START_EVENT. The workflow publishes it only after the entry step committed, claimed the stream and made it live, so the handle's stream is live.
  • The edge: when no event arrives within one getEvent wait (RUN_START_EVENT_TIMEOUT_SEC), or the workflow ends first, awaitEntryRunStart's committedRun() (lifecycle/run-start-event.ts) returns a committed, still running run without checking its stream. An entry step stalled that long between its commit and its claim would hand handleChatStream a stream that is not live yet, answering the SEND with stream_unavailable.
  • Fix direction: CFW's T1-B rule. For a run that owns its stream, committedRun() returns only once the stream holds a chunk stamped with the run (its run_started), or the run left running.

FU-CFW-ENTRY-FAILSAFE-SILENT-HANDLE: the CFW entry fail-safe returns a handle for a run that never started its stream — open ​

Owner: unassigned (reliability orchestrator to route). Pre-existing, now more reachable after T1-B.

  • At entryOutcomeFailSafeMs the executor runs abandonRunStart. A 'committed' verdict returns the handle, even when the run's entry never wrote its run_started, for example when it stalled after its commit.
  • The caller then reads a stream that does not exist (stream_unavailable), with no typed signal.
  • Fix direction: on 'committed' without run_started, return a typed outcome (or let the handle's stream() wait for the run), instead of a handle indistinguishable from a started one.

Completion strategies and preserved thinking (MR 1) follow-ups ​

MR 1 of the completion-strategies spec (2026-10-04-completion-strategies-and-preserved-thinking-design.md) wired capability-driven completion strategies, the early refusal, the prompt-prefix check, ordered parts and the fail-closed skills / memory paths into core, llm-vercel, runtime-js, the CF DO and CF Workflows. These items are the Temporal and DBOS halves, the risks MR 1 could not close, and items for the conformance owner.

FU-COMPLETION-TEMPORAL: Temporal does not resolve completion strategies — open ​

Owner: completion strategies MR 2 (Temporal).

runLLMStep (runtime-temporal/src/activities.ts) passes LEGACY_RESOLUTION to prepareForcedCompletionCall, so every forced call is the legacy named-tool request, and forced completion on Opus 5.5 / Sonnet 5.5 / Fable 5.1 / Mythos 5.1 fails (a 400 per attempt, then framework_forced_completion_exhausted). The workflow classifies a forced call's error with the package-local R1 shim (forced-error-r1.ts: every error step is a recoverable miss, because the activity → workflow JSON boundary loses the error's class, retryable and code), and isTerminalLLMStepError keeps the !isForcedCall exemption (no error chunk). To close: resolve capabilities per call inside runLLMStep (resolveCompletionCapabilities), carry ErrorDetail and the applied strategy across RunLLMStepResult, delete the R1 shim and use resolveForcedCallTransition, route the native_output protocol classification and native finishWith through phase 2, early refusal in the startRunCommit activity (checkCompletionSupport), the prefix check and stampPromptFingerprint in runLLMStep, suppress text_delta on native_output, set completionStrategy / capabilityProfile on hooks and the attempt_started chunk, rename forcedToolMissing (it now also carries framework_completion_strategy_unavailable). Every changed workflow-side branch goes behind wf.patched('helix-completion-strategy-v1') with a runReplayHistory test over a pre-change history (spec §7). Proof: the both-profiles cells of e2e/src/__tests__/forced-completion-parity.integ.test.ts on temporal-*.

FU-COMPLETION-DBOS: DBOS does not resolve completion strategies — open ​

Owner: completion strategies MR 3 (DBOS).

The DBOS gaps match Temporal's (LEGACY_RESOLUTION in workflows/shared.ts, the R1 shim in workflows/forced-error-r1.ts, the forced exemption in steps/call-llm.ts, no refusal, no prefix check), plus the spec §6 / §5.2 prerequisites: serialize.ts keeps only temperature (and a non-existent maxTokens) of LLMConfig, so providerOptions (thinking, effort), maxOutputTokens, headers and topP never reach the model; the systemPrompt function is rendered once with empty state at serialization (render it per step, then the prefix check applies); the system prompt lacks the completion instruction and workspace fragment parity; the directive must move from the workflow body (no adapter) into callLLMStep; early refusal at registerAgent and execute(). Recovery tests must replay old-shape recorded step results through the new bodies.

FU-COMPLETION-DBOS-SKILLS-RECOVERY-STRAND: DBOS keeps a throwing skills provider non-fatal until it can fail closed in a step — open ​

Owner: completion strategies MR 3. Surfaced by: MR 1 Task 13 review; ruling R24 (whole-branch review A1).

Core resolveSkillsCatalog fails closed (framework_skills_unavailable, retryable), and so do runtime-js, the CF DO, CF Workflows and Temporal. DBOS resolves the catalog in the workflow body, not in a step, so failing closed there would reject the body outside any step-retry or terminal path and could strand a recovering run active; and DBOS has no history-bound thinking path in MR 1, so failing closed only adds a failure mode. MR 1 therefore keeps the pre-0.50 non-fatal behaviour on DBOS: both workflow bodies call resolveDbosSkillsCatalog (runtime-dbos/src/workflows/standard-workflow.ts, used by standardAgentWorkflow and persistentAgentWorkflow), which logs a warn (errorId: 'dbos.skills.catalog_unavailable') and runs the turn with an empty catalog (skills-catalog-nonfatal.test.ts). To close (MR 3): resolve the catalog inside a checkpointed @DBOS.step (also FU-SKILL-DBOS-CATALOG-IN-WORKFLOW-BODY), fail closed there with framework_skills_unavailable, and route the failure through finalizeFailure; then delete resolveDbosSkillsCatalog.

FU-COMPLETION-BEDROCK-VERTEX-NATIVE-OUTPUT: native output through Bedrock / Vertex providers is unverified — open (risk) ​

Owner: llm-vercel. Surfaced by: MR 1 final review.

A 5.5-generation model reached through @ai-sdk/amazon-bedrock or @ai-sdk/google-vertex (Anthropic) resolves to anthropic-5.5 → native_output; so do AI Gateway and OpenRouter ids. For every Anthropic-family model the adapter sets providerOptions.anthropic.structuredOutputMode: 'outputFormat'; the extra provider-specific key is added only for @ai-sdk/anthropic-backed models with a custom provider name (its providerOptionsName getter, which Vertex's Anthropic models share). Whether @ai-sdk/amazon-bedrock (and the gateway / OpenRouter providers) read anthropic.structuredOutputMode at all, and whether they send Output.object as native structured output rather than a forced json tool (a 400 on 5.5), is not verified; none of those packages is installed in the repo. Vertex's Anthropic models run @ai-sdk/anthropic's messages model, so they most likely honour the key, but that is unverified too. Documented as a limitation in docs/llm/vercel.md. To close: a wire-level test with each provider and a captured fetch, or a live probe; if a provider forces a tool, map its own structured-output option or resolve such models to unresolved with a clear reason.

FU-COMPLETION-LIVE-NATIVE-CONTINUATION: live proof that a synthetic completion call replays under binding enforcement — done ​

Owner: e2e. Surfaced by: MR 1 final review.

Case 5 of packages/e2e/src/__tests__/completion-strategies-live-anthropic.integ.test.ts continues case 3's session (claude-sonnet-5-5, thinking, prefix_mismatch_behavior: 'error') with a new user turn after a forced completion the model answered with JSON text, and asserts the continuation is accepted (200) with the synthetic tool_use + tool_result in its history and the byte-prefix property intact across the whole session. Its first live run (effort max, $0.24) never reached that path: the model answered every forced call with a real __finish__ call. Ruling R22 moved the case to effort high, with up to three attempts, and a SKIP when none reaches the path. The R22 run (2026-10-05, $0.0263) reached it on attempt 1: the native JSON-text answer was replayed as the synthetic call in turn 2, both turn-2 requests returned 200 under enforcement, and the byte-prefix chain held across all six requests.

FU-COMPLETION-OUTPUT-CONFIG-FORMAT: send output_config.format once @ai-sdk/anthropic supports it — open ​

Owner: llm-vercel.

@ai-sdk/anthropic@3.0.38 (the peer floor) sends native structured output as the beta output_format field, which Sonnet 5.5 accepts (gate G1, 2026-10-04). When a provider release sends the GA output_config.format, raise the peer floor, drop any beta header dependence, re-run the wire-level boundary test (forced-completion-anthropic-boundary.integ.test.ts) and the key-gated live suite. The adapter has no runtime provider-version check (spec §3.3 envisaged one); add it then if older versions must be refused.

FU-COMPLETION-RM35-PARTS: a __finish__ tool-call finish drops the response's parts (RM-35) — open ​

Owner: the RM-35 owner. Related: RM-35 (catalogue).

llm-vercel maps any step whose tool calls include __finish__ to a structured_output result, and core persists no assistant row for it, so that response's ordered parts (and signed thinking) are not kept. Under native_output the model may also answer by calling __finish__ itself (gate G4.5), which takes this path. The next turn's history lacks that response entirely rather than replaying it, so preserved-thinking continuity is lost for it (not a 400). When RM-35's fix persists the assistant row, carry parts (and helixPrefix) on it and share the persistence code with the native-output normalisation in planStepProcessing (spec §15).

The MR !301 Greptile review re-raised this as its G2 (a signed __finish__ response is not persisted). Parked by the R26 brief: it is RM-35 (open catalogue finding, another owner), unchanged by MR !301.

FU-COMPLETION-DO-CLIENT-TYPED-CODES: DO frontend clients drop a refusal's typed code — open ​

Owner: runtime-cloudflare. Pre-existing for every non-409 entry failure; surfaced by MR 1 Task 15.

The DO answers a completion-support refusal with HTTP 400 { error, code }, but runtime-cloudflare/src/do-clients.ts / do/do-clients.ts (execute, resume, submitToolResult) and do/do-executor.ts turn any non-409 failure into a plain Error, so a caller cannot read framework_completion_strategy_unavailable without parsing the message. To close: parse the body with PeerErrorBodySchema and throw a typed DOPeerError / HelixError carrying the code, with a test per client.

FU-COMPLETION-CHAT-SEND-TYPED-REFUSAL: a chat send's completion-support refusal carries no code — open ​

Owner: ai-sdk. Surfaced by the MR 1 final re-review.

agent-server's /start, /resume and submit routes answer a completion-support refusal with a typed 400 { error, code }, as the DO does. A chat send (POST /chat, ai-sdk handleChatStream) instead turns the executor's refusal into a 200 SSE data-resume-rejected that carries only the message (handle-chat-stream.ts), so a useHelixChat caller cannot tell framework_completion_strategy_unavailable from any other rejection. The docs list only the typed routes, so they are accurate. To close: carry the refusal's code on the rejection (as parseHelixChatError reads for a failed stream), with host- and client-tier tests.

FU-COMPLETION-DO-SUBMIT-REFUSAL-CHILD-PATHS: the DO submit refusal's sub-agent paths are untested; a refused root can still stall — open (partly done) ​

Owner: runtime-cloudflare. Surfaced by the R24 re-review.

The DO /submit-tool-result completion-support gate runs in the root's local resolver and in a sub-agent DO's forwarded-submit shortcut, and a root relays an owning sub-agent DO's typed 400. Only the root's local path is tested (completion-strategies-do-entry.cf.test.ts); the sub-agent gate and the relay are not. In the shortcut the gate runs before the unknown-call check, so an unknown call on a refusable child answers 400 instead of 404. Separately, when the ROOT agent becomes refusable but the pending call belongs to a child DO, the forwarded submit is accepted, the child finishes, and the root's wake refuses and leaves the root suspended. To close: test both sub-agent paths, order the gate after the unknown-call check, and fail a suspended root whose wake is refused (as the active branch does).

Done (R26 G4): forwardToOwner now runs the same completion-support gate on the ROOT's agent and adapter before forwarding (typed 400, nothing written, nothing forwarded), so a refusable root no longer lets a child finish into a refused wake. Tested on real DO instances (ForwardSubmitAgentServer in packages/e2e/src/test-worker.ts, cases in completion-strategies-do-entry.cf.test.ts): a refusable root refuses before forwarding (child stays pending, no rows, no model call); a refusable child refuses the forwarded submit in its shortcut and the root relays the 400 verbatim; a control proves the seeded forward lands on the child. The cells seed the root's ownership entry: on the DO runtime a remote sub-agent that suspends on a client tool fails fast at the parent (RemoteSubAgentClientToolUnsupportedError, core/src/errors/agent-errors.ts, GitLab #107) and a child's registerPending writes ownership into its own DO store, so no natural flow reaches forwardToOwner today.

Not done, by decision: reordering the shortcut gate after the unknown-call check. In the shortcut a call with no pending entry is the submit-before-register race: the resolver writes a pre-submit stub and answers accepted (runtime-js/src/client-tool-resolver.ts, the rootSessionId branch of submitOnce); unknown_tool_call is only returned for a missing session, which the shortcut's isSubAgent check already excludes. Moving the gate after the resolver would write a stub on a refusable child, and the 404 it was meant to restore is unreachable there.

Remaining: fail a suspended (paused) root whose wake is refused, as the active branch does.

FU-COMPLETION-COMPANION-RESUME-REFUSAL-REF: a refused companion resume leaves the parent's ref running — open ​

Owner: runtime-js / core companion dispatch. Pre-existing for any resumeChildAgent throw; surfaced by MR 1 Task 14.

When the early refusal rejects a persistent companion's resume, the parent's SubSessionRef for it stays running. Also, on runtime-js a refused ephemeral sub-agent emits subagent_start before its refusal. To close: mark the ref failed (typed) on any child-start throw, and emit subagent_start only after the child start succeeded, with a unit case each.

FU-COMPLETION-CFW-REPLAY-REFUSAL: CF Workflows re-runs the completion-support check on replay — done ​

Owner: runtime-cloudflare. Surfaced by: MR 1 Task 20 review.

entryRefusal (runtime-cloudflare/src/workflow.ts) runs checkCompletionSupport in the workflow body outside a step.do, so it runs again on every replay. A deploy that changes how an outputSchema agent's model resolves (a capability-table change, a removed or changed capabilities override) can fail an instance whose run already committed. Documented (CF Workflows runtime page, migration guide). To close: record the refusal decision for the entry in a step.do (so a replay reuses it) or skip the check once the instance's run-start commit is recorded.

Done (R26 G3): the check runs inside step.do('<agentType>-completion-support') at the top of the body's main try, returning { ok: true } | { ok: false, error: { message, code } } as data (no retries); the typed non-retryable HelixError is rebuilt outside the step and ends through the unchanged pre-commit refusal path. A replay reuses the recorded decision. Proven by two cases in runtime-cloudflare/src/__tests__/workflows/completion-strategies.workflow.test.ts (a recorded { ok: true } runs to completion although the live adapter would refuse; a recorded refusal fails typed although the live adapter would admit). An instance recorded before this step existed checks fresh on replay (the old behaviour).

FU-COMPLETION-DEMOTE-NO-K5: the DO's failed demote writes no K5 checkpoint — open ​

Owner: runtime-cloudflare (DO wake recovery). Pre-existing; surfaced by MR 1 Task 15 review.

demoteStrandedExecution(..., 'failed', ...) (used for auto_resume_exhausted and, since MR 1, for a stranded run whose agent is now refused) ends the session failed without the K5 commitState terminal commit and writes errorDetail after the run status, against the terminal order (K5 → run status → stream). Also, callers of ensureExecutionContinues other than the wake path skip the interrupt-flag check. The recovery test covers only the reset path (run already completed). To close: route the demote through the K5 terminal commit and add the failed-path recovery case.

FU-COMPLETION-CONFORMANCE-WORKER-CAPABILITIES: the conformance worker's ProxyModel has no capabilities — open ​

Owner: the conformance owner (this session does not edit packages/e2e/src/conformance/**).

packages/e2e/src/harness/cf-do/worker/conformance-worker.ts builds new ProxyModel(new BridgeClient(...), c) without a capability profile, so a cf-do conformance cell always runs the legacy profile and cannot exercise native_output, the prefix check or the early refusal on the DO. The CF Workflows harness added by conformance plan 0c MR-1b (packages/e2e/src/harness/cfw/worker/cfw-deps.ts, cfw-conformance-worker.ts) builds its ProxyModel the same way, so a cfw cell is legacy too. The MR 1 e2e harness (ProxyModel, the bridge protocol) already forwards resolveCapabilities / checkOutputFormat and carries outputFormat, minOutputTokens, parts and providerOptions; the worker needs a way to select the profile per cell. Then: forced-completion scenarios for both profiles (spec §11, "conformance").

FU-COMPLETION-DI28-CITATIONS: DI-28 cites lines MR 1 moved — open ​

Owner: the conformance owner (see the catalogue citation-pinning rule).

DI-28's fix evidence (packages/e2e/src/conformance/catalogue/discipline/DI-28.ts:20, mirrored in docs/conformance/findings.md:240) cites e2e/src/__tests__/forced-completion-parity.integ.test.ts:410 and forced-completion-parity.cf.test.ts:240, which MR 1 Task 24 reshuffled. Its contrast citations docs/guide/agents.md:294-296 and docs/guide/finishing-agents.md:243-246 also moved with MR 1's docs. Re-cite with the MR number, pinned to the merge commit.

FU-COMPLETION-WF-NOISO-LONG-PATH: some *.wf-noiso files cannot start from a long worktree path — open (local only) ​

Owner: runtime-cloudflare test infra.

Nine long-named *.wf-noiso.test.ts files fail to start when the repository lives under a long path (for example a ~/.superset/worktrees/... worktree): the workerd temporary directory path exceeds the platform limit. CI checkouts are short and unaffected. To close: shorten the per-file temp dir (hash the file name) in scripts/run-wf-noiso.mjs or the pool config.

FU-COMPLETION-FAILED-RUN-PHASE-STATE: a failed run's still-pending forced phase differs across runtimes — open ​

Owner: runtime-js / runtime-cloudflare. Surfaced by: MR 1 whole-branch review A4 (parked by ruling R24).

When a run fails for a reason other than the forced call's own outcome while its forced-completion phase is still pending (a skills or memory failure on a forced call, a throwing generateStep), CF Workflows fails the phase in its K5 terminal commit and writes the exhausted diagnostic, while runtime-js and the CF DO leave the phase pending for the same failures. retry() is unaffected on both (it restores the phase from the step commit, never from K5). The divergence is stated in docs/internals/concepts.md (Completion strategies). To close: align the runtime-js / DO onFailed K5 with CF Workflows, or document a CF-Workflows-only contract, with a cross-runtime case.

FU-COMPLETION-CFW-EXHAUSTED-CHUNK: CF Workflows writes the K5 exhausted chunk before the K5 commit — open ​

Owner: runtime-cloudflare (CF Workflows). Surfaced by: MR 1 whole-branch reviews A5 + B7 (parked by ruling R24).

CF Workflows writes the exhausted forced_completion chunk for a failed run's pending phase best-effort BEFORE its K5 commit. A step.do retry or a crash re-writes it, and a K5 that fails every attempt leaves the stream showing exhausted for a phase that is still pending in state. The chunk is diagnostic only (no reader acts on it). To close: write it after the K5 commit lands, once per run (e.g. keyed on the K5 checkpoint id).

FU-COMPLETION-PREFIX-REFUSAL-STEP-START: a prefix refusal leaves an attempt's step_start unpaired — open ​

Owner: core (step iterator) / runtime-cloudflare (CF Workflows). Surfaced by: MR 1 whole-branch review A7 (parked by ruling R24).

The prompt-prefix check (framework_prompt_prefix_changed) runs after the attempt's step_start is written (runtime-js step-iterator.ts, around the buildLLMCallbacks call; CF Workflows steps.ts, the prefix check after the attempt's step_start), so the stream holds a step_start with no step_committed / step_discarded and no attempt_started; the run fails right after. To close: move the check above buildLLMCallbacks (before step_start), or write step_discarded for the attempt before failing, with a stream-order test on both runtimes.

FU-COMPLETION-OVERRIDE-SCHEMA-UNSUPPORTED-TEST: no runtime-level test of a per-call schema refusal under an override — open ​

Owner: runtime-js / runtime-cloudflare. Surfaced by the R26 re-review.

With appliesLlmConfigOverride, an agent with llmConfigOverride is not refused at entry, and a forced call whose overridden model is native-only with a schema it cannot take fails with framework_completion_schema_unsupported. That path is covered only in parts (the llm-vercel checkOutputFormat test, the core forced-call-error tests). To close: a runtime-js case (native-only resolver keyed on the override model, an unrepresentable schema) and a DO cell for an override to an unsupported model.

FU-COMPLETION-PREFIX-CHANGED-RECOVERY: no recovery operation for a framework_prompt_prefix_changed session — open ​

Owner: core. Surfaced by: the MR !301 superpowers review (recommendation 3; parked by the R26 brief). Future work.

A session on a model with history-bound thinking that fails with framework_prompt_prefix_changed (its system prompt or tool set changed under signed reasoning blocks) can only be abandoned today: every further call replays the same signed blocks against the changed prefix. To close: a recovery operation (for example a retry / resume option, or a store-level rewrite) that strips the signed reasoning parts from the session's earlier assistant rows so the history is no longer bound to the old prefix, then lets the run continue under the new prompt, with a test on each runtime that has the prefix check.

Redis store: multi-step writes and existence checks (surfaced by RM-50) ​

RM-50 (!269) made createSession and appendMessages single Lua scripts. The audit behind it found the multi-step writes below, where a concurrent reader can see a partial state. Each still needs a fix. "Delete race" means the write runs between a separate existence check and its own commands, so it can recreate a status-less session hash that loadState rejects with InvalidStoredStateError: the old appendMessages shape. The fix pattern is the HEXISTS sessionId guard inside one script, as in APPEND_MESSAGES_ATOMIC_SCRIPT.

FU-REDIS-RM50-01: remaining multi-step Redis writes — open ​

Owner: unassigned (reliability orchestrator to route). All sites are in packages/store-redis/src/redis-state.ts.

SiteShapePartial state a reader can observe
saveStateLua, then createSessionCheckpoint (EXISTS + pipeline), then a separate pointer HSETcheckpointId stays empty until the trailing HSET, so getLatestCheckpoint falls back to stepCount ordering. The trailing HSET can race a delete. The dangling pointer is RM-38
promoteStagingEXISTS/HGET/LLEN reads, then a Lua script with no in-script existence checkdelete race; its no-staging path builds a checkpoint from separate reads
createCheckpoint / createSessionCheckpointEXISTS + pipeline (pointer last)delete race
incrementStepCount / incrementResumeCountEXISTS + pipeline HINCRBYdelete race
mergeCustomStateEXISTS + SMEMBERS + a non-MULTI pipelinehalf-applied ops (e.g. DEL list before RPUSH reads []); mixed old/new keys
cloneSessionloadState + createSession + HSET + appendMessagesthe target is visible without its messages. RM-27, cross-reference only; owned elsewhere
deleteSessionHGETALL + index Lua + pipelines + SCAN + DELthe long window is where every delete race above lands; a same-id create running concurrently has its keys deleted by the later DELs
compareAndSetStatusEXISTS + Lua + EXPIRE + HGET + index EVALthe status:* index lags the hash
createRun / updateRunStatusEXISTS / HMGET + pipelinethe run is briefly missing from listRuns; a delete race recreates a partial run hash
addSubSessionRefs / updateSubSessionRefEXISTS + pipelinethe ref is briefly missing from getSubSessionRefs; a partial ref hash can leak
truncateMessagesLLEN + DEL/LTRIM fixed by B2 (RM-22): one Lua scriptnone: the trim and the messageCount rewrite (from LLEN, only when the hash exists) are atomic, and serialize with the atomic appendMessages
loadState (a reader)two non-transactional pipelinesa concurrent atomic write can yield a mixed version (old array keys, new list contents)

Also: saveState and saveStateAndPromoteStaging never move status:* index entries, so listSessions({ status }) drifts.

FU-REDIS-RM50-02: sessionExists (EXISTS) vs the append / create guard (HEXISTS sessionId) — open (LOW) ​

sessionExists treats any key as a session, but createSession and appendMessages test the sessionId field. They disagree on a legacy status-less hash left by the pre-RM-50 appendMessages. A caller that creates-if-missing on sessionExists then skips the create, and its next append throws Session not found (runtime-js/src/run-loop.ts:~2638, and CF Workflows' saveAgentState create-if-missing). Fix direction: make sessionExists test HEXISTS sessionId. Recorded as evidence under RM-50 / CP-75.

FU-REDIS-RM50-03: persistent-spawn initial-message append error is swallowed — open (LOW) ​

runtime-js/src/run-loop.ts:4754-4761 catches a failed appendMessages of a re-spawned persistent child's initial user message and only logs a warning. The child then runs without its initial message. This is reachable only if the session vanishes right after createSession. Fix direction: rethrow, or fail the spawn with a typed error. Recorded as evidence under RM-50 / CP-75.


Eviction recovery (DO wake auto-resume) follow-ups ​

Deferred items surfaced by the four-lens code review of the eviction-recovery change (DO wake auto-resume + crash-loop guard + companion wait hardening + JS orphan detection).

FU-EVICT-01: ai-sdk HelixChatTransport swallows /resume 409 and drops the user message — open ​

packages/ai-sdk/src/transport/helix-chat-transport.ts (~:243-265) never checks res.ok on submit; a 409 from /resume (likelier now that wake auto-resume keeps stranded sessions active) silently drops the accompanying user message. Pre-existing; surface a typed error or retry.

FU-EVICT-03: JS runtime marks a SUSPENDED persistent child's ref completed — open ​

runtime-js/src/run-loop.ts (~:4338-4350): the spawn IIFE treats a child run that exits at a suspension boundary as completion and writes the ref completed. Consequence: core's paused_awaiting_client ref-status exemption (added for orphan detection) is dead code on runtime-js — nothing ever writes that status outside the store contract tests. Align the ref status with the child's real outcome.

FU-EVICT-04: usage-store counters for wake auto-resume / exhaustion — open ​

Fleet operability of eviction recovery is logs-only. Precedent exists (CLIENT_TOOL_RUNTIME_RESTARTED_COUNTER, SCHEMA_LIMITS_FALLBACK_COUNTER); add counters for auto-resume attempts and auto_resume_exhausted demotes so operators can count them fleet-wide without log aggregation.

FU-EVICT-05: dead-onFire audit of relative-deadline alarm subscribers — open ​

AlarmScheduler.fireExpired(now) can never fire a subscriber whose compute() returns Date.now() + interval (deadline always future). All three current subscribers (heartbeat, interrupt-poll, execution-watchdog) have this shape — their onFire bodies are unreachable under real alarms and their real work happens elsewhere (documented on registerSubscriber). Either give subscribers durable next-fire deadlines or remove the misleading onFire bodies.

FU-EVICT-06: injectable poll cadence for companion waits — open ​

WAIT_FOR_RESULT_POLL_MS (core, 1s) and the DO adapter's 250ms poll are not injectable; the orphan-detection and failure-window tests burn 2-5s of real time each. Make the cadence a dispatch-dep/config so tests run on a fast clock, and consider backoff for the cross-DO 250ms poll.

FU-EVICT-07: residual copy-tests in hibernation-recovery.test.ts — open ​

shouldRejectInterrupt / simulateSuccessPathOnComplete (~:39-115) re-implement production logic locally — the anti-pattern whose simulateOnStart sibling was deleted by the eviction-recovery change. Replace with tests against the real code paths or delete.

FU-EVICT-08: park-failure compensation can still lose a recovered child's output — open ​

When the poll failure window trips on TRANSIENT unreachability, the ref is marked failed/delivered and the child is best-effort aborted. If the abort ALSO fails (child unreachable) and the child later recovers and completes, its output is never delivered (notifier skips delivered refs). Residual accepted for now — the LLM is told the wait failed and can re-spawn — but a reconciliation sweep (compare ref vs live child status on parent wake) would close it.

FU-EVICT-09: crash-loop attempt over-count on a transient recovery CAS loss — open ​

recoverStrandedExecution persists the incremented crash-loop record (nextRecord, attempt+1) BEFORE calling ensureExecutionContinues. If ensureExecutionContinuesImpl's version-pinned active → paused reset CAS then LOSES (e.g. a concurrent /submit-tool-result bumped the version between the two loads), the impl returns early without running an LLM step, yet the attempt was already counted. Repeated pathological interleaving could reach maxWakeResumeAttempts slightly early → a premature auto_resume_exhausted demote. Self-healing (the concurrent producer that won the CAS drives the run → committed message progress → counter resets on the next recovery) and /retry-recoverable, so Minor. A clean fix: have ensureExecutionContinues report whether it actually started a resume, and only persist the incremented record on a real attempt. Deferred to preserve the pure-planner / thin-recovery split.

FU-EVICT-10: unpinned re-arm CAS after a failed post-reset resume — open ​

The paused → active re-arm in ensureExecutionContinuesImpl's resume .catch (fired only when WE reset active → paused and the resume then threw in its cleanup phase) is NOT version-pinned. It is gated on didResetFromActive && stateAfter.status === 'paused', so it cannot resurrect a genuine HITL suspension, and it is self-correcting (the next recovery cycle reads the interrupt flag first and demotes to interrupted, honoring any user interrupt). Verified non-resurrecting; pinning it to the post-reset version would make the intent explicit. Minor.


Remote skill loading (@helix-agents/skill-cli) follow-ups ​

FU-SKILL-CLI-MANIFEST-READ-UNWRAPPED: runCli doesn't wrap manifest read/parse — RESOLVED ​

Status: resolved · Severity: low (UX polish)

runCli now wraps the manifest readFile / JSON.parse in a try/catch that prints a helix-skills: <message> diagnostic and returns exit code 2 (matching the manifest-validation path); the bakeSkills call is likewise wrapped (exit 1 on a resolve/bake throw). A missing or malformed helix.skills.json no longer surfaces as an unhandled rejection. Regression tests: cli.test.ts › runCli error paths (missing manifest → 2; invalid JSON → 2).

FU-SKILL-CLI-MARKETPLACE-NO-ZOD: marketplace/plugin JSON parsed without a Zod schema — open (LOW) ​

Status: open · Severity: low (defense-in-depth)

ClaudeMarketplaceSource parses marketplace.json / plugin.json via JSON.parse + a TypeScript interface cast, not a Zod schema — malformed marketplace metadata yields an opaque error. Add a Zod schema for these structures for defense-in-depth (consistent with the rest of skill-cli, where the manifest and lockfile are Zod-validated).

Also lacking a malformed-marketplace-JSON test: the untrusted-external-data case (a marketplace.json / plugin.json that parses as JSON but has the wrong shape) has no test today. Add one alongside the Zod schema so the hardened validation has teeth.

FU-SKILL-CLI-INTEGRITY-EXCLUDES-FRONTMATTER: lockfile integrity hash omits skill frontmatter — RESOLVED ​

Status: resolved · Severity: low (edge case under mutable tags/branches only)

packages/skill-cli/src/bake.ts previously built the integrity file-map from the PARSED body (files[${meta.name}/SKILL.md] = body), not the raw SKILL.md — so the frontmatter (name / description / allowed-tools) was excluded from skillIntegrity(files), and a frontmatter-only change did NOT move the lockfile integrity hash.

Fix (landed): walkSkills now carries the raw SKILL.md text onto WalkedSkill.raw, exposed via ResolvedSkillSource.getRaw(name); bake.ts builds the integrity file-map's ${meta.name}/SKILL.md entry from resolved.getRaw(meta.name) (the full raw file incl. frontmatter) instead of the parsed body. A frontmatter-only change now moves the hash and is detected by helix-skills sync --check. Regression test: bake.test.ts › integrity hashes the raw SKILL.md — a frontmatter-only change moves the hash. (The codegen still uses the parsed body, unchanged.)

FU-SKILL-CLI-INTEGRITY-PER-MANIFEST-ENTRY: lockfile integrity is per-manifest-entry, not strictly per-skill — open (LOW) ​

Status: open · Severity: low (doc/wording accuracy)

packages/skill-cli/src/bake.ts aggregates ALL selected skills of one manifest entry into a single files map keyed by the manifest entry's localName, then writes one integrity per lock.skills[localName]. A multi-skill entry (e.g. a claude-marketplace plugin that ships several skills, or a git source whose include lists more than one) therefore gets ONE combined tree-hash spanning all its selected skills — not one hash per skill.

Doc wording reconciled: the README + docs/reference/skill-cli.md no longer claim "one tree-hash per skill" — they now say "one integrity hash per manifest entry (over its selected skills' raw files)", which matches the implementation. The integrity.ts docstring was updated likewise. The item is kept open as a design choice: switching the lockfile to truly per-skill lock entries (one integrity per resolved skill name) is still a possible future change, but the doc inaccuracy that motivated this item is now fixed.

FU-SKILL-CLI-TEMPDIR-CLEANUP: degit temp dirs never removed — open (LOW) ​

Status: open · Severity: low (build-time-only resource leak)

packages/skill-cli/src/git-source.ts and packages/skill-cli/src/marketplace-source.ts create mkdtemp directories to materialize the degit fetch but never rm them — each bake leaks a temp dir under the OS tmpdir. Build-time only (the CLI is never imported at runtime), so no production impact, but it accumulates across repeated helix-skills sync runs. Fix: wrap the resolve/bake usage in a finally { await rm(dir, { recursive: true, force: true }) } after the bake reads the materialized files.

Status: open · Severity: low (security-hardening)

fileSystemSkillProvider.readResource (packages/skill-fs/src/filesystem-provider.ts) enforces skill-dir containment with a lexical check (resolve + relative), NOT realpath. A symlink that lives INSIDE a skill dir but points OUTSIDE it is therefore followed by read_skill_file — and the path argument is LLM-supplied (runtime-untrusted input), so a crafted skill resource tree could exfiltrate files outside the skill dir. Disclosed in the docs as operator-trusted-only (v1 limitation #2), so this is hardening, not a live bug for operator-provisioned dirs.

Hardening: after the lexical check passes, realpath BOTH the target file and the skill dir and re-assert containment, returning null on escape. Trade-off to decide first: this would break operators who intentionally symlink skill resources to shared content outside the dir — so decide opt-in (a provider option) vs default-on before implementing.

FU-SKILL-DBOS-CATALOG-IN-WORKFLOW-BODY: DBOS resolves the skills catalog in workflow-body code — open (LOW) ​

Status: open · Severity: low

DBOS resolves resolveSkillsCatalog directly in @DBOS.workflow body code (packages/runtime-dbos/src/workflows/standard-workflow.ts / persistent-workflow.ts), NOT inside a @DBOS.step. For an in-code provider this is fine (deterministic data, no IO). For a fileSystemSkillProvider it runs replay-sensitive node:fs IO in workflow code — a soft determinism/cache-churn issue. It cannot corrupt committed state (the LLM result is checkpointed downstream) and degrades non-fatally, but it is the reason fileSystemSkillProvider is discouraged on DBOS (see the Skills cross-runtime matrix in docs/internals/concepts.md). Fix: wrap the DBOS catalog resolution in a checkpointed @DBOS.step, mirroring Temporal (which resolves it inside the per-step activity).


FU-SKILL-CORE-YAML-IN-WORKER-BARREL: @helix-agents/core barrel drags yaml (CJS) into Worker bundles — open (LOW) ​

Status: open · Severity: low

Lifting parseSkillFile into @helix-agents/core (commit 0a8a94c1) added a yaml dependency to core. parseSkillFile is re-exported from the core barrel (@helix-agents/core) and statically imports yaml, a CommonJS package whose yaml/dist/log.js does require('process'). Any consumer that bundles the core barrel into a Cloudflare Worker without nodejs_compat and without tree-shaking parseSkillFile out now pulls yaml in, and esbuild's __require('process') shim makes workerd fail at startup ("Dynamic require of 'process' is not supported"). Core can't be marked sideEffects: false (it has a required import-time enablePatches() in state/immer-state-tracker.ts), and its dist is a single bundled index.js, so per-file sideEffects won't help either.

This surfaced in @helix-agents/store-cloudflare's test worker (tsup.worker.config.ts, noExternal match-all, no node compat), which is worked around by stubbing yaml there (the store worker never parses SKILL.md). Real Worker apps typically enable nodejs_compat (which provides process) and the other Cloudflare CI lanes pass, so impact is low — but the core barrel is advertised as Workers-safe and a Node-CJS dep now taints it. Fix options: expose parseSkillFile via a dedicated core subpath (e.g. @helix-agents/core/skill-file) kept out of the main barrel and consumed by skill-fs/skill-cli; or move the parser to a Node-only package (skill-cli would depend on skill-fs). Either keeps the Workers-facing barrel yaml-free.


Deep test-coverage review (post-FU-sweep) — deferred breadth gaps ​

A deep 8-audit test-coverage review of the recently-shipped arc (Temporal parallel-tool execution, history-continuation gate, parent_session_id round-trip, DBOS HITL, cross-runtime correctness waves, stream/DO + O5, providerExecuted/lifecycle/usage) found one production bug (D1 saveStateAndPromoteStaging dropped parent_session_id — fixed + contract test added) and a set of coverage gaps. The closeable parity/breadth gaps were closed in the same sweep (CF parallel-tool compose, Temporal activity-retry idempotency, postgres re-stage upsert, legacy providerExecuted round-trip, onStateChange plain-tool parity, JS scalar-LWW exact, Temporal parallel-tool redis/postgres legs, DBOS in-place-push composes + doc correction, DBOS previousRunId unit). The items below are deferred breadth extensions (each documented in its area's test header where applicable).

FU-DBOS-INTEG-SUITE-GATED-IN-CI: runtime-dbos integration suite runs 0 tests in routine CI — open (HIGH) ​

Severity: high — the runtime-dbos integration suite (~49 files / ~88 tests) is gated behind RUNTIME_DBOS_INTEG=1 (GL #108), and the test:integ:dbos CI job does NOT set it, so the suite registers zero tests in MR/main pipelines. DBOS HITL guarantees whose ONLY coverage is in packages/runtime-dbos/src/__tests__/integration/* are unprotected in CI (a regression merges green). Load-bearing arbiter/rollback/timeout contracts DO have non-gated unit + e2e-parity backstops (so this is depth-not-sole- protection for most), but ~6 flows are gated-only: multi-parallel suspended.toolCallIds, persistent multi-turn 2nd submit, persistent SEND-path hooks, F1 resume-time-hooks, sub-agent child-owned suspend/forwarded-concurrent, submit-after-completion-timeout, predicate checkpoint-replay determinism. Action: either fix GL #108 (DBOS-singleton reset / event-driven waits) and flip include, or set RUNTIME_DBOS_INTEG=1 on the test:integ:dbos job (the suite passes locally with it set — verify stability first). Decide deliberately; the gate was added for a reason.

FU-CFW-CORE-HOOK-PARITY: CFW Workflows lacks non-suspend core-hook parity coverage — open ​

lifecycle-hooks-parity.wf-noiso.test.ts covers only the suspend/resume scenario on cfw-workflows-d1; the full core-hook set (onAgentStart, onMessage, onStateChange, beforeTool, afterTool, onAgentComplete) is not asserted on that backend (agent-hooks-parity.integ.test.ts covers JS/Temporal/DBOS; CF-DO has agent-hooks-invocation.test.ts, CFW has no equivalent). Action: add a CFW workflow-pool core-hook invocation test via the recordedHooks bridge.

FU-CF-STREAM-CHUNK-PARITY: 31-variant chunk round-trip + terminal-state idempotency not extended to CF stream managers — open ​

stream-chunk-roundtrip-parity (31 variants) and stream-terminal-state-parity (endStream/failStream idempotency, latestSequence) run only InMem+Redis. The DO-internal / binding-side managers run the G1–G4 concurrency contract but not the chunk-fidelity sweep nor the terminal-state idempotency assertions, so a DO/D1 chunk variant that loses a field through the SQLite JSON round-trip wouldn't be caught. Action: add a workerd stream-chunk-roundtrip-parity.cf over DOStreamManager + extend terminal-state assertions to the DO managers.

FU-EXPIRED-SESSION-D1-COVERAGE: expired-session cleanup has no D1/DO coverage — open ​

expired-session-cleanup-parity.integ.test.ts runs js+temporal (memory/redis/ postgres) only; cf-do-d1 is hard-filtered (Node-context suite). D1's listSessions/expiresAt round-trip + CAS-to-failed sweep is unverified end-to-end (honest gap, acknowledged in the test header). Action: add expired-session-cleanup-parity.cf.test.ts over selectViable(CF_BACKENDS, { only: ['cf-do'] }).

FU-MIXED-STATUS-RESUME-CONTINUE-PARITY: mixed-status resume→continue run-record only on Temporal — open ​

The suspend(turn1)→resume(turn2)→continue(turn3) run-record sequence (3 runs, turns [1,2,3], mixed status vector) is pinned only on Temporal (multi-turn-continuation.integ.test.ts). JS, DBOS, CF have no equivalent; CF continuation run-record-per-turn is documented-as-parity but never directly asserted (CF is also excluded from the e2e continuation floor). Action: add the 3-turn mixed sequence to the e2e continuation suite as a describe.each(getViableBackends({requires:['hitl']})) block + a listRuns() .length === 2 assertion to the CF DO continuation test.

FU-REVIEW-MINORS: assorted low-risk coverage minors — open (LOW) ​

Branch+history: interaction untested (all runtimes keep branch baseMessages semantics — distinct code path); 3rd-turn continuation history-drop relies on the > 0 gate but isn't explicitly tested at turn 3; previousRunId === undefined on a first run not asserted; remote sub-agent resume dimension is JS+memory only (not fanned out across runtimes). Each is cheap to add.


Persistent sub-agents convergence — Phase 4 (CFW) follow-ups ​

Surfaced during the Phase 4 (Cloudflare Workflows) convergence + its code review. None block Phase 4 acceptance — the convergence is correct and verified (18 unit tests + real-workerd cfw-persistent.wf-noiso.test.ts).

FU-CFW-COMPLETION-AT-LEAST-ONCE: completion-message append is not idempotent on step retry — ✅ FIXED ​

The parent's per-step completion injector appends hidden completion messages in one step.do and persists completionDelivered in a SEPARATE step.do (workflow.ts ${agentType}-inject-completions + -mark-completions-delivered). Splitting them (Phase 4 review fix) means a flag-write failure retries WITHOUT re-appending. The residual: if the APPEND step.do itself crashes mid-write, Cloudflare retries it and appendMessages (non-idempotent) double-appends — the framework-wide appendMessages-in-step.do pattern, not specific to this path. Blast radius is one workflow run (CFW has no multi-turn execute yet). Fixed: the inject-completions step.do now re-reads the recent message tail and drops completion messages already present (dedupeCompletionMessages in steps.ts, keyed on content + metadata.source === 'sub-agent-completion') before appendMessages, so a step retry after a partial write no longer double-appends. Pure dedup helper unit-tested in dedupe-completion-messages.test.ts. Residuals (LOW, from the review): (1) the dedup window is sized to the current batch, and the dedup key is the message CONTENT — so in the rare case of ≥3 completions in one batch that includes a same-output explicit re-spawn whose prior identical completion is still within the tail window, the new (legitimate) completion would be wrongly dropped (a lost notification, never a duplicate). Common single-child batches are unaffected (the prior completion is separated by the intervening LLM step + spawn tool-result, outside a window of 1). Fully robust fix = a per-delivery unique key in the completion message metadata (e.g. subSessionId) instead of content — a core buildCompletionMessage change, deferred. (2) Only the pure dedupeCompletionMessages helper is unit-tested; the inject-step wiring (getMessageCount + tail offset/window) is not (a step.do retry is hard to simulate in a unit test). Both are acceptable for this LOW item (CFW has no multi-turn execute yet).

FU-TERMINATE-FORCEFUL-ABORT: durable-runtime terminateChild does not abort the running child workflow — open (LOW, CFW + DBOS + Temporal) ​

terminatePersistentChild is a no-op on all three durable runtimes (CFW steps.ts buildCfwDispatchDeps, DBOS, Temporal activities.ts): termination relies on the durable interrupt flag core's handleTerminateChild sets, which the child workflow observes at its next interrupt-check poll. This matches pre-convergence behavior (flag-only; never aborted the child workflow), so it is NOT a regression and is UNIFORM across CFW/DBOS/Temporal (deliberately — a Temporal-only forceful cancel would make Temporal inconsistent with the others). Gap: a running child blocked inside a long tool/LLM step keeps consuming resources after terminate returns. Action (cross-runtime):

  • CFW: forceful workflowBinding.get(childWorkflowId).abort() — needs the workflow binding threaded into AgentSteps + a durable child-workflow-id registry (the id gains a __respawn__N suffix on re-spawn).
  • Temporal: wf.getExternalWorkflowHandle(childWorkflowId).cancel() — carries a child-workflow-id-reconstruction risk across a parent suspend/resume boundary (the child was spawned under the original parent workflow id; a resumed parent has a __resume-N id), so the id must be tracked durably, not reconstructed.
  • DBOS: the equivalent cancel on the child DBOS workflow handle.

FU-WORKSPACE-FAILFAST-PERSISTENT-CHILD (Phase 6 / C8): persistent-child workspace fail-fast — DBOS fixed; CFW covered; Temporal FIXED (TR-a) ​

Phase 6 closed the C8 gap where a persistent CHILD agent declaring a workspace bypassed the parent-executor assertRuntimeSupportsWorkspaces and failed LATE. Status per runtime:

  • DBOS — FIXED. serialize.ts now captures hasWorkspace on SerializedAgent (the WorkspaceConfig itself isn't JSON-serializable), and handleSpawnAgent asserts fail-fast BEFORE any ref/session/workflow side effects. Unit-tested.
  • CFW Workflows — already covered. The persistent child self-asserts at its own workflow entry (workflow.ts:522), so it fails at child-start (not late at tool-call). Minor follow-up (LOW): a blocking parent currently waits for the step.waitForEvent timeout rather than failing at the spawn; add a spawn-dispatch assert (child config is reachable at workflow.ts:2826 via paConfig.agent.workspace) so the failure surfaces at the parent immediately.
  • Temporal — FIXED (Temporal rewiring TR-a). The child run-start activity initializeAgentState (activities.ts) calls assertRuntimeSupportsWorkspaces, so a persistent child declaring a workspace fails fast at startup with the shared message (the throw surfaces within the maximumAttempts: 3 cap, like other deterministic init failures). The child workflow's top-level catch RETURNS {status:'failed'} (does not throw), so the blocking dispatch in workflow-dispatch.ts inspects the resolved childResult.status to mark the parent's SubSessionRef 'failed'. sub-agent-fail-fast-temporal.integ.test.ts is now un-skipped and green against a real Temporal server (unlike CFW it fails at the parent immediately — no blocking-spawn timeout). NOTE: the earlier claim that Temporal v7 companion dispatch "is not implemented" was STALE; it was largely working before the rewiring (verified by the persistent integ suite).

FU-DBOS-BLOCKING-SPAWN-SEMANTICS: DBOS blocking-spawn blocks-until-idle and mis-reports a failed child as completed — open (MED) ​

Catalogued as CP-69 (runtime conformance catalogue, packages/e2e/src/conformance/catalogue/control-plane/CP-69.ts; see docs/conformance/findings.md) — consequence (1), blocks-until-idle. Consequence (2) is not part of CP-69 and is only partly mitigated: runFinalizeRun's shouldPreservePriorRunStatus (runtime-dbos/src/steps/run-lifecycle.ts) preserves and propagates a failed/interrupted status on the clean-exit finalize only when the child's most recent turn failed or was interrupted. An earlier failed turn followed by a completed turn can still leave the session reported as completed, so (2) remains open for that shape (not re-verified by a test here).

From the Phase 5 review. A persistent child runs every turn with skipSessionStatusUpdate:true, so its SESSION status stays active until the recv loop idles out, and finalizeLoop (persistent-workflow.ts) then hardcodes status:'completed'. Consequences for the C5 blocking spawn (which polls the child's session status): (1) it blocks until the child's persistentIdleTimeoutMs (default 24h) rather than until the child's first turn; (2) a child whose turn FAILED but then idled out is reported completed. Mitigated for now:pollChildSessionTerminal is bounded (returns running instead of hanging forever on a paused/erroring child); the semantics are documented in the handler + design with a "use a short persistentIdleTimeoutMs for blocking children" recommendation. Deeper fix: poll the per-turn RUN record (getCurrentRun, finalized each turn with the accurate status at shared.ts:2007) instead of the session status — turn-based + failure-accurate — and/or make finalizeLoop carry the child's real terminal status. Must be integ-verified.

FU-DBOS-COMPANION-INTEG-GAPS: DBOS persistent companion lacks end-to-end delivery + crash/recover integ tests — open (MED) ​

From the Phase 5 review. (1) No integ test proves the Phase 5e notifier actually DELIVERS a non-blocking child's completion to the parent (unit-proven + integ-non-regression only — companion-non-blocking finishes the parent before the child completes, so the notifier is never the delivering path). (2) No crash/recover integ test exercises the C5 blocking-spawn body-replay (the design said this "MUST be integ-verified before claiming C5"). Add: (a) a parent with a turn AFTER a non-blocking child completes, asserting the hidden sub-agent-completion message appears; (b) a restart-mid-blocking-spawn test asserting no duplicate child + a single consistent tool result. Env: PG on 5434

  • Redis (see the Phase 5 design doc's env note). Pattern available: the Temporal rewiring (TR-b) proves gap (a) with a race-free pre-seed test — seed a terminal-but-undelivered non-blocking child ref + child state, run the parent one turn, assert the hidden sub-agent-completion message + the ref's completionDelivered flip (see persistent-subagent-temporal.integ.test.ts "completion notifier delivers..."). The same pattern ports directly to DBOS and sidesteps the shared-MockLLM concurrent-spawn race.

FU-TEMPORAL-NOTIFIER-LIVE-NONBLOCKING-RACE: live non-blocking spawn not integ-tested on Temporal (shared MockLLM queue) — open (LOW, test-infra) ​

The Temporal completion notifier (TR-b) is proven by a race-free pre-seed integ test (see FU-DBOS-COMPANION-INTEG-GAPS above) + core unit tests (completion-delivery.test.ts). A LIVE non-blocking spawn (parent spawns, child runs concurrently, completes, parent's next turn delivers) is NOT integ-tested: parent + child share one MockLLMAdapter response queue, so the concurrent child races the parent for responses (persistent-subagent-temporal.integ.test.ts:200 stays it.skip-ped for this reason). Fix needs per-agent LLM response queues in the Temporal test harness (or a deterministic child that consumes no LLM responses). Same test-infra limitation as DBOS; not a runtime bug.

FU-REVIEW-CONVERGENCE-FOUR-PASS: deferred items from the (two) four-reviewer convergence passes — open (LOW-MED) ​

TWO four-independent-reviewer passes over the convergence branch. Round 1 (correctness / replay-safety / cloudflare / test-coverage) → fixed 84ad171 (Temporal blocking inline-output R2-I1; CFW waitForResult dedup R3-C1; CF-DO sendMessage terminated guard R3-I6; timeout comment R1-C1) + 54ae452 (R1-C2 paused-child terminate flag ✅; R3-I5 registry-cast removed via the new CompanionParentConfig ✅; R3-I3 D1 V14 comment ✅). Round 2 (concurrency / type-design / fix-verification / store-layer) → review #3 verified all round-1 fixes correct; fixed 63fb801 (streamManager optional → removed 3 as unknown as StreamManager casts; mode union → removed a cast; UpdateSubSessionRef.status += paused_awaiting_client; dropped unused CompanionParentConfig.name; RoutingMockLLMAdapter opts spread; completionDelivered doc fix; new cross-store contract test for status-only-preserves-completionDelivered). Deferred-items round (the user's "fix deferred items"): a7db7cd (R1-C3 dispatcher Zod arg-validation ✅; terminated-ref notifier skip-early ✅), 710213e (R2-C-1 ✅ — extractRootErrorMessage cause-unwrap + C8 assert moved after createSession so a non-blocking workspace child's failure is persisted + delivered), + the parallel-auto-name spawn DOCUMENTED (spawn tool name description + dispatcher comment steer the LLM to name concurrent same-type children — see below). Round 3 (concurrency / edge-cases / parity-audit / comment-accuracy) → review #1 verified the deferred-items fixes correct; fixed 7d9f229: DIV-1 (CRITICAL — waitForResult double-delivery on JS+Temporal: core now sets completionDelivered on the inline-delivered ref) ✅; C4 canonical signal (sendMessage guard now checks non-empty pendingClientToolCalls, not status==='paused') ✅; extractRootErrorMessage hostile-getter guard ✅; Zod hardening (name .min(1), timeout .positive()) ✅; a batch of comment-rot (DBOS waitForChildTerminal/registry/deps JSDoc, Temporal dispatch header, drifted line refs, DIV-4 DBOS no-resumeMetadata documented as intentional) ✅. Accepted/architectural (NOT bugs): DIV-2 — CF-DO listChildren reads refs directly without resolveAndSyncSubSessionRefs (a DO can't loadState a sibling DO), so a just-completed non-blocking child shows running until the parent's next checkCompletedChildren turn; DIV-3 CF-DO sendMessage C4 is best-effort under a transient /status failure; DIV-5 CFW 'spawning' interim status is never surfaced to the LLM. Round 4 (verify-fixes / resource-DoS / public-API-semver / test-quality) → review #1 verified all round-3 fixes correct; fixed e852b08: R4-M1 (orphan ref on a failed spawn — core now compensates the ref to 'failed' so the name stays re-spawnable, mirroring DBOS) ✅; R4-I3 (terminate sets completionDelivered:true so terminated refs aren't re-scanned each turn) ✅; R4#4-I1 (parity-harness defensive store reset between scenarios) ✅. The public-API/semver audit (R4#3) is captured in docs/superpowers/specs/2026-06-06-phase8-changeset-semver-analysis.md (core needs a MAJOR bump). Tracked (disproportionate / known): waitForResult unbounded wait on Temporal (activity-proxy + heartbeat tuning, or a Zod timeout.max() cap); per-turn O(N) notifier loadState (scales with RUNNING children — a store-side non-delivered index would bound it); RoutingMockLLMAdapter unbounded key Map (test-only, negligible). children; store-atomic add-if-absent remains the proper fix).

Still deferred (cosmetic / disproportionate only):

  • R1-minor: stale completedAt on re-spawn. The re-spawn path resets status→running + completionDelivered:false but leaves the prior run's completedAt (UpdateSubSessionRef can't null a field). Add a clear mechanism or set it explicitly. Data-quality only (never read for control flow).
  • R3-I2 (cosmetic): runtimeName: 'CloudFlare Workflows' uses non-canonical brand casing ('Cloudflare'). Left as-is to avoid churning the shared assertRuntimeSupportsWorkspaces message contract + changelog references.
  • R4 test gaps: the C6 core call-order assertion ✅ (now has teeth via the CTD-stub opLog test + the Temporal recording-store test) and the harness blocking-FAILURE scenario ✅ (blocking-child-fails) are DONE under FU-TEST-AUDIT-PERSISTENT-SUBAGENTS. Remaining: non-blocking LIVE completion-delivery e2e (only unit + pre-seed today).

Already tracked elsewhere: R2-I2 = FU-DBOS-BLOCKING-SPAWN-SEMANTICS; R3-I4 = FU-CFW-COMPLETION-AT-LEAST-ONCE.

FU-TEST-AUDIT-PERSISTENT-SUBAGENTS: comprehensive test-coverage audit of the persistent-subagent feature — high-leverage deterministic gaps FILLED; timing-sensitive integ gaps tracked (LOW-MED) ​

A 6-slice deep audit (core-unit / runtime-adapters / integ-cross-runtime / e2e / stores / parity-harness) of the entire persistent-subagent test suite. Established that happy paths + the marquee bugs (C1–C8, DIV-1, R4-M1, terminate- truth) were well covered, but real gaps clustered in core-unit error branches, zero-coverage helpers, and a Temporal-adapter blind spot. Filled (committed):

  • companion/helpers.test.ts (NEW, 14): generateAutoName MAX+1/gaps/prefix- isolation/substring-prefix/ephemeral-skip; resolveChildName auto/provided/ active+interrupted-dup-throw/terminal-reuse; ACTIVE/TERMINAL set membership + paused_awaiting_client in neither.
  • companion-tool-dispatch.test.ts (+21): spawn error branches, sendMessage branches incl. C6 append-before-flag ORDERING (new opLog stub gives it real teeth — the prior test asserted both writes but not order), listChildren/ getChildStatus/waitForResult/terminateChild branches, arg-validation edges.
  • dispatch-companion.test.ts (NEW, 20, runtime-temporal): the CRITICAL gap — Temporal had ZERO direct executeCompanionToolCall unit tests vs CFW's 14. Mirrors CFW + pins Temporal-adapter seams (spawnMetadata/resumeMetadata lifting, error mapping, C6 ordering via recording-store, DIV-1 flag).
  • cross-store contract (sub-session-operations.ts): paused_awaiting_client round-trip — runs on all 5 stores (verified memory + real PG + Redis).
  • parity harness: blocking-child-fails scenario (JS + Temporal) — child runs then errors; parent completes, ref recorded failed.
  • session-types.test.ts (+5): assertPersistentSubSessionRef (was zero).
  • routing-mock-adapter.test.ts (+2): gracefulExhaust:false registered-key throw vs unregistered-key always-graceful.
  • run-loop-persistent.test.ts: replaced the expect(true).toBe(true) placebo companion-guard test with a real one (inject a CompanionTool into an agent with no persistentAgents → hits + asserts the guard).
  • persistent-subagent-js.integ.test.ts: the duplicate-active-name e2e was would-pass-if-reverted (conditional if (isError)); now documents it as a timing-dependent SMOKE test (deterministic teeth live in the 3 unit tests above) + added a CROSS-BRANCH singleton-name invariant (exactly one terminal worker ref) that breaks on a bad revert regardless of timing.

Review pass over the audit additions (4 independent reviewers: test-teeth / fixture-fidelity / adjacent-coverage / cross-runtime-consistency). ZERO critical; all actionable items fixed: core handleSendMessage terminal-child-with-stale- pending-map fall-through (the !childIsTerminal clause — delivered:false, not a C4 throw); Temporal hook-firing (beforeTool/afterTool success+failure — the file claimed to pin hooks but every agent had none) + unresolvable-parent-agentType branch; CFW mock updateSubSessionRef narrowed to the real-store updatable-field whitelist (was a permissive full-spread); generateAutoName descending-order (counter > maxCounter false arm); DBOS failed-child waitForResult inline branch; fixed an overstated e2e comment (the singleton toHaveLength(1) is structural-by- construction, NOT a guard-revert detector) + a wrong helper test title.

Remaining (timing-sensitive / cross-runtime integ — tracked, not blocking): JS live non-blocking delivery between turns + C1 CAS-loser + C4 e2e + waitForResult timeout/abort; Temporal sendMessage-to-interrupted + terminate- running + live non-blocking (see FU-TEMPORAL-NOTIFIER-LIVE-NONBLOCKING-RACE) + dedup-across-2-turns + non-blocking workspace-failure integ; DBOS notifier- delivers + C5-replay + @DBOS.step replay-once + blocking-inline-failure (see FU-DBOS-COMPANION-INTEG-GAPS); CFW at-least-once (FU-CFW-COMPLETION-AT-LEAST- ONCE) + waitForResult handler branches; CF-DO terminate-running + C4 fail-open; Postgres V11 column-on-populated-table. A fully-deterministic duplicate-active- name e2e needs a child-side sync barrier (gated tool) — disproportionate vs the unit teeth that already exist.

FU-PARITY-HARNESS-TABLE-DRIVEN: table-driven cross-runtime persistent-subagent harness — JS+Temporal SHIPPED; DBOS/CF drivers remaining (LOW-MED) ​

The table-driven harness is built: packages/e2e/src/__tests__/persistent-subagent-parity-harness.integ.test.ts. A shared scenario library (runtime-agnostic: per-agentType RoutingMockLLMAdapter scripts + a ref-based assertion) runs against a thin per-runtime DRIVER. Drivers implemented: JS (in-process JSAgentExecutor) + Temporal (docker runner). 4 deterministic blocking/sequential scenarios (blocking-spawn-completes, two-blocking-spawns-auto-named, terminate-truth-preserves-completed-ref, wait-for-result-after-blocking-child) × 2 drivers = 8 tests, green. Assertions normalize on getSubSessionRefs + child loadState (every runtime persists these; Temporal/DBOS don't persist companion ToolResultMessages). The DRIVERS array makes adding a driver one line.

Remaining:

  • DBOS driver — blocked on FU-DBOS-BLOCKING-SPAWN-SEMANTICS. Every current scenario is blocking-spawn-based; DBOS blocking-spawn polls the child SESSION status until idle-timeout, so it would hang/mis-report. Add the DBOS driver once that lands (or add a non-blocking + completion-notifier scenario set that DBOS handles cleanly).
  • CF-DO / CFW drivers — run in the separate workerd vitest pool, so they need adapter wrappers (can't share the Node describe loop directly).
  • More scenarios to fold in: non-blocking spawn + completion-notifier delivery (needs the timing handled — see the 7b pre-seed pattern), sendMessage happy-path, getChildStatus, re-spawn (skip on DBOS — idempotency), and stream-event parity for companion__spawnAgent (tool_start/tool_end) on Temporal/DBOS (JS has it at persistent-subagent-parity.integ.test.ts:1398). Design: docs/superpowers/specs/2026-06-05-persistent-subagents-parity-harness-design.md.

FU-DBOS-ROUTING-LLM-CONVERGE: DBOS persistent integ uses a local routing LLM instead of the shared core one — open (LOW, convergence) ​

Phase 7a added a shared RoutingMockLLMAdapter to core (session/agent-routing mock LLM). persistent-subagent-dbos.integ.test.ts still uses a LOCAL createRoutingLLM (lines ~85-140) with its own AgentScript format, predating the shared adapter. Converging it to RoutingMockLLMAdapter (route by agentType, setResponses(key, MockResponse[])) removes the duplication and lets the DBOS suite share the same exhaust-isolation semantics. Deferred because it's a sizable script-format rewrite on the finicky DBOS integ env (PG :5434 + Redis) with no behavior change — the local adapter works. The DBOS non-blocking-spawn skip is a separate concurrent-DBOS-workflow constraint, not a queue race, so it is NOT unblocked by the routing adapter alone (unlike Temporal, where Phase 7b un-skipped the live non-blocking spawn via the shared adapter — persistent-subagent-temporal-routing.integ.test.ts).

FU-TEMPORAL-HASPERSISTENT-PER-ENTRY-ACTIVITY: notifier gate costs one extra activity per workflow entry — open (LOW, perf) ​

TR-b gates the per-turn completion notifier on agentDeclaresPersistentChildren, resolved via an ACTIVITY once at each workflow entry (fresh + resume) so the gate survives suspend/resume without plumbing a flag through AgentWorkflowInput + both child-spawn sites. Cost: one extra activity round-trip per workflow entry for ALL agents, including non-persistent ones (DBOS reads agentSerialized.persistentAgents?.length inline — no activity — because it embeds the serialized agent in the workflow input). Optimization: thread a hasPersistentAgents boolean onto AgentWorkflowInput (set by the executor for the parent and by dispatchCompanionTools for children — the child's config is reachable there via the registry), eliminating the per-entry activity. Behavior is already correct; this is purely a round-trip saving.

FU-DBOS-DETERMINISM-SCAN-COMPANION: the workflow-body Date.now() scan doesn't follow into core companion handlers — open (LOW) ​

From the Phase 5 review. The routed read/terminate handlers run core's executeCompanionToolDispatch in the DBOS workflow body, and core uses bare Date.now() in resolveAndSyncSubSessionRefs / handleTerminateChild (completedAt). The no-date-now-in-workflow-body.test.ts scan only checks the three local workflow files, so it no longer enforces the invariant for code that genuinely runs in the body. Runtime impact is BENIGN (those timestamps are one-time-idempotent — guarded by terminal-status skips — and never feed control flow), so it's a test-coverage gap, not a behavioral bug. Fix: extend the scan to follow imported companion handlers, or add a core lint rule for dispatcher timestamp helpers.

FU-DBOS-SPAWN-REPLAY-DUPLICATE: auto-named spawn can duplicate a child on workflow body-replay — ✅ FIXED (one residual) ​

From the Phase 5 review (flagged as a pre-existing property, NOT a Phase 5 regression — the old fire-and-forget spawn shared it). Companion handlers run in the DBOS workflow body and re-run on recovery (and inside CFW's retryable step.do). For an AUTO-named spawn, a body-replay re-ran resolveChildName → generateAutoName which, seeing the prior ref still at running, returned the NEXT name (researcher-2) and addSubSessionRefs+startPersistentWorkflow a brand-new child; an EXPLICIT name threw already running. Fixed: spawn now keys idempotency on the deterministic toolCall.id. At handler entry, if a SubSessionRef already carries this toolCall.id, its name is reused so the derived child id is identical across replay and addSubSessionRefs / spawnPersistentChild are idempotent. First-spawn refs store the real toolCall.id as parentToolCallId. Applied to core handleSpawnAgent (JS / Temporal) and the DBOS bespoke handler. CFW Workflows and CF-DO also route through core handleSpawnAgent but had been passing a CONSTANT 'companion-spawnAgent' id — surfaced in code review: with the new id-keyed dedup that constant would collide distinct spawns onto one child, so both call sites (steps.ts handleSpawnAgent, do/persistent-companion-tools.ts) now thread the real per-call toolCall.id. All 5 runtimes covered. Unit-tested on core, DBOS, and the CFW path (distinct-spawns-no-collision + same-id-replay). Residual (LOW): an explicit-name RE-spawn (a new tool call after a terminal run) resets the existing ref via updateSubSessionRef, which doesn't carry parentToolCallId, so a replay of the RE-spawn degrades to a deterministic no-dup error (not a duplicate). Fully closing it needs parentToolCallId in UpdateSubSessionRef across all 5 stores — deferred (explicit-name re-spawn is rarer).

FU-CFW-WAITFORRESULT-STALE-PARK: CFW waitForResult can park forever on an already-completed child — open (LOW, pre-existing) ​

CFW's handleWaitForResult (steps.ts) is NOT routed through core (it must park in the workflow body) and reads getSubSessionRefs verbatim — no lazy ref-sync. If a non-blocking child already completed (its sub-agent-complete event already fired) but the parent ref is stale running, waitForResult returns a blocking descriptor and the body parks on an event that won't fire again → waits until timeout. Pre-existing (unchanged by Phase 4). Action: run the synchronous read-portion of handleWaitForResult through the same lazy-sync used by the other handlers (return the result inline when the child is already terminal, park only otherwise).


Persistent-cacheable-memory + cache-strategy follow-ups (post-MR !214) ​

Surfaced by the eight-review pass after MR !214 merged. Full triage: docs/superpowers/specs/2026-05-31-deferred-items-triage.md. Two of the four deferred items (multimodal recall query, unreachable format branch) were fixed in the memory-recall-followup-cleanups branch. The two below are dedicated efforts.

FU-MEM-FROMSTART-RETRY: retry({ mode: 'from_start' }) double-claims and is cross-runtime divergent — ✅ done ​

Status: done · Closed by: branch from-start-retry-fix

Resolution: Removed RetryOptions.mode and the 'from_start' value entirely across core and all four runtimes. retry() now has a single behavior — restore from a checkpoint (latest or checkpointId) and re-run the triggering message — with a genesis fallback: if no checkpoint exists, restart fresh from the triggering message instead of throwing. from_start was broken on every runtime since introduction and had zero test coverage.

FU-RETRY-GENESIS-PARSE-WEDGE: genesis-state stateSchema.parse({}) can wedge a session in active — ✅ done ​

Status: done · Closed by: branch from-start-retry-fix

Resolution: Wrapped the entire post-CAS retry() body in a try/catch (JS, Temporal, CF) that best-effort rolls the session back active→failed on ANY error (including a ZodError from stateSchema.parse({})), then re-throws the original error. Tests in js-agent-executor-retry-genesis.test.ts and executor-retry-genesis.integ.test.ts pin the rollback and verify a second retry() can succeed after the parse-throw.

FU-RETRY-BOGUS-CHECKPOINTID: explicit-but-unresolvable checkpointId silently wipes state — ✅ done ​

Status: done · Closed by: branch from-start-retry-fix

Resolution: Added an explicit guard after getCheckpoint(checkpointId): if the caller provided an explicit checkpointId and the store returns null, throw Checkpoint '${checkpointId}' not found for session ${sessionId}. instead of falling through to genesis. The general post-CAS rollback (FIX B) covers this error too — the session is returned to failed and no truncation occurs. Tests in both js-agent-executor-retry-genesis.test.ts and executor-retry-genesis.integ.test.ts verify the error message, rollback, and transcript preservation.

FU-CF-DO-RETRY-DOUBLE-CAS: Cloudflare Durable-Object /retry double-claims the session and never runs — ✅ closed ​

Status: ✅ closed — handleRetry now does a read-only status precheck; the single authoritative ['failed']→'active' CAS is owned by JSAgentExecutor.retry(). CF-DO /retry runs again instead of wedging. · Closed by: branch fix/cf-do-retry-double-cas. · Severity: high (correctness — CF-DO /retry was non-functional) · Pre-existing: introduced by the v7 stateless-suspension redesign (d7fab7fb2), NOT by the from-start-retry-fix work. Surfaced during the post-!220 exhaustive coverage audit.

Resolution: Removed handleRetry's own compareAndSetStatus(sessionId, ['failed'], 'active') (durable-object-agent-base.ts:1843) and replaced it with a READ-ONLY precheck: load state, return 400 if no state exists, 409 if status is not 'failed' — without mutating status. The single atomic ['failed']→'active' transition is now owned exclusively by JSAgentExecutor.retry() (js-agent-executor.ts:2547), which is also used directly by the JS runtime. The precheck is advisory — if state flips between the read-only load and the inner CAS, retry()'s own CAS-failure handling + rollback applies, since handleRetry no longer pre-claims. The two post-resolve rollback paths (updateStatus(sessionId, 'failed') on agent-resolution / missing-namespace errors) are now idempotent no-ops (the store is still 'failed' at those points), kept as defense-in-depth. New boundary coverage: packages/e2e/src/__tests__/cloudflare-do-retry-protocol.cf.test.ts (real-workerd DO) — (a) /retry on a failed session reaches 'completed' with materialized output (pre-fix wedged; teeth-verified by reverting the precheck → RED again), (b) non-failed session → 409, session untouched. This closes the prior zero-coverage gap on the CF-DO /retry route.

Symptom: DurableObjectAgentBase.handleRetry (runtime-cloudflare/src/do/durable-object-agent-base.ts:1843) does its own compareAndSetStatus(sessionId, ['failed'], 'active'), then delegates to JSAgentExecutor.retry() (:1940), which performs the samecompareAndSetStatus(['failed'], 'active') (runtime-js/src/js-agent-executor.ts:2547). By the time the inner CAS runs the status is already 'active', so it fails and retry() throws "Cannot retry: agent status is 'active'…" — and that throw happens before the inner post-CAS rollback try-block (:2567), so the session is left wedged in 'active'. The throw propagates inside the un-awaited execution IIFE, so the HTTP /retry response is still 200 with a runId, but the run never starts and no subsequent retry() can succeed (the ['failed'] CAS never matches again).

Why it went unnoticed: the CF-DO /retry boundary has zero test coverage — the cloudflare-do-* suite exercises /start, /status, /interrupt, and persistent-agent flows, but never /retry. (The Workflows-runtime retry path, executor.ts, is a separate code path and is covered.)

Recommended fix: make handleRetry perform a read-only status precheck (load state, return 409 early if not 'failed') and let JSAgentExecutor.retry() own the single atomic CAS — eliminating the double transition while preserving the clean early-409. Then add a DO /retry boundary test (checkpoint retry + non-failed-status 409) to the cloudflare-do-* suite. Small, self-contained; warrants its own MR rather than riding on the from_start removal.

FU-CF-MEM-COUNTER: Cloudflare re-extracts the whole transcript every step — open ​

Status: open · Severity: low (efficiency, not correctness)

Symptom: Cloudflare builds a fresh MemoryManager per durable step (runtime-cloudflare/src/steps.ts:4530, dispatched at workflow.ts:1984), resetting lastExtractionMessageCount (packages/memory/src/memory-manager.ts:57) to 0. Realtime extraction therefore re-walks the entire filtered transcript every step (~O(steps²) extractor + dedup LLM calls; prompt cost grows with conversation length).

Not a correctness bug: there is no store-layer hash key; dedup is the heuristic LLM/similarity pass in storeExtractedMemories, which suppresses re-extracted content (small tail-risk of stray dupes only). The counter is meaningless in Cloudflare's stateless-per-step model; the JS runtime keeps one MemoryManager per runLoop so the counter accumulates correctly there. Temporal does no realtime extraction.

Recommended fix: persist the extraction offset in durable customState (or a per-session memory-store marker) and rehydrate it per step. Touches the MemoryManagerLike contract → warrants a planned change. The same externalization would let Temporal adopt realtime extraction cleanly. Needs a brainstorm before a plan.

FU-REDIS-STAGE-UPSERT-BY-TOOLCALLID: RedisStateStore.stageChanges double-applies a retried toolCallId — ✅ closed ​

Status: ✅ closed — RedisStateStore.stageChanges now upserts by toolCallId; idempotent under Temporal activity retry; parity with Postgres + in-memory.

Symptom (before fix): RedisStateStore.stageChanges did an unconditional rpush, and promoteStaging's Lua applies every LRANGE entry with no dedup by toolCallId. A retried stageChanges for the same toolCallId therefore double-applied its writes at promotion (e.g. an /items/- append duplicated into the array).

Why it became reachable: runtime-temporal now stages per-tool via executeToolActivity — a Temporal activity with at-least-once retry semantics (maximumAttempts: 3). Temporal+Redis is the first per-tool-staging-on-Redis combo; before the parallel-tools split, Redis staging was only ever exercised once per step.

Closure: stageChanges now runs STAGE_CHANGES_UPSERT_SCRIPT — a single EVAL that scans the staging LIST for an entry whose decoded toolCallId matches and LSET-replaces it IN PLACE (preserving its FIFO position), else RPUSHes a new tail entry (TTL applied inside the script). The scan+write is one atomic EVAL, so it can't race concurrent parallel-tool staging on the same step, and a retry collapses to a single applied entry. This is the Redis analogue of Postgres's ON CONFLICT (session_id, step_id, tool_call_id) DO UPDATE and the in-memory findIndex → replace-in-place; all three stores are now idempotent-by-toolCallId. promoteStaging's Lua and the size-limit pre-flight guard are unchanged. Corrupt / old-shape entries that fail cjson.decode are treated as non-matching (can't alias a live toolCallId), mirroring getStagedChanges / the promoteStaging Lua.

Teeth: the cross-store staging contract (packages/core/src/testing/staging-operations.ts) gained three idempotency-by-toolCallId assertions (re-stage replaces / applies append once; distinct ids stay separate FIFO entries; middle re-stage preserves position) — these run against ALL stores incl. Redis, Postgres, and in-memory in CI. Redis-specific coverage lives in packages/store-redis/src/__tests__/redis-staging-fifo.integ.test.ts (Idempotency by toolCallId describe). Both verified locally against a live Redis; reverting only the redis-state.ts impl re-fails the new tests ([t1,t2,t3,t2] duplicate + double-applied append).


TII-C1+SUBAGENT (Cross-runtime sub-agent matrix) follow-ups ​

TII-C1+SUBAGENT: Cross-runtime sub-agent multiplexed stream resume matrix — ✅ landed (with 2 open follow-ups) ​

Status: ✅ landed Closed by: the commit series implementing docs/superpowers/plans/2026-05-12-cross-runtime-subagent-resume-matrix.md (20 commits ending at HEAD on omnara/stateless-suspension-redesign).

Summary: Extended the single-agent cross-runtime-resume-matrix.{integ,cf}.test.ts with parent-with- sub-agent topologies across 6 runtime × store × stream-manager combos (JS+memory, JS+redis, JS+postgres+redis, JS+D1+DO, DBOS+postgres+redis, Temporal+memory, Temporal+redis+redis) plus the CF-DO slice. 30 active tests pass with vitest run; 67 env-gated tests skip cleanly without REDIS_URL / POSTGRES_URL; it.todo markers preserve documented gaps (CF Workflows combo — blocked on Miniflare). S6 doubly-nested HITL is now a negative-invariant test, not it.todo — see FU-MATRIX-DOUBLY-NESTED-HITL (resolved, documented-as-unsupported).

Scenarios covered: single sub-agent baseline (S1), multi-turn with sub-agents (S2 — see FU-MATRIX-MULTITURN), nested γ-cascade (S3), γ-fan-out (S4), child failure (S5), doubly-nested HITL (S6 — negative-invariant, see FU-MATRIX-DOUBLY-NESTED-HITL), cleanupToStep boundary (S7), remote sub-agent (S8), multi-tab race (S9), fromSequence past end (S10), mid-cursor resume (S11), and stream_resync on path-3 resume (S12).

Side fix landed during matrix work: runtime-temporalworkflow.ts was missing the parent's tool_end chunk for subagent__<name> after subagent_end on the success path of local sub-agent dispatch. Mirrors runtime-js's paired-close pattern. The matrix's S1 path-6 tool_end assertion (commit 808f36273) surfaced the bug; the fix lives in commit 9f6723900.

Two follow-ups still open (tracked separately):

  • FU-MATRIX-MULTITURN — no portable cross-runtime API for multi- turn continuation; S2 documents the gap by skipping Temporal rows with descriptive name suffixes.
  • FU-MATRIX-DOUBLY-NESTED-HITL — ✅ resolved (documented-as- unsupported, GL #73): child suspending on its own client-tool returns completed instead of suspended_client_tool on both JS and Temporal. S6 pins this as a negative-invariant assertion.

FU-DBOS-LOGGING-GAPS: Bare catch {} blocks in runtime-dbos lose operator observability — ✅ closed ​

Status: done (post-release follow-up cluster, alongside Path B Phase 2). Surfaced by: pr-review-toolkit:silent-failure-hunter audit of the v6→v7 stateless-suspension MR (commit landing audit performed 2026-05-13).

Closure: all 4 bare catch {} sites now emit a typed log line at the appropriate severity. Control flow is unchanged — the authoritative fallback mechanism (CAS gates, failStream/endStream, TTL fallbacks) still owns user-observable correctness. The log lines are observability-only:

  • packages/runtime-dbos/src/workflows/shared.ts:1122-1133 / 1144-1155 (now widened to include the warn log line): args.logger?.warn?.('dbos.workflow.emit_error_chunk_failed; failStream remains authoritative', {...}) and the analogous emit_output_chunk_failed.
  • packages/runtime-dbos/src/handles/active-handle.ts:295 → this.cfg.logger?.debug?.('dbos.active_handle.poll_load_state_failed; will retry on next interval', {...}). Debug because a single missed poll-iteration is fully tolerable; a sustained pattern is visible at this level.
  • packages/runtime-dbos/src/lifecycle/resume.ts:161, 167 → both .catch() handlers now log warn with structured context. The control flow degrades gracefully to a misleading AgentAlreadyRunningError if the store is genuinely broken, so observability is the only operator hook for diagnosing that case.
  • packages/runtime-temporal/src/workflow.ts:1660-1666 → top-level workflow-failure escape hatch now logs error (not warn) via wf.log.error('temporal.workflow.persist_terminal_state_failed; durable state may be stuck in non-terminal status', {...}). This is the canonical place where durable state can drift from the in-memory result, so it gets the highest severity.

No tests changed; verified runtime-dbos 364/364 + runtime-temporal 243/243 + integration usage-tracking 4/4 still pass after the edits.

Description: Four sites in runtime-dbos swallow errors with a bare catch {} (no log line). The user-observable contract is preserved at each site by a bounded primary mechanism (CAS gates, authoritative failStream/endStream, TTL fallbacks), so these are NOT bugs — but the operator loses signal when transient infra failures occur. Inconsistent with sibling sites in the same package that DO log at debug/warn.

Sites:

  1. packages/runtime-dbos/src/workflows/shared.ts:1112-1123 and 1133-1144 — dispatcher.emitErrorChunk / dispatcher.emitOutputChunk "best-effort" catches. If miniflare D1 / Redis pubsub is degraded, the canary signal is suppressed. Comment ("failStream is authoritative") is defensible but a args.logger?.warn?.(...) log line would aid investigation.
  2. packages/runtime-dbos/src/handles/active-handle.ts:295 — _pollForPendingClientTool's loadState catch. State-store flakes during the poll race are unobservable.
  3. packages/runtime-dbos/src/lifecycle/resume.ts:161, 167 — loadState(sessionId).catch(() => null) and getCurrentRun(sessionId).catch(() => null) in the [CROSS-RUNTIME] active-session guard. If the state store is genuinely broken, hasPendingClientTools reads false and the code throws AgentAlreadyRunningError — misleading classification.
  4. packages/runtime-temporal/src/workflow.ts:1660-1666 — inner persistTerminalState "best-effort" catch in the top-level error path. If persistTerminalState fails, the workflow still returns {status:'failed', error} from in-memory, but durable state may be stuck in 'active', which the next execute(sessionId) would trip on with a misleading "already active" rejection.

Action: Replace each bare catch with catch (err) { args.logger?.warn?.('<operation> failed; <fallback mechanism> remains authoritative', { sessionId, runId, error: ... }); }. No control-flow change — just add the log line.

Effort: ~2 hours. Priority: low — observability polish.


FU-DBOS-RECONNECT-STREAM-UNIT-TEST: Multi-turn stream reactivation lacks focused unit test — ✅ closed ​

Status: done (post-release follow-up cluster, alongside Path B Phase 2). Surfaced by: pr-review-toolkit:pr-test-analyzer audit of the v6→v7 stateless-suspension MR. Rated 6/10 (not blocking).

Closure: added 4 focused unit tests in packages/runtime-dbos/src/__tests__/lifecycle/execute.test.ts that pin the reactivateStream → initStream ordering contract. The shared MockCfg.streamManager now exposes reactivateStream as a first-class mock (was previously absent because no existing test exercised the path), and makeCfg accepts overrides for getStreamInfo, reactivateStream, and initStream.

The tests:

  1. Reactivates when prior stream is 'ended' — captures call order via shared array; asserts ['reactivateStream', 'initStream'] exactly.
  2. Reactivates when prior stream is 'failed' — parity case for the failStream()-produced status.
  3. Skips reactivation when prior stream is 'active' — negative case ensures the guard isn't fired wastefully on an already-live stream.
  4. First-turn execute (no prior stream) — null-stream path; neither reactivate nor unexpected calls should fire on the reactivation branch.

Pins the documented contract so any future refactor that silently drops reactivation in favor of a different mechanism will fail these tests even when the end-to-end coverage (standard-mode-multi-message.integ.test.ts

  • cross-runtime-subagent-resume-matrix-multiturn.integ.test.ts S2 paths 3/4/6) continues to pass for the wrong reason.

Description: packages/runtime-dbos/src/lifecycle/execute.ts:366-388 calls streamManager.getStreamInfo() and conditionally streamManager.reactivateStream() when status is 'ended' or 'failed'. This is the multi-turn continuation reactivation path (without it, the second execute() on a session whose first turn ended would throw Cannot write to stream in 'ended' state).

The path is covered end-to-end by:

  • runtime-dbos/.../integration/standard-mode-multi-message.integ.test.ts:36 — second execute on same sessionId sees prior messages
  • packages/e2e/.../cross-runtime-subagent-resume-matrix-multiturn.integ.test.ts S2 paths 3, 4, 6

…but there is no direct unit test in packages/runtime-dbos/src/__tests__/lifecycle/execute.test.ts that mocks getStreamInfo: () => ({status: 'ended'}) and asserts reactivateStream is invoked before initStream. A refactor that silently drops reactivation in favor of a different mechanism could pass the unit tests while breaking the documented contract.

Action: Add a focused unit test:

ts
it('reactivates ended stream on continuation', async () => {
  const reactivateStream = vi.fn().mockResolvedValue(undefined);
  const streamManager = {
    getStreamInfo: vi.fn().mockResolvedValue({ status: 'ended' }),
    reactivateStream,
    initStream: vi.fn().mockResolvedValue(undefined),
  } as never;
  const stateStore = { /* mock with existing session that triggers
    isContinuation = true */ } as never;
  await executeImpl({ stateStore, streamManager, ... }, agent, opts);
  expect(reactivateStream).toHaveBeenCalledWith('test-session-id');
  expect(reactivateStream).toHaveBeenCalledBefore(streamManager.initStream as never);
});

Effort: ~1 hour. Priority: low — defensive test coverage.


FU-RUNTIME-JS-STATUS-NORMALIZATION: persistViaSaveAndPromote writes AgentStatusValue into SessionStatus column ​

Status: closed — fixed in commit 7c2944b8b (see below). Surfaced by: the DBOS fix agent for FU-DBOS-ERROR-CHUNK (commit 2ff95b0ed, partially reverted in 8f88e5d4d). The agent attempted an in-place mutation fix that broke 4 runtime-js unit tests; reverting left the underlying bug unfixed.

Description: packages/runtime-js/src/run-loop.tspersistViaSaveAndPromote(input) does a structural cast of input.state (typed AgentState<TState, TOutput>) to SessionState<TState, TOutput>. The two types differ on the status field:

  • AgentState.status is an AgentStatusValue ('running', 'waiting_tool', 'completed', 'failed', 'interrupted', 'paused').
  • SessionState.status is a SessionStatus ('active', 'completed', 'failed', 'interrupted', 'paused'). 'running' and 'waiting_tool' both map to 'active'.

The structural cast does NOT convert the field. As a result, saveStateAndPromoteStaging may persist 'running' into the states.status column when the row should hold 'active'. loadState later rejects the value via SessionStatusSchema.parse, producing the ZodError cluster that breaks ~100 of 240 CF workerd tests (see FU-D1-STATUS-SCHEMA-DRIFT).

The prior in-place mutation fix (2ff95b0ed) broke 4 unit tests because sessionState aliases input.state; mutating status also corrupted the live in-memory AgentStatusValue that the runLoop switch reads after the call.

Fix: Build persistedState as a fresh spread object before the saveStateAndPromoteStaging call, converting only the persisted copy's status:

typescript
const persistedState: SessionState<TState, TOutput> = {
  ...sessionState,
  status: agentStatusToSessionStatus(sessionState.status as unknown as AgentStatusValue),
};

Validation:

  • runtime-js: 532/532 pass (all 4 previously-broken completion-retry tests green)
  • store-cloudflare unit: 264/264 pass
  • CF workerd e2e: 225 pass / 6 fail (all 6 are the pre-existing edge-case failures unrelated to this bug)
  • typecheck: 55/55 tasks pass
  • FU-D1-STATUS-SCHEMA-DRIFT also closed by this fix (schema widening from 2ff95b0ed stays; write-side normalization eliminates the primary Zod rejection).

FU-DBOS-ERROR-CHUNK: DBOS failStream doesn't emit AI SDK error chunk ​

Status: ✅ closed — see "Done items" section. Closed by: commit e7506723e.


FU-D1-STATUS-SCHEMA-DRIFT: D1StateStore.loadState Zod schema rejects valid status ​

Status: closed — root cause was FU-RUNTIME-JS-STATUS-NORMALIZATION (write-side bug, not a read-side schema gap). Fixed together in the same commit as that follow-up. The schema widening in state-schemas.ts (from 2ff95b0ed) stays as a defense-in-depth measure to accept both AgentStatusValue and SessionStatus literals in checkpoint state blobs. Surfaced by: npm run test:cloudflare in packages/e2e — most CF integration tests fail with:

ZodError: Invalid option: expected one of "active"|"completed"|"failed"|"interrupted"|"paused"
  at D1StateStore.loadState (packages/store-cloudflare/src/d1-state.ts:2591:37)

The actual rejected value was 'running' — an AgentStatusValue written to D1 by persistViaSaveAndPromote's structural cast without converting the status field. Normalizing at write time (new persistedState object in run-loop.ts) eliminates the rejection.

Verified pre-existing: Reverting packages/store-cloudflare/, packages/runtime-cloudflare/, and packages/ai-sdk/ to baseline 2d5c3d786 reproduces the same failures identically — TII-C1+SUBAGENT matrix work did not introduce this.


FU-MATRIX-DOUBLY-NESTED-HITL: Sub-agent client-tool suspension — documented as unsupported ​

Status: ✅ resolved (documented-as-unsupported; path 3 chosen) — GL #73 Surfaced by: cross-runtime-subagent-resume-matrix-suspended.integ.test.ts S6 scenario in the cross-runtime sub-agent matrix.

Description: When a child sub-agent calls a client-executed tool (via defineTool({ execute: 'client' })), the expectation is that the child's session suspends with status 'suspended_client_tool' and the parent waits for a submitToolResult against the child's sub-session id. In practice, both runtime-js and runtime-temporal return parent status 'completed' immediately — the child's client-tool dispatch does not propagate a suspension up through the ephemeral sub-agent boundary.

Resolution (path 3): Doubly-nested HITL (ephemeral sub-agent calling a client-executed/approval-gated tool) is documented as unsupported in docs/internals/concepts.md (Client-Executed Tools section). The matrix test S6 is now a negative-invariant assertion (not it.todo): it pins the current behavior (parent 'completed', exactly one subagent_start/subagent_end pair) so any future change that starts supporting nested HITL fails the assertion and forces a conscious revisit. Users needing inner-agent human approval should either hoist the client-executed tool to the top-level agent, or use a persistent sub-agent (stable session id, drivable via follow-up turns) instead of an ephemeral one.

Deferred-but-available paths if this is ever scheduled for real support: (1) fix the framework to propagate child suspension across the ephemeral boundary, or (2) add an explicit inheritClientTools option on createSubAgentTool.

Priority: medium — affects users building agents-that-call-agents where the inner agent needs human approval. The single-level HITL case (parent agent's client tool) works correctly across all runtimes; only the doubly-nested ephemeral case is unsupported.


FU-TEMPORAL-CHILD-MAX-STEPS: Temporal child workflows never receive their agent's maxSteps ​

Status: open. Owner: CF-D1 MR 2 (Temporal round). Surfaced by: CF-D1 MR 1a final review (I-5).

Temporal sub-agent / companion child workflows are started from inside the parent workflow (workflow-dispatch.ts spawn / continue / resume sites, workflow.ts:~541), which has no agent registry, so childArgs never carry the child agent's maxSteps. Before DI-28 an absent budget meant "uncapped", so a child's own maxSteps was never honoured; applying the new default there would have capped every child at 50 regardless (a regression for children with maxSteps > 50 or Infinity). MR 1a keeps children uncapped (workflowMaxSteps). Fix: have the activities that plan a spawn / continuation / resume return the child's resolved budget (wireMaxSteps(childAgent.maxSteps)) and put it into childArgs; then drop the child exception in workflowMaxSteps. Add a child-cap parity case.


FU-RESUMECOUNT-WINNER-BUMP: SessionState.resumeCount was not counted on Temporal and DBOS ​

Status: ✅ resolved (S5 snapshot MR). Surfaced by: the S5 rebase onto CF-D1 MR 1a.

resumeCount is documented as the number of resume / retry entries. On origin/main, Temporal and DBOS bumped it with incrementResumeCount BEFORE their instance started (a loser write), for both resume() and retry(); S5 removed those calls ("losers write nothing", workflow ids from the run id) without a winner-path replacement, so it stayed 0. Fix: as on Cloudflare Workflows, the run-start WINNER of an executor resume / retry entry bumps it once in its entry state, never on an activity / step re-execution whose commit already landed, and a losing caller never writes it: Temporal in startRunCommit (TemporalRunStartEntry.countsResume, set for an executor resume(), not for a parent's re-spawn of a parent_suspended child, which main never counted either), DBOS in runStartRunCommit. runtime-js (and the CF DO, which runs its loop) never counted it on main and still does not; docs/guide/querying.md says so per runtime. Tests: runtime-temporal/src/__tests__/resume-count-winner.test.ts, runtime-dbos/src/__tests__/resume-count-winner.test.ts, the real-server Temporal commit-checkpoint-truth.integ.test.ts (concurrent resume / retry, sequential resumes), the DBOS integ dbos-cas-races / c3-single-winner / resume-standard / retry-*, and the e2e resume-count-parity{,-dbos}.integ.test.ts (JS ×3 at 0, Temporal ×3, DBOS).


FU-CFW-ORPHANED-EXECUTE-CLAIM: a crash between a CONTINUATION's claim and create() strands the session ​

Status: ✅ resolved (S5 snapshot MR, rebase onto CF-D1 MR 1a). Surfaced by: CF-D1 MR 1a final review (I-3) and DoD audit (D1).

MR 1a's execute() claimed the session BEFORE workflowBinding.create(), so a Worker cancelled between the claim and create() left a claim with no instance (a new session was reclaimable; a continuation's → active CAS stranded it).

Fix: there is no executor-side claim any more. The single-winner gate is the instance's own run-start commit, and execute() / resume() / retry() write nothing before create() except a write-once createSession for a NEW session (an active row with no run). A cancel there leaves no run, and the next execute() simply runs: its run-start commit finds no live run. Tests: the runtime-cloudflare executor unit case "FU-CFW-ORPHANED-EXECUTE-CLAIM" and the real-engine cfw-terminal-truth.wf-noiso.test.ts "D1" case.


FU-JS-PAUSED-AWAITING-CHILDREN-EXECUTE: runtime-js execute() continues a parent suspended on a sub-agent ​

Status: open. Owner: CF-D1 MR 3 (JS / DBOS round). Surfaced by: CF-D1 MR 1a final review (I-6).

A parent suspended on a sub-agent's pending client tool is paused with an empty pendingClientToolCalls (the wait lives in suspendedAwaitingChildren / suspendedStepId), so runtime-js's paused-with-pending guard lets execute() start a new turn and orphan the child's pending call. CF Workflows refuses this case since MR 1a; add the same check to runtime-js (and so the CF DO) and DBOS.


FU-CFW-E2E-MULTITURN-ROW: add CF Workflows to the e2e multi-turn matrix ​

Status: open. Owner: conformance owner (plan 0c, the CFW executor lane). Surfaced by: CF-D1 MR 1a (CP-88).

CF Workflows multi-turn continuation is proven on the real engine in runtime-cloudflare/src/__tests__/cfw-terminal-truth.wf-noiso.test.ts, but the e2e cross-runtime families (cross-runtime-subagent-resume-matrix-multiturn, client-tool-multiturn-resume) have no CFW combo: the e2e CFW harness adapter (e2e/src/harness/setup-helpers/cfw-workflows-d1.ts) is hand-rolled and builds its own fixed agent__<name>__<sessionId> instance ids instead of driving the real CloudflareAgentExecutor. Add the CFW rows once plan 0c's real-executor CFW lane exists, and retire the adapter's hand-rolled ids.

Unblocked (plan 0c MR-1b): the real-executor env exists: setupCfwMiniflare (e2e/src/harness/setup-helpers/cfw-miniflare.ts, the cfw-workflows-d1 row) runs the real CloudflareAgentExecutor from Node over Miniflare's Workflows binding. The hand-rolled adapter is now the cfw-workflows-d1-pool row. The two families still hand-write their combos, so adding the CFW rows is still open.


FU-MATRIX-MULTITURN: portable cross-runtime multi-turn continuation ​

Status: ✅ resolved for JS / DBOS / Temporal / CF-DO (GL #74) and CF Workflows (GL #109, CF-D1 MR 1a, CP-88/CP-89). Surfaced by: cross-runtime-subagent-resume-matrix-multiturn.integ.test.ts S2 in the cross-runtime sub-agent matrix.

Resolution: AgentExecutor.execute(agent, msg, { sessionId }) now continues a completed session into a new turn portably (Option A — execute() is the continuation verb; no new interface method). A spike refuted the original premise (a Temporal workflow-id collision): Temporal's default WorkflowIdReusePolicy already starts the new run once turn 1 has CLOSED, and initializeAgentState already preserves history. The real blocker was that turn 1's completion left the sessionId-keyed stream in a terminal 'ended' state and the durable runtimes' execute() never reactivated it. Fix: a reactivateStream- on-continuation block, mirroring what runtime-js / runtime-dbos and Temporal's own retry() already do.

Per-runtime status:

  • runtime-js / runtime-dbos: already worked (matrix S2 js/dbos rows).
  • runtime-temporal: fixed — added the reactivation block to execute(); matrix S2 Temporal rows re-enabled.
  • CF-DO: already worked — its DO /start handler delegates to JSAgentExecutor.execute(), which has the reactivation block. Regression-guard test added (proven load-bearing).
  • CFW Workflows: fixed in CF-D1 MR 1a: execute() claims the session before creating the instance, passes entry: 'new' | 'continue', runs each turn in its own write-once instance (…__turn__<n>, keyed on the session's resume counter), reactivates the stream and clears a leftover interrupt flag; getHandle() resolves the live instance. Proven on the real engine under vitest-pool-workers (runtime-cloudflare/src/__tests__/cfw-terminal-truth.wf-noiso.test.ts). The e2e matrix has no CFW row yet: see FU-CFW-E2E-MULTITURN-ROW.

Portable contract floor = completed. A still-running session rejects with AgentAlreadyRunningError (mutex preserved — Temporal via the status === 'running' guard, others via CAS). Continuation from interrupted / failed / paused is not part of the portable contract (use resume() / retry()); runtime-js's broader execute() acceptance set is a documented JS-only superset — tracked separately as GL #110 (decide between dev-warning vs deprecation vs status quo).

Design + plan: docs/superpowers/specs/2026-05-23-cross-runtime-multiturn-continuation-design.md.


FU-JS-EXECUTE-SUPERSET: runtime-js accepts a broader prior-state set than the portable contract — ✅ closed (status-quo) ​

Status: ✅ closed as documented status-quo — GL #110. Decision: do NOT tighten runtime-js (design decision #3); document the superset; pin the portable floor with a cross-runtime test. A dev-warning on non-floor JS execute() continuation remains an OPEN OPTION under GL #110 if ever desired, but is a deliberate UX/behavior change out of scope here. Priority: low Surfaced by: the #74 audit (FU-MATRIX-MULTITURN resolution sweep).

Description: runtime-js's execute() CAS at packages/runtime-js/src/js-agent-executor.ts:1641-1646 accepts prior states ['paused', 'interrupted', 'completed', 'failed'], while the portable cross-runtime contract is completed only. JS code that continues a failed / interrupted session through execute() silently works on JS but breaks when ported to Temporal / CF-DO / DBOS, where the correct verbs are resume() / retry().

Closure (decision: status-quo + floor pinned):

Per the cross-runtime multi-turn continuation design (docs/superpowers/specs/2026-05-23-cross-runtime-multiturn-continuation-design.md, decision #3): runtime-js KEEPS its broader execute() CAS acceptance set as a documented, non-portable superset. We do NOT tighten it — no JS users break — and the framework simply stops promising the broader set cross-runtime. The portable floor is the intersection (completed).

This is now enforced by test, not documentation alone:

  • Portable floor pinned cross-runtime: execute() continuation from a completed prior run runs across every viable backend (js-{memory,redis,postgres}, temporal-{memory,redis,postgres}, dbos-postgres) in the describe.each(getViableBackends({ requires: ['hitl'], excludes: ['cfw-workflows'] })) block at packages/e2e/src/__tests__/session-model-continuation.integ.test.ts ("execute() continuation floor (from completed)"). TEETH: the row fails if any runtime rejected continuation-from-completed.
  • JS superset pinned (JS-only): two JS-runtime unit tests at packages/runtime-js/src/__tests__/js-agent-executor-continuation.test.ts ("execute() continuation SUPERSET (JS-specific, GL #110)") assert JS execute() continues from interrupted (real interrupt) and failed (store CAS transition) prior statuses. TEETH: each fails if JS tightened its CAS source-set — a future tightening becomes a deliberate, test-visible change rather than a silent regression.

The divergence is documented at docs/internals/concepts.md ("Multi-turn continuation" — "runtime-js additionally accepts a broader prior-state set on execute(); that is a JS-specific superset, not a cross-runtime promise") and docs/internals/session-model.md.

Open option (GL #110, NOT in scope here): a dev-mode warning on non-floor JS execute() continuation (gathers real-usage signal before any tighten) — or a full tighten in a major release. Either is a deliberate UX/behavior change; neither blocks real workloads today.


A.2 (Temporal HITL rewrite) follow-ups ​

FU-A2-01: Per-call agent hooks not plumbed through Temporal runtime ​

Status: ✅ done — see "Done items" section. Closed by: commit 2693eb18f (+ changeset da8bd52ea).

FU-A2-02: SUBMIT_TOOL_RESULT_SIGNAL_NAME no-op signal call still in executor.ts ​

Status: ✅ done — see "Done items" section. Closed by: commit 869787e5c.

FU-A2-03: Harden persistent sub-agent edge cases under __resume-N ​

Status: ✅ done — see "Done items" section. Closed by: commit 3ebed60f7.

FU-A2-04: workflow.ts size still 1,589 LOC after rewrite — ✅ closed ​

Status: done (post-release follow-up cluster). Surfaced by: A.2 Task 3.2 implementer's self-review.

Description (historical): Spec target was ~300-400 LOC; original landing was 1,152 LOC, and it had since grown to 1,692 LOC as FU-A2-09 (γ-cascade re-spawn) and other v7-stateless suspension work landed. The bulk of the extra was inlined sub-agent / companion-tool / remote-sub-agent dispatch (~700 LOC across dispatchSubAgents, dispatchCompanionTools, dispatchRemoteSubAgents) plus chunk emission (subagent_start / subagent_end / tool_end) needed to settle SSE consumers, plus the parent-suspended re-spawn loop.

Closure: extracted all three dispatch functions + their context interfaces (SubAgentDispatchContext, SubAgentCall, SubAgentDispatchResult, CompanionDispatchContext, CompanionToolCall, RemoteSubAgentCall) into a new sibling file packages/runtime-temporal/src/workflow-dispatch.ts.

Sizing impact:

  • workflow.ts: 1692 → 1002 LOC (-690 LOC, -41%)
  • workflow-dispatch.ts: 776 LOC (new file, all dispatch + types)

The workflow body itself (mode dispatch + main loop + suspension branches) is now the dominant content of workflow.ts rather than being buried under 700 lines of dispatch helpers.

ESM cycle handling: Both dispatchSubAgents and dispatchCompanionTools need to start child workflows of the same function they were extracted from (agentWorkflow). To avoid a workflow.ts ↔ workflow-dispatch.ts cycle, the dispatchers accept agentWorkflow as a function parameter typed against a structural AgentWorkflowFn = (input: AgentWorkflowInput) => Promise<AgentWorkflowResult> alias declared in workflow-dispatch.ts. The two callsites in agentWorkflow pass agentWorkflow explicitly. dispatchRemoteSubAgents doesn't spawn child workflows so it doesn't need the parameter.

Verification:

  • typecheck: clean across both packages
  • runtime-temporal unit: 245/245 pass
  • runtime-temporal integration (against Docker Temporal): 49/49 pass
  • e2e parity (lifecycle-hooks-parity + client-tool-happy-parity on temporal-memory): 6/6 pass

Behavior-preserving refactor — no test changes, no signature changes on the public surface, no determinism risk introduced. The cycle- breaking parameter pass is the only contract change, and it's purely internal to the runtime-temporal package.

FU-A2-05: Cross-runtime e2e tests need migration from signal-based to durable submit/resume patterns ​

Status: ✅ done — see "Done items" below for closing summary. Surfaced by: A.2 Task 4.4 deeper e2e validation (post-6b5f4b309).

Description: A.2 Task 4.2 migrated packages/runtime-temporal/src/__tests__/ files to the new model. But the cross-runtime e2e files in packages/e2e/src/__tests__/*temporal*.integ.test.ts still depend on pre-A.2 patterns. Notable failures:

  • client-tool-double-submit-temporal.integ.test.ts, client-tool-already-completed-durable-temporal.integ.test.ts — call handle.signal('submitToolResult', ...) directly. Post-A.2 the workflow exits on suspension so the signal hits a completed workflow → WorkflowNotFoundError: Completed workflow. Migrate to use executor.submitToolResult() + durable observation.
  • client-tool-abort-temporal.integ.test.ts — interrupt during HITL test expects status 'interrupted' but gets 'suspended_client_tool'. Same path as C-3 but for a slightly different scenario; needs the same durable-flag-on-resume fix to apply. Possibly fixed already by A.2 commit a907f1698; verify.
  • client-tool-temporal-crash-recovery.integ.test.ts — both tests fail with 'suspended_client_tool' instead of 'completed'. Same root cause as C-5 but in a crash-recovery scenario. The __finish__ terminal-step detection may not be hitting on the recovery path.
  • persistent-subagent-temporal.integ.test.ts — 6 tests failing with expected 'running' to be 'completed'. Persistent sub-agents not completing — likely a workflow-body rewrite issue with persistent sub-agent dispatch (the inlined dispatchSubAgents may have edge cases).
  • client-tool-subagent-ownership-temporal.integ.test.ts — sub-agent ownership routing test fails. Potentially related to sub-agent ownership writes through the new commitSuspendedStep path.
  • remote-subagent-output-persistence-temporal.integ.test.ts — 1 test failing on remote sub-agent + Temporal interaction.

Action: Per-file triage similar to Task 4.2, but for cross-runtime e2e files. Per-file outcomes likely:

  • migrate (replace signal sends with executor.submitToolResult + durable poll)
  • delete (if testing pre-A.2-only behavior covered elsewhere)
  • diagnose-and-fix (if it surfaces a real A.2 regression like persistent-subagent)

Effort: ~3-5 days for triage + per-file migrations + diagnosing the persistent-subagent regression. Priority: HIGH — these are real test failures on the main e2e path. Production functionality may be affected for the persistent-subagent and remote-subagent paths.

FU-A2-07: Sub-agent ownership routing — StaleStateError on parent's commitSuspendedStep ​

Status: ✅ activity-level fix landed (commit landing this entry update). The commitSuspendedStep save now retries on cross-session StaleStateError, reloads, merges clientToolCallOwnership, and re-saves with the fresh version. Confirmed by activities-commit.test.ts "FU-A2-07: retries on StaleStateError and merges with fresh clientToolCallOwnership" + the FU-A2-07 e2e test reaches its later assertions (no longer fails on the first parent commit).

Surfaced by: FU-A2-05 Phase 2 migration (commit 8f3c889bf). File: packages/e2e/src/__tests__/client-tool-subagent-ownership-temporal.integ.test.ts (test still skipped pending FU-A2-09; see below).

Description (closed): When a parent dispatches a sub-agent that suspends with a pending client-tool entry, the parent's commitSuspendedStep save threw StaleStateError: State version mismatch ... expected version 1, but store has version 2. Root cause: the child's own commitSuspendedStep activity calls writeRootOwnership(rootSessionId=parent, ...) which loads + mutates + saves the parent's session row — bumping the parent's version between the moment the parent's commitSuspendedStep activity loaded input.state and the moment it tries to save with expectedVersion: input.state.version.

Fix: bounded retry (5 attempts) inside commitSuspendedStep. On StaleStateError, reload the latest persisted state, merge the clientToolCallOwnership map (associative — child's added entries + parent's locally-computed entries), and re-save with the fresh version.

FU-A2-09: Temporal sub-agent resume after failed:parent_suspended doesn't drive child to completion ​

Status: done — see "Done items" section.

FU-A2-08: Remote sub-agent SubSessionRef metadata not persisted on Temporal ​

Status: ✅ done — see "Done items" section. Closed by: commit 5085756d0.

Effort: ~0.5 day. Priority: medium — affects remote sub-agent observability.

FU-A2-06: AgentWorkflowResult status enum extension is a DTO breaking change ​

Status: ✅ done — see "Done items" section. Closed by: commit e221ac4d3.

FU-A2-38: Re-enable Cloudflare DO real-DO end-to-end tests (G1 + G2) ​

Status: ✅ done — see "Done items" section. Closed by: commit 30432ad34. Surfaced by: v7 stateless-suspension code review #X (expected +0 to be 5 workerd-suite triage).

Description: Two test files under packages/e2e/src/__tests__/ were committed in 9cf39bd66 wip: preserve doc agent work before stash recovery without the supporting test-worker.ts fixtures and wrangler.toml durable-object bindings. They have NEVER passed on this branch:

  • cloudflare-do-persistent-agents.cf.test.ts — references persistentAgentTestLLM import + PERSISTENT_AGENT_SERVER durable object namespace; both undefined at runtime.
  • cloudflare-do-interrupt-protocol.cf.test.ts — references interruptTestLLM import + INTERRUPT_AGENT_SERVER durable object namespace; both undefined at runtime.

Both test files are currently describe.skip with an inline comment explaining the situation (re-grep 9cf39bd66 in those files).

Action:

  1. Add persistentAgentTestLLM + PersistentAgentTestServer DO class exports in packages/e2e/src/test-worker.ts (mirror the subAgentTestLLM pattern).
  2. Add interruptTestLLM + InterruptTestServer DO class exports.
  3. Add the two new DO namespace bindings (PERSISTENT_AGENT_SERVER, INTERRUPT_AGENT_SERVER) to packages/e2e/wrangler.toml.
  4. Remove .skip from the describe(...) blocks in both test files.

Effort: ~1 day. Priority: medium — this is the only real-DO coverage of the v7 interrupt protocol (G2) and persistent companion-spawn (G1) paths. Currently those paths are exercised only via D1-with-JS shims, not the workerd DurableObjectAgentBase directly.

FU-A2-43: Pre-existing DBOS HITL gaps surfaced by lifecycle-hooks-parity + approval-gate-hook-parity + client-tool-happy-parity — ✅ closed ​

Status: ✅ closed — fully subsumed by GL #75 (DBOS approval-gate suspension + onMessage/onStateChange) and GL #111 (Batches A–D), shipped across MRs !201–!206. Every failing row this FU enumerated is now un-skipped on dbos-postgres and passing; see the per-row closure map below. Surfaced by: round-3 CI verification after DBOS_CAPS was tagged with 'hitl' in P1.10 — the parity tests now run for dbos-postgres but fail because DBOS's DBOS.recv() blocking model differs from the durable-state-only suspension model used by JS/Temporal/CF.

Closure — per-row map (verified against the actual test files):

The root cause was structural: DBOS_CAPS lacked the 'hitl' capability tag, so getViableBackends({ requires: ['hitl'] }) silently excluded dbos-postgres from the parity matrix. Tagging it (now at packages/e2e/src/harness/capability-sets.ts:55, DBOS_CAPS) made the three parity files run for DBOS — surfacing the gaps below, all since closed. With the gaps closed, every DBOS-specific skip gate in the three files collapsed to the generic env-only backend.skipReason gate (no runtime-conditional skips remain):

  • lifecycle-hooks-parity.integ.test.ts — client-tool happy path row + tracingContext round-trip row: closed by Batch D (per-call agent.hooks resolution on DBOS, GL #75/#111) + the GL #75 suspension work. Both rows are now gated only by itOrSkip (backend.skipReason ? it.skip : it).
  • lifecycle-hooks-parity.integ.test.ts — approval-gate deny row (previously failed with expected 'completed' to be 'suspended_client_tool' because DBOS did not suspend on approval-gates): closed by FU-DBOS-APPROVAL-GATES (GL #75 — added dispatchApprovalGatedTool). The row now asserts suspended_client_tool and runs on dbos-postgres.
  • approval-gate-hook-parity.integ.test.ts — approve path (all four hooks) + deny path (onMessage only): closed by FU-DBOS-ONMESSAGE-ONSTATECHANGE (DBOS now fires onMessage / onStateChange from the workflow body / executeTool step in the canonical cross-runtime order). The itOrSkipPendingMessageStateGap gate is now backend.skipReason ? it.skip : it (env-only); both rows run on dbos-postgres.
  • client-tool-happy-parity.integ.test.ts — scenario 1 (suspended.toolCallIds result shape) + scenario 2 (exactly-once submit arbiter): closed by GL #111 Batch A. Scenario 3 (client_tool_timeouttool_end chunk): closed by GL #111 Batch B. All three scenarios are gated only by the generic backend.skipReason (the itOrSkipTimeoutOnDbos gate is now identical to itOrSkip).

Orthogonal items NOT part of A2-43's scope (tracked separately): FU-TRACING-LANGFUSE-PENDING-RUNS-OVERWRITE-ON-RECV-WAKE (still open — observability degradation on DBOS Branch 1 recv-wake) and the per-user-message onMessage cross-runtime divergence (FU-USER-INPUT-ONMESSAGE-CROSS-RUNTIME, since closed). These are distinct from the lifecycle/approval/client-tool parity rows A2-43 enumerated and are not gated by A2-43's closure.

Historical detail (pre-closure):

Failing tests (all dbos-postgres only):

packages/e2e/src/__tests__/lifecycle-hooks-parity.integ.test.ts:

  • client-tool happy path: onAgentSuspended + onAgentResumed fire exactly once with linked sessionId — 30s timeout (hooks never observed)
  • tracingContext set by onAgentSuspended persists through resume → completion — 30s timeout (same root cause)
  • approval-gate deny path: onAgentSuspended + onAgentResumed both fire — expected 'completed' to be 'suspended_client_tool' (DBOS does not suspend on approval-gates; resolves immediately)

packages/e2e/src/__tests__/approval-gate-hook-parity.integ.test.ts:

  • approve path fires { beforeTool, afterTool, onMessage, onStateChange } with correct payloads
  • deny path fires onMessage only — beforeTool / afterTool / onStateChange do NOT fire

packages/e2e/src/__tests__/client-tool-happy-parity.integ.test.ts:

  • submitToolResult routes client-tool result back to LLM for next turn — ✅ closed (Batch A, gap 1)
  • exactly one submit wins; conversation reflects the winning result — ✅ closed (Batch A, gap 2)
  • emits tool_end with client_tool_timeout error and continues on next step — ✅ closed (Batch B, gap 3)

Action:

  1. (Short term) The lifecycle-hooks-parity tests are skipped on dbos-postgres via the itOrSkipDbos gate (see packages/e2e/src/__tests__/lifecycle-hooks-parity.integ.test.ts).
  2. (Medium term) Decide whether DBOS should grow a true suspend-on-pending-approval path that mirrors the JS/Temporal/CF semantics, OR refine Capability to expose 'hitl-approval-gate' separately from 'hitl' so DBOS only opts into the subset it can actually serve.
  3. (Long term) Alternatively, add a DBOS-specific test variant that asserts the DBOS.recv() model directly rather than asserting parity with the suspension-based runtimes.

Effort: ~2-3 days for option 2 (capability split) — preferred because it gives DBOS a clean opt-out without losing parity coverage on JS/Temporal/CF. Priority: medium — DBOS HITL works in production via DBOS.recv; the gap is in cross-runtime parity assertions, not user-visible functionality.

Update: the approval-gate suspension portion is now resolved — see FU-DBOS-APPROVAL-GATES below (GL #75). The remaining e2e parity gaps are split to GL #111.

FU-DBOS-APPROVAL-GATES: DBOS approval-gate suspension ​

Status: ✅ resolved for approval-gate suspension (GL #75); e2e HITL parity gaps split to GL #111.

runtime-dbos previously rejected submitToolResult({kind:'approval-response'}) and had no requireApproval dispatch path. Closed by adding dispatchApprovalGatedTool parallel to dispatchClientTool, reusing the shared pendingClientToolCalls map + DBOS.recv(toolCallId) suspension primitive. Function-form requireApproval was initially evaluated inline (fail-closed); it is now step-wrapped/checkpointed as of Batch C (gap 5 — see the Batch C update below). onAgentSuspended/onAgentResumed now fire for approval-gate + client-tool auto-continue. 9 package-level integration tests cover approve/deny/predicate-variants/shape/hooks/persistent-mode.

Deferred (GL #111): the cross-runtime e2e parity tests were blocked by 5 pre-existing DBOS gaps the parity tests assert — suspended.toolCallIds result shape (gap 1, ✅ closed Batch A), exactly-once submit (gap 2, ✅ closed Batch A), client_tool_timeout tool_end chunk (gap 3, ✅ closed Batch B), per-call agent.hooks resolution (gap 4, ✅ closed Batch D), and checkpointing the requireApproval function-form predicate (gap 5, ✅ closed Batch C). All five are now closed.

Batch A update (#111 gaps 1 + 2 closed): the suspended.toolCallIds result-shape gap and the exactly-once-submit arbiter gap are closed in the feat/dbos-hitl-parity-gaps-1-2 MR. ActiveHandle now populates suspended: { toolCallIds } from the durable pendingClientToolCalls map; markPendingSubmittedBestEffort is now the OCC arbiter, so the first concurrent submit observes submittedAt == null and wins (accepted), subsequent ones see the stamped value and return already_completed. The corresponding e2e parity rows (client-tool-happy-parity.integ.test.ts scenarios 1 + 2) are un-skipped on dbos-postgres and pass. Scenario 3 (timeout tool_end chunk) is closed in Batch B (gap 3, see below).

Batch C update (#111 gap 5 closed): the function-form requireApproval predicate is now evaluated inside a @DBOS.step (ApprovalGateStep.evaluateApprovalGatePredicateStep in packages/runtime-dbos/src/steps/approval-gate.ts), so its boolean result is checkpointed in DBOS workflow history. On replay or step retry, DBOS returns the recorded boolean WITHOUT re-invoking the predicate — the suspend-vs-run decision is replay-deterministic even when the predicate is non-pure (clock / random / external reads). Fail-closed semantics (throw → true) live INSIDE the step so the checkpointed value is always a boolean — never a thrown error — and the workflow body never observes the throw. Throws emit a warn with errorId: 'dbos.approval_gate.predicate_threw'. The previous "MUST be pure" caveat on evaluateApprovalGate is lifted; runtime-dbos now matches runtime-temporal's activity-wrapped predicate evaluation. Locked by 7 new unit tests in __tests__/steps/approval-gate-step.test.ts (per-branch return-value contract + fail-closed errorId warn) and 3 new integration tests in __tests__/integration/approval-gate-predicate-checkpoint.integ.test.ts (predicate-called-exactly-once, throw-caught-inside-step, single evaluateApprovalGatePredicateStep entry in DBOS.listWorkflowSteps(workflowID) with the boolean output). The pre-existing GL #75 approval-gate suite (12 tests) passes unchanged.

Batch B update (#111 gap 3 closed): the client_tool_timeout tool_end chunk gap is closed in the feat/dbos-hitl-parity-gap-3 MR. dispatchClientTool now emits a paired tool_end chunk for both outcomes (submit success + timeout failure) via a new emitToolEndStep (@DBOS.step), emitted BEFORE the afterTool hook to match the canonical cross-runtime order (tool_end chunk → afterTool); the timeout tool-result message carries the humanized error text plus the CLIENT_TOOL_ERROR_CODE metadata, matching JS/Temporal/CF. The emitToolEndStep failure path logs errorId: dbos.client_tool.emit_tool_end_failed and never loses the turn outcome. Server-tool failures are unaffected (verbatim error, no error code). The aborted branch in client-tool-resolver.ts is defensive — exercised by the resolver unit tests (client-tool-resolver.test.ts) but unreachable from the DBOS workflow body, because cancellation rejects DBOS.recv rather than firing the abort signal (cleanup runs in lifecycle/interrupt.ts), so dispatchClientTool's tool_end emit never runs for abort. e2e parity scenario 3 (emits tool_end with client_tool_timeout error and continues on next step) is un-skipped on dbos-postgres and passes; all 3 client-tool-happy-parity scenarios now pass on dbos-postgres.

Batch D update (#111 gap 4 closed): per-call agent.hooks (and options.hooks / options.hookManager) now fire on runtime-dbos. Implementation lives in packages/runtime-dbos/src/observability/hook-manager-registry.ts (process-local Map<workflowId, HookManager>) + packages/runtime-dbos/src/observability/merge-hook-manager.ts (mirrors runtime-js/src/js-agent-executor.ts:3065-3108 buildHookManager — constructor → agent.hooks → options.hooks order; options.hookManager parents via createChild()). executeImpl / resumeImpl / retryImpl compute the merged manager BEFORE DBOS.startWorkflow and register under the new workflowId; HookStep.runHookStep and ExecuteToolStep.executeToolStep resolve via HookManagerRegistry.get(DBOS.workflowID) ?? <static fallback> on every step invocation. Cleanup runs in a non-blocking .finally() on dbosHandle.getResult(); start-failure paths delete the entry eagerly. Cross-worker recovery falls back to the constructor-bound static (documented split-brain risk matching Temporal/CF replaceAgent semantics at runtime-cloudflare/src/registry.ts:182-193). Persistent-mode SEND path leaves the existing registry entry alone (subsequent inbox messages still see the first execute()'s per-call hooks); cleanup for persistent workflows is bounded by process lifetime — tracked as FU-DBOS-PERSISTENT-HOOK-CLEANUP below. Locked by 15 new unit tests (__tests__/hook-manager-registry.test.ts, __tests__/merge-hook-manager.test.ts) + 6 new integration tests (__tests__/integration/per-call-hooks.integ.test.ts). Un-skips 2 of 5 newly-viable e2e parity rows on dbos-postgres; the remaining 3 expose orthogonal DBOS gaps (FU-DBOS-RESUME-RUNID-CLIENT-TOOL still open; FU-DBOS-ONMESSAGE-ONSTATECHANGE closed by follow-up — see below — un-skipping the 2 remaining approval-gate-hook-parity rows). All 5 #111 gaps (1, 2, 3, 4, 5) are now closed.

FU-DBOS-RESUME-RUNID-CLIENT-TOOL: resumed runId equals suspended runId on DBOS — ✅ closed ​

Status: ✅ closed — Option D (test relaxation + documented divergence) per the FU's own closure criteria #4. Per-runtime conditional assertion in lifecycle-hooks-parity.integ.test.ts Scenario 1 + documented divergence in docs/internals/concepts.md > "Resume runId semantics".

Symptom: on DBOS, the cross-runtime client-tool resume path detects the existing PENDING workflow (blocked on DBOS.recv) and returns a ResumedHandle wrapping the SAME workflow rather than starting a fresh resume-scoped workflow (lifecycle/resume.ts:139-273 — the documented [CROSS-RUNTIME] branch). The resumed run's runId equals the suspended run's runId. JS / Temporal / CF allocate a new runId per resume. DBOS Branch 2 (resume of an interrupted / paused workflow) DOES allocate a fresh runId, matching the reference runtimes.

Closure:

DBOS Branch 1's previousRunId: runId self-loop at packages/runtime-dbos/src/workflows/shared.ts:1579-1601 is semantically correct: "recv-driven in-place continuation under the same runId". The runId equality is the explicit signal of in-place continuation, not a defect. Closure delivered via:

  • Test relaxation (packages/e2e/src/__tests__/lifecycle-hooks-parity.integ.test.ts Scenario 1): replaced the unconditional resumedCtx.runId !== suspendedCtx.runId assertion with a per-runtime conditional — DBOS asserts equality (in-place continuation), other runtimes assert inequality (fresh resume workflow). The runIds.has(suspendedCtx.runId) AND runIds.has(resumedCtx.runId) invariants on stateStore.listRuns() continue to pass on DBOS via same-key identity. The itOrSkipPendingResumeRunIdGap skip gate was removed.
  • Documented divergence (docs/internals/concepts.md > "Resume runId semantics"): cross-runtime contract for runId / previousRunId semantics. previousRunId is the load-bearing linkage signal for span stitching on both shapes; hook consumers MUST use it (not the runId !== previousRunId inequality).
  • Independent grep evidence: zero downstream production consumers assume runId !== previousRunId. packages/tracing-langfuse/src/langfuse-hooks.ts:521-522 uses previousRunId only as metadata. Parent-span linkage uses tracingContext.lastActiveSpanId / rootSpanId. packages/core/src/hooks/types.ts:375-380 defines previousRunId as optional with no equality contract.

Hook consumer contract:

Tracing-langfuse and audit-hook consumers stitch the resumed span as a child of the suspending span via previousRunId, which is available in the onAgentResumed payload on every runtime:

  • Same-runId case (DBOS Branch 1): Observer sees onAgentResumed({ runId: X, previousRunId: X, sessionId }). Treat as a self-loop "resume of run X in place" — emit a child span keyed on (sessionId, runId) with parentSpan = (sessionId, previousRunId). The self-loop is the explicit signal that this is a DBOS recv-driven auto-continue, not a fresh resume.
  • Distinct-runId case (JS / Temporal / CF, and DBOS Branch 2 — the new-resume-workflow path): Observer sees onAgentResumed({ runId: Y, previousRunId: X, sessionId }). Treat as the canonical "new run resumes prior run" and emit a child span keyed on (sessionId, Y) with parentSpan = (sessionId, X).

The previousRunId field is the load-bearing linkage signal — NOT the requirement that runId !== previousRunId. This contract is now canonical: future hook consumers that depend on the inequality alone should be updated.

FU-TRACING-LANGFUSE-PENDING-RUNS-OVERWRITE-ON-RECV-WAKE — ✅ closed ​

Status: ✅ closed — registerResumedRun now preserves an existing pending entry's accumulated data when one is already registered under params.runId (the DBOS Branch 1 self-loop). Behavior is byte-for-byte identical on the no-collision path (JS / Temporal / CF / DBOS Branch 2 allocate a fresh runId → existing is undefined → fresh entry, exactly as before). Surfaced by: FU sweep #2 brainstorm (revisit of FU-DBOS-RESUME-RUNID-CLIENT-TOOL).

Symptom (pre-fix): on DBOS Branch 1 recv-driven resume (same runId across the suspend/resume boundary — the documented self-loop where runId === previousRunId, see docs/internals/concepts.md "Resume runId semantics"), StatelessSpanEmitter.registerResumedRun did this.pendingRuns.set(params.runId, {...fresh...}), which OVERWROTE the suspended run's existing pendingRuns entry (containing toolInputs, subAgentInputs, the original suspend-time startTime, and original metadata). The suspended entry is still populated at suspension time on DBOS Branch 1 (no finalize happens at suspend; pendingRuns.delete only fires inside emitRootSpan or cleanup()). On JS / Temporal / CF / DBOS Branch 2 there was no collision because params.runId is a fresh key.

Closure: registerResumedRun now reads any existing entry under params.runId and preserves it: startTime: existing?.startTime ?? Date.now(), toolInputs: existing?.toolInputs ?? new Map(), subAgentInputs: existing?.subAgentInputs ?? new Map(), metadata: params.metadata ?? existing?.metadata. sessionId/agentName/input/parentSpanId come from the resume params (the resume is the authoritative source for those). The fix carries a comment citing the concepts-doc "Resume runId semantics" section and this FU name. This preserves the full suspend→resume span-duration window and any in-flight tool/sub-agent input buffered before suspension.

Tests that lock this in: two new unit tests in packages/tracing-langfuse/src/__tests__/suspend-resume-hooks.test.ts (describe('registerResumedRun — DBOS Branch 1 self-loop ...')): (1) same-runId resume preserves accumulated toolInputs/subAgentInputs and original startTime (RED-first against the plain .set(fresh), GREEN after the fix; teeth-verified by reverting to .set(fresh) → preservation test fails); (2) no-collision case — a NEW runId produces a fresh empty entry and leaves the suspended run's entry untouched (behavior unchanged). Package suite green.

Scope note: observability-only (no correctness impact) and only on DBOS Branch 1, but landed as a clean guard to tighten cross-runtime observability parity.

FU-DBOS-ONMESSAGE-ONSTATECHANGE: DBOS does not fire onMessage / onStateChange hooks — ✅ closed ​

Status: ✅ closed — DBOS now dispatches onMessage and onStateChange from the workflow body / executeTool step, matching the cross-runtime canonical order beforeTool → execute → onStateChange → onMessage → afterTool. Surfaced by: packages/e2e/src/__tests__/approval-gate-hook-parity.integ.test.ts approve path + deny path. The previous itOrSkipPendingMessageStateGap gate was lifted on dbos-postgres; both rows now run.

Closure summary: added the missing hook fires in two sites that mirror the JS / Temporal / CF emission points:

  • onStateChange — inside runExecuteTool (packages/runtime-dbos/src/steps/execute-tool.ts), fired when tracker.hasChanges() is true. Payload { previousState, newState, source: 'tool', toolName } matches run-loop.ts:2063 exactly. previousState is a structuredClone of customState captured BEFORE the tool runs. Hook throws are NON-FATAL (parity with run-loop.ts:2072-2081 — logged at warn with errorId dbos.execute_tool.on_state_change_hook_failed, tool result preserved).
  • onMessage — in the workflow body (packages/runtime-dbos/src/workflows/shared.ts), fired AFTER appending: (a) the assistant message, (b) each phase-1 tool-result message (including the approval-gate deny synthetic tool-result), (c) the phase-2 (finishWith) tool-result, (d) the phase-3 (companion) tool-result. Routed through args.dispatcher.runHook so the invocation is checkpointed by HookStep. Hook throws are FATAL (parity with step-iterator.ts:734 / :1235 returning { kind: 'failed' } — the throw propagates out of the workflow body so DBOS retry / recovery surfaces the failure).

Canonical order fix: to land onMessage between onStateChange and afterTool, the afterTool invocation was moved OUT of runExecuteTool and deferred to the workflow body. ToolStepResult now carries durationMs?: number as the deferred-afterTool signal — present on results from runExecuteTool (regular tool / approve-resume tool), undefined on client-tool / deny / unknown-tool results (the client-tool path fires its own afterTool inline; deny never fires afterTool; unknown-tool never reached tool.execute). This mirrors runtime-js's deferredAfterTool pattern (run-loop.ts:2173-2180 + step-iterator.ts:1238-1257) and converges DBOS on the same cross-runtime order as runtime-temporal's in-activity sequencing (activities.ts:545-549).

Locked by: 9 new unit tests (5 in execute-tool.test.ts + 4 in run-agent-loop-one-turn.test.ts, later expanded to 6 in the fix-round with strengthened canonical-order + phase-2 finishWith + phase-3 companion coverage) + 1 new integration test + 2 newly-passing e2e parity rows.

  • Unit: __tests__/execute-tool.test.ts (onStateChange payload + decoupling + non-fatal throw + step-body order); __tests__/run-agent-loop-one-turn.test.ts (assistant onMessage, phase-1 tool-result onMessage, canonical order, fatal throw, phase-2 finishWith onMessage + deferred afterTool, phase-3 companion onMessage).
  • Integ: __tests__/integration/on-message-on-state-change.integ.test.ts (real Postgres; full canonical order assertion with order counter + strictly-monotonic-and-contiguous messageIndex assertion).
  • E2e: approval-gate-hook-parity.integ.test.ts approve + deny rows on dbos-postgres.

GL #111: with this closed, all 5 #111 gaps are resolved AND the DBOS cross-runtime hook surface is at full parity with JS / Temporal / CF for the { beforeTool, afterTool, onMessage, onStateChange } set (for the assistant + tool-result message paths; per-user-message onMessage cross-runtime divergence is tracked separately under FU-USER-INPUT-ONMESSAGE-CROSS-RUNTIME).

FU-USER-INPUT-ONMESSAGE-CROSS-RUNTIME: per-user-message onMessage fires on JS/CF but not on Temporal/DBOS — ✅ closed ​

Status: ✅ closed — joint Temporal + DBOS fix landed on feat/dbos-followups-post-111 (FU sweep #2). No JS / CF code changed — this brought the lagging runtimes up to the reference behavior.

Symptom (before closure): the onMessage hook fires for every appended message on runtime-js and the Cloudflare DO path (via runtime-cloudflare/src/steps.ts) — including each user-input message at the top of a turn — but did NOT fire for user-input messages on runtime-temporal and runtime-dbos. All four runtimes already agreed on assistant + tool-result message fires.

Verified citations (still valid for the reference runtimes):

  • JS fires per-user-message onMessage after the initial append + onAgentStart: packages/runtime-js/src/run-loop.ts:542-551.
  • CF DO fires per-user-message onMessage after the initial append: packages/runtime-cloudflare/src/steps.ts:611-622.

Closure:

  • runtime-temporal (packages/runtime-temporal/src/activities.ts in initializeAgentState — declared at line 4824, fire site immediately after the onAgentStart save block): added a per-user-message onMessage loop using invokeHookWithStateTracking. messageIndex is computed from preAppendMessageCount (read from the durable message log via stateStore.getMessageCount on continuations, derived from conversationMessages.length on fresh sessions) plus the loop index — mirroring the JS reference's state.messages.length - input.newMessages.length formula across multi-message inputs and continuations. Resume entries continue to be handled by applyResultsAndReload (the per-message firing loop ~line 2512 within the function declared at ~line 2086) which fires onMessage for drained tool-result messages only.
  • runtime-dbos (packages/runtime-dbos/src/workflows/shared.ts in runAgentLoopOneTurn, immediately after the onAgentResumed / onAgentStart branch): captured preAppendMessageCount via dispatcher.getMessageCount({ sessionId }) BEFORE appendMessages (replay-safe through LoadStateStep.getMessageCountStep, which delegates to SessionStateStore.getMessageCount — an O(1) indexed count that avoids the getMessages 10k cap), then fires per-user-message onMessage via dispatcher.runHook gated by args.isResume !== true. Persistent-mode coverage is automatic: each runAgentLoopOneTurn call in persistent-workflow.ts (lines 425, 520, 585) is invoked with isResume defaulted (= false), so every inbox-message turn fires per-user-message onMessage — matching how JS fires it on each execute() call. Locked in by packages/runtime-dbos/src/__tests__/integration/onmessage-persistent.integ.test.ts.

Test that locks this in: packages/e2e/src/__tests__/user-input-onmessage-parity.integ.test.ts. The harness-driven matrix runs 4 scenarios (single-message initial, multi-message initial, continuation pre-append-count, resume-path suppression) across 7 Node-context backends (js-{memory,redis,postgres}, temporal-{memory,redis,postgres}, dbos-postgres). Pre-closure: 12 PASS on js-{memory,redis,postgres}, 16 FAIL on temporal-{memory,redis,postgres} + dbos-postgres. Post-closure: 28/28 GREEN. CF DO is the reference implementation (runtime-cloudflare/src/steps.ts:611-622) and is not enumerated as a tested backend here — workerd-context coverage is deferred to the future IMPL_PENDING_O5 lift. Persistent-mode DBOS coverage is locked separately by packages/runtime-dbos/src/__tests__/integration/onmessage-persistent.integ.test.ts.

Changesets: .changeset/temporal-user-input-onmessage.md + .changeset/dbos-user-input-onmessage.md (both patch).

FU-TEMPORAL-HISTORY-PREFIX-PRE-CLAIM-DROP: history-prefix dropped on Temporal fresh sessions because executor pre-claims — ✅ closed ​

Status: ✅ closed — bounded fix in initializeAgentState (Option A from the original fix sketch). The brainstorm-corrected condition is existing !== null && existingMessageCount === 0 (read via stateStore.getMessageCount), not existing.messages.length === 0 — the latter wouldn't typecheck because SessionState has no messages field (messages are stored separately and read through the store's message API).

Surfaced by: FU#2 fix-round (feat/dbos-followups-post-111) attempt at a "fresh session with history: prefix" scenario in packages/e2e/src/__tests__/user-input-onmessage-parity.integ.test.ts. The scenario passed on JS but failed on Temporal with messageIndex === 0 instead of === history.length.

Symptom (pre-fix): when executor.execute(agent, { message, history: [...] }, { sessionId }) ran against a fresh session on runtime-temporal, the history: prefix was silently dropped from the message log AND the per-user-message onMessage fired with messageIndex === 0 (instead of === history.length).

Root cause: Temporal's executor pre-claims the session via stateStore.createSession() at packages/runtime-temporal/src/executor.ts:434-443 BEFORE the workflow starts (this is by design — it closes a startup race with concurrent execute() calls). By the time initializeAgentState activity runs:

  • existing !== null is always true (the session was just created by the pre-claim).
  • The messagesToAppend ternary at activities.ts:5096-5097 takes the newUserMessages-only branch (skipping conversationMessages).
  • The preAppendMessageCount ternary at activities.ts:5115-5118 takes the getMessageCount branch, which returns 0 (the pre-claim created an empty session).

Net effect: conversationMessages (the resolved history prefix) was computed at line 4934 but never used on fresh sessions when the executor pre-claims. The two ternaries were structurally unreachable for the existing === null branch on the user-input entry point.

Closure: runtime-temporal/src/activities.ts initializeAgentState now reads existingMessageCount once via stateStore.getMessageCount (only when existing !== null to avoid an unnecessary round trip) and derives treatAsFresh = existing === null || existingMessageCount === 0. Both messagesToAppend and preAppendMessageCount derive from treatAsFresh — pre-claimed-but-empty rows now take the history-prefix prepend path, with preAppendMessageCount === conversationMessages.length. The existing customState merge / version carry-over branch above (sessionState = { ...existing, ... }) still runs against existing regardless — only the message-log derivation changes. Continuations (real prior messages, existingMessageCount > 0) take exactly the pre-fix path. Side-effect audit: fresh+no-history is unchanged (conversationMessages.length === 0), and the second existingMessageCount read used by preAppendMessageCount is deduplicated against the one feeding treatAsFresh.

Test that locks this in: Scenario 5 in packages/e2e/src/__tests__/user-input-onmessage-parity.integ.test.ts ("fresh session with history prefix"). Asserts (a) the persisted message log starts with the 3-entry history prefix followed by the new user message; (b) onMessage fires exactly once for the new user message with messageIndex === history.length; (c) the run record's startMessageCount (when the store tracks it) is at least history.length + 1. DBOS is skipped pending its sibling FU (FU-DBOS-HISTORY-PREFIX-NEVER-HONORED) because DBOS discards history: entirely at normalizeAgentInput — separate root cause.

Changeset: .changeset/temporal-history-prefix-pre-claim.md (patch bump).

Scope clarification: This was NOT a regression introduced by FU#2 — the history-prefix-drop predated FU#2; FU#2 simply added an onMessage fire site whose messageIndex derivation surfaced the symptom. The onMessage cross-runtime parity (the FU#2 invariant) held on the user-input paths that scenarios 1-4 exercise; scenario 5 was attempting to extend coverage and surfaced this separate bug, which is now also closed.

FU-TEMPORAL-HISTORY-PREFIX-ON-CONTINUATION: Temporal drops caller-provided history: on continuations — ⛔ reframed / superseded ​

Status: reframed — the original "Temporal drops it, CF appends it, match CF" framing was WRONG. Superseded by FU-CROSS-RUNTIME-HISTORY-CONTINUATION-DOUBLING (below). Temporal's drop turns out to be the least-broken behavior; "matching CF" would ship a history-doubling bug into Temporal.

Why reframed (Wave 1 investigation, full 4-runtime read): history: is a fresh-session seeding input (see AgentInputObject.history JSDoc at packages/core/src/types/executor.ts:64-74 — "provide the full conversation history directly" / caller-owns-the-store; and the documented continuation pattern in docs/guide/concepts.md:184-198 passes only { message } and relies on auto-loaded history, with the explicit "stored once per session, not duplicated per run — O(n) not O(n²)" principle). Passing history: on a session that already owns prior messages is undefined / unsupported misuse, and EVERY runtime that honors it on a continuation produces a doubled durable log:

  • JS (packages/runtime-js/src/js-agent-executor.ts:3201-3203, 3246, 3275-3280): on a same-session continuation it appends [...inputHistory, ...newMessages] ON TOP of the already-persisted prior messages → durable log doubled; the in-memory LLM view (:3201 override) also diverges from the store.
  • CF (packages/runtime-cloudflare/src/steps.ts:538-541, 596-597): appends [...providedHistory, ...userMessages] unconditionally; on a continuation-with-history the prior messages are already stored → doubled; onMessage messageIndex (:617) is also miscomputed against the provided-history length.
  • DBOS (post-FU-DBOS-HISTORY-PREFIX-NEVER-HONORED, packages/runtime-dbos/src/workflows/shared.ts): the history-prefix prepend is gated on args.isResume !== true, so a continuation execute() (which is isResume !== true) also prepends → same doubling. (Chosen deliberately for parity with JS/CF on the supported fresh-session path; see that FU's closure.)
  • Temporal (packages/runtime-temporal/src/activities.ts treatAsFresh branch): the ONLY runtime that does NOT double — it drops history: on continuation. Silent, but not corrupting.

So the real defect is cross-runtime history-doubling on continuation, NOT a Temporal-specific gap. Tracked as the FU below.

FU-CROSS-RUNTIME-HISTORY-CONTINUATION-DOUBLING: history: on a continuation doubles the durable message log (JS / CF / DBOS) — ✅ closed ​

Status: ✅ closed — gated the history: prefix on an empty durable log in ALL runtimes. On a non-empty session history: is now a no-op (drop) + uniform <runtime>.history_dropped_on_continuation warn (js. / cf. / dbos. / temporal., carrying { sessionId, droppedCount, existingMessageCount }); the fresh-session seeding path is unchanged. Per-runtime predicates: JS (js-agent-executor.ts) honors inputHistory only when existingMessageCount === 0 on an implicit continuation (the in-memory LLM view now matches the store; branch path unaffected); CF (steps.ts) prepends providedHistory only when the durable log is empty AND fixed the onMessage messageIndex to derive from the durable count (existingMessageCount + conversationMessages.length + i); DBOS (workflows/shared.ts) tightened the gate from !isResume to !isResume && preAppendMessageCount === 0; Temporal (activities.ts) kept its already-correct treatAsFresh drop and added the observable warn. Locked in by the cross-runtime parity test packages/e2e/src/__tests__/history-continuation-noop-parity.integ.test.ts (asserts no doubling + fresh-seed preserved) plus focused per-runtime unit tests (js-agent-executor-continuation.test.ts, runtime-cloudflare/src/__tests__/steps.test.ts, run-agent-loop-one-turn.test.ts, runtime-temporal/.../activities-hooks.test.ts). RED→GREEN teeth-verified: js-memory failed (doubled) before the gate and passes after; temporal-memory passed throughout (it already dropped); reverting the JS gate re-reds the parity test. Changeset: .changeset/cross-runtime-history-continuation-no-double.md. Surfaced by: Wave 1 history-prefix investigation (the read that reframed FU-TEMPORAL-HISTORY-PREFIX-ON-CONTINUATION).

Symptom: executor.execute(agent, { message, history: [...] }, { sessionId }) against a session that ALREADY has persisted messages appends the provided history ON TOP of the existing log on JS, CF, and DBOS — duplicating the prior conversation in the durable store (and, on JS, diverging the in-memory LLM view from the store). Temporal alone drops it (silently). No runtime produces a correct, non-duplicating result. history: is intended as fresh-session seeding (per AgentInputObject.history JSDoc + docs/guide/concepts.md continuation pattern).

Clean cross-runtime fix (the portable contract): gate the history: prefix prepend on "the durable message log is empty" in ALL runtimes — i.e. promote Temporal's already-correct fresh-only predicate (existingMessageCount === 0) into the portable contract, rather than promoting CF's unconditional-prepend (the doubling behavior):

  • JS js-agent-executor.ts continuation path: only seed history when the loaded session has no prior messages.
  • CF steps.ts:538-597: gate the [...providedHistory, ...] prepend on empty existing state; fix the onMessage messageIndex (:617) to use the durable count.
  • DBOS shared.ts: change the prefix gate from !isResume to !isResume && preAppendMessageCount === 0 (DBOS already reads getMessageCount pre-append, so this is a one-line gate tightening).
  • Temporal activities.ts: keep the drop, but make it observable — logger.warn with errorId: 'temporal.history_dropped_on_continuation' when !treatAsFresh && conversationMessages.length > 0 (and, ideally, all four runtimes warn uniformly when history: is passed on a non-empty session).

Lock it in: a cross-store/cross-runtime parity test (in the packages/core/src/testing harness family + an e2e parity row) asserting that history: on a non-empty session does NOT duplicate the durable log on ANY runtime (and either no-ops or warns uniformly). This is the right place to enforce the single portable contract rather than per-runtime conditionals.

Out of scope for Wave 1: this is a focused cross-runtime change touching all four runtimes + a new parity test; it should be its own MR. Wave 1 fixed the supported fresh-session path (FU-DBOS-HISTORY-PREFIX-NEVER-HONORED) and reframed the continuation question to this FU.

FU-DBOS-STARTMESSAGECOUNT-IGNORES-SEEDED-HISTORY: DBOS run-record counts ignore the seeded history: prefix on a fresh session → msg-N UI-id collision on reconnect/SSR — ✅ closed ​

Status: ✅ closed — RECLASSIFIED from "low priority / arguably defensible metadata nuance" to a real UI-correctness bug. The "unverified concern" in the original writeup (does the same pre-run snapshot cause a msg-N id collision?) was VERIFIED to be a real collision. Bounded one-runtime executor change (no workflow-body / replay-determinism change). Surfaced by: Wave 1 — the e2e user-input-onmessage-parity Scenario 5 DBOS row (Assertion 3 on startMessageCount).

Symptom: on a fresh-session execute(agent, { message, history: [...] }, { sessionId }), DBOS recorded run.startMessageCount === 0 AND run.startUIMessageCount === 0, because the executor captures both counts from the store BEFORE DBOS.startWorkflow (packages/runtime-dbos/src/lifecycle/execute.ts, read pre-workflow) while the history prefix + new user message are appended INSIDE the workflow body (runAgentLoopOneTurn, shared.ts). JS/Temporal capture both counts post-append (= history.length + newMessages.length). The persisted message log + the per-user-message onMessage messageIndex ARE correct on DBOS (proven by Scenario 5 Assertions 1+2); only the run-record baseline counts diverged.

Root cause (the real bug): the UI-id derivation keys off startUIMessageCount. On the live/reconnect/SSR path the in-flight assistant bubble's id is deriveExistingMessageId(startUIMessageCount) → msg-${startUIMessageCount} (ai-sdk/src/handler/handle-chat-stream.ts; the same value seeds the SSR partial in snapshot.ts:482-488). The persisted converter assigns msg-${uiIndex} to each UI-visible (non-tool, non-hidden) message in order (helix-to-aisdk-converter.ts). With startUIMessageCount === 0 the bubble id is msg-0, which COLLIDES with the FIRST seeded-history message (also msg-0). AI SDK useChat keys by message id, so the bubble dedupes against the history message and is DROPPED on reconnect/SSR — the same bug class as the already-fixed SF-C1 / Review-1-H2 duplicate-bubble issues; the fresh-session-with-history: path was the remaining hole.

Fix: on the FRESH STANDARD path only (gated on standardStartMessageCount === 0, i.e. genuinely fresh), execute.ts now PROJECTS the about-to-be-seeded history into the run-record counts via a new projectSeededHistoryCounts(input, base, base) helper. The helper parses the input with parseAgentInput (mirroring the body's normalizeAgentInput) and applies the SAME m.role !== 'system' filter as shared.ts's filteredHistory; it returns null (→ keep base counts) when the input doesn't parse as AgentInput or carries no honored history, so the no-history fresh path is unchanged. Both projected counts are base + filteredHistory.length + newMessages.length (JS-parity). The same projection is mirrored on the PERSISTENT fresh-start path (!sessionExisted → its body runs the same runAgentLoopOneTurn), carried ONLY into the START-path initialRunMeta; the persistent SEND/continuation path keeps the unprojected base counts (a live workflow does not re-seed history — its preAppendMessageCount > 0 drops the prefix). The branch path was already correct (it counts messagesPage.total post-clone) and the standard continuation path keeps the SF-C1 UI-merged base counts — neither is touched.

Gate-match + JS-parity reasoning (the careful part): the body seeds iff args.isResume !== true && preAppendMessageCount === 0 && filteredHistory.length > 0. On both fresh executor paths isResume is false and the durable count is 0, so the gate reduces to "history present" — exactly what the helper checks. Value note: the SPEC originally proposed startUIMessageCount = base + history.length (excluding the new user message). That was WRONG and would have left the collision in a different slot: the in-flight assistant bubble's UI index is the count of UI-visible messages that existed when the run STARTED, and the new user message is part of the turn's input (persisted before the run record), so it MUST be counted. JS reads startUIMessageCount = loadUIMessageCount AFTER persisting [...filteredHistory, ...newMessages] (js-agent-executor.ts _initializeForExecute), i.e. history.length + newMessages.length. For a 3-entry history + 1 new user message the correct value is 4 → msg-4, which is collision-free (history = msg-0..2, new user = msg-3) AND matches the committed assistant id. We implemented the JS-parity value (4), NOT the spec's history.length (3), to avoid introducing a NEW divergence. startMessageCount and startUIMessageCount coincide here because none of the seeded rows are tool rows.

Teeth: unit coverage in packages/runtime-dbos/src/__tests__/lifecycle/execute.test.ts (two new tests: "fresh session with history: projects seeded history + new message into runMeta counts" asserts intercepted runMeta.startMessageCount === 3 and startUIMessageCount === 3 for a 2-entry filtered history + 1 message on a fresh session; "…excludes system messages from the projection" asserts a system-bearing 3-entry history still projects to 3, mirroring shared.ts's filter). Verified RED-first (pre-fix → both 0); GREEN after; teeth-verified (reverting the standard-path runMeta wiring back to the base counts → both RED again, expected 0 to be 3; restored → GREEN). The e2e gate in user-input-onmessage-parity.integ.test.ts Scenario 5 un-gated the DBOS row of Assertion 3 (startMessageCount >= history.length + 1) AND added Assertion 4 (no-collision): run.startUIMessageCount === history.length + 1 and the derived msg-${startUIMessageCount} is not any seeded-history msg-N. Verified RED-first on dbos-postgres (pre-fix → expected 0 to be greater than or equal to 4); GREEN after, all 35 backend×scenario runs pass (JS/Temporal already passed the un-gated assertion). Changeset: .changeset/dbos-fresh-history-uimsgcount.md.

FU-TEMPORAL-CONTINUATION-RUN-RECORD: Temporal execute() continuation does not write a run record (listRuns under-reports turns) — ✅ closed ​

Status: ✅ closed (Wave 4) — bounded one-runtime executor change, no workflow-body / replay-determinism change. Surfaced by: Wave 2's execute() continuation floor e2e parity test (packages/e2e/src/__tests__/session-model-continuation.integ.test.ts), whose runs.length >= 2 assertion failed on temporal-{memory,redis,postgres} (got 1).

Symptom: a continuation execute(agent, msg, { sessionId }) on a completed Temporal session does NOT create a second run record. stateStore.listRuns(sessionId) returns 1 run after two execute() turns, vs 2 on JS / DBOS / CF. Root cause: the executor's createRun call is gated behind if (!existingState) (packages/runtime-temporal/src/executor.ts:434, createRun at :488-500), so only the FIRST execute() writes a run record; continuation turns reuse the sessionId-derived workflow id (Temporal's default workflow-id-reuse) and skip createRun. This defeats the F3 intent (observable runs via listRuns) for continuation turns.

Reference behavior: DBOS writes a run record per execute() turn (the DBOS continuation test in session-model-continuation.integ.test.ts asserts runs.length === 2); JS / CF create per-turn run records too. Temporal is the outlier.

Fix sketch: hoist createRun out of the if (!existingState) guard so a continuation execute() writes a run record with a fresh runId and turn = priorRuns + 1 (load prior run count or derive turn from the session). Care needed around: (a) the workflow-body's run-status transition (it currently transitions the run created at start), (b) the runId vs the reused sessionId-derived workflow id, (c) startMessageCount/startUIMessageCount capture for the continuation turn. This is a contained Temporal-runtime change deserving its own MR + 4-battery (turn numbering + workflow-id-reuse interaction).

Scope (historical): Wave 1 gated the floor test's runs.length >= 2 assertion to non-Temporal runtimes (with a comment), keeping it as real teeth on JS/DBOS/CF and excluding Temporal pending this FU. The floor test still asserted the portable contract (message-history continuity) on Temporal, which passed.

Closure (Wave 4): the createRun call in TemporalAgentExecutor.execute() was hoisted out of the if (!existingState) claim guard, so EVERY execute() turn writes a fresh run record — not just the first. The metadata is computed per-branch:

  • Fresh session (!existingState): turn=1, startSequence=0, message counts from the inbound history + newMessages (unchanged behavior).
  • Continuation turn (existingState present): turn = listRuns().total + 1, startSequence = getStreamInfo().latestSequence (live stream head), startMessageCount = getMessages().total, startUIMessageCount = loadUIMessageCount(...) — the same cursors resume() already uses. createSession remains gated to the fresh path (a continuation session already exists).

stateStore.listRuns(sessionId) now reports per-turn runs on Temporal exactly like JS / DBOS / CF (distinct runIds, monotonically-increasing turn).

What this fix restores — and what it does NOT: the hoist restores run-record EXISTENCE + per-turn turn numbering. It does NOT add a terminal run-status transition. The ONLY in-workflow run-status write is commitSuspendedStep.updateRunStatus(input.state.runId, ...), which fires only on a SUSPENSION boundary; it targets THIS turn's runId (threaded via workflowInput.runId → projectSessionToAgent), so a suspending turn transitions its own record independently. A plain completing continuation never hits that path — its run record stays at the initial running status the executor wrote at createRun time. The terminal completion path (persistTerminalState) does not touch run records at all (so there was never a stale-run-transition risk either way). Run records are observability-grade — nothing load-bearing reads the terminal status.

Replay determinism (the careful part): createRun runs in the executor (client side), NOT inside the deterministic workflow body — so there is no workflow-history change and no new non-determinism. The workflow reads input.runId (a plain serialized field) deterministically on every replay. The continuation createRun (and the metadata reads that feed it) is best-effort: failures are logged at warn and do not fail execute(). Because those metadata reads run AFTER startWorkflow, they live INSIDE the same best-effort try/catch as createRun (Waves 3-5 deep-review FIX 1) — a read hiccup degrades to a skipped run record rather than rejecting execute() and orphaning the already-started workflow. This matches the fresh-session and resume() paths.

Teeth: integration coverage in packages/runtime-temporal/src/__tests__/integration/multi-turn-continuation.integ.test.ts ("completes two-turn continuation and sees turn-1 history") now asserts listRuns(sessionId).runs.length === 2, distinct runIds, turn 1 then 2, and runs[1].startMessageCount > 0. Verified RED-first (reverting the hoist → expected 1 to be 2); GREEN after. Unit coverage in packages/runtime-temporal/src/__tests__/executor-run-records.test.ts (the former "does NOT call createRun when continuing an existing session" test, flipped to "calls createRun (turn = priorRuns + 1) on a continuation execute()") pins turn=2 + the live-cursor metadata. The single-turn case is still exactly 1 run (the fresh-session createRun test asserts toHaveBeenCalledTimes(1)). The previously-gated Temporal exclusion in the cross-runtime execute() continuation floor e2e parity test was removed — listRuns(sessionId).runs.length >= 2 now runs on Temporal too. Changeset: .changeset/temporal-continuation-run-record.md.

FU-DBOS-HISTORY-PREFIX-NEVER-HONORED: DBOS executes drop the history: input prefix entirely — ✅ closed ​

Status: ✅ closed — bounded one-runtime fix (initial entry only), no architectural change. Surfaced by: FU sweep #2 brainstorm (Scenario 5 design for FU-TEMPORAL-HISTORY-PREFIX-PRE-CLAIM-DROP).

Symptom (before fix): executor.execute(agent, { message, history: [...] }, { sessionId }) against runtime-dbos silently dropped the history: prefix. normalizeAgentInput (packages/runtime-dbos/src/workflows/shared.ts) extracted only newMessages from parseAgentInput, discarding the history field. The persistent message log started at the new user message; onMessage fired with messageIndex === 0 regardless of history.length. JS / CF / Temporal(fresh) all honored it. The history field IS preserved across the DBOS serialization boundary (plain JSON on StandardWorkflowInput.input) — no serialization gap, purely a normalize-time drop.

Closure: normalizeAgentInput now returns { newMessages, history } (it already had history available from parseAgentInput, which returns { newMessages, initialState, history }). In runAgentLoopOneTurn, on INITIAL entry only (args.isResume !== true), the history prefix is prepended: messagesToAppend = [...historyPrefix, ...userMessages] where historyPrefix = args.isResume !== true ? (history ?? []) : []. The single dispatcher.appendMessages call writes history + new user messages atomically. preAppendMessageCount is still read from the durable log BEFORE the append (timing unchanged) and is now gated on messagesToAppend.length === 0 (was userMessages.length === 0). The resume path is untouched — historyPrefix is empty when args.isResume === true, so resume continues to drain tool-results without re-prepending history.

onMessage scoping (FU-USER-INPUT-ONMESSAGE-CROSS-RUNTIME preserved): the per-user-message onMessage loop fires ONLY for userMessages, never the history prefix (history is prior context, not new user input — matches JS, which fires onMessage only for input.newMessages at run-loop.ts:542-551, not the loaded/provided history prefix). The index math is messageIndex = preAppendMessageCount + historyPrefix.length + i. For a fresh session (preAppendMessageCount === 0) with a 3-entry history, the single new user message fires at messageIndex === 3 === history.length — matching JS's baseIndex = state.messages.length - input.newMessages.length (where state.messages already includes the prepended history, so the formula resolves to history.length). On resume / no-history turns, historyPrefix.length === 0, so the index reduces to the prior preAppendMessageCount + i.

Reference behavior (cross-runtime contract this now matches):

  • runtime-js (packages/runtime-js/src/js-agent-executor.ts:3246): state.messages = [...conversationMessages, ...newMessages].
  • runtime-cloudflare (packages/runtime-cloudflare/src/steps.ts): allMessagesToSave = [...conversationMessages, ...userMessages].
  • runtime-temporal (post FU-TEMPORAL-HISTORY-PREFIX-PRE-CLAIM-DROP): treatAsFresh ? [...conversationMessages, ...newUserMessages] : newUserMessages.

Teeth: unit coverage in packages/runtime-dbos/src/__tests__/run-agent-loop-one-turn.test.ts ("prepends history prefix on initial entry; onMessage fires only for user message at index history.length") asserts (a) the first appendMessages call carries [h1, h2, h3, newUserMsg], and (b) onMessage fires exactly once — for the new user message at messageIndex === history.length, NOT for either historical user entry. Verified RED-first (history dropped → first append carries only the user message → fails); GREEN after the prepend. Reverting the prepend (messagesToAppend = [...userMessages]) re-reds the append assertion; dropping the + historyPrefix.length index offset re-reds the messageIndex assertion. The e2e gate in packages/e2e/src/__tests__/user-input-onmessage-parity.integ.test.ts Scenario 5 (un-skipped on DBOS by this closure) is the cross-runtime contract: history lands in the log, messageIndex === history.length on a fresh session — assertions the unit test mirrors exactly.

FU-DBOS-PERSISTENT-HOOK-CLEANUP: explicit cleanup of HookManagerRegistry entries on persistent-workflow terminate — ✅ closed ​

Status: ✅ closed — single-call fix, no architectural change. Surfaced by: packages/runtime-dbos/src/lifecycle/execute.ts executePersistentImpl start path. Persistent workflows are long-lived (recv-loop until idle-timeout / abort / crash recovery), so the getResult() .finally() cleanup approach used by standard / resume / retry did not apply.

Symptom (before fix): in long-running processes that spawn many persistent agents and let them terminate via idle-timeout, the HookManagerRegistry retained entries for terminated workflows until process exit. Bounded by process lifetime — no production impact for typical short-lived workers; observable as growing-Map memory for long-lived processes managing thousands of persistent agents.

Closure: added HookStep.cleanupPerCallHookManagerStep (@DBOS.step, idempotent because HookManagerRegistry.delete is a Map.delete no-op when the key is absent — safe across replay), called directly from the persistentAgentWorkflow body's try/finally wrapper. The try/finally was hoisted (in FU#4 fix-round H3) to wrap the entire workflow body — beginRun, the LiveModelRegistryHolder lookup, AND runPersistentLoop all sit under the same cleanup umbrella, so the step fires on every terminal path: idle-timeout, shutdown via inbox, shutdown via control, recv cancellation, uncaught errors propagating out of the recv loop, AND pre-loop setup throws from beginRun or the liveModel missing-check. DBOS.workflowID is captured at the top of persistentAgentWorkflow and used directly in the same scope's finally so the cleanup uses the same key that executeImpl used at HookManagerRegistry.set time (no threading through RunPersistentLoopArgs needed). Cross-worker recovery semantics are unchanged: a recovered worker doesn't have the per-call registry entry to begin with and falls back to the constructor-bound HookStep.hookManager, matching the documented split-brain shared with runtime-cloudflare and runtime-temporal.

Teeth: workflow-body assertions live in packages/runtime-dbos/src/__tests__/persistent-workflow-cleanup.test.ts (mocks the DBOS SDK to drive the workflow body and asserts HookStep.cleanupPerCallHookManagerStep is called with the right workflowId on each terminal path), and direct unit coverage for the step body lives in packages/runtime-dbos/src/__tests__/cleanup-step.test.ts (registry delete + idempotency). Commenting out the new await HookStep.cleanupPerCallHookManagerStep({ workflowId }) call in the persistent workflow body locally fails the cleanup-test assertions with the original symptom; restoring re-greens.

Follow-up (cross-runtime, not DBOS-specific): an empty/whitespace submitToolResult({ error: '' }) passes through to an empty tool_end.error string (humanizeClientToolError's default branch returns the code verbatim). This is identical across JS / Temporal / Cloudflare / DBOS, so it should be fixed once in packages/core (either the humanize-error.ts default branch or a .min(1) on the submit error schema). Tracked here; out of scope for Batch B.

Design + plan: docs/superpowers/specs/2026-05-23-dbos-approval-gate-suspension-design.md.

FU-DBOS-PERSISTENT-CT-MULTITURN: persistent recv-loop multi-turn client-tool re-arm — ✅ closed ​

Status: ✅ closed — bounded one-call fix (not architectural). The it.skip is now it and the test passes in ~1s. Surfaced by: packages/runtime-dbos/src/__tests__/integration/client-tools-persistent.integ.test.ts "client tool inside persistent workflow turn — submit resumes, multi-turn flow continues" (was it.skip).

Closure: the original FU framed this as "unbounded design work / architectural" — that framing was overstated. The actual root cause was a single missing reactivateStream call between persistent recv-loop turns; the fix is the persistent-mode analog of the standard-mode reactivate call that already lives at packages/runtime-dbos/src/lifecycle/execute.ts:441-461 (for standard-mode between-execute() continuation).

Actual root cause: turn N's finalizeRun → endStream puts the session-level stream in 'ended' state. Turn N+1's dispatchClientTool then (a) registers the new pending entry at shared.ts:336-344, (b) calls emitToolStartStep at shared.ts:355, which throws Cannot write to stream <id> in 'ended' state from runEmitToolStart (steps/stream-lifecycle.ts:278-298), and (c) the catch block at shared.ts:365-378 rolls back the pending entry via clearPendingClientToolCallStep and re-throws. Net observable: no visible pending for tc-2, no tool_start chunk on the stream — which is what the test's waitForCondition('turn-2 pending tool tc-2') saw as a timeout. (The persist-then-rollback sequence is observationally indistinguishable from a never-persist failure mode, but the fix-round corrected the FU's earlier "never persisted" framing because future debugging must start from the actual call order.)

Fix (commit 0805c755f + 330182bba + fix-round soft-fail in this branch):

  • Added runReactivateStream plain function + StreamLifecycleSteps.reactivateStreamStep @DBOS.step in packages/runtime-dbos/src/steps/stream-lifecycle.ts. The plain function guards on getStreamInfo so it is idempotent and safe to call against an 'active' stream (no-op) — the underlying Redis store throws when called against an active stream, so this guard is load-bearing. Fix-round soft-fails on streamManager.reactivateStream rejection (matches standard-mode .catch(err => warn) at lifecycle/execute.ts:455-460); if reactivate truly fails the next emitToolStartStep throws and the shared.ts:365-378 catch block rolls back cleanly.
  • Added reactivateStream to the StepDispatcher interface (packages/runtime-dbos/src/steps/index.ts) and wired it into all five dispatcher construction sites (standard-workflow + persistent-workflow production factories, plus three test mocks).
  • Called dispatcher.reactivateStream({ streamId }) in packages/runtime-dbos/src/workflows/persistent-workflow.ts immediately before BOTH runAgentLoopOneTurn invocations in the recv loop (drained-message branch at line 436, regular inbox-message branch at line 501). The initial runAgentLoopOneTurn at line 342 does NOT need reactivation — the stream was just initialized fresh by initStreamStep.

Teeth: the integ test now (a) state-store invariants for both turns (pending tc-1 / cleared → tc-2 / cleared, both completedClientToolCalls markers), AND (b) chunk-level invariants: after turn-1 ends, capture streamManager.getStreamInfo(sessionId).latestSequence as the high-water; after turn-2 ends, use createResumableReader(sessionId, { fromSequence: turn1HighWater }) to read only turn-2 chunks and assert a tool_start chunk for tc-2 (the literal write that was failing) plus a text_delta containing 'blue' (turn-2 final text). With reactivate removed locally, the test fails with the original waitForCondition timeout AND would fail the chunk assertion (the tool_start for tc-2 is never written because the throw aborts dispatch). Fix-round also adds call-site assertions on the mock dispatcher.reactivateStream in persistent-workflow.test.ts (no-call on initial turn, one call on drained branch, two calls on regular-inbox 2-message run) with comment-out teeth verification on each.

Cross-runtime symmetry: all three stream-manager backends reject writes to 'ended'/'failed' streams identically (packages/store-memory/src/in-memory-stream.ts:216-226, packages/store-redis/src/redis-stream.ts:303-347 Lua + :602-618 JS-side, packages/store-cloudflare/src/stream-durable-object.ts:328-342). The JS and CF runtimes avoid this DBOS-persistent-mode-only failure because their executors reactivate BEFORE each turn (packages/runtime-js/src/js-agent-executor.ts:1763-1771, packages/runtime-cloudflare/src/executor.ts:985-993) — and DBOS standard-mode does the same at packages/runtime-dbos/src/lifecycle/execute.ts:441-461 (with .catch(err => warn) soft-fail). DBOS persistent-mode was the outlier because it processes inbox-message turns inside the long-lived workflow body, bypassing the executor-level reactivate. The fix adds the reactivate inside the recv loop so persistent-mode now matches the other three runtimes' behavior at the workflow-body level (and matches standard-mode's soft-fail semantics — see fix-round in this branch).

Working coverage (single-turn HITL in persistent mode, retained): persistent-client-tool-interrupt.integ.test.ts, persistent-companion-client-tool.integ.test.ts. Multi-turn standard-mode flows are exercised by standard-mode-multi-message.integ.test.ts and the e2e parity matrix.

FU-DBOS-SEED-BRANCH-PERSISTENT-PATHS: extend seedStateSchemaDefaults integ coverage to branch + persistent-start paths — ✅ closed ​

Status: ✅ closed — dedicated integ coverage added at both call sites.

Surfaced by: integ-fix fix-round (pr-test-analyzer crit 5 follow-on). The helper is wired from THREE call sites in lifecycle/execute.ts — the standard-path entry (~line 494), the branch-from-checkpoint path (~line 227), and the persistent start path (~line 825). The new state-schema-seed.integ.test.ts exercises the standard-path entry comprehensively (spurious-write, dropped-keys warn, legacy-invalid warn); the other two call sites are covered transitively (the helper is the same code), but a dedicated integ test would lock them against future drift if any of the three call sites diverges.

Closed by: two new integ tests in packages/runtime-dbos/src/__tests__/integration/state-schema-seed.integ.test.ts:

  1. Branch-path seeding (branch-path: seedStateSchemaDefaults seeds new schema keys on the branch target (FU#5)) — runs a source session with a narrow stateSchema (only counter), then branches into a new session via executor.execute(branchAgent, ..., { branch: { fromSessionId, checkpointId } }) where branchAgent declares an extended schema (adds newKey: z.string().default('default-value')). Asserts (a) the branched session's persisted customState contains newKey: 'default-value', and (b) an observe tool running INSIDE the branched workflow body sees the seeded default via ctx.getState(). Teeth-verified: commenting out the branch-site seedStateSchemaDefaults call at execute.ts:227 makes the test fail with expected undefined to be 'default-value'.
  2. Persistent-start seeding (persistent-start: seedStateSchemaDefaults seeds defaults visible to tools in the persistent workflow body (FU#5)) — defines a defaultMode: 'persistent' agent whose stateSchema declares persistentKey: z.string().default('seeded-default'), executes it with mode: 'persistent', and asserts an observe tool running INSIDE the persistent workflow body sees the seeded default via ctx.getState(). Uses a 3 s persistentIdleTimeoutMs so the workflow exits cleanly for cleanup(). Teeth-verified: commenting out the persistent-start seedStateSchemaDefaults call at execute.ts:825 makes the test fail with expected undefined to be 'seeded-default'.

Both tests are integration-only and run under RUNTIME_DBOS_INTEG=1; recorded in .changeset/dbos-seed-branch-persistent-coverage.md (patch bump).

FU-GL-113: handle.result() race in client-tools-standard.integ.test.ts timeout test — ✅ closed ​

Status: ✅ closed — auto-resolved by Batch B's standard timeout test rewrite. Tracks GL #113.

Surfaced by: Batch A's _pollForPendingClientTool addition (commit d3fb7b497) made handle.result() return suspended_client_tool at ~50ms before the in-workflow DBOS.recv deadline fired. The pre-existing timeout test asserted result.status === 'completed' directly and failed deterministically against the new poll-driven semantics.

Closed by: Batch B's client-tools-standard.integ.test.ts rewrite (commit e2ba46569, refined in the Batch B fix-round). The test now exercises the canonical cross-runtime suspend→resume pattern:

execute → handle.result() === 'suspended_client_tool'
       → (no submit; deadline elapses inside the live workflow)
       → resume(agent, sessionId) → resumeHandle.result() === 'completed'

Asserts the timeout tool_end chunk (errorCode: 'client_tool_timeout', humanized error text, result undefined) and the tool-result message's CLIENT_TOOL_ERROR_CODE metadata. Header comment at packages/runtime-dbos/src/__tests__/integration/client-tools-standard.integ.test.ts:140-156 documents the "REGRESSION LOCK" narrative. The equivalent e2e timeout-finalization test (agent-server-submit-http.integ.test.ts, Bug #3 in the integ-fix tier-3) was rewritten in the same pattern (commit e2ba46569). All three timeout flows (per-package standard, e2e HTTP, e2e parity scenario 3) now consistently use suspend→resume.

FU-GL-112: v0.7 merge-train pipeline thrash on !199 — ✅ closed ​

Status: ✅ closed — resolved by runner contention easing; no in-code action required. Tracks GL #112.

Surfaced by: the v0.7 merge train when MR !199 (Batches A→B) ran into 6 stuck no_updates_running jobs from runner contention with the parallel MR !200, followed by transient flakes in:

  • test:integ:rest — Redis pub/sub timing under load.
  • test:examples:research-assistant-cloudflare-do — Playwright transient flake.
  • test:unit — turbo task-graph race when ai-sdk#test started before runtime-cloudflare#build finished on a cache miss.

Closed by: runner contention easing (the parallel MR !200 merged; cache-miss/contention conditions cleared). All retries eventually succeeded on the original !199 run. The root cause was CI infrastructure capacity, not an in-code defect — there is no code change that would close it. The individual transient-flake patterns are tracked separately:

  • Redis pub/sub timing is a known intermittent under heavy CI parallelism; absorbed by the standard 3-retry pattern.
  • Playwright DO-example flakes are tracked under the example-app testing discipline (see CLAUDE.md docs/dev/development-workflow.md).
  • Turbo task-graph correctness is enforced by the package dependency graph (CLAUDE.md "Package Dependencies"). The specific ai-sdk#test vs runtime-cloudflare#build race was an order-of-arrival artifact under cache-miss + worker-contention, not a missing ^build declaration.

Out of scope for closure: no individual code change. If similar thrash recurs under future merge trains, options include (a) reducing parallel-MR concurrency on the runner pool, (b) sticky-CPU on the test:unit turbo workers, or (c) bumping the retry count on the affected jobs.

FU-A2-42: Pre-existing runtime-temporal integration test failures (12 tests, mostly conversation continuation + initial-state) ​

Status: ✅ closed — see "Done items" section. Closed by: commit f4a169588.

FU-A2-41: Redis customState pipeline → Lua atomic script ​

Status: ✅ done — landed on branch omnara/stateless-suspension-redesign. Surfaced by: second-round adversarial review #5 finding P2-A. Closed by: atomic SAVE_STATE_ATOMIC_SCRIPT in packages/store-redis/src/redis-state.ts.

Description (closed): RedisStateStore.saveState's customState write previously ran as a non-atomic pipeline AFTER the Lua version check: SADD/DEL/RPUSH for arrays + HSET for scalars + orphan cleanup. A crash mid-pipeline left the session hash at the new version while customState was partially updated. An orphan-recovery loop on the next saveState swept up dangling list entries, but the window was unbounded for paused sessions.

Fix: the version check, main hash field write, and customState replacement now run inside a single EVAL call — SAVE_STATE_ATOMIC_SCRIPT. The script:

  1. CAS-checks expectedVersion against the stored value.
  2. HSETs every main hash field plus the new version in one call.
  3. Applies the session-key TTL.
  4. (When customState is provided) clears the prior customState structures, writes the new baseline as scalars (HSET) + arrays (SADD into the array-keys index, RPUSH into per-key lists), and applies the customState TTLs.

The customState replacement is structural: keys absent from the new baseline are dropped, mirroring the in-memory store's deep-clone-then-apply semantics. The orphan-recovery loop is gone because partially-applied state is now structurally impossible — any crash before the script completes leaves the prior state untouched.

Encoding choice — variadic ARGV instead of cjson: the script uses a flat [expectedVersion, ttl, listPrefix, mainCount, ...kvs, 'apply'|'skip', scalarCount, ...sKvs, arrayCount, ...arrays] layout. cjson would be cleaner but ioredis-mock (used by unit-test files under packages/store-redis/src/__tests__/) ships without a Lua cjson binding. The variadic format works on both real Redis and the mock — eliminating the need for an allowSequentialFallback downgrade on the saveState path.

Verification: packages/store-redis/src/__tests__/integration/save-state-atomic.integ.test.ts covers happy-path round-trip, replacement semantics (absent-key-drop), scalar↔array type transitions, CAS rejection (no partial writes on stale version), 10-way parallel-save serialization (exactly-one-wins; durable state is the winner's complete payload), no-orphan-list-keys after a sequence of array/scalar/delete transitions, and empty-customState clears.

FU-A2-40: CFW Workflows γ-cascade re-spawn ​

Status: ✅ done — see "Done items" section. Closed by: commit 6cbd78808. Surfaced by: second-round adversarial review #6 finding P1.7.

Description: When a parent CFW Workflow suspends mid-sub-agent dispatch, in-flight children are abandoned. Unlike runtime-temporal (which marks children as failureReason: 'parent_suspended' and re-spawns them on parent resume — see runtime-temporal/src/workflow.ts:995-1178, closing FU-A2-09), CFW Workflows has no equivalent path:

  • Suspension commit doesn't mark children with parent_suspended.
  • applyResultsAndReload doesn't detect or surface them.
  • The workflow body doesn't have a re-spawn loop on resume.

After CFW parent resumes, in-flight children are abandoned. The parent loops back to the LLM with stale suspendedAwaitingChildren in state, the LLM produces duplicate sub-agent calls or the session deadlocks waiting for results that won't arrive.

Workaround until fix lands: the workflow body now logs a clear warning when suspendedAwaitingChildren exists after a resume drain (see packages/runtime-cloudflare/src/workflow.ts near "FU-A2-40"). Operators can configure agents that suspend mid-sub-agent dispatch to restart from a checkpoint upstream of the dispatch.

Action:

  1. Add parent_suspended marking in CFW suspension commit (mirror the Temporal commitSuspendedStep flow).
  2. Extend CFW applyResultsAndReload to return childrenToRespawn matching the Temporal contract.
  3. Add a re-spawn loop in runtime-cloudflare/src/workflow.ts resume branch — likely 100-150 LOC adapting the Temporal pattern.
  4. Add a workerd unit test asserting end-to-end re-spawn behavior.

Effort: ~2-3 days. Priority: medium-high — blocks production use of CFW Workflows for any agent that suspends mid-sub-agent dispatch (a common HITL pattern with researcher / writer hierarchies).

FU-A2-39: Lazy-load Node-only harness setup helpers so workerd .cf.test.ts files can reuse harness-driven tests ​

Status: ✅ done — see "Done items" section. Closed by: commit 0d789d40d. Surfaced by: v7 stateless-suspension code review #X workerd-suite triage.

Description: packages/e2e/src/harness/backend-descriptor.ts eagerly imports every setup helper at module top-level:

ts
import { setupJsRedis } from './setup-helpers/js-redis.js'; // → ioredis (Node-only)
import { setupJsPostgres } from './setup-helpers/js-postgres.js'; // → pg (Node-only)
import { setupTemporalRedis } from './setup-helpers/temporal-redis.js';
// ... etc

When a .cf.test.ts file in workerd pulls in the harness (directly or via importing a .integ.test.ts companion), workerd's parser rejects ioredis/pg source at module-load with SyntaxError: Unexpected token ':', before any test code runs.

This blocks the C2-style "re-run the same integ test in workerd via a 1-line import companion" pattern, e.g. lifecycle-hooks-parity.cf.test.ts (currently .skip'd, see file).

Action:

  1. Convert eager imports in backend-descriptor.ts to lazy: register('js-redis', { setup: () => import('./setup-helpers/js-redis.js').then(m => m.setupJsRedis), envRequirements: ['NODE'] }).
  2. Update getViableBackends to honor envRequirements: ['NODE'] the same way it honors ['WORKERD'] — synthesize skipReason for Node-only backends in workerd.
  3. Re-enable lifecycle-hooks-parity.cf.test.ts (remove the describe.skip, restore the import './lifecycle-hooks-parity.integ.test.js').

Effort: ~0.5–1 day. Priority: medium — the .cf.test.ts import-companion pattern is the only way to cover hooks parity for CF DO + CFW Workflows in workerd context. Until this lands, CF backends in cross-runtime parity tests fall back to indirect coverage via dedicated workerd tests (less convenient when adding new hook contracts).

FU-TYPE-SAFETY-2026-05: Eliminate any + forced casts across all packages ​

Status: ✅ closed — see "Done items" section. Closed by: commits 4fd490327 (Stages A + B.1 + B.3), 47e3faa18 (Stages B.2 + B.4 + C + D + ToolContext.getState docs), c94ca7606 (close-out review findings). Surfaced by: four parallel pr-review-toolkit:type-design-analyzer agents (2026-05-10) run after the BC removal pass landed.

Summary: Eliminated any annotations and forced as Type casts across all packages (Stages A through D). The full migration write-up lives in the "Done items" section below; the short version is that ~200+ violations across the framework were resolved through a combination of (a) shared type aliases (AnyTool, AnyAgentConfig), (b) Zod schemas with _DRIFT_ASSERTIONS satisfies _SchemaTypeDriftAssertions compile-time drift guards (so runtime schemas can't drift from the TypeScript types they validate), and (c) typed test-mock helpers replacing as any. See the Done section for per-stage detail.

FU-A2-18: Re-enable usage-tracking.integ.test.ts on DBOS — ✅ closed ​

Status: done (post-release follow-up cluster, alongside Path B Phase 2) Surfaced by: Skipped + TODO audit during exhaustive coverage review.

Description (historical): packages/runtime-dbos/src/__tests__/integration/usage-tracking.integ.test.ts:46 was wrapped in describe.skip by commit be5e793bc (2026-05-07) to silence a "fresh CI node" timing flake — the DBOS workflow's first usage-store entry write intermittently didn't land before executor.result() resolved, surfacing as expected 'failed' to be 'completed' + expected null not to be null assertions.

Closure: the underlying timing flake no longer reproduces. Verified by:

  • Re-enabling describe.skip → describe (no other code changes required).
  • Stress-loop locally: 10/10 runs clean under RUNTIME_DBOS_INTEG=1 against the docker-compose Postgres store.
  • Full DBOS-suite ordering: usage-tracking passes 4/4 in 1.1s even when interleaved with the other 88+ runtime-dbos integration tests in the same vitest process.

⚠ Verification was local-only. The broader runtime-dbos integration suite is paused in CI by the vitest.integ.config.ts RUNTIME_DBOS_INTEG gate (see GL #108). This means CI does NOT exercise this file as a regression guard today. If the original timing flake re-emerges on a CI runner, the test will not catch it until GL #108 is resolved and the suite is un-paused. The script wrapper at packages/runtime-dbos/scripts/run-integ-tests.mjs now prints a loud warning when the suite is silently no-test'd so the pattern doesn't slip past contributors.

The most likely fixers (sequenced earlier in this branch's work train):

  • The #104 ActiveHandle._fetchResult race fix (f9dcabe39, formerly bfa101151) tightened the post-submit / pre-clear ordering in DBOS — adjacent to but not the same as the usage write path.
  • Phase-B Stage-3 / Round-5 DBOS hardening (merged into main via f942fc27f) tightened multiple step-body ordering invariants.
  • DBOS SDK bumps post-2026-05-07.

No code changes were needed to re-enable; the describe.skip was the only blocker.

Action:

  1. Diagnose the flake — likely a beforeAll timeout on cold-start DBOS.
  2. Either bump beforeAll timeout or split heavy setup out.
  3. Replace describe.skip with describe.

Effort: ~1-2 hours (diagnose) + cleanup. Priority: medium — coverage gap on a tier-1 supported runtime.


FU-A2-44: Round-3 P3.R3-CONC partial closure — Redis state-store atomic Lua scripts ​

Status: ✅ partial close — patchMetadata, updateStatus, setInterruptFlag, compareAndSetStatus script hoist all landed. Surfaced by: round-3 review #3 P1.C1 / P1.C3 + P2 cluster.

Closed by these scripts (in packages/store-redis/src/redis-state.ts):

  • PATCH_METADATA_SCRIPT — read + merge + write all in one EVAL. Two concurrent patches with disjoint keys can no longer race (TOCTOU between the JS-side HGET and HSET is gone). Cross-store contract is preserved: per-key last-write-wins; the win is that N parallel patches on N disjoint keys retain all N.
  • UPDATE_STATUS_ATOMIC_SCRIPT — status field write, interrupt-context handling, secondary-index ZREM/ZADD, and TTL refresh all in one EVAL. Concurrent transitions can no longer leave a session indexed in two status:* ZSETs.
  • SET_INTERRUPT_FLAG_SCRIPT — HSET (reason + timestamp) + EXPIRE in one EVAL. Closes the narrow window where a process crash between the HSET and the EXPIRE left the flag without a TTL. Cross-store contract is preserved: last-writer-wins on reason.
  • COMPARE_AND_SET_STATUS_SCRIPT (P3.R3-MISC) — hoisted from a per-call template.join('\n') to a module-level static readonly const. Saves a string allocation per CAS call; makes the script cacheable for EVALSHA if/when we opt in.

Verification: packages/store-redis/src/__tests__/integration/atomic-state-ops.integ.test.ts covers patchMetadata round-trip, additive merging, last-write-wins semantics on conflicts, 20-way parallel-patch retention of all keys, session-not-found error; updateStatus atomic + secondary-index move + interruptContext clear-on-leave; setInterruptFlag TTL atomicity + last-write-wins reason; compareAndSetStatus hoisted-const sequence test + interruptContext preservation on transitions to interrupted/paused.

Round-3 Redis-stream closures included in this batch:

  • INIT_STREAM_SCRIPT — atomic HSETNX status + sequence/createdAt + HSET updatedAt + EXPIRE meta/chunks all in one EVAL (P3.R3-CONC.3). Closes the window where a crash between HSETNX status and EXPIRE left the stream key without a TTL.
  • getStreamCount switched from KEYS pattern to SCAN pattern (P3.R3-MISC). Non-blocking on large keyspaces.
  • NOT_ACTIVE chunk-write rejection now logs at the throw site (P3.R3-OBS) so chunk drops are visible in production traces regardless of caller-side catch behavior.
  • RedisStreamManager constructor accepts a logger option (defaults to noopLogger) and warns at construction when maxChunks: 0 is configured (P3.R3-PERF.6 — operators see the unbounded-stream cliff at startup, not buried in production traces).
  • chunkParseCache moved from a single module-level Map shared across every RedisStreamManager to a per-instance ChunkParseCache with proper LRU eviction (P2 closure). Test isolation improved; steady-state cache hit rate also improved (LRU vs FIFO).

Round-3 batch 4 closures (this commit):

  • deleteSession TOCTOU fix (P3.R3-CONC): the cleanup ZREM list now enumerates ALL valid SessionStatus values (not just the one read back from HGETALL hash.status), so a concurrent updateStatus between the HGETALL and the index Lua can't leave the sessionId orphaned in a status index we didn't touch. ZREM on absent members is a no-op, so over-enumerating is safe. The same fix extends to the per-status run indexes added by P3.R3-PERF.1 below.
  • listRuns({ status }) per-status secondary index (P3.R3-PERF.1): added sessionRunsByStatusKey() helper. createRun ZADDs into the initial 'running' index; updateRunStatus ZREMs old + ZADDs new when status changes. listRuns with a status filter now does an O(log N + page) ZRANGE on the per-status index instead of loading every run and filtering in JS. deleteSession cleans up all per-status run indexes.
  • cleanupOrphanedStagingData pipelining + observability (P3.R3-PERF.5 + P3.R3-OBS): each SCAN page now batches its LINDEX 0 calls into one pipeline and DELs into a single multi-key DEL — turns 2*pageSize round-trips into 2. On completion logs a structured redis-state.cleanup_orphaned_staging.summary with { scannedKeys, deletedExpired, deletedUnparseable, totalDeleted, durationMs, maxAgeMs }.
  • RedisStateStoreOptions.logger?: Logger added (P3.R3-OBS) — defaults to noopLogger. Currently consumed by the cleanup summary log; reserved for future operator-visible warnings.

Still open under P3.R3-CONC (deferred):

  • RedisLockManager fencing-token INCR not atomic with the lock acquisition (low priority — Redlock's mutex still bounds the race per getNextFencingToken call). Tracked here for completeness; not blocking production.

A.2 — Coverage + demo + streaming items migrated from deleted branch-scoped review docs ​

The following items were surfaced during the branch-scoped audit / review / verification cluster (branch-audit-review.md, test-state-deep-review.md, exhaustive-coverage-review.md, exhaustive-fixes-report.md, code-review-v7.md, demo-review-v7.md, doc-audit-v7.md, resumable-streaming-deep-review.md, and the six sub-project-*-verification.md files) — all of which were deleted as part of branch cleanup. These items remained open at the time of deletion; recording them here preserves the trail.

Also tracked in GitLab issues (each FU code below maps into one of the three umbrella issues; the issue body has the full context):

FU-A2-16: CFW Workflows real workerd HITL e2e ​

Status: ✅ closed (real-workerd workflow-body HITL coverage landed; a residual harness-matrix gap remains tracked under O5 — see "Residual" below). Severity (historical): medium-high — tier-1 supported runtime with 100% mocked-step tests Surfaced by: exhaustive-coverage-review.md §2.2. Closed by: Wave 4 test-only commit (this branch).

Description (historical): All CFW Workflows tests used mocked workflow steps; full-execution.workflow.test.ts had every real-workflow test marked describe.skip, and hitl-cycle.workflow.test.ts only exercised the activity surface. The stated blocker — "full workflow-body invocation requires @cloudflare/vitest-pool-workers v0.9.0+ features that aren't available yet" — was stale: v0.9.14 is installed and introspectWorkflowInstance IS available. The real blocker was the miniflare compatibilityDate pin (2024-01-01), which predates the RPC compat surface that backs introspectWorkflowInstance's unsafeGetInstanceModifier binding (workerd warns: needs the rpc flag or compat date ≥ 2024-04-03).

Closure (real coverage now landed):

  1. Bumped packages/runtime-cloudflare/vitest.workflows.config.ts miniflare compatibilityDate to 2024-09-23 (a Workflows-era date), which enables the introspection RPC binding. No regressions in the existing workflow suites.
  2. Added a same-isolate test-injection seam to packages/runtime-cloudflare/src/test-workflow.ts (testWorkflowInjection) so a workerd-context test can register a real agent config + scripted MockLLMAdapter responses BEFORE env.AGENT_WORKFLOW.create(...); the REAL TestAgentWorkflow.run → runAgentWorkflow body then resolves them. (The test runner + main worker share one isolate under @cloudflare/vitest-pool-workers, so module-level state is shared.)
  3. Rewrote packages/runtime-cloudflare/src/__tests__/workflows/hitl-cycle.workflow.test.ts into a real-workerd HITL suite: it drives a REAL Workflow instance via env.AGENT_WORKFLOW.create(...) + introspectWorkflowInstance(...).waitForStatus('complete'), reads the workflow body's returned AgentWorkflowResult, and verifies the full v7 stateless-suspension cycle end-to-end against real D1:
    • client-tool: execute → real LLM step → client-tool register → suspended_client_tool + pending persisted to D1 + onAgentSuspended (1×) → durable submit → create({mode:'resume'}) → drain → completed
      • pending cleared + onAgentResumed (1×).
    • approval-gate deny: execute → suspended_client_tool → deny submit ({approved:false}) → resume → completed.
  4. Un-skipped the real-workflow block in full-execution.workflow.test.ts (real completion path + unregistered-agent error path) and removed the stale "v0.9.0+ not available" note. The only remaining describe.skip there is the legitimate partyserver-StreamServer-headers one.

Teeth-checked: neutralizing the client-tool LLM response flips the result to completed (suspend assertions fail); corrupting the previousRunId expectation fails against the real suspended-run UUID — both assertions are load-bearing, not vacuous.

Residual (tracked under O5, NOT this FU): the coverage above lives in the runtime-cloudflare package and choreographs the two-instance suspend/resume by writing D1 directly (what executor.submitToolResult() does). It is NOT the harness-driven cfw-workflows-d1 BackendEnv row in the e2e matrix — that still requires the full O5 lift: a workerd-resident worker exposing an HTTP/RPC surface for CloudflareAgentExecutor's execute/submitToolResult/resume, wired into setupCfwWorkflowsD1. Until O5 lands, getViableBackends continues to hard-filter cfw-workflows-d1 (IMPL_PENDING_O5) so the harness-matrix tests (lifecycle-hooks-parity.{integ,cf}.test.ts, etc.) skip that row cleanly. See the O5 entry under "D (Test infrastructure overhaul) follow-ups".

FU-A2-17: CF DO persistent sub-agents real-DO test expansion ​

Status: ✅ done — see "Done items" section. Severity: medium — persistent sub-agents only unit-tested with mocks Effort: ~1 day Surfaced by: exhaustive-coverage-review.md §2.3.

Description: cloudflare-do-persistent-agents.cf.test.ts (closed via FU-A2-38 G1) exists but the coverage is narrow. Expand to exercise the full companion-spawn matrix (blocking + non-blocking modes, child failure, parent interrupt during child run, cascading sub-agents) against a real DO instance.

Note: CF DO interrupt protocol test is already closed via FU-A2-38 G2 — do not duplicate.

FU-A2-19: DBOS bug-as-test in persistent-companion-client-tool — ✅ closed ​

Status: done (post-release follow-up sweep) Severity (historical): medium — test asserted buggy behavior Effort: ~30 min (fix + 4 new unit tests + assertion inversion) Surfaced by: exhaustive-fixes-report.md §304 (deleted).

Description (historical): persistent-companion-client-tool.integ.test.ts:600 asserted the buggy DBOS behavior where the persistent-loop's clean exit overwrote a prior turn's status:'failed' run record with status:'completed' + completionReason:'idle_timeout'. The failure trace was lost; no programmatic record (session-level OR run-level) of the LLM error survived once the persistent loop exited.

Closure:

  1. Root cause in packages/runtime-dbos/src/steps/run-lifecycle.ts's runFinalizeRun. The function unconditionally wrote args.status to updateRunStatus, so a later finalize call with status='completed' overwrote a prior status='failed'.

  2. Fix: runFinalizeRun now detects the specific failed→completed downgrade. When the incoming status is 'completed', it reads the current run record via stateStore.getRun(runId). If the current status is 'failed', the call returns without re-writing — preserving the failure trace and the original completionReason.

  3. Why only failed→completed: other transitions are unaffected:

    • failed → failed: fine (refresh metadata)
    • completed → completed: fine (idempotent — re-write run)
    • running → completed: canonical happy path (re-write)
    • running → failed: real failure write (re-write)

    The narrowing keeps the fix scoped to the exact pre-fix bug class.

  4. Test inversion: the integration test at persistent-companion-client-tool.integ.test.ts:600 now asserts status === 'failed' (the correct post-fix outcome).

  5. New unit tests: 4 new tests in packages/runtime-dbos/src/__tests__/run-lifecycle.test.ts's runFinalizeRun describe block pin the new behavior:

    • failed→completed preservation (the fix)
    • completed→completed idempotent path (not gated)
    • failed→failed metadata refresh (not gated)
    • running→completed canonical happy path (not gated)

    runtime-dbos: 372 → 376 unit tests.

FU-A2-31: Migrate 9 client-tool families to shared harness ​

Status: open Effort: ~24 hours Surfaced by: exhaustive-coverage-review.md §6.

Description: Only 4 of 152 e2e files (~2%) use the shared matrix harness. Nine client-tool-family integration suites are still hand-rolled per-runtime — migrate them to harness-driven matrix tests to gain cross-runtime parity for free and reduce hand-maintained boilerplate. Post-merge work.

FU-A2-CF-MISLABELED-FILES: mislabeled CF test files — ✅ closed ​

Status: done (Wave 5 test-infra sweep) Effort: ~2 hours total Surfaced by: exhaustive-fixes-report.md §300.

Description (historical): Test files named cloudflare-*.integ.test.ts (and cloudflare-*.cf.test.ts) claimed runtime-cloudflare coverage but instantiate JSAgentExecutor via the cloudflare-pool-setup.ts / cloudflare-executor-setup.ts helpers (which wire JSAgentExecutor + D1 store, NOT CloudflareAgentExecutor). Root-cause: those setup helpers made JS-executor usage implicit. The established resolution (begun in the v7 stateless-suspension redesign) renames each mislabeled file to a *-d1-with-js.{integ,cf}.test.ts twin and rewrites its header docstring + describe() label to state explicitly that the agent loop runs on JSAgentExecutor and only the store/stream layer is Cloudflare-specific.

Closure: The earlier rename pass created the corrected *-d1-with-js twins but left the original cloudflare-* files in place as committed duplicates — both copies ran (double test runs) and the originals kept perpetuating the exact mislabel this FU is about. This sweep removed the 7 stale duplicates (their *-d1-with-js twins carry identical test bodies + the corrected naming/docstrings — verified by diff: only docstring / describe / comment label lines differed, identical it() titles and counts):

  • usage-cloudflare.integ.test.ts → usage-d1-with-js.integ.test.ts
  • multi-message-execute-cloudflare.integ.test.ts → multi-message-execute-d1-with-js.integ.test.ts
  • tool-attachments-cloudflare.integ.test.ts → tool-attachments-d1-with-js.integ.test.ts
  • ai-sdk-cloudflare.integ.test.ts → ai-sdk-d1-with-js.integ.test.ts
  • ai-sdk-cloudflare-do.cf.test.ts → ai-sdk-d1-with-js.cf.test.ts
  • cloudflare-edge-cases.cf.test.ts → edge-cases-d1-with-js.cf.test.ts
  • finishwith-tools.cf.test.ts → finishwith-tools-d1-with-js.cf.test.ts

Count reconciliation: the FU estimated "8 remaining"; the actual count of mislabeled-files-with-a-corrected-twin was 7 (all deleted here). The remaining cloudflare-do-*.cf.test.ts files are correctly named — they exercise real Durable Objects without the JS-executor helper. Live code-comment cross-references in runtime-cloudflare / runtime-temporal finishwith/workflow tests and usage-subagent-parity.integ.test.ts were repointed at the surviving *-d1-with-js twins (the stale (CloudflareAgentExecutor) annotation on the deleted usage-cloudflare reference was also removed — it was always wrong; that suite used JSAgentExecutor). Historical docs/superpowers/ plan/spec snapshots were intentionally left as-is (point-in-time records).

FU-A2-CF-SETUP-EXPLICIT: Make cloudflare-pool-setup.ts JS-executor usage explicit — ✅ closed ​

Status: done (post-release follow-up sweep) Effort: ~30 min (doc comments only — no rename) Surfaced by: exhaustive-fixes-report.md §301 (deleted).

Description (historical):packages/e2e/src/helpers/cloudflare-pool-setup.ts and packages/e2e/src/helpers/cloudflare-executor-setup.ts both instantiate JSAgentExecutor from @helix-agents/runtime-js, but their file names imply Cloudflare DO / Workflows execution. A reader looking for the CFW Workflows test wiring would land in the wrong place.

Closure: added prominent EXECUTOR CHOICE sections at the top of each file's docstring explaining:

  1. The "cloudflare" in the name refers to the STORE/ENVIRONMENT, not the executor runtime.
  2. Why the JS-executor + Cloudflare-store combination exists: tests that exercise store/stream-manager contracts don't need the workflow-orchestration runtime; the JS executor keeps timing deterministic and failures localized to the store layer.
  3. A "when in doubt" decision tree pointing readers to the right helper for their test's needs (cloudflare-pool-setup for D1+DO under workerd-pool, cloudflare-executor-setup for D1 under Miniflare, or runtime-cloudflare/src/__tests__/ for actual workflow-step semantics).

A rename was considered but rejected — the existing names match the naming convention of other helpers in the same directory (temporal-setup.ts, redis-setup.ts, etc.) and renaming would churn ~30 test files for cosmetic gain. Explicit doc comments are sufficient.

FU-A2-AI-SDK-V7-MIGRATION: AI SDK e2e tests still use v6 API — ✅ closed (stale) ​

Status: done — closed as STALE after an independent re-audit (Wave 5 test-infra sweep). No migration work was required. Effort: ~half-day (estimated) — N/A, premises no longer hold. Surfaced by: exhaustive-fixes-report.md §249.

Description (historical): packages/ai-sdk/src/__tests__/e2e/* was said to still use the deprecated v6 API; the ask was to migrate to v7 handleChatStream so the suite exercises the current public surface.

Audit / closure: every premise of this FU is now false:

  1. The referenced directory does not exist. There is no packages/ai-sdk/src/__tests__/e2e/. The ai-sdk tests live under __tests__/ (unit) and __tests__/integration/ (*.integ.test.ts).
  2. The suite already exercises the current v7 surface. 28 test files import/drive handleChatStream, StreamTransformer, createReplayEvents, and createHelixChatTransport directly — the canonical v7 chat-stream/transformer API. The pre-v7 frontend-handler API (FrontendHandler / createFrontendHandler / handleFrontendRequest) was removed in the v8 redesign (see CHANGELOG) — it no longer exists to migrate off of. The pre-v7 frontend-handler API (class/factory/request-handler) has no definition or usage in ai-sdk source or tests; the name FrontendHandler survives only in JSDoc describing the new replacements.
  3. The remaining v6 / useChat / tool-input-end references are descriptive or negative, not real deprecated-API usage:
    • "AI SDK v6" mentions refer to the Vercel AI SDK v6 UIMessage format (orthogonal to the Helix v6→v7 redesign this FU implied), all in JSDoc.
    • useChat mentions describe the framework's own useHelixChat as a "drop-in replacement for useChat", and transport-submit-tool-result-server-side.test.ts specifically validates the non-useChat server-side dispatcher path.
    • tool-input-end appears only in comments/assertions confirming the transformer emits the NEW tool-input-available event instead of the deprecated tool-input-end — the deprecated event is never emitted or asserted-present.

No genuine deprecated v6 call exists to migrate; the suite is already on the v7 public surface. Closed with no code change.

FU-A2-HARNESS-SMOKE-WORKERD: harness-smoke.integ.test.ts:69-71 workerd assertion bug — ✅ closed ​

Status: done (post-release follow-up sweep) Effort: ~20 min (context guard + comment) Surfaced by: exhaustive-coverage-review.md §321 (deleted).

Description (historical): The cf-do-d1 and cfw-workflows-d1 are absent from Node context results assertion in packages/e2e/src/harness/__tests__/harness-smoke.integ.test.ts holds in Node context (where the file currently runs via the .integ.test.ts extension), but would fail if mirrored to a .cf.test.ts companion — under workerd those backends ARE present because their setup helpers can load there.

Closure: added a runtime context-guard so the assertion only runs when typeof process.versions.node === 'string' (i.e., Node context). A future .cf.test.ts mirror would now see the test correctly skip rather than produce a misleading green result or a mystery failure. Also expanded the surrounding comment to make the context-specificity explicit.

The fix is intentionally defensive — there's no immediate workerd mirror, but the brittle assumption that "this file is Node-only" no longer lives implicitly in the file extension.

FU-A2-LIFECYCLE-HOOKS-E2E: onAgentSuspended/Resumed per-runtime e2e ​

Status: ✅ closed — all runtimes now have lifecycle-hook coverage, including the previously-missing CFW Workflows leg. Surfaced by: exhaustive-fixes-report.md §299. Closed by: Wave 4 test-only commit (this branch), alongside FU-A2-16.

Description (historical): The lifecycle hooks were silently broken on runtime-js and runtime-temporal until this branch (see fixed items #9 and #10 in the deleted exhaustive-fixes-report). Per-runtime e2e coverage was needed across all runtimes to prevent regression.

Closure (per-runtime status):

  • JS / Temporal / DBOS / CF-DO: covered by the harness-driven matrix in packages/e2e/src/__tests__/lifecycle-hooks-parity.integ.test.ts (getViableBackends({ requires: ['hitl'] })) + the .cf.test.ts companion for the CF-DO workerd leg. onAgentSuspended / onAgentResumed are asserted exactly-once with the client_tool discriminator, the previousRunId linkage (load-bearing on every runtime), and a runtime-conditional runId-equality assertion (DBOS Branch 1 self-loop vs fresh-runId runtimes).
  • CFW Workflows: NOW covered by packages/runtime-cloudflare/src/__tests__/workflows/hitl-cycle.workflow.test.ts (real workerd via @cloudflare/vitest-pool-workers + introspectWorkflowInstance — see FU-A2-16 for the mechanism). Both hooks fire exactly-once across the real two-instance suspend/resume cycle, and onAgentResumed.previousRunId === suspendedCtx.runId is asserted against the real suspended-run UUID (teeth-checked).

The CFW Workflows leg is NOT wired into the shared getViableBackends matrix (the harness cfw-workflows-d1 BackendEnv row stays IMPL_PENDING_O5 — see FU-A2-16 "Residual" + the O5 entry), because the matrix needs a CloudflareAgentExecutor HTTP/RPC surface inside workerd. The package-level workflow test exercises the same hook contract directly against the real workflow body, so the contract is covered even though the matrix row remains gated.

FU-A2-SUB-AGENT-USAGE-PARITY: Sub-agent usage scoping parity (Temporal/CF/DBOS) ​

Status: ✅ closed (matrix parity test added; the one runtime divergence that was temporarily gated has since been root-caused + fixed under FU-TEMPORAL-POSTGRES-SUBAGENT-OVERSTEP, so all viable backends now run). Effort: ~half-day Surfaced by: exhaustive-fixes-report.md §299 (P2 list).

Description: Sub-agent token-usage scoping (whether child usage is attributed to parent's run vs reported separately) is verified on runtime-js but not on Temporal/CF/DBOS. Add parity tests.

The portable contract (the "Residual E" usage-scoping model from packages/runtime-js/src/__tests__/usage-scoping.integ.test.ts): when a parent agent spawns a child sub-agent and the child records usage,

  1. the parent's DIRECT rollup (getRollup(sessionId, { includeSubAgents: false })) shows the sub-agent dispatch but NOT the child's tool call / custom metrics (those are scoped to the child's session);
  2. the child's usage lives under the child session, discoverable via the parent's subagent usage entry subSessionId; and
  3. the AGGREGATED rollup (getRollup(sessionId, { includeSubAgents: true })) surfaces the child's custom metrics summed exactly once (no double-count, no omission).

Closure: new matrix-driven parity test packages/e2e/src/__tests__/usage-subagent-parity.integ.test.ts, driven by getViableBackends({ requires: ['sub-agents'] }), pins all three invariants against a single shared script. Verified locally (Docker Redis + Postgres + Docker Temporal dev server) on six backends — js-{memory,redis,postgres}, temporal-{memory,redis}, dbos-postgres: 6 passed. The scoping contract held identically on every passing backend (the read path goes through usageStore.getRollup(...) directly, mirroring the per-runtime usage suites, to avoid the executor-wiring asymmetry where some harness setup helpers don't thread usageStore into the executor they build, so handle.getUsage() would return null). Teeth-checked: neutralizing the child's recordUsage() calls makes the child-rollup + aggregated-rollup assertions fail (expected undefined to be 1), proving the aggregation assertions are not vacuous. cf-do-d1 / cfw-workflows-d1 are in the matrix but skip in Node (workerd-only, IMPL_PENDING_O5); CF DO sub-agent usage scoping is exercised via usage-d1-with-js.integ.test.ts (D1UsageStore driven by the JS executor).

One backend gated: temporal-postgres is in the matrix but SKIPPED — its sub-agent → parent-finish step loop genuinely diverges (the parent over-steps past its FINISH output). This is NOT a usage-scoping bug (where the run completes, the child rollup is correctly scoped); it's a Temporal × Postgres-state step-loop divergence tracked separately as FU-TEMPORAL-POSTGRES-SUBAGENT-OVERSTEP below.

FU-TEMPORAL-POSTGRES-SUBAGENT-OVERSTEP: parent over-steps past FINISH after a sub-agent completes (Temporal + Postgres state) — ✅ closed ​

Status: ✅ closed Severity: medium — a parent → sub-agent → parent-finish flow did NOT reach a clean completed status on the temporal-postgres backend. Surfaced by: the usage-subagent-parity.integ.test.ts matrix (FU-A2-SUB-AGENT-USAGE-PARITY).

Description: On the temporal-postgres backend (TemporalAgentExecutor + PostgresStateStore), after a child sub-agent completed and its result was recorded on the parent session, the parent's next LLM step threw and the run failed (status: 'failed') instead of completing. Confirmed deterministic and store-specific — temporal-memory / temporal-redis completed the SAME script in four LLM calls, and js-postgres / dbos-postgres also completed cleanly.

Root cause (NOT a usage-scoping or result-reload defect):PostgresStateStore.saveState silently dropped parent_session_id. Both saveState UPDATE branches (the skipCheckpointCreation path and the checkpoint-creating path) wrote root_session_id but omitted parent_session_id — only saveStateAndPromoteStaging had been fixed in an earlier cross-store-parity pass (code review #8 finding I1); saveState was the remaining outlier (the in-memory, Redis, and D1 stores all round-trip the field).

The Temporal child sub-agent's initializeAgentState persists its parentSessionId via saveState. With the dropped column the child loaded back from Postgres with parentSessionId: undefined, so when the child completed, persistTerminalState's !state.parentSessionId guard mistook the child for a ROOT session and finalized (endStream) the SHARED parent stream. The parent's next runLLMStep then threw Cannot write to stream … in 'ended' state from the LLM stream callback (after generateStep had already consumed the FINISH response). Because runLLMStep runs as a Temporal activity with a 3-attempt retry policy, Temporal re-ran the step — consuming an extra mock LLM response — and the run failed. temporal-memory / temporal-redis round-trip parentSessionId, so their children correctly skip shared-stream finalization; js-postgres / dbos-postgres drive the child stream lifecycle differently and never depended on the loaded child parentSessionId here.

Fix: add parent_session_id to the UPDATE column list (and parameter array) in BOTH saveState branches of packages/store-postgres/src/state/postgres-state-store.ts, mirroring the already-correct saveStateAndPromoteStaging.

Teeth:

  • Cross-store contract test (packages/core/src/testing/state-operations.ts → "should round-trip v7 suspension fields through saveState") now also asserts parentSessionId round-trips through saveState. Verified RED on Postgres without the fix (expected undefined to be 'parent-sess-xyz'), GREEN with it; memory / Redis / D1 already passed.
  • The isTemporalPostgres gate in usage-subagent-parity.integ.test.ts was removed — temporal-postgres now runs with full teeth (all 7 viable backends GREEN).

Changeset: .changeset/temporal-postgres-subagent-overstep.md (@helix-agents/store-postgres patch).

FU-A2-DO-STREAM-MANAGER-PARITY: DOStreamManager round-trip parity test — ✅ closed ​

Status: ✅ closed (Wave 3) — reframed: the runtime-cloudflare DOStreamManager ALREADY had a full contract-harness parity test (packages/runtime-cloudflare/src/__tests__/do-stream-manager-contract.cf.test.ts, real workerd via cloudflare:test). The genuine remaining gap was the SEPARATE binding-side store-cloudflare DurableObjectStreamManager (stub.fetch + SSE consumer), which had bespoke unit tests but never ran the shared streamManagerContractTests factory. Wave 3 added that parity test (packages/store-cloudflare/src/__tests__/durable-stream-contract.integ.test.ts, real StreamServer DO in Miniflare), which both fills this parity gap AND verifies the G4 fix (FU-G4-STORE-CF-STREAM-DO). All G1–G5 contract guarantees are now asserted against the binding-side manager. Effort: ~2 hours Surfaced by: exhaustive-fixes-report.md §299 (P2 list).

Description: No round-trip parity test exists for DOStreamManager (write → read → assert chunk sequence integrity). Add one mirroring the in-memory/Redis parity tests.

FU-A2-EXPIRED-SESSION-CROSS-STORE: expiredSessionCleanup extension to Temporal+CF ​

Status: Temporal portion ✅ closed; CF DO portion open (gated on the O5 workerd-harness lift). FU-A2-21 closed the cross-store parity for memory/Redis/Postgres/D1. Effort: ~2 hours Surfaced by: exhaustive-fixes-report.md §299 (P2 list).

Description (original): Extend the expiredSessionCleanup parity test to cover Temporal + CF DO runtimes — previously only memory + the three remote stores had it (per FU-A2-21).

Why this is portable: expiredSessionCleanup (packages/agent-server/src/cleanup.ts) is runtime-agnostic — it touches ONLY the SessionStateStore abstraction (listSessions → loadState → compareAndSetStatus) and never invokes the executor. So the contract ("identify past-expiry sessions, CAS them to failed/session_expired, leave future-expiry and terminal sessions untouched") is the same on every backend the harness can stand up; the runtime is just the vehicle that constructs (and, on Temporal, writes through) the store.

Closure (Temporal): the parity test (packages/e2e/src/__tests__/expired-session-cleanup-parity.integ.test.ts) broadened its backend filter from getViableBackends({ only: ['js'] }) to getViableBackends({ excludes: ['dbos', 'cfw-workflows'] }). In Node context this resolves to six backends — js-{memory,redis,postgres} + temporal-{memory,redis,postgres} — so all five scenarios run on Temporal (30 logical runs). Verified locally against Docker Redis + Postgres + the in-process Temporal test environment: all 30 pass. No runtime divergence was found — the store-abstraction assertion is fully portable, and the Temporal harness's store wiring persists expiresAt / status identically to the JS path.

Remaining gap (CF DO): this suite is not yet wired to run on cf-do-d1 in the workerd pool. O5 Phase 1 landed real CF setup helpers and made cf-do-d1 a live parity participant for the lifecycle-hooks-parity suite, but it did so via a context-split: the Node and CF backend registries are now separate (NODE_BACKENDS vs CF_BACKENDS) so the workerd bundle never statically references Node-only loaders. This integ file drives its matrix via getViableBackends, which reads NODE_BACKENDS only — so it can never surface cf-do-d1.

A previous .cf.test.ts companion that merely re-imported this integ file relied on the pre-O5 assumption that cf-do-d1 lived in the single shared registry with an IMPL_PENDING_O5 skip marker (and would auto-activate once the marker dropped). The context-split invalidated that assumption: after O5 the re-import resolves to zero viable backends in workerd, registers zero suites, and fails collection (No test suite found). That dead companion was removed (fix-forward after the O5 Phase 1 merge); expiredSessionCleanup cross-store coverage continues to run Node-only via the integ file (six js/temporal backends).

When this suite is CF-wired in a later O5 phase, it will follow the lifecycle-hooks-parity pattern — a Node-import-free shared scenario module plus a dedicated CF entrypoint over selectViable(CF_BACKENDS, …) — NOT a re-import companion. That also requires the cf-do-d1 harness SessionStateStore shim to implement listSessions / compareAndSetStatus (the O5 Phase 1 shim only covers the lifecycle-hooks scenario's needs). The DOStateStore cross-store contract is exercised today by the dedicated workerd-context .cf.test.ts suites under packages/store-cloudflare / packages/e2e, just not yet through this parity harness. Tracked here + under O5 in docs/dev/test-infrastructure-roadmap.md.

Scope note: dbos and cfw-workflows are excluded by intent. DBOS would only re-exercise the postgres store already covered by js-postgres / temporal-postgres, and CFW Workflows shares CF DO's not-yet-CF-wired harness gap for this suite with no distinct store to add. The FU's explicit ask was Temporal + CF DO.

FU-A2-CF-TEST-MIRROR: C-1..C-5 contract tests need *.cf.test.ts mirrors ​

Status: ✅ closed (Wave 5 — two genuinely-missing mirrors landed; two already covered; one reframed as redundant, see mapping below). Effort: ~half-day Surfaced by: exhaustive-fixes-report.md §299 (P2 list).

Description: Five contract tests (C-1..C-5 in the deleted coverage review) run in Node context only. Mirror them to .cf.test.ts for real workerd-context coverage.

Closure. The "C-1..C-5" labels map to the five shared-contract-factory suites that the Cloudflare store packages run. Mapped against what already had real-workerd coverage:

#Contract suite (shared factory)Implementation under testNode-context suiteReal-workerd .cf.test.ts?
C-1stateStoreContractTestsD1StateStorestore-cloudflare/.../integration/state-store-contract.integ.test.ts (Miniflare)ADDED — store-cloudflare/.../state-store-contract.cf.test.ts
C-2usageStoreContractTestsD1UsageStorestore-cloudflare/.../integration/usage-store-contract.integ.test.ts (Miniflare)ADDED — store-cloudflare/.../usage-store-contract.cf.test.ts
C-3streamManagerContractTestsDOStreamManager (in-process DO)—already real-workerd: runtime-cloudflare/.../do-stream-manager-contract.cf.test.ts (Wave 2)
C-4streamManagerContractTestsDurableObjectStreamManager (binding-side)store-cloudflare/.../durable-stream-contract.integ.test.ts (Miniflare, real StreamServer DO)reframed redundant — see below
C-5(no distinct CF store)——n/a

What was added (genuinely missing). store-cloudflare had NO cf-pool harness at all (no vitest.cf.config.ts, no wrangler.cf.toml, no test:cloudflare script). Added that infra (mirroring runtime-cloudflare's exactly, NOT a parallel invention) plus the two D1 mirrors. Each runs the SAME shared factory from @helix-agents/core/testing against the D1Database bound by workerd itself (env.AGENT_DB via cloudflare:test), not Miniflare's Node-side D1 proxy. runMigration / table-clearing are inlined in the .cf.test.ts (the Node helper miniflare-setup.ts imports miniflare, unavailable in the workerd pool). A D1Database binding is NOT DO-IO-restricted, so — unlike the in-process DOStreamManager mirror — the store + every assertion run directly in the worker context with no runInDurableObject proxy. Result: 157 tests (137 state-store + 20 usage-store) green on first run in real workerd, no skipKnownFailures, no opt-outs.

Teeth verified. (1) Skipping runMigration makes all 20 usage tests fail with no such table: __agents_usage: SQLITE_ERROR from workerd's internal cloudflare-internal:d1-api — proving the test hits the live workerd D1, not a no-op binding. (2) Wrapping the store in a Proxy that returns a wrong queryUsage total (999) makes 20/20 fail with meaningful diffs (expected 999 to be 3, wrong lengths, wrong pagination caps) — proving the contract assertions catch a semantic violation against the workerd-bound store. Both sabotages reverted; suite green.

Why C-4 is NOT given a separate .cf.test.ts (reframed, not skipped). The binding-side DurableObjectStreamManager mirror (durable-stream-contract.integ.test.ts) already runs the contract against the real partyserver StreamServer Durable Object inside Miniflare (createRealStreamContext() loads the bundled worker; the inline test-double is not used). For a DO accessed via its HTTP/SSE binding surface, that path is real-workerd-equivalent: the DO body executes in workerd, and the manager talks to it over the same fetch/SSE surface it uses in production. A separate .cf.test.ts would re-host the identical StreamServer + contract with no new code path exercised — a vacuous duplicate. The distinct in-process DOStreamManager surface (the part that genuinely benefits from runInDurableObject workerd hosting) is already covered by C-3's do-stream-manager-contract.cf.test.ts. This mirrors the Wave-3 FU-A2-DO-STREAM-MANAGER-PARITY reframing discipline (reframe rather than blindly produce a no-value duplicate). The push-SSE late-attach G4 nuance on this binding path is separately tracked by FU-G4-BINDING-AFTER-ENDSTREAM-RETURNS-DONE.

C-5 has no distinct Cloudflare store implementation to mirror (the fifth slot in the deleted review did not correspond to a CF store contract); nothing to add.


Demo (Round-3 web-app-testing) follow-ups ​

FU-DEMO-01: remote-agents-temporal/approve.ts v7 migration — ✅ closed (verified migrated) ​

Status: done (verified during post-release follow-up sweep — the migration already landed in an earlier work train and the example typechecks + builds clean). Severity: WAS P0 for that example; the original audit predated the actual migration. Effort: ~5 min (verify + clean up stale comment). Surfaced by: demo-review-v7.md §232-242 (deleted; original audit).

Closure:

  1. examples/remote-agents-temporal/src/approve.ts — already uses the v7 durable-submit pattern. The file's own docstring (lines 14-17) explicitly calls out the v6→v7 migration: "v6 used handle.signal('submitToolResult', payload) — that signal was deleted in v7's stateless-suspension redesign (Task 3.3) so this script HAD to migrate." The body now uses executor.submitToolResult({ kind: 'client-tool-result', ... }) followed by executor.resume(), with proper handling of unknown_tool_call and already_completed responses.

  2. examples/remote-agents-temporal/src/client.ts — the only residual issue was a stale comment block (lines 95-101) that still described the old v6 behavior ("The signal-based submitToolResult flow used here continues to work in v7..."), which was false post-A.2 stateless-suspension. Cleaned up to correctly describe the v7 surface (Temporal now emits 'suspended_client_tool' etc., approve.ts uses durable submit).

  3. example-remote-agents-temporal — npm run typecheck and npm run build both pass cleanly. The example is not fail-fast anymore; running the end-to-end flow requires a Temporal + Postgres+Redis stack but the code itself is correct.

Nothing else to ship for this entry.

FU-DEMO-02: client-tools README v6 snippet + nextjs-redis snapshot helper — ✅ closed (verified outdated) ​

Status: done (verified during post-release follow-up sweep — the two pending sub-items have already been resolved by intervening work; this entry stays in the docs as a back-reference). Effort: ~1 hour (audit). Surfaced by: demo-review-v7.md §245-251 (deleted; original audit).

Closure:

  1. examples/client-tools/README.md — re-audited 2026-05-18 post-release. The README no longer contains any v6 snippet; the current "How the client-tool flow works" walkthrough documents the v7 useHelixChat / createHelixChatTransport / submitToolResult path end-to-end. No diff needed. Searches for "v6" / "deprecated" return zero matches in the README.

  2. examples/nextjs-redis snapshot fallback — the framework helper the demo was waiting on is buildSnapshot (exported from @helix-agents/ai-sdk at src/index.ts:199 as export { buildSnapshot, type BuildSnapshotDeps } from './handler/snapshot.js'). examples/nextjs-redis/src/lib/agent-client.ts:141 already uses it directly:

    ts
    export async function getSnapshot(
      sessionId: string
    ): Promise<FrontendSnapshot<ChatState> | null> {
      return buildSnapshot<ChatState>(deps, { sessionId });
    }

    No fallback path is needed because the helper exists; the "snapshot fallback path still expects a v7 framework helper that isn't shipped yet" sub-item in the original audit pre-dated the buildSnapshot export landing.

Nothing to ship for this entry — included here only so the original audit trail stays auditable.

FU-DEMO-04: opennext-cloudflare-do — coordinator UI hardcoded child name — ✅ closed ​

Status: done (verified during May 18 post-release follow-up audit for GL #100). Severity: medium (demo polish, not framework) Effort: done via option 1 (system prompt constraint — the cheapest fix in the original issue). Surfaced by: Round-3b web-app-testing live-LLM verification.

Closure: the CoordinatorAgent system prompt explicitly instructs the LLM NOT to pass a name argument to companion__spawnAgent, so auto-naming produces the canonical researcher-1 that the UI polls. See the in-code rationale at examples/opennext-cloudflare-do/src/app/coordinator/[sessionId]/ChatClient.tsx:83-97. A future framework enhancement could add GET /api/coordinator/<id>/children for dynamic name discovery via getSubSessionRefs, but that's an enhancement, not a closure requirement.

FU-DEMO-05: opennext-cloudflare-do — unauthenticated /debug route — ✅ closed ​

Status: done (verified during May 18 post-release follow-up audit for GL #100). Severity: medium (information disclosure if demo deployed without auth) Effort: done via env-var gate (option 2 in the original issue). Surfaced by: Round-2 web-app-testing curl probes.

Closure: both /api/chat/[sessionId]/debug and /api/debug/[sessionId] routes are gated by gateDebugRoute() from examples/opennext-cloudflare-do/src/lib/debug-gate.ts. Default behavior is 404 (no signal that the route exists, so scrapers don't get a "look harder" hint). Set ENABLE_DEBUG_ROUTES=true on the worker (typically via .dev.vars for local development) to enable. Strict === 'true' comparison documented to prevent future "helpful" refactors from accepting '1' / 'TRUE' / 'yes'.


Streaming hardening (v7.0.1+) ​

FU-STREAMING-01: Concurrent writer + reader race in core — ✅ closed ​

Status: done (closed by Sub-project A streaming-hardening, MR !184). Severity: 7/10 Surfaced by: resumable-streaming-deep-review.md §261.

Closure: the cross-impl streamManagerContractTests harness in @helix-agents/core/testing codifies the no-gap + global-monotonic invariant via G1.concurrent-writers (5 writers × 10 chunks, asserts no dups, no gaps, monotonic-as-yielded). Wired into all three implementations (in-memory, Redis, DOStreamManager). The race-detection test in packages/e2e/src/__tests__/streaming-races.integ.test.ts exercises the same shape against every viable backend under adversarial load.

FU-STREAMING-02: Redis getChunksFromSequence array-index assumption — ✅ closed (verified already resolved) ​

Status: done (verified during post-release follow-up sweep; the original audit predated the current implementation OR the fix landed in passing without back-reference). Severity (historical): 8/10 Surfaced by: resumable-streaming-deep-review.md §262 (deleted).

Description (historical): getChunksFromSequence was reported to assume array indices map 1:1 to sequence numbers in the Redis chunk list — fragile against producers that skip sequences (e.g., retry- then-resume).

Closure: the current implementation at packages/store-redis/src/redis-stream.ts:1212-1231 does NOT make the array-index assumption. It:

  1. Loads the entire list (LRANGE 0 -1).
  2. Filters by parsed._sequence > fromSequence per entry.
  3. Sorts the filtered set by sequence before returning.

The class-level comment at line 231 explicitly states "we filter by _sequence (not list index) in getChunksFromSequence." The sort step's comment at line 1228 explicitly handles "out-of-order list insertion (concurrent writers)."

Both alternatives the original audit proposed (invariant verification, or a (seq, chunk) Hash) are unnecessary — the filter-then-sort already provides the same robustness without the storage-shape change. The cost is one full-list scan per getChunksFromSequence call, which is acceptable because the JS-side parse-cache (mentioned at line 1206) amortizes parsing across reads.

FU-STREAMING-03: DO SSE late-attach race vs concurrent writes — ✅ closed ​

Status: done (closed by Sub-project A streaming-hardening, MR !184). Severity: high (CF DO production path) Surfaced by: resumable-streaming-deep-review.md §263.

Closure: root cause identified as a lost-notification race in DOStreamManager.createResumableReader's iterator wait loop: a write landing between the status check and waiter registration fired notifyWaiters() before the reader's waiter was registered, so the reader hung up to 5 seconds on the safety timer. Fixed by reordering the loop to register-before-poll (waiter registered BEFORE the storage SELECT; any write either lands before the register and is caught by SELECT, or fires the waiter). The G2.late-attach-during-burst contract test and the FU-STREAMING-03 repro race-detection test both exercise the scenario and verify sub-100ms delivery.

FU-STREAMING-04: cleanupToStep + attached reader interaction — ✅ closed ​

Status: done (closed by Sub-project A streaming-hardening, MR !184). Surfaced by: resumable-streaming-deep-review.md §264.

Closure: behavior pinned as G4 of the StreamManager concurrency contract: cleanupToStep(N) that deletes chunks past a reader's cursor MUST cause the next iterator.next() call to throw StreamTruncatedError (new typed class, exported from @helix-agents/core). Implemented across all three impls via a per-stream truncated_at_step marker that the reader consults on each iteration. Contract tests: G4.truncation-throws + G4.truncation-after-endStream-returns-done + G4.caught-up-on-active-stream-does-not-throw + G4.resetStream-clears-truncation-marker. Two known gaps tracked separately:

  • FU-G4-STORE-CF-STREAM-DO: binding-side DurableObjectStreamManager path lacks G4 (requires new SSE event type).
  • FU-G4-PERF-POLISH: ✅ closed — sticky O(N) reader rebuild while the truncation marker stays set (bounded with a per-reader marker cache on redis + DO; the clear-on-reset sub-issue predated this FU and was already done).

FU-STREAMING-05: Inline mock DO scaffolding divergence from real Phase 8 semantics — ✅ closed ​

Status: done (closed by Sub-project A streaming-hardening, MR !184). Surfaced by: resumable-streaming-deep-review.md §265.

Closure: the 240-LOC packages/runtime-cloudflare/src/__tests__/helpers/mock-sql.ts helper (whose header explicitly admitted "NOT a full SQLite implementation") was DELETED entirely. The 5 dependent test files (do-stream-manager-resumable.test.ts, -paused, -pause-cas, -cleanup-to-step, -info-metadata) migrated to .cf.test.ts and now run inside real workerd via cloudflare:test's runInDurableObject. 26 tests preserved with all original assertion intent. The shared helpers/stream-manager-cf-helpers.ts provides the runInDurableObject boilerplate to keep the 5 files thin.

FU-G4-STORE-CF-STREAM-DO: StreamDurableObject lacks G4 truncation enforcement — ✅ closed ​

Status: ✅ closed (Wave 3) — the binding-side DO stream path now surfaces G4 truncation. Severity: medium (parity gap, not data loss) Surfaced by: Batch 3 spec review of streaming-hardening MR (Sub-project A).

Closure (Wave 3): wired a truncated event end-to-end through the binding-side (store-cloudflare) DO stream path, which previously deleted orphaned chunks on cleanupToStep but never surfaced StreamTruncatedError to attached readers:

  • packages/core/src/store/stream-event.ts — added StreamTruncatedEvent ({ type: 'truncated', truncatedAtStep, atSequence? }) to the StreamEvent union + Zod StreamEventSchema + a type guard (with a new core unit test in store/__tests__/stream-event.test.ts).
  • packages/store-cloudflare/src/stream-durable-object.ts — handleCleanupToStep now writes a truncated_at_step marker and broadcasts the truncated event to attached WS + SSE connections; handleReset / handleReactivate clear the marker (matching memory/redis/DOStreamManager).
  • packages/store-cloudflare/src/durable-stream.ts — the binding-side SSE reader (createSSEReaderCore) recognizes truncated and throws StreamTruncatedError (close-before-throw discipline; subsequent next() returns done).
  • packages/ai-sdk/src/cloudflare/sse-parser.ts — parseDOSSEStream surfaces the new truncated SSE event (ParsedStreamTruncated + isParsedTruncated) instead of dropping it.

Best-effort nuance (documented): the binding-side reader is a push-SSE consumer with no per-next() marker poll, so G4 truncation is event-driven — the throw fires only if the reader is attached when cleanup broadcasts. A reader that attaches AFTER cleanup (replay /sse?fromSequence=X) won't get a truncated event. The G4.truncation-throws contract case attaches before cleanup, so it passes (active-stream throw enforced). One sibling contract case — G4.truncation-after-endStream-returns-done — IS opted out for this manager via skipKnownFailures with a documented push-SSE justification (not faked): see FU-G4-BINDING-AFTER-ENDSTREAM-RETURNS-DONE below. Its portable safety property is pinned by a dedicated binding-side test instead.

Locked by: the new binding-side contract-parity test packages/store-cloudflare/src/__tests__/durable-stream-contract.integ.test.ts (runs the shared streamManagerContractTests against the real StreamServer DO in Miniflare — this also closes FU-A2-DO-STREAM-MANAGER-PARITY) plus the binding-side push-SSE behavior tests in packages/store-cloudflare/src/__tests__/integration/durable-stream-late-attach-g4.integ.test.ts.

Residual — handleResume (paused→active) marker-clear has no direct read surface (Waves 3-5 deep-review L1, accepted-no-test): handleResume (stream-durable-object.ts:841) clears the truncated_at_step meta row with an UNCONDITIONAL DELETE FROM meta WHERE key = 'truncated_at_step' on the paused→active transition (mirroring the handleReactivate:877 / handleReset:1119 DELETEs). The deep-review noted this clear is only "trivially" tested — the existing reactivate contract case passes for a different reason (the push-SSE reader does not poll the marker). We investigated adding a DIRECT "marker is gone after resume" assertion and found the DO exposes NO read surface for the truncated_at_step row:

  • /info deliberately excludes it from its SELECT … WHERE key IN (…) allow-list (handleInfo:912).
  • The DO never reads the row back for any behavioral decision — the only consumer of G4 truncation is the binding-side reader, which reacts to the broadcast truncated EVENT emitted at cleanupToStep time (durable-stream.ts SSEReaderState.truncated), NOT the meta row.
  • The SSE attach path (setupSSEConnection:462) does not read the marker or emit a truncated frame on attach, so a fresh post-resume reader can't observe its presence/absence.
  • Miniflare 3.x exposes no public API to query an arbitrary user DO's internal SQLite (meta) table from the test (only getDurableObjectNamespace + HTTP stubs).

Asserting the row directly would require EITHER a production code change (add a test-only read endpoint to the DO) OR reaching into Miniflare internals in a way that does not map to the real meta table — both contortions the deep-review explicitly said to avoid for this harmless-on-push-path Low. The clear is therefore verified by code-read + the unconditional SQL DELETE, and is structurally identical to the handleReactivate / handleReset DELETEs that ARE exercised (the shared contract's G4.resetStream-clears-truncation-marker case + the binding-side late-attach tests). No test added; tracked here. Re-evaluate if a future read surface for the marker is added (e.g. surfacing truncatedAtStep on /info), at which point the direct assertion becomes feasible.

Description: The G4 contract ("cleanupToStep MUST throw StreamTruncatedError on attached readers when chunks past the reader's cursor are deleted") was wired into InMemoryStreamManager, RedisStreamManager, and runtime-cloudflare/DOStreamManager in the streaming-hardening MR. The binding-side DurableObjectStreamManager (in packages/store-cloudflare/) delegates via stub.fetch to a SEPARATE in-DO impl — StreamDurableObject — which the original plan conflated with DOStreamManager. As a result, the binding-side path does NOT enforce G4.

The binding-side reader is a pure SSE consumer (createSSEReaderCore in packages/store-cloudflare/src/durable-stream.ts). Closing the gap requires adding a new event type to the StreamEvent discriminated union (currently chunk | end | fail | status) — call it StreamTruncatedEvent — and wiring it across:

  1. packages/core/src/store/stream-event.ts — new event type + schema
  2. packages/store-cloudflare/src/stream-durable-object.ts — handleCleanupToStep writes truncated_at_step marker AND broadcasts the new event to attached SSE connections.
  3. packages/store-cloudflare/src/durable-stream.ts — createSSEReaderCore.getNext recognizes the new event and throws StreamTruncatedError.
  4. packages/ai-sdk/src/cloudflare/sse-parser.ts — recognize the new event so it doesn't drop into the "unknown frame" branch.
  5. Add .cf.test.ts infra to packages/store-cloudflare/ so the cross-impl contract test from @helix-agents/core/testing.streamManagerContractTests can run against DurableObjectStreamManager and verify the fix.

Effort: ~80-120 LOC across 3 packages + test infra. Real wire- protocol change.

Mitigation in the meantime: The streaming-hardening release notes (and the corresponding changeset) should call out that G4 is NOT enforced on the binding-side DurableObjectStreamManager path in v7.0.1; consumers using that import directly should be aware that cleanupToStep will not surface as StreamTruncatedError on attached SSE readers until this follow-up lands.

FU-G4-BINDING-AFTER-ENDSTREAM-RETURNS-DONE: push-SSE binding reader can't satisfy the poll-reader "first next() after post-endStream cleanup is done" mechanic — ⏳ accepted divergence ​

Status: ⏳ accepted divergence (documented + opted out, portable safety property pinned by a dedicated test). Severity: low (best-effort push-SSE timing; the REAL G4 safety guarantee — never throw after terminal — holds). Surfaced by: running the binding-side Miniflare contract suite during the Wave 3 four-battery fix-round (the original Wave 3 commit had never actually run this suite green — G4.truncation-throws was hanging 60s and G4.truncation-after-endStream-returns-done was failing; both were wrongly dismissed as Miniflare flakes).

The divergence: the shared contract case G4.truncation-after-endStream-returns-done asserts the FIRST iterator.next() after a post-endStream cleanupToStep returns done. This is a POLL-reader mechanic — InMemoryStreamManager, RedisStreamManager, and the in-process DOStreamManager re-read the store / marker on each next(), so a just-deleted chunk vanishes immediately. The binding-side DurableObjectStreamManager is a push-SSE consumer: when a reader attaches, the DO eagerly replays ALL surviving chunks into the reader's local buffer BEFORE a later cleanupToStep deletes them server-side. The reader therefore delivers the already-pushed chunks (which genuinely existed and were streamed) and then terminates with done — it satisfies the real guarantee (it never THROWS StreamTruncatedError after a terminal status) but not the poll-reader "first next() is done" timing.

Resolution (Wave 3):

  • Opted the case out for this manager via skipKnownFailures['G4.truncation-after-endStream-returns-done'] with an in-line push-SSE justification (durable-stream-contract.integ.test.ts). NOT faked — the active-stream throw case (G4.truncation-throws) remains enforced and is NOT opted out.
  • Pinned the portable safety property (drain-to-done after endStream+cleanupToStep with an attached reader NEVER throws StreamTruncatedError) with a dedicated binding-side test in integration/durable-stream-late-attach-g4.integ.test.ts.

Also fixed in the same fix-round (the real bug behind the 60s hang): the binding-side reader (durable-stream.ts) set only state.truncationThrown before throwing StreamTruncatedError, not state.closed. A consumer's catch-then-retry (G4.truncation-throws asserts the next next() returns done) on a still-active stream therefore blocked indefinitely instead of returning done. Now sets state.closed = true before both throws, mirroring do-stream-manager.ts / redis-stream.ts.

If we ever want full parity: the binding reader would need to NOT eagerly buffer (i.e., re-fetch survivors lazily per next()), which defeats the push-SSE design. Not worth it — the divergence is delivery timing of already-valid chunks, not a missing guarantee.

FU-G4-PERF-POLISH: G4 truncation marker check is O(N) per next() + not cleared on resetStream ​

Status: ✅ closed Severity: low (perf polish; no correctness issue) Surfaced by: Code-quality review of Batch 3 streaming-hardening commit (32dbdea76).

Original description (predates the fixes below): the G4 truncation-mark mechanism stores a per-stream marker the reader iterator consults on every next(). Two improvements were proposed: (1) cache the marker on the reader closure to avoid re-scanning the chunk list on every yield; (2) clear the marker on resetStream.

As-scoped reframe (Wave 3):

Sub-issue (2) — clear the marker on reset — was already DONE before this FU was actioned. All three impls clear the marker on the transition back to active:

  • memory: in-memory-stream.ts resetStream + reactivateStream
  • redis: RESET_STREAM_SCRIPT HDEL __truncated_at_step (redis-stream.ts) + the CAS-to-active script used by resumeStream/reactivateStream
  • DO: resetStream + reactivateStream DELETE ... WHERE key = 'truncated_at_step' (do-stream-manager.ts)

Guarded by the cross-store contract test G4.resetStream-clears-truncation-marker (core/src/testing/stream-manager-contract.ts). Nothing to do here.

Sub-issue (1) — cache the marker / avoid O(N) per next() — was partly stale, now bounded on redis + DO:

  • Memory needed NO change. Its marker check is already O(1) (a field read on the in-memory stream object) and its chunk lookup is amortized-O(1) via the [R3C-C1] hint index (findNextChunk). There is no per-next() list scan to eliminate.
  • Redis + DO had the real cost. While the truncation marker is SET, the resumable reader re-ran a FULL buffer rebuild on every next() (redis: LRANGE chunksKey 0 -1; DO: SELECT ... WHERE sequence > cursor). This is wasted work in the steady state where a long-lived reader survived a cleanup and keeps draining surviving chunks below the marker.

Fix (closing sub-issue 1 on redis + DO): added a closure-local lastObservedTruncStep to both resumable readers. The expensive rebuild is now SKIPPED when the marker value is unchanged AND the local buffer still has undrained survivors past the cursor — the exact steady state the FU flagged. The rebuild deliberately STILL runs (a) once the buffer drains and (b) when a new cleanup advances the marker, because the on-drain rebuild is what surfaces any chunks the writer appended after the cleanup while the marker stayed set, and confirms emptiness for the G4 throw gate. This bounds the perf win to the safe path and cannot trade G4 correctness for it. The cache is reset when the marker is cleared (reset/reactivate).

Verification: the full G4 contract suite stays green on redis (redis-stream-contract.integ.test.ts, requires Redis) and DO (do-stream-manager-contract.cf.test.ts, requires workerd). A focused regression test (redis-stream.integ.test.ts → FU-G4-PERF-POLISH: marker-cache skip still delivers post-cleanup writes) pins the correctness-critical path: it FAILS with a blanket marker-unchanged skip (spurious StreamTruncatedError) and PASSES with the buffer-non-empty guard. The optional store-read-count micro-test was skipped as brittle per the FU.

Touched:

  • packages/store-redis/src/redis-stream.ts (+ integ test)
  • packages/runtime-cloudflare/src/do/do-stream-manager.ts
  • (packages/store-memory/src/in-memory-stream.ts — no change needed)

A.1 (Stateless purity cleanup) follow-ups ​

FU-A1-01: runId generator pattern duplicated in executor.ts ​

Status: ✅ done — see "Done items" section. Closed by: commit 869787e5c.


O5 (workerd parity harness) follow-ups ​

FU-O5-CF-DO-APPROVAL-STREAM-READ: approval-gate approvalId not readable via env.streamManager on cf-do-d1 ​

Status: open — Scenario 3 (approval-deny) is gated OFF on cf-do-d1 in the shared lifecycle-hooks-parity suite via ParityOpts.gateApprovalStreamRead.

What's gated. The shared lifecycle-hooks parity scenario module's Scenario 3 (approval-gate deny path) recovers the approvalId synchronously from env.streamManager.createReader(sessionId) (it needs it to POST the approval-response submit). On the cf-do-d1 backend this is NOT presentable:

  • The DO base persists its event stream to its OWN INTERNAL DOStreamManager (over ctx.storage, wired in durable-object-agent-base.ts > ensureInitialized), NOT the binding-side DurableObjectStreamManager over env.STREAMS that the cf-do-d1 BackendEnv.streamManager exposes. So env.streamManager.createReader(sessionId) reads a DIFFERENT (empty) stream than where the DO emitted the tool_approval_request chunk — it returns no reader and the scenario throws stream reader missing.
  • The approvalId is generated at approval time and lives ONLY on the tool_approval_request chunk; it is NOT persisted on state.pendingClientToolCalls (see PendingClientToolCall in core/src/types/state.ts — no approvalId field), so there is no portable way to recover it from the harness BackendEnv (state store / status route).

Why it's a gate, not a fake. Scenarios 1 + 2 (the portable client-tool suspend/resume + previousRunId linkage + tracingContext round-trip contracts) DO EXECUTE on cf-do-d1 and pass — the approval path is the only genuinely runtime-divergent leg. Per the spec's gating discipline (refinement #5: "synchronous mid-suspend handle internals" are CF-non-portable), Scenario 3 is skipped on cf-do-d1 rather than weakened. The EXECUTING CF-DO approval-gate coverage continues to live in the dedicated standalone DO tests (which read the DO's /sse route directly).

To close. Either (a) surface the approvalId on the DO /status (or a new route) / on pendingClientToolCalls so the scenario can recover it portably, or (b) give the cf-do-d1 BackendEnv.streamManager a reader that proxies the DO's internal /sse stream so createReader returns the real approval chunk. Then drop gateApprovalStreamRead: true from the cf-do entrypoint and confirm Scenario 3 EXECUTES.

Where: gate in e2e/src/parity/scenarios/lifecycle-hooks.ts (ParityOpts.gateApprovalStreamRead + itOrSkipApproval); applied in e2e/src/__tests__/lifecycle-hooks-parity.cf.test.ts.


FU-O5-CFW-TRACING-CONTEXT-PERSISTENCE: tracingContext set in onAgentSuspended does not survive resume → completion on cfw-workflows-d1 ​

Status: open — Scenario 2 (tracingContext round-trip) is gated OFF on cfw-workflows-d1 in the shared lifecycle-hooks-parity suite via ParityOpts.gateTracingContextPersistence. (Since plan 0c MR-1b the workerd row this suite runs on is named cfw-workflows-d1-pool; cfw-workflows-d1 is now the Node-driven conformance lane.)

What's gated. The shared lifecycle-hooks parity scenario module's Scenario 2 writes state.tracingContext inside onAgentSuspended (via env.stateStore.saveState, mirroring the langfuse-hooks pattern) and asserts it round-trips through loadState() AFTER the resumed run COMPLETES. On JS / Temporal / DBOS this holds — tracingContext is a durable top-level SessionState field that survives a completion save. On cfw-workflows-d1 it does NOT:

  • The CFW D1 store (store-cloudflare/src/d1-state.ts) packs tracingContext into the suspension_context JSON column (V8/V9 schema), alongside suspendedAwaitingChildren / suspendedStepId / expiresAt. The suspend-leg hook's saveState writes it correctly (confirmed — the column round-trips on load).
  • BUT the resume-leg completion's saveStateAndPromoteStaging rebuilds suspension_context from the now-COMPLETED (non-suspended) in-memory state, which carries no tracingContext, so it writes suspension_context = NULL — clobbering the suspend-time write. The post-resume loadState().tracingContext is therefore undefined.

This is a genuine CFW-runtime (D1-store data-model) divergence, NOT a harness adapter artifact: the setupCfwWorkflowsD1 adapter only does AGENT_WORKFLOW.create() + poll + the durable submit; it never touches state persistence. The same workflow body + D1 store are exercised by the standalone hitl-cycle.workflow.test.ts.

Why it's a gate, not a fake. Scenario 1 (the portable client-tool suspend/resume + exactly-once cardinality + previousRunId linkage contract) and Scenario 3 (approval-deny) BOTH EXECUTE and pass on cfw-workflows-d1 — the tracingContext persistence assertion is the only genuinely runtime-divergent leg. Per the spec's gating discipline (refinement #5), Scenario 2 is skipped on cfw rather than weakened.

To close. Promote tracingContext out of the suspension-scoped suspension_context column into a durable top-level column so it survives non-suspended saves like the JS/Temporal/DBOS stores — then drop gateTracingContextPersistence: true from the cfw entrypoint and confirm Scenario 2 EXECUTES.

The fix must cover BOTH CF stores. packages/runtime-cloudflare/src/do/do-state-store.ts (CF-DO — note: this DO store lives in runtime-cloudflare, NOT store-cloudflare) uses the identical buildSuspensionContext packing + NULL-on-non-suspended-save as packages/store-cloudflare/src/d1-state.ts (CFW), so at the store level CF-DO clobbers tracingContext on a completion save exactly the same way. CF-DO's Scenario 2 passes today NOT because its store preserves the field, but because of resume topology: CF-DO resumes IN-PLACE inside the live DO (the in-memory SessionState still carries tracingContext through to the completion save), whereas CFW resumes via a COLD-LOADED fresh workflow instance whose completion save rebuilds suspension_context from a state that already dropped it. So the top-level-column migration must be applied to BOTH d1-state.ts and do-state-store.ts — and when CF-DO later joins a scenario that resumes via a fresh load (not in-place), it would hit the same clobber without it.

Where: gate in e2e/src/parity/scenarios/lifecycle-hooks.ts (ParityOpts.gateTracingContextPersistence + itOrSkipTracing); applied in runtime-cloudflare/src/__tests__/lifecycle-hooks-parity.wf-noiso.test.ts.


D (Test infrastructure overhaul) follow-ups ​

See test-infrastructure-roadmap.md for the full list. Key items still open at A.2 closure:

  • F6 (workspaces on Temporal/CFW Workflows) — explicit fail-fast, needs design work for cross-runtime workspace support.
  • F7 (DBOS resume() contract bug fixes) — being addressed in sub-project A.3 (next).
  • O5 (workerd-context CF DO + CFW Workflows harness setup helpers) — blocked on sub-project B. Scope sharpened by FU-A2-16 / FU-A2-LIFECYCLE-HOOKS-E2E (closed): real-workerd CFW Workflows workflow-body execution + HITL suspend/resume + lifecycle hooks are now directly covered at the package level (runtime-cloudflare/src/__tests__/workflows/*.workflow.test.ts, driving env.AGENT_WORKFLOW.create(...) + introspectWorkflowInstance). What O5 still owes is the harness-matrix leg: a setupCfwWorkflowsD1 (and setupCfDoD1) that returns a real BackendEnv whose executor is a workerd-resident CloudflareAgentExecutor reachable over an HTTP/RPC surface, so the shared getViableBackends-driven tests (lifecycle-hooks-parity, cross-runtime-resume-matrix, etc.) can run the cfw-workflows-d1 / cf-do-d1 rows instead of hard-filtering them under IMPL_PENDING_O5.

Cross-service remote agents (2026-07-09) ​

Deferred items from the cross-service remote-agents effort (spec: docs/superpowers/specs/2026-07-09-cross-service-remote-agents-design.md; guide: docs/guide/cross-service-remote-agents.md). Each item below was verified against the source at the time of writing — file references are the place to start.

FU-REMOTE-01: worker-handler workspace validation + /workspace bridging — open ​

createRemoteAgentWorkerHandler (runtime-cloudflare/src/do/worker-handler.ts) performs no construction-time workspace wiring validation — the file does not mention workspaces at all — so a DO producer whose agent declares workspaces surfaces the misconfiguration as a runtime WorkspaceFailedError instead of a startup error, and the handler does not bridge /workspace introspection. Parity target: AgentServer's construction-time validation. Spec §4.3 non-goals.

FU-REMOTE-02: migrate legacy snapshot surfaces to checkpoint-pinning — open ​

Three snapshot surfaces still use the live two-read pattern that round-2 review proved unsound (state and sequence read non-atomically → skipped or duplicated in-flight-step patches for local frontends):

  • createServerSnapshot (core/src/stream/resumable.ts:62-73) — loadState() then getSnapshotSequence().
  • ai-sdk buildSnapshot (ai-sdk/src/handler/snapshot.ts:163-199) — loadState() then getStreamInfo().latestSequence; its getLatestCheckpoint read only supplies checkpointId/stepCount metadata, it does not pin state.
  • DO-native handleSnapshot (runtime-cloudflare/src/do/durable-object-agent-base.ts:3650-3683) — loadState() then streamManager.getLatestSequence().

The remote-protocol handlers already checkpoint-pin; port the same discipline to these three. Spec §3.1.

FU-REMOTE-03: toolCallId stamping on state_patch chunks — open ​

Optional §7 in-step re-ordering aid: state_patch chunks carry no toolCallId today (verified — no such field on the chunk type). Stamping the originating toolCallId would let consumers attribute same-step parallel-sibling patches. Additive wire field; v2. Spec §7 item 4.

FU-REMOTE-04: oldestRetainedSequence on RemoteStatusResponse — resolved (decided no) ​

Spec §11 open item 7 asked whether /status should also carry the retention floor. Decided no: the floor is already carried by RemoteSnapshotResponse.oldestRetainedSequence (core/src/types/remote-protocol.ts:208, §3.4) and by the error{code:'TRUNCATED', floor} frame on /sse, so RemoteStatusResponse stays a cheap liveness/recovery poll without duplicating it — a consumer that needs the floor calls getSnapshot. Closes spec §11 item 7.

FU-REMOTE-05: remote completion fold degrades on a bare catch — no typed-error branching — open ​

processRemoteSubAgentCompletion (core/src/orchestration/remote-sub-agent-completion.ts:187-218) wraps getSnapshot/getUsage in bare try/catch blocks with ZERO instanceof branching, so a genuine 500, a schema-drift error, a missing route on an old peer (RemoteAgentFeatureUnavailableError), a RemoteAgentNotFoundError, and a raceWithTimeout timeout all emit the IDENTICAL warn. The typed errors are constructed by the transport (core/src/transport/http-transport.ts:330-344) but the fold never looks at them. Cheap improvement: a reason classification field on the warn — separates the cases in logs without changing degrade semantics. Open design question: should the CODE match the documented contract (branch on typed errors, e.g. treat feature-unavailable as expected version skew and a 500 as an incident), or should the docs keep matching the code (all read failures degrade identically)? Decide before adding the field.

FU-REMOTE-06: fold read timeouts bound the parent, not the request (no AbortSignal on the transport) — open ​

raceWithTimeout in the completion fold bounds how long the PARENT waits; it does not cancel the underlying request. RemoteAgentTransport (core/src/types/remote-protocol.ts:366-397) only accepts an AbortSignal on stream() — getStatus/getSnapshot/getUsage/interrupt take none — so a timed-out fold read keeps running in the background. Relatedly, HttpRemoteAgentTransport.fetchWithRetry (core/src/transport/http-transport.ts:357-396) passes NO signal and NO timeout to fetchFn, so those three reads are unbounded at the fetch layer and a 5xx retry loop can outlive the fold. The complete fix widens the transport interface (a breaking change for third-party transports) — hence deferred.

FU-REMOTE-07: fold timeouts are not configurable — open ​

GET_STATUS_TIMEOUT_MS (5s) / GET_SNAPSHOT_TIMEOUT_MS (10s) / GET_USAGE_TIMEOUT_MS (10s) are module constants (core/src/orchestration/remote-sub-agent-completion.ts:62-64). The helper's only per-call knob (maxFoldStateBytes) originates on RemoteSubAgentConfig and is hand-threaded through four runtime DTOs, so adding a second knob is a cross-runtime change, not a local edit. Too-tight a bound degrades a healthy-but-slow producer (silently dropping its state/usage fold), so the knob matters if operators ever hit it. Deliberately deferred; the rationale is documented inline at the constants.

FU-REMOTE-08: remote stream-proxy dispatch loop is implemented four times — open ​

The core helper executeRemoteSubAgentDispatch (core/src/orchestration/remote-sub-agent-dispatch.ts:222) is used ONLY by Temporal (runtime-temporal/src/remote-sub-agent-activity.ts:43). Three runtimes hand-roll their own for await (transport.stream(...)) loop with the same concerns (remoteSource stamping, REMOTE_REF_PERSIST_CHUNK_INTERVAL persist cadence, dead-producer observations, ALREADY_RUNNING → attach, terminal fold via processRemoteSubAgentCompletion):

  • runtime-js/src/execution/remote-subagent-executor.ts:483
  • runtime-dbos/src/steps/execute-remote-subagent.ts:460
  • runtime-cloudflare/src/steps.ts:2745

They are behaviorally equivalent today and pinned by the cross-runtime remote parity battery (packages/e2e/src/__tests__/remote-agent-*.integ.test.ts), but a future fix to one can silently miss the other three. Either converge the three onto the core helper (each has a real durability reason not to — DBOS step boundaries, CFW step boundaries, JS in-process retry/timeout) or document that per-runtime rationale inline at each loop so the duplication is a decision rather than drift.

FU-REMOTE-09: sub-agent CALL-COUNT usage stats are at-least-once under durable retry — open ​

The TOKEN fold is idempotent at read — selectEmbeddedRemoteRollup (core/src/usage/rollup-utils.ts:260) picks ONE entry per subSessionId by the §6.2 total order. The CALL-COUNT stats are not: processSubAgentEntry (core/src/usage/rollup-utils.ts:129-131) increments subAgentStats.totalCalls/byType[...] for EVERY entry, and the D2 rule records recordSubAgent on both anchor branches, so a durable retry or the dual fast/recovery recording path inflates the call counts. The comment at runtime-cloudflare/src/steps.ts:2257-2258 ("re-recording is idempotent-at-read") overstates this. Fix: either narrow the comment ("token fold is idempotent at read; call-count stats are at-least-once") or dedup the stats by subSessionId at rollup time (groupSubAgentEntries already groups them, so the plumbing exists).

FU-REMOTE-10: computeTerminalUsage is a permanent stub — JS producers never attach end.usage — open ​

computeTerminalUsage (runtime-js/src/run-loop.ts:3955-3961) returns {} unconditionally: RunLoopInput carries usageContext (record-only), not a rollup-capable UsageStore handle. Consequence: a JS-runtime producer's end frame never carries usage, so every consumer of a JS producer makes a getUsage() recovery round-trip per remote-child completion. Correctness-neutral (the fold still happens via the recovery path, tagged remoteRollupOrigin: 'recovery'), but the end.usage fast path never materializes for JS producers. Fix: thread a UsageStore (or a getRollup callback) through RunLoopInput.

FU-REMOTE-11: EndFrameSchema strictly validates end.state / end.usage — open ​

EndFrameSchema (core/src/transport/http-transport.ts:43-48) validates state as z.record(...) and usage as UsageRollupSchema, so ONE malformed field fails the WHOLE end frame → malformedFrameEvent → a recoverable: true error event (:473-479). This is asymmetric with the deliberately-opaque chunk passthrough (:424-430), where the payload is not validated at all.

Corrected impact (the original write-up claimed the child's output is discarded — it is not): every consumer treats the recoverable error as a stream drop and falls through to its post-drop getStatus recovery branch (e.g. runtime-js/src/execution/remote-subagent-executor.ts:628-641, core/src/orchestration/remote-sub-agent-dispatch.ts:505-518), which harvests a completed child's output and folds via the recovery path. So the real cost is a wasted stream re-attach plus a degrade from the end fast path to getStatus+getSnapshot+getUsage — output is lost only if that recovery ALSO fails. Suggested fix: parse state/usage as z.unknown().optional() in the envelope and validate them at the fold step, so a bad optional field can't invalidate a valid terminal frame.

FU-REMOTE-12: e2e cross-service workers hand-redeclare SubAgentNamespace with a stale (wrong) rationale — open ​

packages/e2e/src/cross-service/workers/producer-worker.ts:17-31 and consumer-worker.ts:19-27 locally re-declare the SubAgentNamespace interface and blame "a rollup-plugin-dts quirk [that] drops the export keyword". That diagnosis is wrong: the root re-export simply did not exist, and it was ADDED in commit d93a2a773e (runtime-cloudflare/src/index.ts:187 now exports type SubAgentNamespace). Both workers should import the real type from @helix-agents/runtime-cloudflare and delete the misleading comments.

FU-REMOTE-13: examples/research-assistant-temporal is BROKEN — missing two activities — open ​

Its hand-written worker registration list (examples/research-assistant-temporal/src/activities.ts, the createActivities return object) omits agentDeclaresPersistentChildren and injectPendingCompletionNotifications. agentWorkflow calls activities.agentDeclaresPersistentChildren UNCONDITIONALLY on every workflow entry (runtime-temporal/src/workflow.ts:670), so any run of this example dies at the first activity call with Activity function agentDeclaresPersistentChildren is not registered on this Worker — the exact failure that hit examples/remote-agents-temporal before it moved to prototype-derived dynamic binding. Fix the example (see FU-REMOTE-14 for the systemic fix).

FU-REMOTE-14: no production-facing canonical Temporal activity binder — open ​

Root cause behind FU-REMOTE-13: the SDK ships a complete, canonical activity list ONLY as createTestActivities() — a public but TEST-named subpath export (@helix-agents/runtime-temporal/testing → runtime-temporal/src/__tests__/testing/test-activities.ts). Production/example worker authors therefore hand-roll the registration object, and it ROTS silently the moment the framework adds an activity: agentWorkflow proxies the whole GenericActivities surface, so there is no compile-time link between the proxied methods and the registered names. examples/remote-agents-temporal/src/activities.ts works around this by binding the prototype's own methods dynamically (see its bindActivities doc comment). Ship a production-facing binder (e.g. createAgentActivities(activities) from the package root) so no consumer has to choose between a test-named import and a hand-rolled list.

FU-REMOTE-15: the demo frontend's §7 reconstruction logic has NO runtime test — open (REQUIRED) ​

examples/remote-agents-cloudflare-do/consumer/public/app.js (1511 lines: three reconstruction modes, the heldThrough rewind guard, boundary re-snapshot, late-join, provenance dedup) is the reference implementation of the spec §7 consumer contract and is currently CODE-REVIEWED ONLY — six review rounds each found a real bug in it. Task 6's smoke test (examples/remote-agents-cloudflare-do/tests/smoke.cf.test.ts) runs in workerd via cloudflare:test and cannot execute browser DOM code, so it does not cover any of this. The repo already ships the right harness pattern for exactly this — real wrangler dev + Playwright + a record/replay OpenAI mock, used by examples/opennext-cloudflare-do and examples/research-assistant-cloudflare-do (see example-app-testing.md). Recording a fixture requires an OPENAI_API_KEY — hand-crafting fixtures is explicitly forbidden by that doc, which is why this is deferred rather than done. Treat as REQUIRED, not nice-to-have: this is the only executable check the §7 contract can get.

FU-REMOTE-16: unify createRemoteAgentWorkerHandler (DO) and AgentServer routing behind one protocol router — open ​

Both the Cloudflare DO worker handler and the Node AgentServer are "a fetch(Request) → Response that speaks the remote-agent protocol" — the same eight-route table (start / resume / sse / status / snapshot / usage / interrupt / abort), auth gate, agentType allowlist, projection, and error envelopes. Today they are SEPARATE standalone impls: createRemoteAgentWorkerHandler imports nothing from agent-server, the projection logic is duplicated (worker-handler-projection.ts vs AgentServer's inline mapper), and the auth posture diverges (worker handler = warn-and-serve; AgentServer = fail-closed-throw).

Better design: a shared RemoteAgentProtocolRouter core — the route table + auth + allowlist + projection + envelope handling — parameterized by a resolveTarget(agentType, sessionId) seam. The only irreducibly DO-specific piece is "resolve sessionId → DO stub + forward"; the Node side resolves to its in-process executor. Everything else is shareable.

Now unblocked by the "no backwards compatibility" directive (free to reshape both public surfaces). Nontrivial refactor; deferred. NOT a correctness issue — both handlers work and are independently tested; this is a duplication/consistency cleanup.

(For the record: FU-REMOTE-17 — the de-backcompat CODE sweep that was the sibling of this item — is DONE; see the remote-agents-de-backcompat.md changeset and the required status/chunk/awaitingClientTool tightening + REMOTE_PROTOCOL_VERSION removal.)

FU-REMOTE-18: consumer-side TRUNCATED → re-snapshot recovery is unimplemented — open (low priority; currently UNREACHABLE) ​

The spec §3.4 consumer contract — "on a TRUNCATED fail frame, refetch the snapshot and resume from its checkpoint-pinned streamSequence" — is documented and half-built but NOT wired on the consumer:

  • Producers emit it correctly: agent-server/src/http/handler.ts:379 and runtime-cloudflare/.../durable-object-agent-base.ts:4087 emit a fail/error frame with errorDetail.code === 'TRUNCATED' + a floor when a consumer requests fromSequence below the retention floor.
  • The transport surfaces it correctly: do-stub-transport.ts maps it to a {type:'error', errorDetail:{code:'TRUNCATED'}, recoverable:false} event.
  • But no consumer dispatch loop acts on it. remote-sub-agent-dispatch.ts:504 (+ the runtime-js / runtime-dbos / runtime-cloudflare twins) returns {kind:'failed', failureReason:'remote-failed'} on ANY !recoverable error, and TRUNCATED is recoverable:false — so getSnapshot() is never called and a TRUNCATED fails the sub-agent even though its output/state is recoverable. (A non-recoverable TRUNCATED even bypasses the existing post-drop getStatus harvest.)

Reproduction test (committed, teeth-having): packages/e2e/src/__tests__/cross-service-do-do-truncated.integ.test.ts — a describe.skip known-gap guard that drives the real shared dispatch through a mid-stream TRUNCATED and asserts the consumer re-snapshots + completes gapless. It flips green the instant a consumer implements the branch.

Why low priority / currently UNREACHABLE in production: the retention floor (oldestRetainedSequence) is effectively always 0 today — nothing advances it. resetStream() (the only floor-raiser) is called only in tests; the run-scoped cleanupToStep(minSequence) deletes a partial step's RECENT chunks (above the floor), not old ones, and has no production caller; and there is no automatic retention cap / TTL / max-size trimming on stream chunks (streams retain everything). So a consumer's fromSequence (≥ 0) is never below the floor, and TRUNCATED is never emitted in normal operation. This is a designed-ahead §3.4 resilience signal for a FUTURE with bounded stream retention; implement the consumer recovery when (if) that retention is wired.

The fix (when needed): in each dispatch loop's case 'error', branch on event.errorDetail?.code === 'TRUNCATED' BEFORE the !recoverable fail — call transport.getSnapshot(childSessionId), adopt its state, re-attach stream({ fromSequence: snapshot.streamSequence }), then un-skip the test above.


CI health (MR !272) ​

MR !272 (branch ci-health) bumps wrangler/Miniflare, passes the test env vars through turbo, fixes the Cloudflare Node-proxy flake family, relaxes two wall-clock bounds and stops .integration_base retrying script_failure. Entries here are either new, or record the resolution of an entry that another MR files (so that entry is not duplicated). When that other MR lands, mark its entry fixed and point it here.

Resolutions of entries filed on other branches ​

  • FU-FLAKE-OPENNEXT-WRANGLER-DEV-CRASH (reported on !269, redis-create-session-atomic; this is the only entry for it): fixed by !272. wrangler 4.141.0 no longer dies under the refresh-spam tests. The in-worker Uncaught Error: Network connection lost / Unable to enqueue errors are still logged (see FU-OPENNEXT-UNCAUGHT-ON-CLIENT-ABORT below) but are no longer fatal to wrangler dev. The example jobs now fail the readiness loop when the server never answers, and after the run print an explicit "wrangler dev exited during the Playwright run" line if the process died. Evidence: over the 45 pipelines before !272, 3 of 36 test:examples:opennext-cloudflare-do runs crashed (16755525964 on main, 16756530064 and 16756897246 on !269), all on 4.63.0; locally 0 crashes in 7 full runs on 4.141.0 (2 of them under a 10-process yes CPU stress) and 0 in 3 full runs on 4.63.0 on the same machine (the crash does not reproduce on macOS arm64). The CI runs on !272 are the evidence for the Linux amd64 runner; see the MR description. Correction to the entry: research-assistant job 16756390703 did not crash and has no 503. It failed status should transition from started/running to ended because the replayed mock finished the run before the first status poll; !272 fixes that assertion (completed is a valid first observation).
  • FU-FLAKE-CF-D1-PROXY-FETCH-FAILED (filed on !264, cf-multirun-flake-fix): fixed by !272. Root cause: Miniflare 3's Node proxy (undici 5.29) drops loopback connections under concurrent load (ECONNRESET, surfacing as TypeError: fetch failed). Transport A/B, 4 runs of concurrent D1 + DO proxy calls: Miniflare 3 22 of 13,440 calls failed, Miniflare 5 0 of 13,440. The same failure explains the CAS-race losers that saw D1StateError instead of StaleStateError, the large-state appendMessages failure, the G1.concurrent-writers DO failures and the durable-lock acquire failure. concurrent-subagent-fold.integ.test.ts also gets Workflows-style bounded retry-on-throw in its mock step.do (runStepWithWorkflowRetries) and prints result.error in every status assertion. With one injected fetch failed per test, the old mock fails 3 of 6 tests and the new one passes 6 of 6. No transport-level retry was added: the writes are not idempotent, and the bump removes the transport failure.
  • FU-FLAKE-DURABLE-LOCK-RACE-RETRY (filed on b8-dod-catchup): fixed by !272. The race test now uses a fresh resource per attempt and releases in finally via Promise.allSettled; its first failure (a Miniflare 3 proxy ECONNRESET) is also gone with the bump.
  • FU-FLAKE-MINIFLARE-SYNC-PROXY-DESYNC (on main): marked fixed in place, in "Known CI flakes" above.
  • FU-FLAKE-REMOTE-DUP-START-RACE (on main, fixed by !271): listed in the exposure table below only.

FU-FLAKE-CF-DO-STREAM-WALLCLOCK-BOUND: wake-latency bounds failed under CI load — ✅ fixed (CI health 2; !272's bound raise did not fix it) ​

Reopened and fixed (CI health 2): !272's 2000ms bound still failed (expected 4146 to be less than 2000, !288 16900882197), and the same class hit six more contract cases (G1 drain 5s, CU.empty-backlog 1s, WL.* / RF.* 1s endsWithin; 7 sightings after !272). The contract now has no positive wall-clock window: each case awaits the event, and DOStreamManager / RedisStreamManager run it with the new internal safetyNetPollMs: Infinity, so only the write's or transition's wake can deliver. Red-first: removing endStream's / failStream's wake makes the DO contract hang (test timeout), and dropping Redis's end pub/sub handling fails 6 Redis cases the same way. The poll itself keeps a test per manager (do-stream-manager-resumable.cf.test.ts "safety-net poll", redis-stream-safety-net.integ.test.ts).

Tests: G3.bounded-staleness in packages/core/src/testing/stream-manager-contract.ts (< 400ms; runs for store-memory, store-redis, store-cloudflare and the DO stream manager) and iterator returns done on endStream well before any safety-net tick in packages/runtime-cloudflare/src/__tests__/do-stream-manager-resumable.cf.test.ts (< 500ms).

Symptom: expected 404..588 to be less than 400/500 with the wake working. Jobs: main 16755525962 (588 vs 500) and 16754640436 (407 vs 400, 531 vs 500); 16756029846 (404); 16755485929 (434); 16749367148 (556 vs 500); 16756390700 (411); store-memory test:unit 16756727163 (579). On !272's first pipeline, before this fix landed, test:cloudflare 16758170964 failed the same way (822).

Fix: both bounds are now < 2000ms, and G3's nextWithTimeout is 4s. The bounds only need to separate a working wake (~50ms) from the 5s safety-net poll; the old values measured scheduler latency under load. Local: store-memory contract 20/20 green on the new bound; the old bound did not fail locally (0/20 even under CPU stress), so the CI jobs above are the red evidence.

FU-CF-DO-RESUMABLE-READER-DROPS-FRAMES: live DO continuation truncated after the first frame of a read — ✅ fixed (MR !272) ​

Class: product bug (silent data loss on the live stream), latent on main, exposed deterministically by the wrangler 4.141 / workerd 1.20260925 bump.

Symptom: in the opennext example, the confirmation after a client-tool turn streamed live as "The page" and stopped; a reload showed "The page background has been changed to blue." refresh-exact-messages.spec.ts C failed 3/3 on 4.141 (3/3 green on 4.63). Probably also behind the long-running comprehensive-refresh.spec.ts:472 flake ("turn 2 has 0 text parts after refresh"), which Playwright's in-run retry has been hiding on main.

Root cause: DOStreamManagerClient.createResumableReader (packages/runtime-cloudflare/src/do/do-clients.ts) split each network read into lines and returned from inside the line loop on the first chunk, discarding every later frame of that read. Newer workerd coalesces back-to-back DO SSE writes into one read.

Fix: complete lines are queued across next() calls. Unit test yields every frame when one read() carries several SSE frames fails before (1 of 4 deltas) and passes after. Changeset: @helix-agents/runtime-cloudflare patch. Conformance finding RM-51 (S1), pending the owner's catalogue batch 4. No scenario observes it yet (the owner's pending plan is 0c, a cf-do host cell), so the record's evidence exception or its flip to fixed is the conformance owner's to make.

FU-CI-TURBO-ENV-PASSTHROUGH: turbo's strict env mode dropped test env vars — ✅ fixed (MR !272) ​

The root pnpm run test:integration dropped RUNTIME_DBOS_INTEG (the DBOS suite collected 0 files and exited 0) and TEMPORAL_ADDRESS (Temporal tests used localhost:7233, which worked in CI only because the Kubernetes executor exposes service containers on localhost). Fixed with globalPassThroughEnv in turbo.json and a CI guard, pnpm run check:turbo-env (lint job), that fails when test code reads a variable turbo would drop. The conformance owner was filing this; no ID existed on main or conformance-catalogue-snapshot-truth when !272 was written, so this is the ID. The DBOS gate itself is removed by !259 (dbos-integ-ci-unpause), which also owns the test:integ:dbos job.

FU-OPENNEXT-UNCAUGHT-ON-CLIENT-ABORT: opennext worker throws on client disconnect mid-stream — open ​

Owner: the B8 session. Under the refresh-spam tests the worker logs Uncaught Error: Network connection lost and Uncaught TypeError: Unable to enqueue from OpenNextNodeResponse._internalWrite (writing into a stream the browser already closed). Non-fatal on wrangler 4.141, but it is noise in every run and was the lead-in to the 4.63 crashes. Fix direction: cancel the upstream reader when the Node response closes, or catch the enqueue after close in the route's streaming bridge.

FU-FLAKE-POOL-WORKERS-STACKED-STORAGE-LOST: test:cloudflare isolated-storage teardown loses the connection — ✅ fixed (CI health 2) ​

Fixed (CI health 2): a keep-alive race on Miniflare's Node loopback server, upstream cloudflare/workers-sdk#14848. vitest-pool-workers 0.9.14 pins miniflare 4.20251011.0, whose loopback server keeps Node's keep-alive timeout (5s + Node 24's 1s buffer). workerd pools the /storage push/pop sockets and closes idle ones itself at 5s, so an idle machine never races; on a starved runner its timer fires late, Node closes first, and the pop after a ~6s test goes out on a dead socket. In the sweep the failure rate by test duration was 0% under 2s, 1.3% at 5–6s and 15.1% at 6–8s; 96% of the failed tests ran ≥5s against 28% of passing ones. Fix: upstream's own change (#14850, server.keepAliveTimeout = 0) as patches/miniflare@4.20251011.0.patch; upstream only ships it in pool-workers releases that need vitest 4. Guard: runtime-cloudflare/src/__tests__/pool-workers-loopback-keepalive.test.ts (red on the stock Miniflare: "closed an idle keep-alive socket"). Drop the patch with the pool-workers bump.

Owner: the B8 session. filestore-provider.cf.test.ts (runtime-cloudflare) and CloudflareFileStoreFileSystem grep tests fail with Error: Network connection lost from WorkersTestRunner.updateStackedStorage (@cloudflare/vitest-pool-workers@0.9.14). Jobs: !272 16758170964, !270 16757391364, !264 16754841608. Same family as FU-FLAKE-WORKERD-INTROSPECT-TEARDOWN (pool-workers teardown). pool-workers stays on ^0.9 in !272 because newer releases need vitest 4.

More sightings (2026-09-28), same signature (Error: Network connection lost out of WorkersTestRunner.updateStackedStorage / ensurePoppedActiveTryStorage, workerd stream disconnected prematurely), now spanning more runtime-cloudflare files — atomic-suspend-write, do-stream-manager-paused, do-stream-manager-cleanup-to-step, do-stream-manager-contract (G1), forced-completion-do, and the filestore-filesystem / filestore-provider grep and R2 tests, several failing together in one job: main 16766713719; !281 16766313974; !282 16767233852 and 16767365277 (!282 changes no runtime-cloudflare or store-cloudflare file — only core's ScriptedModel test helper, which no cf test imports; a retry of 16767233852 passed as 16767441949).

FU-FLAKE-CFW-FORCED-COMPLETION-CONNECTION-LOST: forced-completion.workflow.test.ts fails with Network connection lost — ✅ fixed (CI health 2) ​

Fixed (CI health 2): the same updateStackedStorage loopback race as FU-FLAKE-POOL-WORKERS-STACKED-STORAGE-LOST (vitest.workflows.config.ts uses isolated storage), fixed by the same patch. Its later 120s timeout ("D-fix", main 16882767277) ran while the job rebuilt five Next.js apps (FU-CI-TEST-JOBS-REBUILD).

Test: packages/runtime-cloudflare/src/__tests__/workflows/forced-completion.workflow.test.ts (test:workflows, @cloudflare/vitest-pool-workers@0.9.14). Signature:Error: Network connection lost., with a burst of jsg.Error: Instance dispose (broken.outputGateBroken) in the log. That is the FU-FLAKE-WORKERD-INTROSPECT-TEARDOWN family, but that entry names other files. Jobs: !272 16758970357 ("runs terminal-text → forced finish"); before !272: !264 16750600721, !262 16754754857 and !266 16755705444 ("resumes in forced mode on a cold start"). !272 does not change the pool-workers version, the Workflows runtime or this test.

FU-FLAKE-REMOTE-AGENTS-EXAMPLE-USAGE-FETCH: remote-agents example logs a failed usage fetch after a workerd connection reset — open (owner: B8 session) ​

Update (CI health 2): seen on main 16765146699, 16765362858 and 16825373949 (2026-09-27 → 09-30, the same "no console/page errors" assertion after Uncaught Error: Network connection lost in wrangler dev), then 38 green runs in a row with no retry (this job has retries: 0). Not reproduced; left open, owner orchestrator → CI-health-3 (it is an example-app job).

Test: examples/remote-agents-cloudflare-do/tests-e2e/cross-service.spec.ts:22, the assertion at :103 ("no console/page errors across the full flow"). Symptom: the page logs [demo] usage fetch failed: TypeError: Failed to fetch right after workerd reports Connection reset by peer (os error 104) (wrangler 4.141) or Network connection lost (wrangler 4.63). Every earlier assertion passes, and the server stays up. Jobs: !272 16759512532; !268 16757428815 (the old wrangler, so pre-existing on main's tooling). Not this entry: the job's other four failures in the same window (16756390704, 16756192324, 16755700166, 16756092320) are a toBeTruthy state-fold stall. That is CP-77, owned by S5.

FU-FLAKE-REMOTE-AGENTS-CF-DO-SMOKE-TIMEOUT: the remote-agents cf-do example smoke times out at 60s under a loaded runner — open (owner: B8 session) ​

Test: examples/remote-agents-cloudflare-do/tests/smoke.cf.test.ts ("cross-service remote agents (DO → DO over a service binding) > completes the tree, folds child state + usage, snapshots the child, and projects the external mount"), in test:unit. Symptom: Error: Test timed out in 60000ms — no assertion failed. Job: !282 16767233838 (2026-09-28 00:15), in the same minutes the same runner's test:cloudflare (16767233852) was losing workerd connections (FU-FLAKE-POOL-WORKERS-STACKED-STORAGE-LOST). !282 changes no example, runtime or store code, and the smoke passed 3/3 locally at the same head. No earlier sighting of this exact signature was found in the recent failed test:unit jobs; if it recurs, check whether it co-occurs with the pool-workers connection losses (a shared workerd/runner load cause) before raising the timeout.

FU-FLAKE-D1-CAS-RACE-LOSER-D1STATEERROR: a D1 CAS-race loser rejects with D1StateError, not StaleStateError — open (owner: B8 session) ​

Test: packages/store-cloudflare/src/__tests__/integration/atomic-suspend-write.integ.test.ts:515 ("all losers throw StaleStateError when UNIQUE(messages.sequence) fires"), assertion :548. Jobs: main 16758460796 (e0aae15185); earlier !259 16756105477 and 16756470721, and !264 16756272942. Not proven to be the Miniflare proxy transport failure. D1StateError is the wrapper D1StateStore puts around any unexpected error, and the proxy's fetch failed surfaces that way elsewhere. But CI truncates the wrapped cause, and it did not reproduce locally in 0/20 runs on !272 (Miniflare 5, 10 of them under CPU stress) or 0/15 on main (Miniflare 3, under CPU stress). Next step: print r.reason and its cause in the assertion message, so the next CI occurrence says whether it is transport (resolved by the Miniflare 5 bump) or a UNIQUE-error classification gap in saveStateAndPromoteStaging.

FU-FLAKE-INTERRUPT-RESUME-FIXED-SLEEP: interrupt-resume e2e interrupts after a fixed 500 ms sleep — ✅ fixed (CI health 2) ​

Fixed (CI health 2): "should handle checkpoint restoration on resume" interrupts after the run's second step_committed marker and reads the checkpoints once after result() (no sleep, no poll). Seen again on main 16825373934 (2026-09-30). The file's other sleep-then-interrupt cases assert nothing that depends on how far the run got, so an early interrupt cannot fail them. Since S5 a run-start commit also writes a checkpoint, so the old "any checkpoint" assertion held even for an instant interrupt; the test now also asserts the latest checkpoint's stepCount is at least 2. With the old fixed sleep forced to 0ms it fails on that assertion (expected 0 to be greater than or equal to 2).

Test: packages/e2e/src/__tests__/interrupt-resume-d1-with-js.integ.test.ts, "Interrupt/Resume E2E - JS runtime + D1 store › Edge Cases › should handle checkpoint restoration on resume". It fails at :659 with expected 0 to be greater than 0: listCheckpoints is still empty after the interrupt.

Job: !269 job 16760751939 (test:integ:e2e 3/4, pipeline 2886547356). It passed 5/5 isolated locally.

Root cause: a test defect, the fixed-sleep proxy pattern. The test sleeps 500 ms, then calls interrupt(), and expects at least one checkpoint to exist by then. Under CI load the interrupt can land before step 1 commits. The old script_failure auto-retry hid this until !272 removed it.

Fix direction: wait for a real signal before interrupting, not a fixed 500 ms. Either poll for a persisted checkpoint, or hold a tool step on a gate and interrupt while it is held.

Related: !259's FU-DBOS-E2E-FIXED-SLEEP-PROXIES is the DBOS sibling. The other interrupt-resume-* lanes probably share the pattern and should be fixed together.

FU-CI-WORKERS-TYPES-5-PEER: wrangler 4.141 wants @cloudflare/workers-types@^5 — open ​

Owner: the B8 session. pnpm peers check reports wrangler 4.141.0's optional peer @cloudflare/workers-types@^5.20260925.1 unmet (the repo is on ^4.2026…, and the published packages declare ^4.0.0 as their own optional peer). Harmless today (optional peer, types only). Moving to v5 changes the published peer range, so it needs its own MR with a typecheck pass.

Exposure: integration shards that were green only through the removed script_failure retry ​

Sweep of the 45 pipelines before !272 (include_retried=true), integration shards whose first attempt failed with script_failure and whose automatic retry passed. With the retry gone, these show up as red until their fixes land.

Failing test (first attempt)Jobs (failed attempt)Cause / owner
e2e remote-agent-flow "duplicate start throws ALREADY_RUNNING"16755525951 (main), 16757763574, 16757391353, 16756192309, 16756128346, 16756105472, 16755700151FU-FLAKE-REMOTE-DUP-START-RACE, fixed by !271 (on main)
runtime-cloudflare concurrent-subagent-fold (append/modify/removal/failed-child)16757391357, 16756390693, 16756192313, 16754841601, 16750450972FU-FLAKE-CF-D1-PROXY-FETCH-FAILED, fixed by !272
e2e ai-sdk-redis-temporal "Non-streaming mode" (InvalidStoredStateError, status undefined)16755699318, 16755485917, 16754864347, 16754754845RM-50 (Redis createSession not atomic), fixed by !269
store-cloudflare durable-lock "only one holder when racing"16754754849, 16750857812FU-FLAKE-DURABLE-LOCK-RACE-RETRY + Miniflare 3 ECONNRESET, fixed by !272
e2e ai-sdk-d1-with-js (every later D1 call (message?.id === id))16755572339FU-FLAKE-MINIFLARE-SYNC-PROXY-DESYNC, fixed by !272
e2e multi-message-execute-dbos-persistent "idle-timeout exit then new message"16754864348CP-74 (DBOS persistent send into an exiting workflow), owned by B2-core
store-cloudflare atomic-suspend-write CAS race "all losers throw StaleStateError"16758460796 (main), 16756105477, 16756470721, 16756272942FU-FLAKE-D1-CAS-RACE-LOSER-D1STATEERROR (owner B8); not proven covered by the transport fix, see the entry
e2e client-tool-happy-parity dbos-postgres16758483177FU-FLAKE-DBOS-CLIENT-TOOL-PARITY-TIMING, owner B8, filed on !266

Repeats after the sweep window: CP-74 on main 16759539277 and on !266 16759542438 (owner B2-core); the fold failure 16758483181 (fixed here); durable-lock 16758324556 (fixed here); duplicate start 16758324552 (fixed by !271); RM-50 16757995539. RM-50 is fixed by !269, which is not merged yet, so ai-sdk-redis-temporal shards can go red until !269 lands (it merges right after !272).

Not in the table: CP-77 (unowned) and FU-CONF-CP62-DBOS-SETTLE-TIMING (conformance owner) did not appear as auto-retried integration shards in this window. The conformance jobs never auto-retried script_failure; !271's test:conformance:dbos 16756806141 and conformance:verify 16756806142 were retried by hand.


CI health 3 ​

Open items CI-health-3 found while chasing CI reds. They are not fixed in its MRs.

FU-FLAKE-CONF-EXECUTOR-DRIVER-DISPOSE-WALLCLOCK: conformance executor-driver.test.ts dispose bound is a wall-clock upper bound — ✅ fixed (!297) ​

Fixed (!297, conformance owner batch 10): the dispose bounds are now driven on vi.useFakeTimers() / advanceTimersByTimeAsync, and the test asserts which call hit which bound, the order, and the fake-clock settle instant (first call inside its bound, then drainBoundMs, then the hung after-drain call cut at cellWorkBoundMs) instead of elapsed milliseconds. The sibling wall-clock upper bounds in the same describe (the closed-but-unsettled reap case and the hung-first-call case) got the same treatment. The code under test reads only setTimeout, so no test seam was needed.

The original report follows.

Owner: the conformance owner. Surfaced by: !293 pipeline 2911240084.

  • Where: arm64 conformance:verify job 16923705133, gate step unit-tests, on jmif-dev-srv between 19:09 and 19:15 UTC.
  • Retry: job 16923979153, green, disclosed in !293.
  • What was running on the runner:
    • !293's other jobs;
    • no other helix-agents pipeline;
    • no CI-health-3 load runs;
    • other projects' jobs and local sessions: unknown.
  • Test: packages/e2e/src/conformance/__tests__/executor-driver.test.ts › "dispose() with env.terminateCellWork — 'stopped' judged from backend truth › I3/N2: the SLOWEST dispose path takes disposeWorstCaseMs — first call just inside its bound…" (from !290, 4a82e0b347).
  • Failure: expected 501 to be less than or equal to 490 at :3033, i.e. expect(took).toBeLessThanOrEqual(worst + TOLERANCE_MS). A positive wall-clock window that a loaded runner overshoots.
  • Plan (owner, done): drive the dispose bounds with fake timers and assert which call hit which bound, not elapsed ms.

FU-FLAKE-CFW-SPIKE-ITEM34-SLEEP-RACE: cfw spike.integ.test.ts items 3+4 raced its own 2 s sleep — ✅ fixed (!303) ​

Fixed (!303): the cell stays gating (its assertions are the evidence that engine state survives dispose and recreate) and now makes its timing window explicit:

  • all four instances are created first and awaited together, so they pass step a nearly at once;
  • the sleep is 4 s, not 2 s, and the "past the deadline" wait is relative to the latest recorded step-a time (max(tStepA) + SLEEP_MS + 1_500), not a fixed delay after the recreate;
  • a window guard records when each instance was first seen running with one step output and when evict() resolved. If any instance's sleep deadline had already passed by then, the cell throws a HarnessError naming the instance and the measured gap. That is a harness timing fault, never a pass and never an engine failure;
  • the title no longer says "recorded, not gating".

The cell's runtime went from about 5.9 s to about 9.7 s.

The original report follows.

  • Test: packages/e2e/src/harness/cfw/__tests__/spike.integ.test.ts › "cfw T0 spike › items 3+4 (recorded, not gating): engine state across dispose + recreate, and the re-init trigger".
  • Signature: control: expected { status: 'complete', stepOutputs: 1 } to deeply equal { status: 'running', stepOutputs: 1 }.
  • Where: job 16942478435, on !299's pipeline 2914044694.
  • Cause: the cell created its four instances one after another, each with a 2 s sleep. On a loaded runner the time from control passing step a to the end of the dispose exceeded 2 s, so control's sleep finished before the dispose and the engine correctly reported complete. The premise (every instance still asleep at the dispose) was never checked.

FU-FLAKE-CF-VITEST-ONTASKUPDATE-RPC-TIMEOUT: store-cloudflare test:cloudflare exits 1 on a vitest worker RPC timeout with every test green — open ​

Owner: CI-health-3.

  • Where: !292 pipeline 2910411150, arm64 test:cloudflare job 16919213090; retried as job 16919266809.
  • Signature: store-cloudflare 283 passed, 1 skipped, yet vitest exits 1 on the unhandled Error: [vitest-worker]: Timeout calling "onTaskUpdate". It hit right after state-store-contract.cf.test.ts, a 254-test file that ran 67.7 s, while 4 MR pipelines shared the runner.
  • Lead: vitest 2.1.9's worker → main birpc call (onTaskUpdate) has a fixed timeout. A long single file on a starved host misses it.
  • Fix direction: measure, then either split the contract file or fix the starvation. Don't raise the timeout.
  • Unrelated to !292's diff: store-cloudflare is untouched.

FU-STALE-ENTRY-RUNS-ON-COMPLETED-SESSION: a stale resume / retry entry starts a second run after the winner completed (CP-96) — done (!295) ​

Resolution (!295). Fixed in the store, on every runtime and store: the run-start commit carries two required StartRun fields, expectCurrentRunStatus (the run CAS covers the id AND the status) and entryStatuses (a session gate against the status the store reads inside the commit). A run-status mismatch is the C3 loser (AgentAlreadyRunningError); a session-status mismatch is RunStartRejectedError('entry_status_changed'). The rule and the per-runtime sources are in execution-flow.md §Entry gates; the design is 2026-10-04-cp96-entry-status-gate-design.md. Temporal, DBOS and CF Workflows were probed on the real engines: the interleave was already refused by the executor's owner check or the run-id CAS (green before and after); the new gates are proven there by between-plan-and-commit tests. The text below is the original report.

Catalogue: CP-96 (reserved; owner batch 10). It first surfaced as the e2e redis-cas-races.integ.test.ts › "Concurrent Resume Races › should allow exactly 1 executor to succeed when 5 race with 10ms stagger" (expected 2 to be 1, !290 job 16919048157), but it is not Redis-specific.

Mechanism. The status gate of a resume() / retry() entry is not atomic with its run-start commit, and the winner writes its active entry state only AFTER its commit (R29):

  1. Entry A's run-start commit makes run R1 current (running). The session still reads interrupted / failed until A's post-commit entry-state write.
  2. In that window, entry B reads seenRunId = R1 and passes its status gate on the stale status (runtime-js: resumeSession's friendly checks; the assertNoLiveOwner check runs only for a running session).
  3. A's run completes, so R1 is idle.
  4. B's run-start commit expects R1. The CAS matches and the owner guard passes (R1 is not live), so B wins and runs on a session A already finished.

Observed. runtime-js + InMemoryStateStore, deterministic: a gate on appendMessages({ startRun }) holds A after its commit until B reaches its commit, then holds B until A completes. The repro and logs are in the reliability dossier, handoffs/ci-health-3-data/cp96-repro/.

  • resume (continue): runs [interrupted, completed (A), completed (B)]. B appends a second finish turn after the terminal one, and the output is overwritten (a → b). A real agent would run its tools twice.
  • resume (with_message): the same; B's message and turn land on the completed session.
  • retry: worst. B's retry rewinds A's successful turn: the transcript keeps only B's [user, assistant, tool], so A's completed result is LOST.
  • execute: benign. B is a later new turn, which execute allows on a completed session anyway.

Scope.

  • runtime-js, every store: confirmed live.
  • Temporal (activities.ts plan gate ~:1084 / :1135; entry state written after the commit at ~:1278), DBOS (run-commit.ts :650 / :706; ~:862) and CF Workflows (steps.ts :948 / :1005; ~:1149): the same structure, plan-time gate plus post-commit entry state. A live probe is pending.
  • The CF Durable Object is safe: its claim block re-checks allowedStatuses inside the same transactionSync as the commit and writes the entry state there (R31, durable-object-agent-base.ts :630-650, entry_status_changed).

Fix options.

  1. Runtime-level.
    • Each entry takes the current run RECORD with its id. If that run is running, the entry applies the live-owner check whatever the session status says (JS: the lease; Temporal / DBOS / CFW: their owner probes).
    • A live winner then refuses the stale entry. An entry that saw an older run already loses on C3's expected_mismatch_idle no-retry rule.
    • Small, but it must be repeated on each runtime.
  2. Store-level (recommended).
    • StartRun carries the entry's accepted session statuses, and every store's run-start commit checks them atomically with the CAS: RunStartRejectedError('entry_status_changed'), which the DO already does.
    • Alternatively, the commit writes the entry's active status itself, as the DO's R31 does, so the window closes.
    • One rule in core's run-start-operations contract covers every runtime. It touches memory, Redis, Postgres, D1 and the DO store.

The choice between them is the user's.

FU-D1-MIGRATIONS-CONCURRENT-ADD-COLUMN: concurrent D1 migrators fail with "duplicate column name" — open ​

Owner: the CF-D2 store-integrity bucket (routed by the orchestrator). Surfaced by: the !304 audit (FU-PG-MIGRATIONS-TABLE-CREATE-RACE's D1 sibling), verified in source.

  • packages/store-cloudflare/src/d1-migrations.ts: applyMigrationAtomic batches a version's DDL and its version row atomically, but two isolates that both read version N-1 both batch it. V2/V3 are plain ALTER TABLE … ADD COLUMN, and addColumnsIfMissing checks pragma_table_info outside the batch, so the second batch fails whole with duplicate column name.
  • Impact: transient. The losing isolate's runMigration throws once; a retry finds the version recorded. No schema damage (the batch is atomic).
  • Fix direction: treat a duplicate-column failure whose version is now recorded as success (re-read the version), or guard the batch on the version row.

FU-MEMORY-CF-MIGRATIONS-NOT-ATOMIC: memory-cloudflare applies a migration's DDL and its version row separately — open ​

Owner: the CF-D2 store-integrity bucket (routed by the orchestrator). Surfaced by: the !304 audit, verified in source.

  • packages/memory-cloudflare/src/migrations.ts runMemoryMigration: a version's DDL runs in one batch and its version row in a second statement, a plain INSERT with no ON CONFLICT. V2's ALTER TABLE … ADD COLUMNs are not idempotent.
  • Concurrent isolates: the second fails on the version primary key or on duplicate column name.
  • Worse, a crash between the two writes: V2's columns exist but V2 is unrecorded, so every later runMemoryMigration fails on duplicate column name, permanently.
  • Fix direction: one atomic batch per version (DDL plus INSERT … ON CONFLICT(version) DO NOTHING), as store-cloudflare's applyMigrationAtomic does, and idempotent ADD COLUMNs.

CI health 2 ​

CI health 2 removed the last genuine-failure retries: test:cloudflare's retry: when: script_failure, the config-level vitest retry in runtime-cloudflare / store-cloudflare / memory-cloudflare vitest.integ.config.ts, and 23 per-test { retry: 2 } options (22 in e2e *.cf.test.ts, 1 in a runtime-cloudflare .wf-noiso file). It fixed the flakes those retries hid. It also stopped every test job rebuilding the repo (see development-workflow.md, "CI test jobs").

The sweep (the evidence base). 164 pipelines updated 2026-09-26 → 2026-10-03 (139 MR, 25 main), every job attempt (include_retried=true); traces of all 120 failed attempts behind a red-then-green job and of 198 example jobs. After !272, test:cloudflare's first attempt was red in 36 of 78 pipelines (7 of 14 on main), every one turned green by its retry:

FamilySignaturePost-!272CauseFix
Loopback keep-alive raceNetwork connection lost from WorkersTestRunner.updateStackedStorage19 of 36Miniflare 4.20251011.0's loopback server kept Node's keep-alive timeout (5s + 1s); on a starved runner Node closed a socket workerd still pooledpnpm patch (upstream workers-sdk#14850) — FU-FLAKE-POOL-WORKERS-STACKED-STORAGE-LOST
StreamManager contract wall clockexpected [] to deeply equal [1,2,3], expected 4146 to be less than 2000, drain: timed out after 5000ms, 'pending' ≠ 'resolved', endsWithin7positive wall-clock windows on DO proxy readsevent awaits + poll disabled — FU-FLAKE-DO-STREAM-CONTRACT-TRIMMED-AFTER-ATTACH, FU-FLAKE-CF-DO-STREAM-WALLCLOCK-BOUND
Fixed sleep before interrupt'Filter Run 1 content…' not to contain 'Run 1'; Timeout waiting for stream … 'ended' after 5000ms4 (+5 before)150ms sleep, then a 5s status deadlineinterrupt on the tool's tool_start — FU-FLAKE-CF-D1-GETCHUNKS-AFTER-RESUME, FU-FLAKE-CF-D1-DO-SEQUENCE-PRESERVE
Self-inflicted load[vitest-worker]: Timeout calling "resolveId" / "onTaskUpdate", CFW forced-completion 120s timeout2each test job rebuilt every package (five next builds) beside its teststest jobs run turbo --only — FU-CI-TEST-JOBS-REBUILD
Shared mock FIFO race (non-blocking child)parent 'failed' ≠ 'completed' in persistent-subagent-d1-with-js(e2e shard, !287 + !289)parent took the child's scripted finishper-agent scripts — FU-FLAKE-PERSISTENT-NONBLOCKING-SHARED-MOCK
CP-52 lifecycle-hooks parity (cf-do-d1)'suspended_client_tool' ≠ 'completed'3 (last 09-27)product bugfixed by S5 (!288)

Other jobs: test:integ:e2e interrupt-resume-d1-with-js checkpoint restore (FU-FLAKE-INTERRUPT-RESUME-FIXED-SLEEP, fixed here); the remote-agents example late-join console error (FU-FLAKE-REMOTE-AGENTS-EXAMPLE-USAGE-FETCH: last seen 2026-09-30, then 38 green runs with no retry; not reproduced, left open); conformance history.companion-continue|temporal-postgres "failed for a different reason" (main 16853177130; routed to the conformance owner).

Hidden Playwright retries (retries: 1, 2 on CI in resumable-streams-nextjs) are out of this MR's scope; see the two entries below (owner: CI-health-3).

FU-FLAKE-PERSISTENT-NONBLOCKING-SHARED-MOCK: persistent-subagent-d1 non-blocking spawn ends failed — ✅ fixed (CI health 2) ​

persistent-subagent-d1-with-js.integ.test.ts › "non-blocking spawn persists ref with spawned status to D1" (expected 'failed' to be 'completed'; !287 pipeline 2905237226, !289 job 16916000765). The PARENT run failed: the non-blocking child runs concurrently with it, and on the shared global-FIFO MockLLMAdapter the parent's second model call could take the child's scripted finish({ result }), which its output schema rejects (invalid_type at summary; call order parent, parent, child). 4 of 25 runs failed under CPU load. The test now uses per-agentType RoutingMockLLMAdapter scripts, as the file's other concurrent cases already did; 25 of 25 green under the same load.

FU-E2E-DBOS-LAUNCH-RACE: concurrent first DBOS.launch() on a fresh database — ✅ fixed (CI health 2) ​

Reproducing FU-E2E-DBOS-LOAD-SENSITIVITY on a fresh Postgres under load, two forks' first launches ran DBOS's schema migration concurrently and failed (duplicate key value violates unique constraint "dbos_migrations_pkey", then deadlock detected exhausting the 8 transient retries); CI's integration shards start on a fresh Postgres too. createDBOSE2EContext now serializes launches with a Postgres advisory lock, taken by polling pg_try_advisory_lock: a blocked pg_advisory_lock call is an open statement, and DBOS's migration runs CREATE INDEX CONCURRENTLY, which waits for every open statement, so the two waited on each other (caught live: virtualxid vs advisory). 15 of 15 fresh-database runs green (under CPU load) after, against 3 of 4 failing before.

FU-CI-TEST-JOBS-REBUILD: every CI test job rebuilt the whole repo beside its tests — ✅ fixed (CI health 2) ​

turbo's test tasks depend on build, and turbo has no cache in CI, so pnpm run test:cloudflare (and test, test:integration, test:workflows, typecheck, lint) rebuilt every workspace package in each job, concurrently with the tests: about 40 tasks, five next builds and the docs among them. Tests that take ~10ms locally took 1–13s; vitest's own worker RPC timed out (60s) twice. Every test job now runs turbo with --only on the build job's artifacts, and the Next.js example jobs build only their own app. Evidence: the MR's before/after job-duration table.

FU-FLAKE-OPENNEXT-REFRESH-PLAYWRIGHT: opennext refresh tests pass only on their in-run retry — open (T1) ​

Owner: orchestrator → CI-health-3. Measured: 16 of 48 test:examples:opennext-cloudflare-do jobs reported a flaky test: comprehensive-refresh.spec.ts:691 "rapid refresh during tool transitions" 5× (incl. main), :617 "very early refresh" 2× (incl. main), :600 "AGGRESSIVE refresh after follow-up completion" (flaky on !288 16898206128, red on !287 16913998804), mid-text-reload-scoring #1/#3/#4/#5/#10/#13, refresh-exact-messages A, api-consistency :300/:344 and stream-resume-consistency :273 (several on S5's own !288 before it merged). Needs a post-S5 re-measure at --retries=0.

Sub-project pipeline ​

Sub-projectStatusDescription
DdoneTest infrastructure overhaul (shared harness + capability sets)
A.1doneStateless purity cleanup (re-port to core helpers)
A.2doneTemporal HITL rewrite (workflow exits on HITL boundary)
A.3doneDBOS resume() contract bug fixes (F7) — closed 7d9498b8e
B (B3 cleanup)doneFU-A2-01 + FU-A2-03 closed; F6 → future-work
CdroppedOriginally listed as "generic transport adapter"; the transport surface (RemoteAgentTransport interface + HttpRemoteAgentTransport + DOStubTransport) is already transport-agnostic with no surfaced demand for additional implementations. C had no spec, plan, or concrete deliverables — placeholder only. New transports (WebSocket, gRPC, etc.) get fresh specs when concrete demand surfaces.
EpendingDoc reconciliation

Done items ​

(closed items are moved here with their closing-PR reference)

FU-PG-MIGRATIONS-TABLE-CREATE-RACE: concurrent PostgresMigrations runs crashed at start on a fresh schema — done ​

Closed by: !304 ("store-postgres: concurrent migrators are safe").

  • Where it failed: main pipeline 2914016750 (41cd3fbbc0), test:integ:e2e 2/4, job 16942254650. cross-runtime-subagent-resume-matrix-suspended S6 (js + postgres + redis) died in setup: duplicate key value violates unique constraint "pg_type_typname_nsp_index" at ensureMigrationsTable.
  • Cause (store):
    • run() / getCurrentVersion() created the migrations table with a bare CREATE TABLE IF NOT EXISTS, outside any lock. Concurrently, both sessions pass the existence check and the loser fails on the catalog's unique index.
    • Behind it, overlapping migrators (only phases 1 and 3 were serialized):
      • a CREATE INDEX CONCURRENTLY waits for every older transaction, so it deadlocked (40P01) against another migrator's phase-1 transaction waiting for the build's table lock (a no-op ADD COLUMN IF NOT EXISTS still takes ACCESS EXCLUSIVE), and against a second migrator's build of the same index. Under host load, 8 migrators re-collided on every attempt;
      • the INVALID-index cleanup dropped a peer's live build (a build in progress also reads indisvalid = false), and could drop a peer's rebuild under the same name from a stale scan;
      • a migrator whose IF NOT EXISTS no-oped on a peer's unfinished build recorded the version anyway.
    • A tablePrefix over 34 characters got truncated identifiers; from 39 on, two index names collided.
  • Cause (e2e): createPostgresTestContext() named its prefix e2e_${testId ?? Date.now()}_. The six cross-runtime-subagent-resume-matrix-* files build a postgres combo with no testId in parallel workers, so two that built in the same millisecond shared tables, and either one's cleanup() dropped the other's.
  • Fix (user decision, option B):
    • Migration lease: one row in <prefix>migration_lease, taken with an atomic upsert, renewed every leaseTtlMs / 3, released at the end, taken over after a crashed holder's ttl. Phases 1 and 3 check it with a renew-and-check as each phase transaction's last statement (clock_timestamp(), so a long phase never blocks its own renewal; MigrationLeaseLostError, retried). One migrator at a time: neither deadlock can form.
    • The migrations and lease tables are created under pg_advisory_xact_lock.
    • peerWaitMs (default 600 s): one run's total wait for the lease and for a build still running on its tables (found via pg_stat_progress_create_index, scoped through pg_locks, so another role's unrelated build never stalls it). Out of budget: MigrationLeaseTimeoutError / MigrationIndexNotReadyError; the timed-out attempt records nothing.
    • Cleanup: each INVALID candidate is retired under SHARE UPDATE EXCLUSIVE on its table (no live build of any role lets it be taken; lock_timeout means busy, skip), re-checked by OID and renamed to idx_<prefix>dead_<oid>; only that name is dropped.
    • Phase 3 records a version only when every concurrent index it builds is valid (MigrationIndexNotReadyError).
    • run() re-enters on 40P01, MigrationIndexNotReadyError or MigrationLeaseLostError (5 attempts, jittered backoff, warn postgres.migration.retry); applied accumulates over attempts. PostgreSQL < 12 is refused.
    • MAX_TABLE_PREFIX_LENGTH = 34, enforced by every store (breaking for longer prefixes; upgrade note in the changeset). The repo's DBOS e2e prefixes and tii_h6_… were shortened.
    • e2e: the prefix gets a random suffix; cleanup() lists also drop abandoned_run_starts and migration_lease.
  • Red-first (each fails on the code before its fix):
    • migrations-concurrent-start.integ.test.ts: concurrent getCurrentVersion() and run() (the CI signature on main; then 40P01 exhaustion under load until the lease; the test now asserts zero retries); another role's unrelated build (waited out the whole budget); a stalled holder's stale scan (dropped its successor's index: OID changed); the 34 / 35-character prefixes (index shapes vs a short prefix; the old 58-character test passed vacuously on colliding names).
    • New behaviour: the lease contract (crashed-holder takeover, a live holder never taken, renewal past the ttl, release), an orphan build waited for, and a waiting migrator that completes.
    • Regression guard (passes on main too): a cancelled build replaced with a valid index.
    • postgres-migrations.test.ts: the peer-wait budget spans the whole run (52 s for a 10 s budget when it was per attempt); the prefix bound and a guard that fails any future migration identifier over it.
    • postgres-test-context-isolation.integ.test.ts: same-millisecond contexts shared e2e_1700000000000_, and cleanup() left abandoned_run_starts.

FU-PG-MIGRATIONS-SESSION-LOCK: serialize a whole PostgresMigrations.run() on one pinned connection — superseded ​

Closed by: !304. The migration lease serializes a whole run without a pinned connection (a session-scoped pg_advisory_lock needed a pin-a-connection method on PostgresAdapter, an interface change for every custom adapter). Nothing is left to do.

FU-CI-NATIVE-BUILD-CPU-FEATURES-HANG: a CI job hung 90 min inside pnpm install on a native build — done ​

Closed by: !296.

  • Where it failed: main pipeline 2911492581 (7bca6187bd), test:integ:e2e 2/4, job 16925087186. It hit server_timeout_running after 5400 s.
  • Signature: the trace ends inside pnpm install at cpu-features install: node-gyp rebuild … gyp http GET https://nodejs.org/download/release/v24.21.0/node-v24.21.0-headers.tar.gz. Another job's gyp ERR! not found: make is the same family.
  • Cause:
    • pnpm-workspace.yaml allowBuilds let cpu-features and ssh2 run their install scripts (!252 kept npm's set).
    • So every CI install compiled them with node-gyp. With a fresh ~/.cache/node-gyp per job, that downloads the Node headers from nodejs.org with no timeout of its own.
    • That GET stalled on the amd64 runner.
  • Fix:
    • Both are now false in allowBuilds. They are optional native accelerators reached only through workspace-docker → dockerode → docker-modem's SSH transport, and ssh2 falls back to pure-JS crypto.
    • The root node-gyp devDependency, which served only those builds, is removed.
    • CI's .pnpm_install runs timeout 900 pnpm install --frozen-lockfile.

FU-FLAKE-APPROVAL-GATES-PLAYWRIGHT: approval-gates Playwright tests passed only on their in-run retry — done ​

Closed by: CI-health-3's second MR.

  • Measured: 18 of 50 test:examples:approval-gates jobs (2026-09-29 → 10-03) had a flaky test.
  • Signature: send-button … element is not enabled right after input.fill(…).
  • Cause: the tests typed before React hydrated the page under next dev. The controlled input's state stayed empty, so Send stayed disabled.
  • Fix:
    • The page exposes a data-hydrated app-ready marker (set in a useEffect), and every spec waits for it before typing.
    • tests/hydration-guard.spec.ts makes hydration deterministically late (JS chunks held back, goto at domcontentloaded). Without the wait it fails with the CI signature 2/2; with it, it passes.
    • retries is now 0.
  • Proof: the suite passed 15/15 at --retries=0 beside a forced example rebuild, at load average 23–121.

FU-B2-05: CF Workflows cross-turn proofs limited by CP-67 / DI-18 — done ​

Closed by: CF-D1 MR 1c (!291). B2's CF Workflows history, from_checkpoint and retry cases now run across turns on the real engine: runtime-cloudflare/src/__tests__/history-b2.wf-noiso.test.ts, describe('FU-B2-05 cross-turn') (C1x: a second execute() sees turn 1's user message and tool activity; C3x: retry() of a failed turn 2 re-runs turn 2's trigger once with turn 1 kept; C7x: from_checkpoint to the turn-2 input checkpoint rewinds within turn 2 only). They cover existing behaviour, so they pass on main as well; 5× green on head.

Update (CF-D1 MR 1a): a second execute() turn now runs in its own instance and reactivates the ended stream (CP-88), and handle.resume() reactivates it too (CP-67); the real-engine suite runtime-cloudflare/src/__tests__/cfw-terminal-truth.wf-noiso.test.ts proves cross-turn runs. B2's history / from_checkpoint / retry cases can now be re-run as cross-turn tests.

Update (CF-D1 MR 1b, !287): DI-18 is fixed: handle.result() waits for a real outcome (no 5-minute cap, waits through an operator pause) and never caches a rejected poll, so it no longer limits these assertions. The B2 cross-turn re-run itself was not part of !287; owner: CF-D1 MR 1c.

Originally: CFW never reactivated an ended root stream for a plain follow-up turn (CP-67), so a second turn hung on real workerd. B2's CFW history, from_checkpoint and retry proofs are therefore multi-step single-instance tests, not cross-turn ones. DI-18 (handle.result() after retry()/resume() can memoize a bogus {failed, 'Unknown error'} under local-dev Workflows) also limits what the workerd tests can assert. Re-run the CFW cross-turn cases when CP-67 is fixed.

FU-CFW-SESSION-ID-REUSE-AFTER-DELETE: a reused session id collides with its old instances — done ​

Status: done. Closed by: CF-D1 MR 1c (!291) for the sub-agent half; MR 1b (!287) closed the root half. Surfaced by: CF-D1 MR 1a DoD audit (D4).

Resolution (CF-D1 MR 1c, !291, design spec C7). Every CF Workflows child instance id is now keyed on the parent run (respawnSuffix(runId)): a sub-agent spawn is …__spawn__<run suffix>, a companion spawn …__spawn__<run suffix>-<callId> (the call id too, because a persistent child's session id repeats within one run when a failed child is re-spawned), a companion continuation / resume …__continue__<run suffix>-<step>-<callId> / …__resume__<run suffix>-<step>-<callId>; respawns already were. A re-executed child-create step whose create() already landed adopts its own instance (already_exists, one shared createChildInstance helper for all five sites). Proof: the real-engine case cfw.child-id-unique-after-session-reuse (C7) in runtime-cloudflare/src/__tests__/cfw-stop-conditions.wf-noiso.test.ts (execute → deleteSession → execute with a repeated tool-call id; red on main: the second run hung on the old child id), plus workflow.test.ts "child create is idempotent across a step re-execution". The child SESSION still survives the parent's delete: FU-DELETE-SESSION-ORPHANS-CHILDREN.

  • Done (snapshot MR + CF-D1 MR 1b). Root turn / resume / retry instance ids are keyed on a random key minted by the call (__turn__<key>), never on a counter that restarts with the session, so a reused ROOT session id always builds fresh ids. CF-D1 MR 1b pins it on the real engine: cfw-terminal-truth.wf-noiso.test.ts "FU-CFW-SESSION-ID-REUSE-AFTER-DELETE" runs execute → deleteSession → execute on the same id three times, each a fresh completed session. The planned per-session-lifetime nonce is therefore not needed for root ids.
  • Was open (closed by MR 1c). A sub-agent's base instance id is agent__<type>__<childSessionId>, and the child session id is derived from the parent session id and the tool call id. If a reused parent session's model emits a tool call id it emitted before the delete, the child's base id is already taken. Real providers mint unique tool call ids, so this needs a deterministic model (a scripted test model, or a provider that numbers calls per conversation). Fix: key child instance ids on the parent run's key (as respawns already are), or on a nonce stored on the session row at creation.

FU-FLAKE-CP52-ALARM-SUSPEND-HOOK-WINDOW: runtime-cloudflare CP-52 run-guard tests raced a timed onAgentSuspended window — done ​

Closed by: !292 (CI-health-3 EVICT-02).

  • Where it failed: main went red after !291 merged (pipeline 2911027602, test:integ:rest job 16922542276): do-run-guards.integ.test.ts › "CP-52 durable wake: the execution loses its pending wake (eviction) — the alarm, not a request, completes the run", waitFor(suspension-settled) timed out, last state completed.
  • Not an !291 regression: an A/B of 28427e5281 against 5bb1eaaf2b passed 5/5 each.
  • Not a product bug: the run completed correctly, with 2 runs and the tool result delivered once.
  • Cause: a wall-clock test window.
    • /test/suspend-hook-delay {ms: 500} held onAgentSuspended for 500 ms.
    • Inside it, the test had to poll, then submit, then evict.
    • On a loaded runner the window expired first, so the still-registered pending wake continued the run and the intermediate suspended_client_tool state the test waited for never appeared.
    • A 700 ms pause between submit and evict reproduces the CI signature.
  • Fix (test-only):
    • The helper worker holds onAgentSuspended on a gate (/test/hold-suspend-hook, /test/release-suspend-hook) that the test releases after its submit / evict / wake-hold.
    • All four timed-window tests in the file use it: the alarm test, the CP-52 race, the submit-during-hook test and the R42 partial batch. Each could otherwise fail, or quietly degrade into its own control case.

FU-FLAKE-JS-LEASE-BLOCKED-RENEWAL: runtime-js lease "condition 1" took the session over under load — done ​

Closed by: !292 (CI-health-3 EVICT-02).

  • Where it failed: runtime-js/src/__tests__/lease-takeover-fencing.test.ts › D2 "condition 1: a model call blocked for > 3 lease periods keeps the run live" failed in !290's arm64 test:unit (job 16919048155): expected { …(12) } to be an instance of AgentAlreadyRunningError.
  • Cause (test design, not a product bug):
    • The test ran a real 150 ms lease with real-time renewals (every 50 ms) across a 600 ms sleep and the intruder's whole entry path.
    • An event-loop stall longer than ttl − ttl/3 (~100 ms) on a loaded host lapses the lease on the store clock, so the takeover is legitimately accepted. That is the documented JS lease contract.
    • A 2-TTL busy-wait before the intruder reproduces the exact CI assertion.
  • Fix:
    • Date (the store's lease clock and the lease's local clock) is frozen.
    • It advances by ttl/3 only after a renewal is observed landing (a spy on renewRunLease), across 12 periods (4 TTLs).
    • The same injected stall now passes.

FU-EVICT-02: a wake-recovered DO child never ends its stream (the "wake race" e2e) — done ​

Closed by: !292 (CI-health-3 EVICT-02). Catalogue: MA-22 (reserved; catalogued in owner batch 10), open with fix evidence until a cf-do cell flips it.

What it was. The e2e cross-service-do-do-recovery › "FU-EVICT-02 wake race" failed 4 of 5 runs on main after a heavy local run (expected 'running' to be 'completed', the consumer session) and passed when the machine was idle. The original hypothesis (a 409 on the ephemeral remote /start surfacing as a failed result) was not the cause. There were two problems:

  • A product bug (MA-22).
    • A DO child (a remote child, or a sub-agent in its own DO) has a parentSessionId, so the JS run loop never ends, fails or pauses its stream: it assumes the stream is the parent's.
    • DurableObjectAgentBase ended the child's own stream only after a COMPLETED run on /start, /resume and /retry (the old "CC58" block). The wake auto-resume drive (ensureExecutionContinuesImpl) did nothing, and a run that ended failed was never failed on any drive.
    • So an evicted child recovered to completed (status, output and every chunk written) with no end frame, and its parent's attached stream stayed running until the remote timeout (10 min).
    • Fix: one finalizeOwnStreamAfterDrive after every drive, wake included. It ends on completion, fails with the run's errorDetail on failure and pauses on suspension, each fenced to the run (CP-85).
  • A hollow test.
    • The wake-race and evict-single-step scripts put slow_tool in the same step as __finish__. Since RM-29 C2 a call co-emitted with __finish__ is never executed, so there was no held-open step, and the child finished in ~15 ms.
    • The lost-POST re-POST then landed on a COMPLETED child and started a second turn instead of the designed 409 → attach. That is CP-95, /start not idempotent: a contract decision for the cross-service bucket.
    • The eviction landed after everything had finished when the machine was idle (green), and inside that second turn when it was loaded (red).
    • Both scripts now hold the step open with a test-released gated_tool (producer worker /__test/release), and __finish__ is its own step. The tests evict and release on observed chunks, never a timer.
    • The D7 lost-POST test also asserts exactly one run_started on the producer.

Evidence.

  • runtime-cloudflare do-eviction-recovery.integ.test.ts: 3 new tests (woken child completes → ended; woken child fails → failed; live child fails → failed) and an ended assertion on the companion child. 4 red on main, 14/14 green with the fix.
  • e2e wake race: red on main (the consumer running at 40 s, CF-D1c's exact signature), green with the fix.

FU-OPENNEXT-EXAMPLE-FLAKY-REFRESH-TESTS: opennext Playwright suite is green only through its in-run retry — done ​

Closed by: the S5 snapshot MR (branch reliability-example-refresh-flake, design spec docs/superpowers/specs/2026-09-26-consistent-snapshot-checkpoint-truth-design.md).

Resolution (S5, plan Task 24): the snapshot is one checkpoint-anchored read and the example's session pages use the canonical client (snapshotResumeOptions(snapshot): seed without the in-flight message, resume from the cursor with replayPrefix; useHelixChat with resync; the count-unchanged refetch workaround deleted). All 7 test.fixme markers are removed and each test passed --repeat-each=15 --retries=0 (15/15 each, on the paced replay mock the suite now runs with, and again on the unpaced mock). The new mid-text-reload-scoring.spec.ts scores no-loss / no-duplication / cursor-and-id-match per reload. Red on origin/main (a118f8e17e, example built from base, unpaced, 6 workers × 50): 24 of 1450 runs lost the in-flight text permanently (rendered "" vs the final text). The retries: 1 in playwright.config.ts is unchanged (out of scope here).

Re-measured after the rebase onto 5bd30e20da (CF-D1 MR 1a): each of the 7 tests again passed --repeat-each=15 --retries=0 (15/15 each), as did the CP-52 and CP-74 cases (15 separate vitest run --retry=0 processes each). Stress (--workers 6 --repeat-each 50 --retries 0): comprehensive-refresh 650/650 (paced mock); mid-text-reload-scoring 1150/1150 (1000 mid-run reloads scored); the unpaced files of the base red run 1449/1450, the one failure a workerd Network connection lost on the first send POST (no session started; see FU-S5-WORKERD-FIRST-SEND-CONNECTION-LOST), zero content losses. The full opennext suite passed 64/64 with retries 0. Logs: tmp/s5-evidence/un-quarantine-*{,-rebased}.log and green-stress-6x50-branch-*.log (local, not committed).

Quarantine (2026-09-27, user decision): these 7 tests fail on the RM-08 torn snapshot often enough to block unrelated MRs. GitLab merges only on a green pipeline, so they are marked test.fixme. Each has a comment that names RM-08, this entry and the owner:

  • examples/opennext-cloudflare-do/tests/comprehensive-refresh.spec.ts "refresh mid-stream: resumes and completes correctly" (was :370)
  • comprehensive-refresh.spec.ts "refresh during follow-up: turn 1 preserved, turn 2 completes, no cross-pollination" (was :472; see also FU-OPENNEXT-REFRESH-DURING-FOLLOWUP-LOSES-TURN2-TEXT)
  • interleaved-content.spec.ts "preserves all text content on refresh mid-stream @slow" (was :223)
  • interleaved-content.spec.ts "text content survives refresh after first tool appears @slow" (was :345)
  • api-consistency.spec.ts "snapshot should have consistent tool count after stream completes" (was :161)
  • interleaved-content.spec.ts "aggressive refresh during tool execution preserves content @slow" (was :287; reported as :293; added later on 2026-09-27 after it flaked on !279 and !276)
  • comprehensive-refresh.spec.ts "AGGRESSIVE refresh loop: state converges to stable completion" (was :400; added later on 2026-09-27 after it flaked on !279 and !276)

Exit condition: S5's snapshot MR removes all 7 test.fixme markers. The same MR shows all 7 green at --repeat-each=15 --retries=0 or more. This entry stays open until that MR merges. The other tests named below are not quarantined.

Owner: S5 (reliability-example-refresh-flake). Every test listed below is a refresh-during- or refresh-after-stream consistency test, the same family as :472, whose root cause is RM-08 (torn snapshot); S5's checkpoint-anchored snapshot work covers them. playwright.config.ts sets retries: 1. Over the 36 opennext jobs in the 45 pipelines before !272, 23 had 1-3 tests that failed once and passed on the retry. Most frequent: comprehensive-refresh.spec.ts:472 (refresh during follow-up, 14 jobs; now its own entry, FU-OPENNEXT-REFRESH-DURING-FOLLOWUP-LOSES-TURN2-TEXT), interleaved-content.spec.ts:287, comprehensive-refresh.spec.ts:520 and :400. !272's first pipeline had two: api-consistency.spec.ts:161 ("assistant message has no content") and comprehensive-refresh.spec.ts:348 (120s timeout). FU-CF-DO-RESUMABLE-READER-DROPS-FRAMES is a likely cause of the refresh-during-follow-up family; re-measure after !272 lands before chasing the rest.

FU-OPENNEXT-REFRESH-DURING-FOLLOWUP-LOSES-TURN2-TEXT: refresh mid-stream in turn 2 leaves turn 2 with no text — done ​

Closed by: the S5 snapshot MR (branch reliability-example-refresh-flake, design spec docs/superpowers/specs/2026-09-26-consistent-snapshot-checkpoint-truth-design.md).

Resolution: the root cause was RM-08 (a torn snapshot): a reload while turn 2 streamed could read a snapshot whose cursor was past turn 2's committed text, or that was ended without its last step, so the resume delivered nothing for that text. The snapshot is now one checkpoint-anchored read (messages = the committed rows up to the checkpoint, streamSequence = that checkpoint's cursor), and the example resumes with the canonical client (snapshotResumeOptions, the C-1 replay prefix). A second, independent cause (FU-CF-DO-RESUMABLE-READER-DROPS-FRAMES) was fixed by !272. Evidence: the test (no longer quarantined) passed 15/15 at --repeat-each=15 --retries=0, before and after the rebase onto 5bd30e20da, and comprehensive-refresh.spec.ts passed 650/650 under 6 workers × 50 repeats with no retries (tmp/s5-evidence/un-quarantine-comprehensive-refresh-refresh-during-follow-up{,-rebased}.log, green-stress-6x50-branch-paced.log). Base a118f8e17e lost text permanently in 24 of 1450 stress runs. See FU-OPENNEXT-EXAMPLE-FLAKY-REFRESH-TESTS.

Owner: S5 (reliability-example-refresh-flake), ruling by the reliability controller, 2026-09-26. Root cause: RM-08 (torn snapshot); S5's checkpoint-anchored snapshot work is the fix. Test: examples/opennext-cloudflare-do/tests/comprehensive-refresh.spec.ts:472 "refresh during follow-up", assertion :517 (getTextParts(turn2Msg).length >= 1, received 0).

Evidence:

  • Local, same machine, --repeat-each=15 --retries=0: 10/15 fail on main's wrangler 4.63.0 and 2/15 on !272 (wrangler 4.141.0 plus the FU-CF-DO-RESUMABLE-READER-DROPS-FRAMES fix). Every failure is at :517.
  • CI: flaky (failed once, then passed on Playwright's in-run retry) in 14 of the 36 opennext jobs in the 45 pipelines before !272, including main 16757438431 and 16754640438. On !272 it failed both attempts in job 16758515639.

What it is: after a reload while turn 2 is streaming, and after the stream ends, the second assistant message has no text part. Turn 1 is intact. It is a content-loss bug on the resume/refresh path, not a test-timing issue: the test waits for waitForStreamEnd before it asserts. Next step: capture the snapshot plus resume SSE (the zz-diag network capture used for the reader bug works) on a failing iteration.

FU-CP81-CF-DO-LIFECYCLE-HOOKS-QUARANTINE: cf-do-d1 lifecycle-hooks-parity Scenario 2 quarantined on CP-52 (filed as CP-81, withdrawn as its duplicate) — done ​

Closed by: the S5 snapshot MR (branch reliability-example-refresh-flake, design spec docs/superpowers/specs/2026-09-26-consistent-snapshot-checkpoint-truth-design.md).

Resolution (S5, plan Task 24): the CP-52 fix landed in S5's snapshot MR (a submit landing inside onAgentSuspended is never dropped; durable wake, bounded by the crash-loop budget). The gateTracingContextPersistence: true line and its comment are removed from lifecycle-hooks-parity.cf.test.ts; the cf-do-d1 Scenario 2 case passed 15/15 separate vitest run --retry=0 processes. Removed from the quarantine register.

Owner: S5 (reliability-example-refresh-flake). The user assigned the CP-81 fix to S5's snapshot MR. CP-81 was withdrawn as a duplicate of CP-52 (same cause and fix site), so the fix and the un-quarantine cite CP-52.

Quarantined (2026-09-27, user decision): only the cf-do-d1 case of packages/e2e/src/__tests__/lifecycle-hooks-parity.cf.test.ts › "tracingContext set by onAgentSuspended persists through resume → completion". The entrypoint sets the existing gateTracingContextPersistence: true option, which skips Scenario 2 for that entrypoint only. The JS, Temporal and DBOS cases (lifecycle-hooks-parity.integ.test.ts) and cf-do-d1 Scenarios 1 and 3 are unchanged.

Failure: expected 'suspended_client_tool' to be 'completed'. CP-52 (was CP-81): the CF DO drops a client-tool submit that arrives while the onAgentSuspended hook is running. ensureExecutionContinues returns early on isExecuting (packages/runtime-cloudflare/src/do/durable-object-agent-base.ts:653), and the run exits suspended with nothing to wake it. CI jobs: 16758514869, 16759542448, 16760323616, 16764206430.

Exit condition: S5's snapshot MR fixes CP-52 and removes the gateTracingContextPersistence: true line (with its CP-52 comment) from the cf entrypoint. The same MR shows the case green at 15 or more repeats with no retries.

FU-DBOS-PERSISTENT-SEND-INTO-EXITING-WORKFLOW: DBOS persistent message sent during idle-timeout exit is lost (CP-74); test quarantined — done ​

Closed by: the S5 snapshot MR (branch reliability-example-refresh-flake, design spec docs/superpowers/specs/2026-09-26-consistent-snapshot-checkpoint-truth-design.md).

Resolution (S5, plan Task 24): the CP-74 fix landed in S5's snapshot MR (R43: the run id is the single dedupe key, and a late inbox read follows the idle finalize). The it.skip and its comment are removed; "idle-timeout exit then new message starts a fresh persistent workflow" passed 15/15 separate vitest run --retry=0 processes on Postgres. Removed from the quarantine register.

Owner: the DBOS persistent-mode session (CP-74 fix). Finding: CP-74 (packages/e2e/src/conformance/catalogue/control-plane/CP-74.ts). A message sent into a persistent workflow that is exiting on its idle timeout is lost. The session already reads completed, but the workflow is still DBOS-PENDING, so execute() sends into an inbox that nothing reads again.

Quarantined (2026-09-27, user decision): only packages/e2e/src/__tests__/multi-message-execute-dbos-persistent.integ.test.ts › "idle-timeout exit then new message starts a fresh persistent workflow", as it.skip (vitest 2 has no fixme) with a CP-74 comment. It failed both attempts on !275. The rest of the file still runs.

Exit condition: the CP-74 fix's MR removes the it.skip and its comment, and shows the test green at 15 or more repeats with no retries.

FU-CONF-STOPWHEN-PARITY: stopWhen honoured only on JS and the CF DO — done ​

Closed by: conformance catalogue batch 6 (branch conformance-catalogue-batch-6), which verified the lead and catalogued it as DI-29 (S1, open). See DI-29 in docs/conformance/findings.md.

The batch-5 reviewer's F9 source lead held: Temporal and DBOS call shouldStopExecution with maxSteps only, and CF Workflows never calls it. A probe with stopWhen: () => true stopped JS after 1 LLM call, while Temporal and DBOS made 3 and CF Workflows ran to maxSteps 5 and failed. The docs now carry DI-29 caveats (the agents guide, execution-flow, and the Temporal, DBOS and CF Workflows runtime pages). The runtime fix itself is tracked by DI-29, not here.

FU-CONF-TEMPORAL-REAP-LATE-CHILD — a child start the server records only after the reap's lookup bound — done ​

Closed by: MR !275 (branch conformance-temporal-drain-children), in its Greptile P1 fix.

Gap:

  1. A parent's final history can record StartChildWorkflowExecutionInitiated with no ChildWorkflowExecutionStarted or StartChildWorkflowExecutionFailed after it.
  2. The reap looked such a child up for 1s, then dropped it and forgot the closed parent.
  3. A child started later escaped every later reap, dispose included.

Fix:

  • StartedWorkflows.terminateTree now KEEPS each start it could not resolve.
  • Every later call looks every kept start up again first ("pass 0"). That includes the next cell's reap about other sessions, and dispose.
  • Once the child exists, pass 0 walks and terminates it, attributed to its own session.
  • A start still unresolved at dispose is reported (CellWorkTermination.unresolvedChildStarts, plus a console.warn) and is never dropped.

Temporal semantics checked:

  • temporalio/temporal's processStartChildExecution returns early when the parent is no longer running and the child has not started, for every ParentClosePolicy including ABANDON. Its comment reads: "ideally for case 3, we should continue to start child … we can't do that".
  • So a late start is possible only from a start task already in flight when the parent closed.
  • The Docker dev server runs this same server code.

Testing:

  • There is no live test. The in-flight window is too narrow to force, and the harness has no hook to pause the server's start task.
  • It is covered by fake-history unit tests (temporal-client-adapter.test.ts, "P1: …"). They failed before the fix: the late child was not reaped, and pendingChildStarts did not exist.

FU-CONF-TEMPORAL-DRAIN-GROUND-TRUTH — judge a conformance cell's Temporal work "stopped" from backend truth, and end the whole workflow tree — done ​

Closed by: branch conformance-temporal-drain-children, the conformance harness MR "reap a cell's whole Temporal workflow tree at cell end; judge drain from backend truth", plus its fix rounds 1 and 2.

Problem: the executor driver judged "stopped" from whether a product handle's promise settled, and ended only the root workflows its client had started. On the temporal-* lanes, a persistent-companion child workflow could outlive its cell. A parent stop never reaches a blocking companion (CP-23), and a non-blocking one is ABANDONed. That orphan kept calling the model and writing to its torn-down parent stream. history.companion-continue (B2-core) produced 20–21 (test server, across two before-runs) / 24 (Docker) unhandled "Cannot write to stream … in 'failed'/'ended' state" rejections from runtime-temporal activities.ts on the full temporal file. test:conformance exited 1 although every cell passed.

Fix:

  • New optional hook BackendEnv.terminateCellWork, implemented only by the Temporal helpers (StartedWorkflows.terminateTree).
    • It walks the workflow tree through each parent's history, never through visibility (the Java test server can't list) and never through runtime-temporal's child-id naming.
    • It gives running executions a short close grace, terminates what is still running top-down, re-describes each execution, and attributes terminations to sessions.
  • New RunLog.reapSince(mark): runCell reaps at EVERY cell end.
    • An execution still running after termination is a genuine leak.
    • A run whose executions closed but whose handle never settled is reported in two disjoint classes: terminatedUnsettled (the harness terminated an execution of its session; session-level, may mix causes) or closedUnsettled (closed on its own; product-side only).
  • stopAndDrain does the same at dispose, replacing terminateInFlight there. A handle-less undrained run still fails dispose, as on JS and DBOS. The worst case is held under DISPOSE_TIMEOUT_MS by the budget test.
  • JS and DBOS lanes are unchanged.
  • The real-infra test is packages/e2e/src/harness/__tests__/temporal-cell-reap.integ.test.ts. It passes on the test server and on the Docker server.

FU-CONF-TEMPORAL-REAP-LATE-CHILD (above) was closed in the same MR. The candidate product finding it surfaced is FU-CONF-TEMPORAL-RESUME-HANDLE-EXTERNAL-TERMINATE.

FU-A2-03 — Harden persistent sub-agent edge cases under __resume-N ​

Closed by: commit 3ebed60f7. Summary: Added focused e2e regression test persistent-subagent-suspension-temporal.integ.test.ts exercising the parent-suspend + persistent-blocking-child interaction. Scenario: parent dispatches a persistent blocking child → child completes → parent calls a client tool (workflow exits with 'suspended_client_tool') → consumer submits result + resumes parent → parent finishes via __finish__. Verifies the persistent SubSessionRef survives the __resume-N workflow boundary intact (no duplication, no status reset). Result: A.2's design is correct — no production fix needed. Non-blocking persistent + parent-suspend is covered structurally by the same code path; the e2e test for it is blocked on shared MockLLMAdapter response-queue races between parent and non-blocking child (a pre-existing test infrastructure limitation documented in persistent-subagent-temporal.integ.test.ts).

FU-A2-01 — Per-call agent hooks not plumbed through Temporal runtime ​

Closed by: commit 2693eb18f (+ changeset da8bd52ea). Summary: Added AgentRegistry.replace(config) API to runtime-temporal and runtime-cloudflare. Replaces an existing static registration with a new agent config, returning whether a prior registration existed. Wired replaceAgent capability through the e2e harness's BackendEnv (no-op on JS/DBOS; calls registry.replace() on Temporal). approval-gate-hook-parity tests un-skipped on Temporal; both approve and deny paths pass. Pattern: tests call await env.replaceAgent(agent) before each executor.execute(agent, ...) to swap the registered reference for one carrying the test's hooks. JS/DBOS continue to honor inline per-call hooks without registry intervention.

FU-A2-02 — SUBMIT_TOOL_RESULT_SIGNAL_NAME no-op signal call ​

Closed by: commit 869787e5c. Summary: Removed the no-op handle.signal('submitToolResult', ...) call in applySubmitToOwner and the orphan resolveWorkflowId helper. Post-A.2 the workflow exits on every HITL boundary so the signal hit a completed workflow and only generated temporal.signal_failed warn logs. The durable pendingClientToolCalls write IS the canonical signal; consumers drive the next loop iteration via executor.resume(). Updated 6 unit tests; deleted 2 tests covering the now-impossible signal-failure scenarios.

FU-A1-01 — runId generator pattern duplication ​

Closed by: commit 869787e5c (batched with FU-A2-02). Summary: Extracted duplicated `run-${Date.now()}-${Math.random().toString(36).slice(2, 10)}` into a module-level generateRunId() helper. Two call sites in execute() and resume() now share one source of truth.

FU-A2-06 — AgentWorkflowResult status enum extension changeset ​

Closed by: commit e221ac4d3. Summary: Added a MINOR changeset for @helix-agents/runtime-temporal documenting the A.2 Temporal HITL rewrite per the project's "MINOR = observable behavior change" versioning policy. Documents three new 'suspended_*' status values, durable-only submitToolResult, __resume-N workflow ID convention, and the five removed signals (submitToolResultSignal, childSuspendedSignal, childWokeSignal, runResumedSignal, RUN_RESUMED_SIGNAL_NAME).

FU-A2-08 — Remote sub-agent SubSessionRef metadata not persisted on Temporal ​

Closed by: commit 5085756d0. Summary: The workflow body's dispatchRemoteSubAgents called recordSubSessionResult (core helper) which only UPDATES existing SubSessionRefs — it doesn't CREATE them. For ephemeral sub-agents the ref is created via recordSubAgentResult's addSubSessionRefs call (mode='ephemeral'); for remote sub-agents, no ref-creation path existed.

Fix: added addRemoteSubSessionRef activity that calls addSubSessionRefs with mode: 'ephemeral' + remote: { streamId, lastSequence } metadata. Workflow body's dispatchRemoteSubAgents calls it BEFORE recordSubSessionResult so the ref exists for the helper to update on the result.

Result: remote-agent-temporal.integ.test.ts now 4/4 PASS (was 3/4 + 1 skip). Cross-package typecheck 55/55 PASS.

FU-A2-05 — Cross-runtime Temporal e2e migration ​

Closed by: commits 13075bc45 (Phase 1: persistent sub-agent fix) + 332c0b2a0–646615b2a (Phase 2: 11 file migrations). Spec: docs/superpowers/specs/2026-05-08-temporal-e2e-migration-design.mdSummary: A.2's Temporal workflow body rewrite caused 18 e2e tests to fail with Completed workflow errors (signal-pattern hits exited workflow) plus a real persistent sub-agent regression (expected 'running' to be 'completed').

Phase 1 fixed the persistent sub-agent regression with two complementary changes: (a) Option A — added resolveAndSyncSubSessionRefs lazy-sync helper to packages/core/src/orchestration/companion-tool-dispatch.ts mirroring DBOS's pattern, syncing stale 'running' refs to terminal status when companion tools read them; (b) Option B — added markPersistentChildCompleted activity called from the workflow body after await childHandle.result() resolves on blocking spawn paths.

Phase 2 added a TemporalAgentExecutor to createTemporalTestContext() (with lazy-startup background worker to avoid contention) and migrated 11 test files from handle.signal('submitToolResult', ...) to the v7 stateless pattern (executor.submitToolResult + executor.resume).

Test results: 13/13 migrated tests pass (was 12/12 fail). 2 tests remain skipped pending FU-A2-07 (sub-agent ownership StaleStateError) and FU-A2-08 (remote SubSessionRef metadata) — both real underlying regressions documented for follow-up rather than band-aided.

FU-A2-09 — Temporal sub-agent resume after failed:parent_suspended ​

Closed by: commit landing FU-A2-09 fix. Summary: After FU-A2-07 fixed the parent's commitSuspendedStep StaleStateError, the parent's resume reached applyResultsAndReload but failed to drive the failed:parent_suspended child to completion. Two interacting issues:

  1. applyResultsAndReload left failed:parent_suspended children in remainingChildren, setting hasMoreSuspended = true. The workflow body's resume path early-exited with 'suspended_awaiting_children', skipping the main loop where dispatchSubAgents could γ-cascade re-spawn.
  2. The main loop calls runLLMStep, which doesn't return sub-agent calls on resume (the LLM has already produced the prior step). So dispatchSubAgents wouldn't re-fire even after lifting the early-exit gate.

Fix (variant of option (c) + workflow-body re-dispatch):

  • applyResultsAndReload now returns a separate childrenToRespawn field listing failed:parent_suspended children. These do NOT count toward hasMoreSuspended (the early-exit gate).
  • The workflow body's resume path re-dispatches them DIRECTLY via wf.startChild (using __resume-N workflow IDs per spec §5.2 mit#1) BEFORE the early-exit gate. Each re-spawn awaits the child workflow's result; on completion, recordSubSessionResult appends the tool- result message to the parent's history (emitting subagent_end + tool_end chunks exactly once). On re-suspension, the parent re- suspends with the child still in suspendedAwaitingChildren.
  • New activity commitChildRespawnResults performs the durable bookkeeping: clears re-spawned children from suspendedAwaitingChildren and re-inserts any that re-suspended.

This is option (c) from the original analysis but cleaner — instead of re-entering the main loop (which requires runLLMStep to fire empty and the dispatcher to re-fetch state), the workflow body does the re-dispatch directly and the main loop only runs when there's actual LLM work to do.

The executor.resume(child) short-circuit (issue #1 in the original description) turned out to be moot — the parent re-spawns the child via wf.startChild from inside the workflow, never going through executor.resume(child).

Un-skips packages/e2e/src/__tests__/client-tool-subagent-ownership-temporal.integ.test.ts.

A.3 — DBOS resume() contract fixes (F7) ​

Closed by: commits 8ec773936, 16e4acb99, 7d9498b8e. Summary: Single truthy-object pitfall on resume.ts:79 caused 3 reproducible DBOS resume() contract failures (resume on completed, active, or concurrent suspended sessions all incorrectly succeeded). Fix: single multi-status CAS with explicit .ok check, mirroring runtime-temporal's post-A.2 pattern. Added 5 regression-guard unit tests so the bug can't return silently.

FU-A2-39 — Lazy-load Node-only harness setup helpers (workerd compat) ​

Closed by: commit 0d789d40d. Summary: packages/e2e/src/harness/backend-descriptor.ts previously imported every setup helper at module top-level. Workerd's parser rejected ioredis / pg source on module-load, blocking the C2 "re-run the same integ test in workerd via a 1-line companion import" pattern. Two coordinated changes:

  1. backend-descriptor.ts: converted eager imports to a lazy SetupLoader pattern using a lazyImport<T>(path: string) helper that defeats the bundler's static analysis. Each backend now exposes setupLoader (lazy) + setup (resolved accessor).
  2. get-viable-backends.ts: switched workerd-context detection from the broken D1Database global probe (not exposed by @cloudflare/vitest-pool-workers) to a WebSocketPair probe (verified empirically). Added a hard-filter for cross-context backends rather than just marking them with a skipReason.
  3. lifecycle-hooks-parity.cf.test.ts: replaced its describe.skip placeholder with a direct import of the integ companion so the harness re-run works under workerd.

The descriptors cf-do-d1 and cfw-workflows-d1 carry an IMPL_PENDING_O5 env requirement marking the deferred O5 work (real-DO + CFW Workflows setup helpers under workerd's vitest pool).

FU-A2-17 — CF DO persistent sub-agents real-DO test expansion ​

Closed by: Wave 4 test-only commit (this branch). Summary: Expanded cloudflare-do-persistent-agents.cf.test.ts beyond the original two FU-A2-38 G1 scenarios (blocking spawn + waitForResult) to exercise the full companion/sub-agent matrix against real workerd DOs. Four new scenarios (all green on workerd):

  1. Non-blocking spawn — parent does not wait. companion__spawnAgent (non-blocking) returns status: 'spawned' and the parent finishes on its own. Teeth: reads the parent's persisted spawn tool result via the /messages route and asserts { status: 'spawned' } with no drained output — a blocking-mode regression (silently waiting) would surface status: 'completed' + output and fail.
  2. Child failure observed by parent (blocking spawn). A failer child whose LLM step errors with recoverable: false lands in status: 'failed'. Teeth: the parent's blocking spawn tool result is { status: 'failed' } (the parent observes the failure, NOT a silent success) AND the child DO's /status reports failed. Verified the assertion fails if the child status is read as completed.
  3. Parent interrupt during (non-blocking) child run — CF DO divergence. Spawns a non-blocking child (so the parent releases workerd's input gate), then POST /interrupts the parent. Asserts the v7 contract (200/400, NEVER 503) and a consistent parent terminal state (completed | interrupted, never failed), and that the independent child DO is not collaterally corrupted (completed | running | not_found, never failed/interrupted). Documented divergence: the CF DO blocking spawnAgent poll loop runs inside the tool's execute(), does NOT break on the parent abort signal, and does NOT write a durable interrupt flag to the child — the parent observes its own flag only at the NEXT step boundary after the tool returns. This matches the per-runtime interrupt table in docs/internals/subagent-execution.md (CF DO: per-step checkInterruptFlag).
  4. Cascading nested sub-agents (3 levels). New cascade-l1 → cascade-l2 → researcher chain; each level runs in its OWN DO of the self-bound class and the DO base rewrites persistentAgents per started agent type, so a child DO spawns a grandchild. Teeth: asserts both the root's and the mid-level child's spawn tool results are completed (proving the chain drained bottom-up across 3 DO instances).

Fixtures added to packages/e2e/src/test-worker.ts: a failer blocking persistent agent and the cascade-l1/cascade-l2 coordinators (all registered on the existing PersistentAgentTestServer registry — no new DO class or wrangler binding needed; the self-bound class hosts every level).

Harness caveat (not a runtime bug): all DO instances share ONE persistentAgentTestLLM FIFO response queue, and non-blocking children run fire-and-forget, so a racing background child can drain a response the parent expected. The deterministic blocking scenarios serialize via the spawn poll loop; the non-blocking scenarios make BOTH finish responses schema-compatible with both agents' output schemas ({ done, text, summary } — Zod strips extras) so whichever agent grabs whichever finish, completion-retry never fires. This mirrors the documented limitation in the JS persistent-subagent-parity.integ.test.ts. No production code changed; no runtime bug found.

FU-A2-38 — Re-enable Cloudflare DO real-DO end-to-end tests (G1 + G2) ​

Closed by: commit 30432ad34. Summary: packages/e2e/src/__tests__/cloudflare-do-persistent-agents.cf.test.ts (G1) and cloudflare-do-interrupt-protocol.cf.test.ts (G2) were previously describe.skip because the supporting test-worker.ts fixtures and wrangler.toml durable-object bindings hadn't been wired up. Both files now run end-to-end against real DOs.

Wired up:

  1. packages/e2e/src/test-worker.ts:
    • persistentAgentTestLLM + PersistentAgentTestServer DO class (self-binding via subAgentNamespace: env.PERSISTENT_AGENT_SERVER, two persistent agent configs — researcher (blocking) + writer (non-blocking) — declared via parent-with-persistent's persistentAgents field).
    • interruptTestLLM + InterruptTestServer DO class with an askUserTool (execute: CLIENT_TOOL_EXECUTE) for the HITL paused-state branch of the interrupt-route contract test.
  2. packages/e2e/wrangler.cloudflare.toml:
    • New PERSISTENT_AGENT_SERVER binding (class_name = PersistentAgentTestServer).
    • New INTERRUPT_AGENT_SERVER binding (class_name = InterruptTestServer).
    • Migrations v6 + v7 for SQLite-backed DO storage.
  3. Both .cf.test.ts files: removed the describe.skip and the inline FU-A2-38 placeholder comments. The G1 tests cover persistent sub-agent companion-spawn under real workerd-DO; G2 pins the v7 interrupt-route contract (200 / 400 / NEVER 503 — the v6 INTERRUPT_NOT_LOCAL 503 contract is gone).

FU-A2-40 — CFW Workflows γ-cascade re-spawn ​

Closed by: commit 6cbd78808. Summary: Mirrored runtime-temporal's γ-cascade closure (FU-A2-09) in CFW Workflows. Three coordinated changes plus dedicated coverage:

  1. packages/runtime-cloudflare/src/steps.ts — commitSuspendedStep now marks each suspendedAwaitingChildren entry's child session as failed:'parent_suspended' with CAS retry. applyResultsAndReload gained a childrenToRespawn: Array<...> return field, populated by a private scanChildrenForRespawn that filters durable child states for status === 'failed' && failureReason === 'parent_suspended'.
  2. packages/runtime-cloudflare/src/workflow.ts — added a re-dispatch loop in the resume branch. For each surfaced child: step.do('respawn-...', () => workflowBinding.create({ id: 'agent__<type>__<id>__respawn__<attempt>', params: {...} })) (the format was __respawn-<attempt> until CF-D1 MR 1a), then a bounded poll (max 5 min, exp backoff) on the child's durable state until terminal, then recordSubSessionResult to append the synthetic tool-result message to the parent's history.
  3. New drain-clear step removes drained children from suspendedAwaitingChildren and resets the suspension discriminators (status: 'active', clear suspendedStepId, clear failureReason) when fully resolved. Without this reset, the main loop's commit path would re-route through commitSuspendedStep and produce a spurious 'suspended_step_partial' final status when the LLM step completes the run.

Coverage:

  • New subagent-respawn-on-resume.integ.test.ts — 3 D1+Miniflare tests pinning the full γ-cascade contract (mark on suspend, respawn on resume + drain + clear awaiting, no-respawn fallback when no marker present).
  • 220/220 runtime-cloudflare integ + 1048/1048 unit pass.

P3.R3-PERF.2 — D1 promoteStaging single targeted read ​

Closed by: commit 388691974. Summary: Pre-fix, D1StateStore.promoteStaging issued THREE reads against D1: a minimal SELECT step_count, custom_state from __agents_states, a COUNT(*) against __agents_messages (for messageCount), and a full loadState (SELECT * FROM __agents_states) for the checkpoint payload. The minimal SELECT and the full loadState BOTH read __agents_states for fields that overlap heavily — wasted work on the hot path (every step boundary calls promoteStaging).

Fix: single targeted SELECT against __agents_states fetching exactly the columns needed for the checkpoint payload (session_id, agent_type, stream_id, step_count, status, output, error, custom_state, interrupt_context, version, resume_count). The COUNT(*) against __agents_messages is unchanged. The loadState call is gone.

Coverage: new d1-state-promoteStaging-reads.test.ts pins the read shape with a prepare()-spy decorator over MockD1Database: exactly 1 SELECT against __agents_states (was 2 pre-fix), 0 SELECT * queries against __agents_states (regression smoke), exactly 1 SELECT (COUNT) against __agents_messages, and exactly 1 SELECT against __agents_staging (getStagedChanges, unchanged). Plus a behavioral assertion that checkpointId is still returned. 242/242 store-cloudflare unit + 402/402 integ pass.

P3.R3-PERF.3 — Postgres three-phase migration runner with concurrent DDL ​

Closed by: commit 1d2bb369c. Summary: PostgresMigrations.run() previously wrapped every migration's statements in a single transaction. CREATE INDEX CONCURRENTLY is forbidden inside transactions, so V7's partial expiresAt index either had to use a regular CREATE INDEX (taking ACCESS EXCLUSIVE on __agents_states for the build duration) or rely on a documented operator manual workaround.

The runner now uses a three-phase model: (1) transactional schema phase (statements — ALTER TABLE / CREATE TABLE), (2) concurrent phase (concurrentStatements — CREATE INDEX CONCURRENTLY etc., runs outside any transaction), (3) version-row commit phase (INSERT ... ON CONFLICT (version) DO NOTHING). Phases are ordered so concurrent indexes can reference columns added in the same migration. INVALID indexes from prior crashed CONCURRENTLY builds are dropped before retry. V7 retagged: partial expiresAt index moved into concurrentStatements.

Concurrency: dropped the session-level pg_advisory_lock (broken under connection pools — different connections can't release it). Replaced with xact-scoped locks inside phases 1/3 + IF NOT EXISTS idempotency for phase 2 + ON CONFLICT for the version-row insert. Safe across multiple migrator processes.

Coverage: postgres-migrations.test.ts gained 9 new tests pinning the contract: lock NOT taken (smoke against the pre-fix pattern), ON CONFLICT DO NOTHING on version insert, concurrent statements run via adapter.query (NOT inside transaction), invalid-index cleanup before CONCURRENTLY rebuild, prefix-aware index scoping, error propagation, logger event surface. The applyUpTo integ helper was updated to mirror the runner's phase model so test schemas match production.

Tests: 197/197 unit + 327/327 integ for store-postgres.

Postgres-LISTRUNS-IDX — V8 composite (session_id, status, turn) runs index ​

Closed by: commit c596e67ab. Summary: Closes Postgres parity with the Redis closure in 296ad28e7. Pre-fix, PostgresStateStore.listRuns(sessionId, { status }) filtered with only idx___agents_runs_session_id available, doing a per-session run scan + in-memory status filter on every status query. V8 adds idx___agents_runs_session_status_turn ON __agents_runs(session_id, status, turn) so the composite supports both the WHERE and the ORDER BY without a separate sort step. Built via the concurrentStatements phase (P3.R3-PERF.3) so the index build doesn't block writes on the hot-path runs table. Coverage includes a behavioral test (5 runs, mixed statuses, turn ordering) and an EXPLAIN-based plan assertion (SET LOCAL enable_seqscan = off to force the index path on small test data).

P3.R3-PERF.4 — message_count denormalization (Postgres V9 + D1 V10) ​

Closed by: commit 054194def. Summary: Pre-fix, every listSessions call evaluated (SELECT COUNT(*) FROM messages m WHERE m.session_id = s.session_id) for every row in the WHERE-matched set — O(N × M) cost where N is matched session count and M is per-session message volume. Both Postgres and D1 paid this cost.

Post-fix: a denormalized message_count INTEGER NOT NULL DEFAULT 0 column on __agents_states. V9 (Postgres) + V10 (D1) ALTER TABLE

  • idempotent backfill UPDATE (recomputes from the messages table; re-runnable for drift recovery). The mutating code paths (appendMessages, saveStateAndPromoteStaging, cloneSession) maintain the column in lockstep with the messages table. listSessions reads it directly — no per-row subquery work.

saveStateAndPromoteStaging's checkpoint payload now reads the message_count from the states row directly (one less SQL round-trip on every step boundary).

Coverage: 5 new tests in each store's list-query integ file covering appendMessages increment, multi-call accumulation, saveStateAndPromoteStaging atomicity, zero-message sessions, and EXPLAIN-based proof that listSessions does NOT touch __agents_messages post-fix. Cursor pagination (the second half of P3.R3-PERF.4) is deferred — denormalization captures the dominant perf win.

D1-DELETESESSION-TOCTOU — orphan-message race in appendMessages ​

Closed by: commit b4e393670. Summary: Pre-fix, D1StateStore.appendMessages did the existence check and the message INSERT batch in SEPARATE D1 round-trips. A concurrent deleteSession could land between them, leaving orphan messages with no states row.

Post-fix: each message INSERT is gated by WHERE EXISTS (SELECT 1 FROM __agents_states WHERE session_id = ?) in an INSERT...SELECT form (SQLite's INSERT doesn't accept WHERE). If the gate fires, meta.changes = 0; we sum changes across the batch and throw StateNotFoundError if the total is less than the message count. The mock D1 was extended to recognize the EXISTS-gated INSERT...SELECT pattern so unit tests reproduce the contract.

Coverage: 3 new tests in d1-state.integ.test.ts covering the canonical race scenario, the happy path regression guard, and a partial-deleteSession edge case.

CLEANUP-ORPHANED-STAGING — cross-session orphan sweep (Postgres + D1) ​

Closed by: commit cab06588e. Summary: Mirrors the runtime-redis closure in 296ad28e7. Pre-fix, Postgres and D1 had only the per-session cleanupOrphanedStaging(sessionId) — operators had no way to sweep abandoned staging rows without enumerating every session ID first. Storage on __agents_staging grew unboundedly for paused / abandoned sessions on production deployments.

Post-fix: both stores expose cleanupOrphanedStagingData(maxAgeMs?: number): Promise<number> that DELETEs staging rows older than now - maxAgeMs (default 1 hour) and logs a structured cleanup_orphaned_staging.summary event. NOT on the SessionStateStore interface — per-store operator method, same convention as Redis.

Coverage: 3 new tests in each store's list-query integ file covering the age filter, zero-deletes case, and multi-session sweep.

P3.R3-BC-MISC (D1 chain collapse) — collapsed initial schema for fresh-DB deployments ​

Closed by: commit b23000542. Summary: Pre-fix, fresh deployments ran the full V1..V10 migration chain — 10 sequential db.batch() round-trips on every fresh-DB worker boot, adding ~100-500ms of dead time.

Post-fix: runMigration() detects fresh databases (no __agents_states table present in sqlite_master) and applies a single SCHEMA_COLLAPSED_INITIAL constant — the post-V10 schema in one batch, plus all incremental version rows for audit-trail parity. Existing deployments at any partial version (1..9) continue down the V1..V_current incremental chain unchanged.

Detection contract uses sqlite_master so a partial schema from a prior crashed run falls through to the safe incremental path (no silent CREATE TABLE IF NOT EXISTS no-ops mis-reporting the version). Audit-trail parity: the collapsed path records EVERY incremental version row, not just CURRENT_SCHEMA_VERSION, so operator queries like SELECT * FROM __agents_migrations WHERE version = 7 behave consistently across deployment paths.

Coverage: new d1-collapsed-schema-parity.integ.test.ts runs both paths in isolation and compares PRAGMA-reported schemas structurally (table column lists + index definitions). 5 tests covering version stamp, per-table column parity (sorted by name to absorb ALTER TABLE column ordering), per-table index parity (name + columns + uniqueness + partial flag), and version-row history parity.

Maintenance contract: when adding V_{n+1}, also append the new statements to SCHEMA_COLLAPSED_INITIAL. The parity test catches missed updates.

Tests: 242 unit + 420 integ for store-cloudflare. Full turbo typecheck 55/55.

FU-TYPE-SAFETY-2026-05 — Eliminate any + forced casts across all packages ​

Closed by: commits 4fd490327 (Stages A + B.1 + B.3), 47e3faa18 (Stages B.2 + B.4 + C + D + ToolContext.getState docs), c94ca7606 (close-out review findings). Surfaced by: four parallel pr-review-toolkit:type-design-analyzer agents (2026-05-10) run after the BC removal pass landed.

Summary: ~200+ any annotations and forced as Type casts removed across all packages. The fix spanned four stages:

Stage A — Fast wins (~110 violations): Introduced shared aliases AnyTool = Tool<z.ZodType, z.ZodType> and AnyAgentConfig = AgentConfig<z.ZodType, z.ZodType> in core; bulk- replaced the <any, any> epidemic across 40+ runtime files. Tightened HookManager.invoke<K extends keyof AgentHooks> so temporal + cloudflare + client-tool-workflow-helper no longer need triple-casts. Switched helix-to-aisdk-converter.ts to use its own isToolPart guard (now validates state + toolCallId, not just the discriminator). Replaced type DurableObjectState = any test stub with a structural shape. Dropped 11 tracker.getState() as any sites in favor of typed generic bridges. Replaced as any as Redis test mocks with a stubRedis(...) helper. Replaced statements: any[] in d1-mocks with MockD1PreparedStatement[].

Stage B — Zod schema introduction (~80 violations): Authored packages/core/src/types/state-schemas.ts exporting SessionStatusSchema, RunStatusSchema, CompletionReasonSchema, SubSessionStatusSchema, SubSessionRefStatusSchema, SubSessionRefModeSchema, AgentStatusSchema, plus discriminated- union schemas for Message, UsageEntry, Memory, the content- part variants (SystemMessageSchema, UserMessageSchema, AssistantMessageSchema, ToolResultMessageSchema), and parseAgentStateCheckpoint / AgentStateCheckpointView helpers. Every schema ships a _DRIFT_ASSERTIONS satisfies _SchemaTypeDriftAssertions compile-time guard binding the schema to its TypeScript counterpart — the previous bare-type guards were decorative and silently passed under drift; the new value-bound guard catches optional-field drift via StructuralMatches<A, B> + KeysMatch<A, B>. Applied at every JSON.parse(row.x) as Y site across store-postgres, store-cloudflare, store-redis, memory-redis, and the DO state / usage / stream stores. Tightened DO HTTP wire-protocol schemas: StartAgentRequestV2Schema.history is now z.array(MessageSchema); SubmitToolResultResponseSchema validates cross-DO fetch responses; AgentStatusResponseSchema.status is the 8-literal enum instead of z.string().

Stage C — Public API surface (~30 violations): Typed DOFrontendExecutor.execute(_agent: AnyAgentConfig, input: AgentInput<unknown>, ...) and getHandle() with consumer-side docs explaining why they're not AgentExecutor implementations. Typed agent-server/types.ts agents record to Record<string, AnyAgentConfig>. Replaced the metadata?: any HTTP-handler parameter with the same Zod-validated shape used elsewhere. Typed do-clients.ts public-API surfaces via generics over <TState, TOutput> where possible.

Stage D — Test mock helpers (~90 violations): Authored packages/llm-vercel/src/__tests__/helpers/make-chunk.ts with a makeChunk<T> factory; ~60 as unknown as TextStreamPart<ToolSet> casts replaced across chunk-mapper.test.ts and chunk-mapper-edge-cases.test.ts. Authored packages/memory/src/__tests__/helpers.ts with typed mockLanguageModel + mockMemoryStore stubs; ~30 {} as any sites swept across 8 memory test files. Rewrote ai-sdk snapshot stability test's (p: any) => callbacks with a LoosePart structural type. Replaced production any patterns in packages/ai-sdk/src/handler/handler-factory.ts (~7 sites), packages/llm-vercel/src/vercel-adapter.ts (~4 sites), and packages/runtime-js/src/{js-agent-executor,run-loop}.ts (5 sites) with typed structural bridges or generic <T>() => accessors.

ToolContext.getState<T>() — caller-driven contract: the <T>() => T signature is intentional, not a violation. The persisted state is opaque at the runtime boundary; the runtime returns whatever the caller asks for and relies on the caller to validate via Zod Schema.parse(ctx.getState()). JSDoc at the API surface (core/src/types/tool.ts) documents the contract end-to-end with code examples.

Lingering documented casts (intentional, not bugs): a handful of as <Type> bridges remain at runtime/SDK boundaries where cross-package type drift is fundamental (Zod 4 framework schemas vs Zod 3 in the Vercel AI SDK's ai package generic constraints, partyserver's unknown[] template-literal return, the DO ResumeMode vs core ResumeOptions literal-union mismatch). Each cast is documented at the site with the specific rationale and is confined to a single bridge call rather than scattered.

Close-out review findings (commit c94ca7606): 3 critical, 5 high, 4 medium issues caught by two superpowers code reviews of the migration work — all fixed inline. Highlights:

  • C1: continue regression in js-agent-executor.ts drain loop was dropping telemetry; scoped to skip only the afterTool dispatch.
  • C2: silent try { } catch {} in DO checkpoint message-rebuild replaced with per-skip warn + summary log.
  • C3: missed JSON.parse(row.chunk) as StreamChunk casts in packages/store-cloudflare/src/stream-durable-object.ts:945, 1051 → applied StreamChunkSchema.parse(...); updated 31 test fixtures with the required step: 0 field.
  • M1: StructuralMatches<A, B> didn't catch optional-field drift → added KeysMatch<A, B> enforcing keyof A and keyof B match exactly. Mutation test (introducing introducedDrift?: string into UserMessageSchema) confirmed the guard now rejects what the previous form silently accepted.
  • M2: AgentStatusResponseSchema.status was z.string() → tightened to z.enum([...]) with the 8 specific literals.
  • M4: DO /resume silently fell through to continue semantics for 'retry' / 'branch' resume modes → explicit reject with warn log + thrown error at the DO boundary.

FU-DBOS-ERROR-CHUNK — DBOS failStream now emits AI SDK error chunk ​

Closed by: commit e7506723e. Summary: Pre-fix, DBOS's runAgentLoopOneTurn on the LLM-terminal failure path called dispatcher.failStream({...}) at packages/runtime-dbos/src/workflows/shared.ts:1098 without first emitting an error chunk into the stream. AI SDK transformers running over the SSE bytes saw the stream go from active → ended without an error event, so consumers couldn't render error state. Runtime-js and runtime-temporal already emit an error chunk before failStream. Fix added dispatcher.emitErrorChunk(...) (new EmitErrorChunkArgs + runEmitErrorChunk + emitErrorChunkStep) immediately before failStream on the failed path, mirroring the other runtimes' pattern. Verified by ai-sdk-dbos.integ.test.ts > should handle DBOS workflow errors as AI SDK error events.

FU-A2-42 — Pre-existing runtime-temporal initialState/branch/history failures (12 tests) ​

Closed by: commit f4a169588. Summary: Pre-fix, packages/runtime-temporal/src/__tests__/integration/temporal.integ.test.ts had 12 deterministic failures because the Temporal workflow body's 'fresh' mode never read AgentWorkflowInput.initialState, branch, or history. When callers started workflows directly (the integration test harness via runner.executeWorkflow(), the agent-server, or any consumer bypassing the executor's pre-claim createSession({initialState}) shortcut), these inputs were silently dropped — customState arrived empty, branched workflows lost source history, etc. Fix routed the three inputs through to the initializeAgentState activity with the same precedence as runtime-js / runtime-cloudflare: load base from branch.fromSessionId (missing source throws 'session not found'), history overrides branch messages (system messages filtered), schema defaults applied to initialState ?? branchState, new-session path prepends resolved conversation history before the user messages. Also deduped the LLM-error chunk: runLLMStep's onError callback writes an error chunk, and then the LLM-terminal failed path was calling persistTerminalState({status:'failed'}) which wrote ANOTHER error chunk. Added skipErrorEmit flag for the LLM-terminal branch. Additionally, top-level catch RETURNS {status:'failed', error} (matching runtime-cloudflare) instead of re-throwing, so callers can await structured failures via handle.result() rather than catching WorkflowFailedError. All 12 tests now green; runtime-temporal unit suite 242/242; runtime-temporal integ 51/51; cross-runtime e2e (temporal combo) 71/71.

FU-AISDK-1: getChunksFromStep does not respect run boundaries ​

Status: CLOSED 2026-05-17 via Option A. Both getChunksFromStep and getAllChunks fallback branches removed from loadPartialContentAsMessage in packages/ai-sdk/src/handler/snapshot.ts. When startSequence is unavailable the loader now returns null (no partial content) — safe degradation, the snapshot uses persisted messages only.

Background: The whitelist-inversion fix (commit 049a19e21, 2026-05-16) made the step-based loading branch unreachable for the originally-reported Bug 4 scenarios (a run with startSequence defined now always uses sequence-based loading). The follow-up fix (commit 3f098277a, 2026-05-17) ALSO closed the settled-run leak path where failed/superseded runs with status === 'paused' (race scenarios: process crash before stream cleanup, CF DO eviction) would incorrectly load chunks via the step-based fallback. The final remaining problem — the currentRun === null case where getChunksFromStep(N) returned chunks across all runs at the same step number — is now closed by the fallback removal.

A first attempt at the same fix (commit 5c4fbde42, included in MR !166 pipeline 1) was reverted because a Playwright test in examples/opennext-cloudflare-do failed during that CI run. Subsequent investigation showed the failure was NOT caused by the fallback removal: running the same Playwright test locally with the fix re-applied (on top of intervening main commits including round-5 review findings, v0.7 StepWrites unification, and providerExecuted persistence) passes reliably. The intervening main commits resolved a separate state- persistence path that had been masking the leak via consistent duplication on both live and refresh sides.

The deferred runtime-cloudflare cleanups identified during investigation (remove ?? 0 default in DOStateStoreClient.getCurrentRun; return startSequence for all non-null currentRun in DO /snapshot) are NICE TO HAVE for type clarity but no longer required — the ai-sdk-side safe-degradation closes the leak vector. Track those as separate follow-ups if/when someone touches that code.

Conformance suite 0a — follow-ups ​

Deferred triage items from the plan-0a whole-branch review. Items the 0a.1 plan ("Conformance Suite 0a.1 — Review Fixes and Driver-Contract Reshape", docs/superpowers/plans/2026-09-24-conformance-suite-0a1-review-fixes-and-contract.md) closed — and those the two post-completion review-fix passes (PR-1: harness code; PR-2: catalogue / CI / docs) closed — have been removed, each verified against the current code rather than taken on a plan's word (the plan document above and the fix commits' messages record what closed what). Each surviving line names the file the issue lives in and why it's deferred rather than fixed now.

  • T1 packages/core/src/llm/scripted-model.ts — the gate/hold mechanism (whenHeld/release) tracks only one held record per gateId; a second concurrent hold on the same id silently overwrites the first's waiters.
  • T1 packages/core/src/llm/scripted-model.ts — getCallCount undercounts calls that resolved via a scripted fail response; it's a public core API used outside conformance too, so this needs either a documented contract or a fix plus a semver note, not a quiet behavior change.
  • T1 packages/core/src/llm/scripted-model.ts — a call that's both pre-aborted (abort signal already set) and scripted to fail is classified 'aborted' regardless of the scripted failure, which can mask a scenario bug that scripts a fail but never actually races the abort.
  • T2 packages/e2e/src/conformance/__tests__/docs.test.ts — the file-level check is stricter than its name suggests (it also asserts cross-file consistency, not just doc-file shape); rename or split it.
  • T3 packages/e2e/src/conformance/judge.ts — describeVerdict and the empty-ledger-entry rendering path still lack direct unit coverage (0a.1's drift.ts now REJECTS a ledger entry with no assertions before it can ever reach judge.ts, so the live risk is lower, but describeVerdict's own handling of that shape is still only exercised indirectly via suite-lane.test.ts / cell.test.ts, not judge.test.ts directly).
  • T4 packages/e2e/src/conformance/observation/persisted.ts — readPersisted's 10k-page pagination guard (for (let guard = 0; guard < 10_000; ...)) still falls through silently on the bound instead of raising a clearly-labeled "truncated" outcome; a session with more than 2,000,000 messages (200/page) would fall out of the loop and get judged by whatever the length-vs-getMessageCount comparison happens to produce (usually a ShortReadError, but with no indication the CAUSE was the guard, not a real corrupt-row skip).
  • T5 packages/e2e/src/conformance/judge.ts — a single root cause can produce duplicate "problem" lines (once per affected assertion id) instead of being collapsed into one line naming all affected ids.
  • T6 packages/e2e/src/conformance/drivers/executor-driver.ts (wrap) — outcome() conflates a throw from starting the call (get() rejecting) with a throw from awaiting its result; both collapse to the same terminal: 'threw' shape, losing which phase actually failed. (0a.1's RunOutcome reshape did not change this — thrown still doesn't distinguish "never started" from "started and threw resolving the result".)
  • T8 packages/core/src/llm/scripted-model.ts (waitForAbort) — accepts and reports true for an abort signal that was already set BEFORE stop() was even called, which can make a stop-latency assertion pass vacuously.
  • T8 packages/e2e/src/conformance/findings.ts (CP-25, CP-26) — evidence provenance is source-only with a live: ... note embedded in the evidence string rather than a first-class live-repro provenance entry; worth reconciling once a live probe scenario exists for these.
  • T10 (generateStep / mock restore) — a mocked step generator is restored via delete on the module namespace rather than vi.restoreAllMocks() / re-assignment, which is fragile against import-order changes; needs the exact call site identified and normalized.
  • packages/e2e/scripts/conformance-verify.ts / packages/e2e/src/conformance/verify.ts — the printed line index in a mismatch report doesn't correspond to the actual .jsonl line number (off by the header/merge step), and a result line with an unexpected/unknown key is silently ignored rather than surfaced as a schema drift.
  • packages/core/src/llm/scripted-model.ts — two concurrent calls against the SAME unscoped .script() (global) queue can have their responses swapped if the queue is consumed out of call order; needs either per-call correlation or an explicit "no concurrent calls without gateId or scriptFor" invariant. (0a.1's scriptFor(sessionId, ...) fixed this for SESSION-scoped queues only — two interleaved sessions' scriptFor responses are proven not to swap — the shared global queue this item describes is untouched.)

New follow-ups from the 0a.1 plan (owning plan noted per item) ​

  • DONE (Phase B increment 1, 2026-09-26) — DRIVER_CAPABILITIES has per-channel observe-persisted / observe-llm / observe-hooks / observe-stream / observe-client (plus observe-host-status and observe-listed-count), every scenario's needs lists the channels it reads, and the per-cell driver view (drivers/cell-driver.ts) makes an undeclared read a HarnessError and an unobservable channel a declared gap (authoring guide §2, "Observation channels are capabilities too"). Original item: 0b packages/e2e/src/conformance/drivers/session-driver.ts / observation/observation.ts — observation CHANNELS are not driver capabilities: observe() is one capability, but a host/client driver may be able to observe persisted and not llm/hooks (or vice versa), and today nothing lets it say so. Add per-channel observe-* capabilities (e.g. observe-persisted, observe-llm, observe-hooks) to DRIVER_CAPABILITIES and to each scenario's needs, so an unobservable channel makes the cell a DECLARED GAP instead of a crash or a vacuous pass (0a.1 final review M6).

  • DONE (Phase B increment 1, 2026-09-26) — the host and client drivers (makeHostDriverCfDo / makeHostDriverJs in drivers/host-driver.ts, makeClientDriverCfDo / makeClientDriverJs in drivers/client-driver.ts) are all built with defineMakeDriver, and defineConformanceSuite refuses a lane whose MakeDriver capabilities differ from its lanes.ts entry. The cf-do agent install is a 'harness' seam (the bundled agent registry, fingerprint-checked in prepare()); js-chat-host registers the wired agent on its chat-route map. Original item: 0b packages/e2e/src/conformance/drivers/session-driver.ts — the host and client tier drivers plan 0b adds MUST be built with defineMakeDriver(capabilities, fn), not a bare function assigned a capabilities property by hand — buildBackendInstance rejects a driver whose own capabilities disagree with its MakeDriver's static set, and defineMakeDriver is the only way to satisfy that invariant correctly (flagged as a concern in 0a.1's Task 3 fix-round-1 report). Also: follow the executor driver's SPLIT, not a blanket wrap — BackendEnv.wireAgent (pure harness wiring: hook recorder, test seams) fails as a HarnessError (:crash), while BackendEnv.installAgent propagates unwrapped and its meaning is declared where it is built — agentPreparation's InstallKind. Use 'product' ONLY for a real product registration API (Temporal's AgentRegistry.replace(); in 0b e.g. agent-server's agents map), so a typed rejection at registration is judged by checkRejection like any other runtime answer. A harness seam (the CF DO module slot, the CFW test injection registry, a no-op) is 'harness', and its failure is a HarnessError that can never match a ledgered .rejected-typed (PR-1 item 4, PR-3 item 3; supersedes the older "don't copy the blanket prepare() wrap" note).

  • 0c packages/e2e/src/harness/setup-helpers/cf-do-d1.ts / cfw-workflows-d1.ts — still silently ignore SetupOptions (including llmAdapter) entirely, unlike the JS/Temporal/DBOS helpers, which now all either honor or reject (rejectUnsupportedSetupOptions) every option (0a.1 Task 3). Plan 0c, which is what actually calls these helpers via runtime-cloudflare's workflow pool, must either honor SetupOptions or reject it the same way. Partly done (Phase B Task 3): the new cf-do backend (setup-helpers/cf-do.ts, which replaced the cf-do-d1 registry row) honours llmAdapter (a ScriptedModel only) and rejects everything else via rejectUnsupportedSetupOptions; the legacy workerd cf-do-d1 and cfw-workflows-d1 helpers still ignore SetupOptions. Done for every Phase B backend (increment 1, 2026-09-26): cf-do honours or rejects every field, and js-chat-host delegates to setupJsMemory, which does too. What stays open is only the two legacy workerd-pool helpers, which no conformance lane uses; plan 0c (CF Workflows lanes) owns them. Plan 0c MR-1b: the CF Workflows conformance lane's new helper, setupCfwMiniflare (cfw-miniflare.ts), honours llmAdapter (a ScriptedModel only) and rejects every other field via rejectUnsupportedSetupOptions. The legacy cfw-workflows-d1.ts helper now backs the cfw-workflows-d1-pool row; MR-2 T9 retires it or keeps it with a stated reason (FU-CONF-CFW-POOL-ROW-RETIREMENT).

  • 0c docs/superpowers/specs/2026-09-23-runtime-conformance-suite-design.md §11 / packages/core/src/llm/scripted-model.ts — plan 0c must first VERIFY whether a workerd isolate is actually shared and reused across a cell's calls the way the Node-pool inProcessControl() assumption relies on (a DO/Workflow eviction can reset module-level state). The SEAM is ready (0a.1 fix wave): ScriptedModelControl methods may return promises, record() assigns the index atomically, every post-record mutation (gateId, abortedAt, outcome, endedAt) goes through update() keyed by callId, and the model has async twins (callsAsync, callsForAsync, remainingScriptedAsync, ready(); the executor driver's observe() already uses callsForAsync) — proven by a core unit test with an async, snapshot-storing control. What remains for 0c: implement the store/RPC-backed control itself; its own gate-signalling mechanism (gates are still in-process — a held call and the release()/reset() that unblocks it may live in different isolates); and reset() isolation for a call whose record()/dequeue() is still in flight on an async control (today only safe once runs have settled). The same question applies to harness/recorded-hooks.ts's hooks bridge (0a.1 Task 4), which also currently assumes the agent runs in the SAME process as the harness. Resolved by the Node bridge (Phase B for cf-do; plan 0c MR-1b reuses it for CF Workflows): the ScriptedModel, its gates and the hook recorder stay in Node (harness/cf-do/bridge-server.ts). Workerd runs a ProxyModel that sends each model call to the bridge, and hooks reach it as hook.fired, so no control state lives in a workerd isolate (plan 0c spec §3.1).

  • 0a.1-final-review packages/e2e/src/conformance/drivers/executor-driver.ts (stopAndDrain) — skips (does not wait on) a run whose RunHandle hasn't arrived yet when dispose() is called; the skip is loud (surfaces via poison/teardown failure, not silent), but it's still a gap in the drain guarantee worth closing. PR-1 widened the window slightly: wrap now starts the call through one extra microtask (Promise.resolve().then(get), so a synchronous throw becomes a threw outcome), so a handle arrives one tick later than before.

  • 0a.1-final-review packages/e2e/src/harness/recorded-hooks.ts (wireAgentHooks) — the re-wire-detection WeakSet tag is on the hooks BUNDLE object; rebuilding an equivalent hooks bundle with fresh wrapper functions (rather than reusing the tagged one) bypasses the "already wired" check silently. Tag the wrapper FUNCTIONS instead, or the agent itself.

  • 0a.1-final-review packages/e2e/src/harness/recorded-hooks.ts and packages/e2e/src/harness/setup-options.ts — both import HarnessError from packages/e2e/src/conformance/types.ts, a layer inversion (harness/ is meant to be a lower layer than conformance/); move HarnessError to a shared module both can import from without crossing that boundary.

  • 0a.1-final-review packages/e2e/src/harness/__tests__/recorded-hooks.test.ts / packages/e2e/src/conformance/__tests__/executor-driver.test.ts — the recorder's "no hooks fired" and "two agents sharing one bundle" cases, and executor-driver's "prepare() called twice" case, are only weakly exercised (mock env.wireAgent / env.installAgent, not the real wireAgentHooks path through a backend); worth stronger, more explicit tests (0a.1 Task 4 fix-round-1, deferred by instruction rather than fixed).

  • 0a.1-review-fix (PR-1) packages/core/src/llm/scripted-model.ts — a failed ScriptedModelControl call is sticky and EPOCH-attributed: its ScriptedModelError.origin names the call, session and reset epoch, and ready() reports a failure from an EARLIER epoch as "a leftover call of a previous cell". That keeps it loud and correctly labeled, but the cell that CRASHES on it is whichever cell runs next on that model (the rebuild its crash triggers then clears it), not the cell that owned the leftover call — so that cell's own verdict may already have been recorded as a pass. PR-3 narrows this: a cell whose OWN driver-started runs are still unfinished at its end (CellRun.leftRunning, read through SessionDriver.runs) now triggers a rebuild, so the next cell gets a fresh model, and a control call made during the leftover grace is caught by that cell's own ready(). What remains is a call from work the driver does not track (e.g. a runtime-spawned child run outliving its parent's result). Attributing that back to the owning cell would need the lane to hold results until the next epoch's ready(); deferred.

  • 0a.1-review-fix (PR-1) packages/e2e/src/conformance/drivers/executor-driver.ts / cell.ts — an installAgent (product registration) failure is judged only on the rejection path (checkRejection). On a POSITIVE path it has no named id to land on, so the cell crashes (<scenario>:crash): loud, but UNLEDGERABLE — a runtime whose registration starts failing for a supported agent can't be ledgered as a finding without a registration assertion id. Consider a synthetic-but-ledgerable <scenario>.installed id.

  • 0a.1-review-fix (PR-1) — FIXED in PR-2, recorded for context:heldOrEnded (scenarios/stop.ts) left whenHeld's timer (up to 2 × HELD_BOUND_MS = 60 s) pending after the run's outcome won the race. ScenarioModel.whenHeld now takes an optional AbortSignal that clears the timer and drops the waiter, and heldOrEnded aborts it once the race is decided (core scripted-model.test.ts, e2e stop-scenario.test.ts).

  • 0a.1-final-review packages/e2e/scripts/benign-teardown.mjs (STRICT_KNOWN_RACE_HEADLINE_PATTERNS) — neither anchor has been verified against a genuine CI capture of the race firing end-to-end (4 live attempts across 0a.1's Task 7 fix rounds, all clean). Capture the REAL DBOS "pool-after-end" race and the REAL Temporal IllegalStateError teardown rendering from CI the first time either fires, add them as strict-forgive samples. (The header comment's old "miniflare D1" attribution was corrected in the 0a.1 fix wave: the literal string is thrown by @temporalio/worker's connection.ts as an IllegalStateError, not a plain Error; the strict anchor was sourced from an existing hand-built test sample, not a real capture, and may need to move to the IllegalStateError/"...to it" form once a real one is seen.) Until then the predicate fails closed by design, which is the safe direction, not a bug.

  • 0a.1-final-review packages/e2e/src/__tests__/helpers/vitest-teardown.ts — its pkill -f workerd teardown runs UNSCOPED (kills every workerd process on the machine, not just this run's own children) and is therefore still NOT wired into any config's globalSetup (0a.1 Task 7 reverted an earlier, unsafe activation attempt). Scope it to the run's own child processes (tracked PIDs, or a run-specific marker) before enabling it for real.

  • 0d (repo-wide cleanup) the dead forceExit/globalTeardown vitest config keys 0a.1's Task 7 found and removed from packages/e2e's configs exist, equally silently no-op, in ~13 other packages (grepped for forceExit/globalTeardown across vitest*.config.ts): packages/agent-server/vitest.integ.config.ts, packages/ai-sdk/vitest.integ.config.ts, packages/memory-cloudflare/vitest.integ.config.ts, packages/memory-redis/vitest.integ.config.ts, packages/runtime-cloudflare/vitest.integ.config.ts, packages/runtime-dbos/vitest.integ.config.ts, packages/runtime-temporal/vitest.integ.config.ts, packages/store-cloudflare/vitest.config.ts and vitest.integ.config.ts, packages/store-postgres/vitest.integ.config.ts, packages/store-redis/vitest.integ.config.ts, packages/tracing-langfuse/vitest.integ.config.ts, and the root vitest.config.ts. Same fix shape as packages/e2e got: drop forceExit; rename a globalTeardown file's export from default to a named teardown and wire it via globalSetup where a real teardown file exists.

  • (untriaged — environmental, unverified) examples/research-assistant-cloudflare-do has at least one red Playwright/workerd test attributable to the same Miniflare temp-dir path-length limit documented in docs/dev/development-workflow.md ("Sandbox / environment gotchas": mkdirat: File name too long, triggered by a deeply-nested checkout path) rather than a real product defect. Not yet confirmed whether it is purely environmental (this session's sandbox path depth) or reproduces from a normal-depth checkout too; needs a run from a short path to settle it before it's either fixed or written off.

Conformance Phase A — follow-ups (not findings) ​

Consumer-code / docs observations surfaced while re-verifying the Phase A bucket-prep catalogue (2026-09-24-conformance-phase-a-bucket-prep). None of these are catalogue findings (they're not framework defects reachable through the conformance suite's driver surface) — they're tracked here so they don't get lost.

FU-CONF-OPENNEXT-DO-ID-PREFIX: the opennext-cloudflare-do example uses two DO-id conventions — open ​

examples/opennext-cloudflare-do/src/lib/agent-client.ts:28 and brief-agent-client.ts:18 (+ README.md:431) derive the DO id via idFromName('session:' + sessionId). The chat handler's own clients derive it via idFromName(sessionId) (no prefix) — packages/runtime-cloudflare/src/do/do-clients.ts:97, 342, 816 — so the example's debug/admin routes built on getDOStub read an empty DO for a session created through the normal chat path. packages/runtime-cloudflare/src/do/do-executor.ts:383, 427, 455, 508, 535 uses the session: prefix convention too — so the runtime-cloudflare package itself carries two conventions, not just the example.

Owner: examples/opennext-cloudflare-do maintainers (align getDOStub's idFromName prefix with do-clients.ts, or document why the two conventions are intentionally different).

FU-CONF-EXAMPLES-RESUME-ONLY-ACTIVE: seven example ChatClients resume only on status === 'active' — open ​

examples/opennext-cloudflare-do/src/app/{chat,coordinator,brief-flow,reasoning}/[sessionId]/ChatClient.tsx (lines 40, 36, 38, 39 respectively) and examples/nextjs-redis/src/app/{chat,reasoning,brief-flow}/[sessionId]/ChatClient.tsx (lines 131, 91, 94 respectively) all gate shouldResume on initialSnapshot.status === 'active' only, so a paused session (suspended on a client tool / HITL approval) never reattaches its stream on page reload — the user sees a stalled page until they act blind or navigate away and back. examples/resumable-streams-nextjs/src/app/chat/[sessionId]/ChatClient.tsx:111 does it right: status === 'active' || status === 'paused'.

Owner: examples/opennext-cloudflare-do and examples/nextjs-redis maintainers (port the resumable-streams-nextjs pattern to the other seven ChatClients).

FU-CONF-RESUMABLE-STREAMS-GUIDE: docs/guide/resumable-streams.md contradicts itself — open ​

  • Lines 113 vs 402: resume: snapshot.status === 'active' vs initialSnapshot.status === 'active' || initialSnapshot.status === 'paused' for what is presented as the same recipe.
  • Lines 149 vs 859: content replay is documented as defaulting to false, then later as contentReplay: { enabled: true } (default) — a shape that does not exist in the code; the real knobs are deps.contentReplayEnabled (on the deps bundle) and the handleChatStream param contentReplay: true (a boolean, not an { enabled } object).
  • Lines 450 vs 453-459: content replay is described as opt-in, then a few lines later as something you "disable globally" — inconsistent framing of the same default.
  • Lines 352-357: the v6 HelixChatTransport class really is gone (new HelixChatTransport(...) does throw), so the callout isn't factually wrong about that. The actual problem is the heading and the advice: "was deleted in v7" reads as "the whole thing is gone," but an export interface HelixChatTransport and an export function createHelixChatTransport producing that same shape still ship at packages/ai-sdk/src/transport/helix-chat-transport.ts:57, 93 — and the callout tells every reader to migrate to DefaultChatTransport + prepareHelixChatRequest/prepareHelixReconnectRequest instead, without mentioning that a much smaller migration (swap new HelixChatTransport(...) for createHelixChatTransport(...), same shape) is also available for whoever was using it as a transport, not as the in-framework chat-handler recipe.

Owner: docs maintainers (docs/guide/resumable-streams.md) — reconcile the two resume-gate recipes, correct the content-replay knob names/defaults, and reword the "HelixChatTransport was deleted" callout so it doesn't read as "there is no replacement" when createHelixChatTransport still ships.

FU-CONF-DBOS-RESUME-REPLAY: hypothesis for bucket B1 — NOT a finding — open ​

packages/ai-sdk/src/handler/handle-chat-stream.ts:772-773 documents that DBOS never reaches the auto-resume race branch (runtimes that auto-resume on submit take a different, earlier success path). But packages/runtime-dbos/src/lifecycle/resume.ts:152-192 (case (a), a CAS failure because the session is already 'active' AND pendingClientToolCalls is non-empty) returns a handle to the EXISTING workflow — still blocked on DBOS.recv, in the SAME run — rather than starting a new one, which is exactly the "continues in the same run" shape the cited comment assumes DBOS doesn't have. (Case (b), same 'active' check but no pending client tools, throws AgentAlreadyRunningError at resume.ts:304; a status that is neither 'active' nor 'interrupted' / 'paused' — e.g. 'completed' — falls through to a THIRD, separate branch outside the 'active' check, resume.ts:307-311, which throws Session "<sessionId>" is not in a resumable state (current: <status>); neither of those two throwing branches is the mechanism this hypothesis is about.) Because DBOS continues after DBOS.recv in the SAME run, startSequence is that run's own start, and streamLiveResponse may replay the whole run into the existing message id. A DO variant of the same risk: packages/runtime-cloudflare/src/do/do-clients.ts:303 defaults startSequence to 0 when the run row is no longer 'running', so the resumedRun?.startSequence !== undefined guard at handle-chat-stream.ts:799 passes trivially with 0. Bucket B1 must probe both of these on the client tier before this becomes a catalogue finding; this entry exists so the reproduction isn't lost in the meantime.

Owner: Phase B bucket B1 (host/client-tier conformance probe).

FU-CONF-DO-WOKEN-RUN-PAUSED-WINDOW: the woken-run ownership exit must not trust paused on the DO — open (Phase B) ​

The executor responder ends its ownership of a run woken by a submit (clientToolSubmitWakesRun) once that run was seen active and is now any non-active status, including paused with the answered entry still pending (packages/e2e/src/conformance/drivers/executor-driver.ts, awaitOwnedRunSettled). That is sound on DBOS, which never writes paused for a live run. On the CF DO it is not: ensureExecutionContinuesImpl (packages/runtime-cloudflare/src/do/durable-object-agent-base.ts ~736-829) CASes a stranded active session to paused and then resume()s it, so a live, self-recovering run passes through paused. A poll inside that window would end ownership early, and the run could leak into the next cell. No lane exercises this today (the cf-do adapter is not in the Node matrix). Before Phase B wires a cf-do lane with this responder, make the end condition per-backend (e.g. a declared pausedMayBeTransient, or waiting for a terminal status on the DO) and add a test for the reset-to-paused window.

Blocks enabling the executor-tier responder on cf-do (Phase B Task 3). The cf-do backend (packages/e2e/src/harness/setup-helpers/cf-do.ts) declares clientToolContinuation: 'auto' and clientToolSubmitWakesRun: true from source, but it has no executor-tier lane: its executor refuses every call with a HarnessError, so the responder never runs there. Resolve this item before giving cf-do an executor lane (or any lane that uses the woken-run ownership exit).

FU-CONF-DO-MINIFLARE-ALARM-ACROSS-EVICT: under Miniflare dispose+recreate, an evicted DO's durable alarm does not wake it — partly resolved by Miniflare 5 (the alarm now fires at its deadline); the on-demand fire's differences remain (Phase B) ​

Under Miniflare 3 (history; see the Miniflare 5 update below): a running DO arms a durable alarm about 30s ahead (DOAB interrupt-poll / execution-watchdog subscribers, packages/runtime-cloudflare/src/do/durable-object-agent-base.ts ~237, ~869-883, ~938-952). In production that alarm survives an eviction and wakes a fresh instance, whose onStart runs wake recovery. The cf-do conformance backend evicts with mf.dispose() + a new Miniflare over the same persist dir, and there the alarm does NOT fire on its own. Measured in packages/e2e/src/harness/cf-do/__tests__/persisted-remote.integ.test.ts ("the durable alarm across an eviction"): the alarm was armed about 30.1s after /start; after the eviction, with no request, the DO stayed asleep until 10s past that deadline; the next request woke it. Whether the stored alarm value itself survives is not observable this way, because the waking request's onStart recovery re-arms the alarm before any read. Consequences: (a) on cf-do only a request (or an explicit alarm fire) wakes an evicted DO, so wake scenarios (Task 7 wake(), Task 8 materialize.wake-hooks / MA-09) must wake it explicitly and cannot rely on the alarm; (b) the product's alarm-driven wake path is not reproduced by the harness. To reproduce it, add a /__conformance/fire-alarms route that calls the DO's real onAlarm() (spec §3 "Wake = the next request, or an explicit alarm fire"), and note which wake each scenario uses.

Task 7 update: how the harness covers the alarm path now. driver.wake(sid, { via: 'alarm' }) (BackendEnv.wake, packages/e2e/src/harness/setup-helpers/cf-do.ts wakeDo) POSTs the worker's entry route /__conformance/fire-alarms/:sid, which calls ConformanceAgentDO.fireAlarmsForConformance() by DO RPC (packages/e2e/src/harness/cf-do/worker/conformance-worker.ts). That method reads ctx.storage.getAlarm(), deletes the alarm (workerd consumes an alarm when it invokes the handler), and calls PartyServer's own alarm(): the entry workerd calls, which runs onStart (DOAB's 'wake' recovery) on a fresh instance and then DOAB's onAlarm() ('watchdog' recovery + the scheduler's reschedule). Nothing is reimplemented. It is RPC and not a DO fetch because PartyServer's fetch() runs onStart before any route handler, which would make every alarm wake a request wake first. Measured live (packages/e2e/src/harness/cf-do/__tests__/timing-controls.integ.test.ts, host and client drivers): the stored alarm DOES survive the dispose+recreate — the fire reads exactly the value armed before the eviction, proving nothing re-armed it first — and firing it re-drives the stranded run to completion. With no alarm armed, nothing is fired (fired: false) and the handler does not run. via: 'request' (the default) is a real product request: the chat snapshot route (the page load).

Which recovery an alarm wake exercises (fix round 1). WakeResult.instance reports it. After an eviction the fire meets a fresh instance ('fresh'): PartyServer's alarm() runs onStart → recoverStrandedExecution('wake') — the SAME hook-less recovery a request wake runs — and onAlarm's 'watchdog' recovery then has nothing to do. So an evicted-DO alarm wake is not independent coverage of the 'watchdog' recovery; that path is reached only by firing on a WARM (non-evicted) DO ('warm', proven live in timing-controls.integ.test.ts). Task 8 (materialize.wake-hooks, MA-09) must not claim alarm-path ('watchdog') coverage for an evicted-DO wake.

Miniflare 5 update (2026-09-27, main 186b0c6a48 bumped miniflare 3 → 5.20260925.0-alpha). The persisted alarm now DOES fire on its own after dispose + recreate, at its deadline, as in production. Measured in persisted-remote.integ.test.ts ("the durable alarm across an eviction", now pinning this; 3/3): with no request, the recreated DO stays asleep until 1s before the deadline, then the alarm alone wakes it, onStart's wake recovery completes the stranded run, and the run is completed. So "the DO woke by itself" is now observable on cf-do. It also means a wake scenario must wake an evicted DO well before the ~30s deadline, or the alarm runs the recovery instead. materialize.wake-hooks wakes within seconds, and the evict/wake self-tests check the DO is still asleep first.

What still differs from production (why this stays open): (1) resolved by Miniflare 5 (above); (2) the harness fires it on demand, not at its deadline (time is compressed, so behaviour that depends on the alarm arriving ~30s later is not reproduced); (3) a throwing handler is reported (handlerError), not retried with workerd's backoff; (4) the handler gets no alarmInfo (retryCount / isRetry; PartyServer's alarm() takes none). Closing this needs Miniflare/workerd to deliver persisted alarms after a restart, or a different eviction model (spec §10).

FU-CONF-TRUE-MID-TEXT-HOLD: no scenario can hold a model call in the middle of a text part — open (Phase B; the hold primitive landed 2026-09-27) ​

The Phase B spec §4 timing controls list driver.reload(sessionId) "mid-text". That is not delivered yet. ScriptedModel's hold (packages/core/src/llm/scripted-model.ts) blocks BEFORE the call streams anything, so no step can stop after some of its text deltas. The Task 7 self-test ("reload mid-message (tool-gated)", packages/e2e/src/harness/cf-do/__tests__/timing-controls.integ.test.ts) streams a step's text and holds the run in a tool gate instead, so the reload lands mid-MESSAGE, between steps. (Until DI-20's fix, !268, the client also errored on the product's finish-step frame before the reload. Since that fix the first instance ends client:unmounted and the reloaded instance resumes an intact stream, but still only between steps.) To close: add a ScriptedModel hold that fires after N text chunks of a response (streaming the first N, then blocking until released), then a true mid-text reload test and scenario.

Progress (2026-09-27). The hold exists: holdAfterDeltas(n, gateId, then) (core ScriptedModel, timing control B in docs/dev/conformance-scenario-authoring.md "Timing controls"). It pauses after exactly n text deltas on every backend that runs the ScriptedModel, cf-do included (the bridge's delivery barrier makes whenHeld mean the deltas reached the DO). The live self-test (harness/cf-do/__tests__/hold-controls.integ.test.ts) shows the caller holding exactly the partial text until release on cf-do and js-chat-host, host and client tiers. What remains is the scenario side: a true mid-text reload (S5's resume.replay-exact mid-text cursor).

Owner: the conformance owner.

FU-CONF-DI20-MASKED-CLIENT-CELLS: re-attribute the client-tier submit/smoke cells when DI-20 is fixed — done (Phase B, 2026-09-27) ​

At the client tier the product Chat rejected the finish-step chunk (DI-20) before anything else those scenarios judge happened, so smoke.hello-completes|{cf-do,js-chat-host}|client, submit.race-a-client|cf-do|client and submit.no-race-client|{cf-do,js-chat-host}|client were ledgered to DI-20.

Resolution. DI-20's fix (!268, 0e618772ee) landed on main first. The Phase B increment-1 branch rebased onto it, re-ran the cells and re-attributed every id from the observed run (identical 5/5):

  • smoke.hello-completes (both backends) and submit.no-race-client|js-chat-host|client now pass. DI-20's record cites them as corroboration; its provenBy is only the 4 stream.ui-chunk-schema{,-client} cells, which require step frames.
  • submit.race-a-client|cf-do|client now reaches the host. With .submitted and .window-open passing, it fails the same five ids as the host cell (the auto-send gets only data-resume-rejected, and the run is stranded paused), so it is ledgered to CP-52.
  • submit.no-race-client|cf-do|client fails .submit-accepted and .continuation-reaches-caller (data-resume-rejected while the woken run completes), so it is ledgered to CP-53.

DI-20 is fixed.

FU-E2E-DBOS-PERSISTENT-IDLE-RESTART-FLAKE: duplicate of CP-74 (a product defect, not a test flake) — closed ​

packages/e2e/src/__tests__/multi-message-execute-dbos-persistent.integ.test.ts › "idle-timeout exit then new message starts a fresh persistent workflow" fails intermittently at :436, which counts USER messages (expected [ { role: 'user', …(1) } ] to have a length of 2 but got 1). The second user message is lost. That is CP-74 (catalogued in MR !267): the idle-timeout exit marks the session completed while the DBOS workflow is still PENDING in its finally, so execute() routes a send into the exiting workflow, and the message stays unconsumed in dbos.notifications. It is not a test race; an earlier version of this item said so, wrongly. Phase B Task 9 saw it on origin/main 59d47f1e0f (1 of 3 isolated runs) and on conformance-phase-b-do-slice (which changes nothing it imports). On MR pipelines the test:integ:e2e job's retry on script_failure (.gitlab-ci.yml, .integration_base) hides it. Track the fix under CP-74 (and FU-DBOS-PERSISTENT-SEND-INTO-EXITING-WORKFLOW).

FU-INFRA-TURBO-STRICT-ENV: turbo run test:integration drops HELIX_TEMPORAL_BACKEND / RUNTIME_DBOS_INTEG — closed (duplicate of FU-CI-TURBO-ENV-PASSTHROUGH, fixed by !272) ​

Phase B Task 9's DoD matrix (2026-09-26) found that turbo's strict env mode dropped RUNTIME_DBOS_INTEG / HELIX_TEMPORAL_BACKEND / TEMPORAL_ADDRESS from test:integration. Before the fix, a local pnpm run test:integration ran the runtime-dbos integration suite gated off (0 tests) and pointed Temporal at localhost:7233. The same defect was fixed on main as FU-CI-TURBO-ENV-PASSTHROUGH (MR !272, 186b0c6a48): globalPassThroughEnv in turbo.json, guarded by pnpm run check:turbo-env. --env-mode=loose is no longer needed.

FU-CONF-RACE-A-RELEASE-ON-HEADERS: race A's hold bracket releases on the submit's response headers — open (Phase B) ​

submit.race-a holds the DO in onComplete and releases it once the auto-send is answered (.submit-returned, packages/e2e/src/conformance/scenarios/submit.ts: responded is true when the response status arrives). A CP-52 or CP-53 fix that withholds the auto-send's response until the DO is idle would deadlock against that hold: the release waits for the response, and the response waits for the release. It fails visibly (a .submit-returned timeout), never silently. If a fix MR hits it, change the bracket, not the fix: release when the request has reached the host (for example a bridge-recorded arrival of the handler's /resume or submit call), and keep the window proof (isInside + isExecuting). Recorded in CP-52's record and the authoring guide's race A section. Also for that MR: race A's caller-side ids must be re-attributed from the observed run before submit.race-a|cf-do|host can appear in CP-52's provenBy (CP-52's conformance: notes).

Owner: the conformance owner (or the CP-52 fix MR, with the owner's review).

FU-CONF-CLIENT-SESSION-REUSE-FENCE: only cf-do enforces single-use session ids — open ​

The cf-do bridge refuses traffic for a session a previous cell closed (bridge-server.ts, THE CELL CONTRACT step 5). js-chat-host has no such fence; it relies on cellId-prefixed session ids. A scenario that hard-codes a session id would therefore leak state across cells there without a HarnessError. To close: record the session ids each js-chat-host cell used (its composed cellContract, cellContractFor in suite.ts) and refuse a later cell's request for one of them.

Owner: the conformance owner.

FU-CONF-CP62-DBOS-SETTLE-TIMING: stop.mid-tool-honors-abort|dbos-postgres|executor had two outcome shapes — resolved (Phase B final fix wave, 2026-09-26) ​

The cell was nondeterministic, with both shapes caused by CP-62. Usually stop() landed after both of the step's tools (slow, held on its gate; note, instant) had started, and the outcome never settled within 30 s (.outcome-settles and its dependents failed; the ledgered shape). Sometimes the note tool's DBOS step had not started yet when stop() cancelled the workflow: that step then threw "Workflow … has been cancelled" through its three retries (1 s, 2 s backoff) and the run failed after ~3 s, so a different set failed (.outcome-interrupted, .transcript-paired, .continuation-completes). Seen in CI job 16756806141 (MR !271): https://gitlab.com/helix-ai/open-source/agents/-/jobs/16756806141. It was not the settle bound straddling retry exhaustion (that cell took 3.7 s), so no bound change could fix it.

Fix: .tool-reached (packages/e2e/src/conformance/scenarios/stop.ts, both stop.mid-tool-* scenarios) now waits until slow is running AND note has run before calling stop(), so the stop always lands at the same point of the step. The assertions are unchanged. Every runtime runs a step's tools in parallel, so this is reachable everywhere. CP-62's fragments were re-checked from fresh runs on a PRIVATE Postgres/Redis (no other session's load on the database; host load average ~40–160 from other sessions): with this change stop.mid-tool-honors-abort|dbos-postgres gave the ledgered 6-id set in 16/16 runs and stop.mid-tool-ignores-abort|dbos-postgres its ledgered 3-id set in 15/16; the pre-change commit gave the usual shape 10/10 locally (the retry-exhaustion shape was CI-only). The three Temporal cells were also identical 5/5 (the same timing applies to them; unchanged ledger). The ledger needed no change.

Residual (not the CP-62 shape; loud, never silent): in 1 of 16 private runs, the FIRST DBOS launch on a brand-new database, the lane rebuild after the stranded honors-abort cell could not stop the run within TERMINATE_BOUND_MS (20 s; the DBOS handle's abort() = cancelWorkflow + status CAS), so the backend was poisoned and the next cell reported :crash (GATED). It did not recur in 20 further warm runs (10 on the final head, 10 on the pre-change commit). If it shows up in CI, look at DBOS cold-start latency against that bound before touching the scenario.

FU-E2E-CROSS-SERVICE-BUNDLES-IN-SRC: leftover cross-service bundles make the e2e typecheck run out of memory — open (pre-existing on main) ​

packages/e2e/src/cross-service/harness.ts (~92) writes each test run's esbuild output to mkdtemp(src/cross-service/dist/bundle-*), and runs leave those directories behind. packages/e2e/tsconfig.json includes src/**/* with allowJs, so a later tsc --noEmit type-checks every leftover consumer.mjs / producer.mjs. After a local test:integration, the e2e typecheck needs ~6.5–8.8 GB and dies at Node's default heap (JavaScript heap out of memory, exit 134). Reproduced on origin/main 59d47f1e0f with the same leftovers (26 bundle dirs: 8.8 GB, rc=134); without them main needs 3.6 GB and this branch 4.0 GB. CI is not affected (fresh checkouts). Found in Phase B Task 9's final fix wave (2026-09-26). To close: write the bundles outside src/ (the OS temp dir, or a gitignored .cross-service-dist/) and/or add src/cross-service/dist/** to the tsconfig exclude, and remove each run's bundle dir in the harness teardown.

Owner: the cross-service remote-agents area owner (unassigned; reported to the reliability orchestrator).

FU-CONF-PHASE-B-SPEC-EXECUTOR-RESPONDER: Phase B spec §4 predates the executor-tier responder — done ​

Done (Phase B increment 1, 2026-09-26): Phase B spec §4 now opens with "Executor-tier responder already landed" (MR !258), and its §11 "As built" records the host and client implementations of the same option.

The conformance harness-API additions (branch conformance-harness-api, plan docs/superpowers/plans/2026-09-25-conformance-harness-api-additions.md) landed the executor-tier client-tool responder (StartOptions.clientTools, the 'submit-tool-result' driver capability) and the BackendEnv.clientToolSubmitWakesRun declaration (with its wake wait and woken-run wait — CP-64) EARLY, ahead of Phase B. Phase B spec §4 should note that both already exist at the executor tier. The spec lives on branch conformance-phase-b-do-slice; update it there on rebase.

Owner: the conformance owner (Phase B spec, on rebase of conformance-phase-b-do-slice).

FU-CONF-CF-BRANCH-ADAPTER-CRASH: smoke.branch-completes crashes on the Cloudflare executor lanes — open (CF Workflows: adapter path gone; branch cell pending MR-2) ​

Plan 0c MR-1b: the CF Workflows executor lane no longer uses a hand-rolled adapter. No branch cell has run on the lane yet: it is not-yet-run until plan 0c MR-2, whose first runs show whether smoke.branch-completes passes. Its cfw-workflows-d1 row is the real CloudflareAgentExecutor (setupCfwMiniflare) and declares branching (CFW_CAPS), so a branch reaches the product's own path. The hand-rolled adapter kept as the cfw-workflows-d1-pool row declares no branching (CFW_POOL_CAPS), and no executor lane runs over it or over cf-do-d1. What follows describes the state before MR-1b.

No Cloudflare backend declares the branching product capability, so on cf-do-d1 / cfw-workflows-d1 the runner substitutes smoke.branch-completes' REJECTION path, which calls driver.branch(...). The hand-rolled CF harness adapters refuse options.branch with a HarnessError (rejectBranchInHarnessAdapter, packages/e2e/src/harness/setup-options.ts — deliberately, so a branch is never silently run unbranched), and a HarnessError is never judged: the cell crashes (smoke.branch-completes:crash) instead of observing the product's typed framework_not_supported. This must be resolved before any CF executor lane runs: Phase B replaces the adapters with real drivers that reach the product's own branch path and surface its typed rejection. The CF lanes do not run in the executor matrix today, so no cell is affected yet.

Owner: the conformance owner (Phase B CF drivers).

FU-CONF-CP64-JUDGE-REASON: the judge matches failed ids, not failure reasons — open ​

A ledgered expected failure is matched by assertion ID only (judgeCell), so smoke.client-tool-completes.completes on dbos-postgres would stay expected-fail (CP-64) even if it started failing for a DIFFERENT cause. By design the judge compares ids; the cause is carried in the failure message instead (.completes renders the outcome with describeOutcome, incl. thrown.message — e.g. "Agent session … is already running"), so a reviewer can see it in the cell output. Revisit if the burn-down needs reason-level matching (e.g. an optional per-ledger-entry message pattern checked by conformance:verify).

Owner: the conformance owner.

FU-CONF-HOLD-CHECK-HIDES-UNDELIVERED-ABORT: a cell-end hold crash may be a product's undelivered abort — open (triage rule, non-blocking) ​

Since MR !282, runCell fails any cell that ends with a ScriptedModel call still blocked in a scripted hold (a plain hold or a paused holdAfterDeltas, via ScriptedModel.outstandingHolds()). The failure is a cell-contract <scenario>:crash, a HARNESS verdict, and the cell's observation is discarded.

Some scenarios rely on the PRODUCT to end their hold: lifecycle.ts's pause-for-resume setup waits for the stop's abort, and materialize.ts waits for the eviction to abort the held call. If a runtime fails to propagate an abort to the held model call (the stop never reaches the LLM adapter's abortSignal), the hold stays blocked to cell end. That product defect then shows up only as a harness :crash naming the hold, never as a ledgerable product finding, and the discarded observation hides the ids that would have shown it.

Harness-cut descendant sessions are excluded — nothing else. When the cell-end REAP (Temporal's terminateCellWork) terminates a DESCENDANT execution of the cell's own roots — a sub-agent or companion child workflow — whose session is NOT the session of any run this cell's RunLog tracks, that child's work was cut by the HARNESS alone: a Temporal workflow terminate never delivers an abort to the activity running a held model call, so the hold stays blocked through no fault of the product. runCell does not count such a hold as a leak; it records it in CellRun.reapedHolds (the results line's reapedHolds, informational, never a verdict) and the model reset aborts it. The own session is read from the child's start input (CellWorkTermination.terminatedDescendantSessions, then minus the tracked run sessions → ReapReport.terminatedSessions). This surfaced on !282's first pipeline (temporal-cell-reap.integ.test.ts on Docker Temporal: the reaped companion child's held call crashed the cell).

NOT exempt, and still a cell crash: a hold in the session of a ROOT or any other TRACKED run, even when the reap terminated that run (a leaked hold, or an undelivered product abort, keeps its root running to cell end — exactly what this check exists to catch; re-review 2 of !282); and a late child kept from an EARLIER cell (it descends from that cell's root, not this one's). So the triage rule below applies to every case but the harness-cut child.

Triage rule (until the fix below): a :crash whose message names ScriptedModel hold gate '…' still held at cell end (or a holdAfterDeltas pause) is a possible product bug — abort not propagated — NOT a harness flake. Check whether the scenario's path relied on an abort (the guide's "Mid-text model holds" cell-contract bullet lists which do) and whether the model record has abortedAt, before dismissing it.

To close: make the scenarios whose hold waits on a product abort record that as a product assertion before cell end (e.g. a .abort-reaches-model id from waitForAbort, released in a finally either way), so an undelivered abort fails a named, ledgerable id and the hold check only ever fires on a genuine scenario leak. Kept separate from FU-CONF-REAP-ALLOWLIST: that one is a verdict-neutral cleanup hiding a leak; this one is a verdict that is attributed to the wrong side (harness instead of product).

Owner: the conformance owner.

FU-CONF-REAP-ALLOWLIST: gate the conformance reap's reaped list on an allowlist — open (non-blocking) ​

The Temporal cell-end reap (RunLog.reapSince, MR !275) records every execution it terminates in CellResult.reaped. The list is informational only, so it never changes a verdict. The trade-off:

  • A NEW product leak is now cleaned up silently and the cell passes. Example: a child or companion left running after its parent completes (the CP-50 or CP-23 shape).
  • Before the reap, such a leak eventually made test:conformance exit red, by accident.

The fix is to gate reaped in conformance:verify against a checked-in allowlist of expected (scenario, backend) reaps. Each entry names the finding that explains it; today that is CP-23 for history.companion-continue|temporal-*. An entry not on the list fails the gate, and a stale allowlist entry (a cell that no longer reaps) is reported like a stale declared gap.

Owner: the conformance owner.

FU-CONF-TEMPORAL-TERMINATED-ACTIVITY-WRITES: an activity still running after its workflow is terminated can still write to the stream — open (non-blocking) ​

Terminating a Temporal workflow does not stop an activity that is already running on a worker. The activity is not told about the terminate until it heartbeats or completes, and it keeps running until it finishes or times out. So after the conformance reap (or dispose) terminates a cell's executions, a runtime-temporal activity that was mid-flight can still:

  • emit to the parent stream (activities.ts emit);
  • write to the store;
  • call the shared ScriptedModel.

That is the "Cannot write to stream … in 'failed'/'ended' state" shape, now bounded by the activity's own duration instead of the workflow's.

The reap does not observe running activities: it reads execution status, not pending activities. The cell-end model reset aborts held model calls, which ends most such activities quickly, and none has been observed after the reap in the matrix or in B2's runs.

Fix direction: after terminating, wait (bounded) until every terminated execution's describe().raw.pendingActivities is empty, or until the harness's background worker reports no in-flight activity for those workflows. Report the ones that don't drain.

Owner: the conformance owner.

FU-CONF-TEMPORAL-RESUME-HANDLE-EXTERNAL-TERMINATE: candidate finding — a Temporal resume handle never settles when its workflow is terminated externally — open ​

This is a candidate for the catalogue, NOT a finding yet. resume() returns the stream-observed v7 handle (runtime-temporal/src/handle.ts:115-146). It resolves only on a terminal, suspension or interrupt chunk, or when the stream drains. A workflow terminated from outside (an operator, a deploy script, the conformance reap) writes no chunk and does not end the stream, so the handle's result() may wait forever.

The conformance harness sees exactly this shape whenever it terminates a __resume-N workflow. It reports that under terminatedUnsettled (harness-caused), never under a product finding. The shape resembles CP-65, but CP-65 is a different cause: a resume re-suspends without a marker.

Next step: a conformance scenario that terminates a resumed run's workflow through the backend and expects the handle to settle, as failed or terminated, within a named bound.

Owner: the conformance owner (catalogue decision).

B8 (transcript validity, RM-29 / DI-17) follow-ups ​

FU-B8-01: Phase B host-tier .heal-spares-pending assertion — open ​

The C5 continuation-heal claim "a call registered as pending (client tool or approval) is left untouched" is proven today only by runtime unit tests. The executor tier cannot seed an interrupted session with a registered pending client tool (seedMessages writes messages only, and a real client-tool suspension leaves the session paused, not interrupted), so the conformance scenario cannot exercise it yet. Needs a host-tier harness capability (spec §2, C5).

Owner: conformance owner (Phase B host tier).

FU-B8-02: RM-10 masks Temporal .llm-input-paired in transcript.* — unblocked (re-run pending) ​

On Temporal, every transcript.* scenario's .llm-input-has-turn1-calls failed because the model input carried no history at all (RM-10), which also masked whether .llm-input-paired would otherwise pass there. B2-core fixed RM-10 (runLLMStep now loads the complete history in the activity). Re-run these ids on the B2 matrix — .llm-input-paired becomes a live RM-29 proof on Temporal for the first time.

Owner: B2-core (RM-10).

FU-B8-03: transcript.finish-with-client-sibling scenario for RM-29 shape (d) — open ​

Shape (d) — a finishWith call deferred by a same-step client tool / suspending sub-agent, executed once after resume — is proven today by runtime unit tests only (JS/DO/CFW/Temporal/DBOS). Add a conformance scenario once Phase B's executor tier supports submitting a client tool result mid-scenario.

Owner: B8 follow-up (needs Phase B executor submit-tool-result capability).

FU-B8-04: transcript.finish-cosiblings.cosibling-not-executed — done ​

Done: transcript.finish-cosiblings now declares .cosibling-not-executed (the co-emitted note tool's in-process execution witness never records the turn-1 cosibling) with its positive control .cosibling-witness-observable (the same tool executed in turn 2 is witnessed), plus .cosibling-state-not-written / .cosibling-marker-observable (its customState marker is absent after turn 1 / persisted after turn 2, read via PersistedObservation.customState). All four passed pre-B8 and post-B8 on all 7 Node backends (the cosibling never ran); RM-29 carries a conformance: note. The Vercel path, where the co-emitted call is never produced, is RM-35.

Owner: B8 follow-up.

FU-B8-05: DI-17 fixed-flip — done ​

Done in !257 (rebased onto !258): DI-17 is fixed with provenBydiscipline.failed-tool-marked (JS ×3 + Temporal ×3) and transcript.finish-with-fails-then-retries (JS ×3); DBOS is contrast; CF DO / CFW / abort branch / Temporal companion dispatch are unit-test-only, re-open on Phase B.

FU-B8-06: synthetic not-executed results are invisible to hooks, the stream and AI SDK UI parts — open ​

Synthetic not-executed tool results (C4/C5) are detectable only on persisted helix Messages (isNotExecutedToolResult). Three parity gaps:

  • onMessage: synthetic results and the results runPhase2FinishWith appends on Temporal are not routed through the onMessage hook the way regular tool results are, so a hook-based observer misses them.
  • Live stream: llm-vercel's chunk mapper emits tool_start for every model tool call, losers included, but no runtime emits a tool_end for a synthetic result, so a useChat part stays in input-available.
  • AI SDK converter: helix-to-aisdk-converter.ts renders a persisted synthetic result as a generic output-error and does not carry metadata.helixSynthetic onto the part, so a reloaded UI cannot tell "not executed" from a real failure.

Needs a parity decision (fire hooks? emit a tool_end error chunk wherever synthetic results are appended? surface the marker on the converted part?) before implementing. The guides (finishing-agents.md, sub-agents.md) state the current limit. Also covers MR !257 review A m4.

Owner: needs parity decision (core/runtime + ai-sdk maintainers).

FU-B8-07: pendingClientToolCalls handling duplicated across JS/Temporal/CFW — open ​

loadAgentState/AgentState don't carry pendingClientToolCalls directly, so each of JS, Temporal and CFW re-derives or re-plumbs the same workaround to pass pending ids into the C5 heal's excludeIds. Consolidate into one shared helper or a field on the loaded state so the three runtimes stop duplicating this.

Owner: open (core/runtime maintainers, cross-runtime cleanup).

FU-B8-08: CFW deferred phase-2 loop duplicates the normal phase-2 loop — open ​

workflow.ts's deferred-finishWith-after-resume path re-implements the same phase-2 (finishWith execution + terminal-result synthesis) logic as the normal in-step phase-2 loop, instead of sharing one helper. Extract a common function so the two loops can't drift.

Owner: open (runtime-cloudflare maintainers).

FU-B8-09: DBOS CAS→startWorkflow crash window widened slightly by the heal write — open ​

DBOS's continuation path has a pre-existing crash window between the completed → active CAS and startWorkflow; the C5 heal write (pairing trailing unpaired calls before the new turn's input) adds one more durable write inside that window, marginally widening it. The same applies to the heal in retry() (failed → active) and resume() from interrupted (MR !257 review F1). Pre-existing risk, not introduced by B8, but worth folding into whatever fixes the underlying crash window.

Owner: open (runtime-dbos maintainers).

FU-B8-10: DBOS dispatcher.appendMessages is at-least-once across a crash — open ​

dispatcher.appendMessages runs the store append inside a DBOS step. If the process crashes after the store write lands but before DBOS checkpoints the step's completion, recovery re-runs the step and appends the same messages again. A duplicated tool_result is an invalid transcript (RM-29). This is pre-existing and affects every dispatcher.appendMessages call site in runtime-dbos/src/workflows/shared.ts (phase-1 tool results, phase-2 finishWith results, the per-step terminal heal), including the new RM-29 superseded-results append (shared.ts ~2518). Temporal and CFW were hardened in the B8 final review: their phase-2 / superseded appends only append results whose toolCallId is still an unpaired trailing call in the persisted history (core filterUnpairedTrailingResults over getAllMessages). Apply the same helper inside the DBOS append step for role: 'tool' messages.

Owner: open (runtime-dbos maintainers).

FU-B8-11: Temporal main-loop no-output phase-2 drops the finishWith tools' customState (RM-34) — open ​

In runtime-temporal/src/workflow.ts's main-loop phase-2 branch, when no finishWith call produces output, phase2.nextState.customState (a failing finishWith's partial writes and its hooks' writes) is dropped; JS persists it via commitStep. Catalogued as RM-34 (pre-existing) and owned by that bucket; the site carries an RM-34 comment. The deferred-finishWith-after-resume path added in MR !257 does NOT repeat it: it commits the deferred phase-2 customState (review F12).

Owner: RM-34 bucket.

FU-B8-12: expectPairedTranscript test helper copied into six files — open ​

The adjacency-aware pairing assertion is duplicated across the RM-29 runtime tests (JS, Temporal, DBOS, CFW). Move it to core's testing utilities (@helix-agents/core/testing) and import it everywhere.

Owner: open (test-infrastructure cleanup).

FU-B8-13: runDeferredFinishWith bare catch {} and run-loop's double buildIteration() — open ​

runDeferredFinishWith (core step-iterator.ts) swallows a buildEffectiveTools failure with a bare catch {} (the normal step then surfaces it), and runtime-js/src/run-loop.ts calls buildIteration() twice per iteration (once for the deferred check, once for runStepIteration). Log the swallowed error and build the deps/hooks once when nothing is deferred.

Owner: open (core / runtime-js maintainers).

FU-B8-14: the deferred-finishWith path skips the maxSteps pre-check — open ​

A finishWith call deferred by a phase-1 suspension runs after resume without the maxSteps pre-check the normal step performs. This matches the pre-fix behaviour (the call belongs to an already-counted step) but is undocumented as a decision. Decide and pin it with a test.

Owner: open (core maintainers).

FU-B8-15: HELIX_SYNTHETIC_METADATA_KEY vs COMMON_METADATA_KEYS — open ​

The synthetic marker key lives in its own constant (transcript.ts) instead of a COMMON_METADATA_KEYS.SYNTHETIC entry beside TOOL_FAILED (types/state.ts). Consider moving it for consistency (a public-export change).

Owner: open (core maintainers).

FU-B8-16: Temporal main-loop interrupt emits run_interrupted before persisting the status — open ​

The main-loop interrupt branch (the finishInterrupted closure in runtime-temporal/src/workflow.ts) emits run_interrupted before persistTerminalState('interrupted'), so a handle can resolve interrupted while loadState still reads active for a moment. The post-drain resume branch and the deferred-finishWith interrupt path (persistFirst: true) already persist first. Make persist-first the only order (pre-existing; changing it reorders the main loop's activity sequence).

Owner: open (runtime-temporal maintainers).

B2-core (LLM history completeness) follow-ups ​

Deferred from the B2-core bucket (spec ../superpowers/specs/2026-09-25-llm-history-completeness-design.md). None of these blocks the B2 contract; each names the finding that owns it where one exists.

FU-B2-01: DBOS checkpoints the full history every step (O(N²) system-DB growth) — open (MED) ​

getMessagesStep is a @DBOS.step, so the complete history it returns (now the WHOLE log, since RM-11 removed the 10k cap) is written to the DBOS system database once per LLM step. A session of N messages over S steps records O(N·S) bytes, O(N²) over a long session. Fix: move history load, message build, the onMessage hook fires and the LLM call into callLLMStep, so the history never becomes a step output. Accepted cost for B2 (spec §3.3).

FU-B2-02: DBOS from_checkpoint / retry edge cases (B6 minors) — open (LOW) ​

  • M1 A stranded session (workflow terminal or missing, session still active) cannot be rewound: resume({ mode: 'from_checkpoint' }) throws AgentAlreadyRunningError.
  • M2 Between the claim CAS and DBOS.cancelWorkflow, the suspended workflow can still complete a step; the rewind then races that write.
  • M3 A sub-agent's JSON input that happens to have a message key is still parsed as an AgentInput (only a content-less message falls back to the JSON form).
  • M4 A rewind that fails after the suspended workflow was cancelled rolls the status back to interrupted but leaves the old pendingClientToolCalls in place.
  • Kept pending calls after a DBOS rewind are not re-dispatched; they rely on the next continuation's trailing heal.

FU-B2-03: DBOS checkpoint state is the pre-step snapshot — open (LOW) ​

DBOS inter-step and terminal checkpoints store state loaded at the top of the iteration, not the state after the step's tool writes (messageCount is live, state is not). A rewind to such a checkpoint restores customState one step early. Other runtimes snapshot the committed post-step state. (B12 review M-1.)

FU-B2-04: CP-68 residue (B13 minors) — open (LOW) ​

  • No dedicated test for the terminate-write race in recordPersistentChildNonCompletion (the re-assert is idempotent and self-heals).
  • CF Workflows: no workerd test for a DIRECT child interrupt (parent-driven is covered by unit tests; the workerd harness cannot drive it, CP-23).
  • Temporal and DBOS never record 'interrupted' on a companion's SubSessionRef, so Temporal's companion__waitForResult on a stale running ref can wait until its timeout (pre-existing; sent to the conformance owner).
  • Temporal __resume__${stepCount} child workflow ids can collide across turns (pre-existing; sent to the owner).
  • A re-interrupt that lands before the resumed child's first step commits rolls back the message sendMessage delivered (root resume({ mode: 'with_message' }) has the same window; sent to the owner).
  • JS: resuming a companion from interrupted does not run the trailing unpaired-call heal (RM-29 territory; sent to the owner).

FU-B2-06: Temporal resume with nothing to drain never settles (CP-65) — open ​

resume() on a session that is still suspended with nothing to drain early-returns without a suspension_marker, so handle.result() hangs. B2's from_checkpoint path does not hit it (it rewinds first), but other resume-on-suspended cells do. Owned by CP-65.

FU-B2-07: CF Workflows minors (B8 / B12 reviews) — open (LOW) ​

  • retry() still reads the triggering message with an unchecked getAllMessages (executor.ts), not loadLLMHistory.

  • The lenient loadAgentState reads the full history even though bookkeeping callers never use it.

  • markAgentFailed no longer refreshes updatedAt; persistFromCheckpointHistoryFailure does a plain saveState after its CAS (same as markAgentFailed).

  • run_resumed.step for a rewind reports the pre-rewind stepCount.

  • A rewind's truncation is not undone when a later post-claim step throws (only the status is rolled back).

  • writeTerminalCheckpoint does extra store calls inside the retried executeAgentStep step; a step retry writes one more (equally truthful) checkpoint.

  • Temporal and CFW retry() run truncateMessages before throwing "No user message found to retry with". Benign: the truncate target is the latest checkpoint, so nothing committed is lost, but the throw should come first.

FU-B2-08: runtime-js minors (B7 review) — open (LOW) ​

  • Cross-turn from_checkpoint stream cleanup is floored at the current run's start sequence, and stream_resync is emitted only when chunks were removed.
  • loadFromSession reads the history twice.
  • retry()'s history-failure seam skips the FIX-B rollback when its CAS is lost.
  • Loading the companion history before the CAS widens the window in which a sendMessage append can land between the read and the resume.
  • No dedicated stream-readback test for the retry() seam; initializeSession seam continuation/checkpoint-branch cases are untested.

FU-B2-09: runtime-temporal minors (B4 / B5 reviews) — open (LOW) ​

  • resume() leaves the session active if the rewind or the with_message append throws after the CAS.
  • extractApplicationFailureDetail accepts ANY ApplicationFailure whose details[0] parses as an ErrorDetail; a future activity that puts a message-only object there would be misread as classified.
  • The branch activity wraps only HelixErrors as non-retryable (the history-load site wraps everything).
  • A branch with both an invalid source and an invalid checkpoint reports state_checkpoint_not_found before state_session_not_found.
  • No unit tests for branch messageIndex / checkpointId-over-messageIndex precedence.

FU-B2-10: core primitive minors (B1 / B2 / B2b reviews) — open (LOW) ​

  • loadLLMHistory reports read: 0 when paging throws midway (the partial count is lost).
  • rewindToCheckpoint has no guard that checkpoint.sessionId === sessionId (callers resolve first); a negative branch.messageIndex is not validated.
  • A restored pending entry keeps isPreSubmitStub verbatim from the snapshot.
  • No direct test that a currently pending entry wins over the snapshot entry.

FU-B2-11: store minors — open (LOW) ​

  • Redis appendMessages is RPUSH then HINCRBY (two round trips); a truncate racing an append can skew listSessions().messageCount until the next write. Resolved by RM-50 (!269): appendMessages is now one Lua script (RPUSH + HINCRBY messageCount, guarded on the session existing), and truncateMessages is one Lua script that rewrites the count from LLEN, so the two serialize and the hash count cannot skew.
  • The executor-tier e2e cases for forced-completion checkpoints cannot fail on stores whose saveState mints a checkpoint (memory / Redis / Postgres). The non-minting-store family in session-model-branching.integ.test.ts covers Temporal and runtime-js; DBOS has no such cell (it runs on Postgres only), and D1 / CFW catch it on workerd.

FU-B2-12: out-of-scope findings surfaced by B2 — tracked by the conformance owner ​

  • CP-30 DBOS retry() ignores RetryOptions (checkpointId, message).
  • CP-34 The DO host /resume drops checkpointId, so from_checkpoint cannot reach the DO host tier.
  • CP-41 CFW / Temporal treat the parent_message flag as an interrupt, so sendMessage to a RUNNING child can stop it with the message unread.
  • CP-58 retry() without options.message after a first-step failure throws "No user message found to retry with" on runtime-js. CFW now behaves the same after a forced-completion failure, whose tail writes a terminal checkpoint (RM-33); an exception exit writes no extra one (R28); retry truncates to the latest checkpoint.
  • CP-69 DBOS blocking spawn waits for the child's idle timeout (FU-DBOS-BLOCKING-SPAWN-SEMANTICS).
  • CP-70 The DO base swallows a typed pre-claim resume failure: the session stays paused with no errorDetail.
  • RM-13 A store that skips a corrupt row on read now makes loadLLMHistory fail the run (state_history_incomplete) instead of silently dropping the row; corrupt-row paging itself stays with RM-13.

FU-B2-14: CF Workflows whole-log reads that only need scalars — open (LOW) ​

All pre-existing on main, performance only:

  • runtime-cloudflare/src/workflow.ts ${agentType}-check-existing returns the full AgentState from the file-local loadAgentState, which loads ALL messages (unchecked getAllMessages) into the step result. A long session risks the Workflows step-result size limit. The step only needs status / stepCount / a few scalar fields.
  • CloudflareAgentExecutor.loadAgentState reads messages with unchecked getAllMessages; it feeds resume() bookkeeping, never the model, so no messages are needed there.

FU-B2-17: a non-owner write can fail a pinned from_checkpoint rewind — done (superseded by S5) ​

Superseded by the S5 snapshot MR: no runtime claims a rewind with a version-pinned CAS any more. A from_checkpoint rewind (and a retry's rewind) is part of the entry's run-start commit (StartRun.truncateTo, planned by planRewindToCheckpoint), which is guarded by the run CAS (expectCurrentRunId), not by the session version, so a non-owner version bump (a submit on a kept call, a child's root-ownership update) no longer fails it. The text below describes the old model. rollbackLostRewindClaim is unused (FU-S5-CORE-UNUSED-REWIND-CLAIM-HELPERS).

Owner: orchestrator (re-brief).

A claim-pinned rewindToCheckpoint (claimedVersion) fails with AgentAlreadyRunningError on ANY version bump between the caller's claim CAS and its rewind save, and nothing is written or truncated. Some bumps come from writers that do not own the session:

  • a submitToolResult on a call the rewind would keep;
  • a child session's root-ownership update (updateRootOwnership) when the rewound session is that child's root.

Resolved: the stranded claim. A claim from a NON-active source (interrupted / paused / completed / failed) that now has no pending calls cannot have a competing owner, because on every runtime:

  • assertRewindable rejects an active session with no pending calls;
  • continue / with_message reject it as running. runtime-js and DBOS always did this; Temporal and CFW now do too, and before this change they would start a second run;
  • execute() rejects it.

So the lost claim was a non-owner bump. Core rollbackLostRewindClaim moves the session back to its source status with a pinned CAS, on JS, Temporal, CFW and DBOS. Each runtime has a forced test: an interrupted source with a non-owner bump ends interrupted and rewinds again.

What remains:

  1. The caller sees a spurious AgentAlreadyRunningError and must retry.

  2. A lost claim is still not rolled back when either:

    • the source was active (a suspended session), or
    • the session now has pending calls.

    A competing claimer may own the session in both cases. The session is left active with its pending calls, which reads as suspended, so it stays rewindable on every runtime and continuable on Temporal and CFW. For a paused / interrupted source with pending calls, the status change to active is the only lasting effect.

A precise fix, one that re-bases across non-owner bumps, needs a claim token that the claim itself writes, which means a store-level field. Owner: B2 follow-up.

FU-B2-16: root test:integration runs two DBOS suites against one system DB — open (LOW) ​

Locally, RUNTIME_DBOS_INTEG=1 pnpm run test:integration runs the e2e DBOS cells and the runtime-dbos integ suite in two concurrent processes against the same DBOS system database, and DBOS e2e cells then fail (e.g. a turn ending failed, or agent-hooks-parity DBOS "expected [] to include 'beforeLLMCall'"). Pre-existing on main (same family, different test names); CI runs these as separate jobs, so it does not hit this. Suspected cause: the process-local hook registry / workflow ownership across two DBOS processes (distinct executor ids, so unproven). Workaround: run the DBOS suites package-direct. Owner: ci-health (B8).

FU-B2-15: pre-existing observations from the B16 fix round — reported to the conformance owner ​

  • CF Workflows non-checkpoint resume() truncates before an unpinned claim, so it can race a live run. Pre-existing on main; reported to the conformance owner (CP-35). B2 later narrowed it: executor.resume() continue / with_message reject a live active session before any write and claim with a version-pinned CAS before any cleanup or truncate, and from_checkpoint claims with a version-pinned CAS before it rewinds. CP-35 stays open for handle.resume() (no guard, truncates before an unpinned claim) and retry()'s cleanupToStep without startSequence; the live-run half is CP-83.
  • At maxSteps, an agent without outputSchema ends completed on JS and Temporal but failed on CF Workflows. Pre-existing on main; reported to the conformance owner (candidate).
  • Temporal handle.result() can resolve (from the stream) before the workflow's last activities persist the final rows and terminal status. Pre-existing on main; reported to the conformance owner (candidate).

S5 (consistent snapshots + checkpoint truth) follow-ups ​

Deferred from the S5 snapshot MR (branch reliability-example-refresh-flake; spec ../superpowers/specs/2026-09-26-consistent-snapshot-checkpoint-truth-design.md, §9 "Not in scope" and the review rounds). None blocks the S5 contract. Owners marked "orchestrator" are routed by the reliability orchestrator.

FU-S5-B1-CONFLICT-REACTION: what a client does after a refused entry — open (owner: B1) ​

A live run-start conflict surfaces as RunStartConflictError (named AgentAlreadyRunningError). A chat send refused that way answers 409 {code: 'state_already_running', retryable: true}; resumes (handleResumePath, the DO's ensureExecutionContinues) and requests with no new user content still attach to the winner's stream. The spec leaves the client-side reaction to B1: whether useHelixChat should retry a refused send once the winner settles, attach to the winner, or keep surfacing the error (today: chat.error + parseHelixChatError, and chat.regenerate() retries).

FU-S5-LIVE-SUPERSEDED-RUN-CHUNKS: a superseded executor's chunks before the new run's claim — open (LOW, owner: orchestrator) ​

Once a newer run has claimed the session stream, every stream manager drops an own-session chunk stamped with another runId (chunkFenceRunId), and endStream / failStream / pauseStream are owner-fenced (CP-85). Between the new run's run-start commit and its claimStream, the stream still names the old owner, so a superseded executor's chunks pass. On runtime-js the lease gate covers that window (gateWriterOnLease). On the CF DO, Temporal, DBOS and CF Workflows the old executor stops at its next fenced write, which can be a whole LLM step of streamed text. A live client can show that text until it reloads; a snapshot never contains it.

FU-S5-LIVE-ATTEMPT-DUPLICATION: a retried attempt's text stays on a live screen — open (LOW, owner: orchestrator) ​

A Temporal activity retry or a CF Workflows step.do retry re-streams the step under a new attemptId. The AI SDK protocol cannot retract rendered text, so a live client shows both attempts until it re-anchors. data-attempt-superseded is a resync trigger: a client wired with useHelixChat({ resync }) (or useAutoResync with stream control) re-anchors on it. Remaining gaps: a client without resync wiring keeps the duplicate until reload, and a run that FAILS with a superseded attempt on screen gets no re-anchor (the failed stream ends before a resync lands).

FU-S5-AUTO-SEND-RECHECK: re-check the spurious auto-send after a cursor resume (B1 / B6) — open (owner: B1) ​

Out of scope for S5 (spec §9). It was expected to disappear because snapshots now anchor at step boundaries and replayed tool calls are never re-dispatched (useHelixChat filters replayed ids; a replayed PENDING client tool runs once after the replay). Re-check on the host and client tiers with the sendAutomaticallyWhen wiring the examples use.

FU-S5-SUBAGENT-INNER-CONTENT-ON-RELOAD: a sub-agent's inner stream content is not in the snapshot — open (owner: orchestrator) ​

Out of scope for S5 (spec §9). A snapshot holds the parent session's committed rows; a sub-agent's inner-stream content (its own text and tool parts rendered inside the parent's message) is not rebuilt on reload. The child's content is durable in its own session.

FU-S5-WORKERD-FIRST-SEND-CONNECTION-LOST: workerd Network connection lost on a first send under load — open (LOW, owner: B8) ​

2 of about 3250 opennext stress runs on the S5 branch (6 workers × 50) failed on the very first send POST with workerd Network connection lost inside miniflare's entry worker: the session never started. 0 in about 1900 base runs, so not proven unrelated. Same local-workerd class as FU-OPENNEXT-UNCAUGHT-ON-CLIENT-ABORT and the "worker restarted mid-request" flake (spec §9, out of scope).

FU-TEMPORAL-DBOS-HOST-TIER-REFRESH-PROOF: no Temporal- or DBOS-hosted chat-handler refresh proof — open (owner: the Temporal round / B6) ​

The char-exact refresh e2e runs on JS ×3, Temporal ×3, DBOS and CF Workflows against the executor, and the host / client conformance cells run on cf-do and js-chat-host. There is no Temporal- or DBOS-hosted handleChatStream refresh spec, and the executor tier has no snapshot route, so resume.replay-exact has no Temporal / DBOS host cell. Add a Temporal-hosted (and DBOS-hosted) chat host plus host cells.

FU-DBOS-HARDCRASH-REPLAY-NONDETERMINISM: DBOS restart-recovery replay fails determinism after an in-process hard crash — open (LOW, owner: orchestrator) ​

An integ restart-recovery test (approval-gate, client-tool-recovery-dbos) failed DBOS's determinism check in about 1 of 2 full-suite runs during S5 Task 13: the orphaned pre-crash body, still running after the harness's in-process hardCrash, recorded a step at the function id the recovered body expected. The S5 ownership guard (an execution that lost its workflow writes no durable step; workflows/execution-guard.ts) addresses that class. Re-measure over 20 or more runs; close if clean, otherwise confirm which executor recorded the step (listWorkflowSteps).

FU-S5-PATH5-FOREIGN-TAB-SEND: a new message from another tab can attach to a live stream — open (owner: B1) ​

handleChatStream answers a refused SEND with 409 state_already_running, but the active-stream attach probe (path 5) can still attach a request whose trailing user message is new (sent from another tab while a run streams) instead of refusing it. Fix direction: compare the trailing user message id with the live run's input before attaching.

FU-S5-STREAM-OUTAGE-SWEEPER: a permanent stream-backend outage leaves the stream active — open (LOW, owner: orchestrator) ​

If the stream backend stays down through a run's terminal path, the terminal claim and the stream transition fail and are logged; the run status is terminal but the stream stays active. The snapshot derives ended / failed from the run, so readers are correct, but nothing sweeps the stream.

FU-S5-DO-BACKLOG-BOUNDARY-LATE: the DO client's caughtUp can resolve late — open (LOW, owner: orchestrator) ​

DOStreamManagerClient's resumable reader learns the end of its replayed backlog from the server's SSE replay; the boundary can be late (never early). A late boundary only delays the switch from the backlog filter to the live filter. An exact boundary needs a DOAB change.

FU-S5-OPENNEXT-DEBUG-ROUTES-DO-ADDRESSING: opennext debug routes read the wrong DO — open (LOW, owner: orchestrator) ​

getDOStub (debug routes only) addresses the DO as session:<id>, but the chat handler's DO clients use the bare id, so the gated debug routes read an empty DO. The same split was fixed in research-assistant-cloudflare-do by S5.

FU-S5-DBOS-OWNERSHIP-GUARD-SDK-COUPLING: the DBOS ownership guard keys on DBOS.tracer identity — open (LOW, owner: orchestrator) ​

workflows/execution-guard.ts detects a replaced DBOS executor by the identity of DBOS.tracer (one per DBOSExecutor). That holds on the DBOS SDK version the package pins; an SDK upgrade must re-verify it.

FU-S5-RUN-ERROR-DETAIL: run records carry no errorDetail — open (LOW, owner: orchestrator) ​

A failed run's classified ErrorDetail is on the session (SessionState.errorDetail) and the failed stream chunk, not on its RunMetadata. Run-level detail needs a schema change in all five stores.

FU-S5-CORE-UNUSED-REWIND-CLAIM-HELPERS: core still exports the !279 claim / hold helpers — open (LOW, owner: orchestrator) ​

Every runtime now rewinds inside its run-start commit (StartRun.truncateTo, planned by planRewindToCheckpoint). No runtime calls rewindToCheckpoint, rewindOutcomeOf, rollbackLostRewindClaim, rollbackFailedResumeClaim, failRewoundSession, rewindHoldOf, isRewindHoldTakenBy or rekeyRewindHold any more. Remove them (no back-compat) or document them as building blocks for a custom loop.

FU-CORE-REWIND-HOLD-JSDOC-STALE: PendingClientToolCall.rewindHeld* JSDoc describes the removed hold machinery — open (LOW, owner: CF-D2 or the next core-touching MR) ​

The JSDoc on PendingClientToolCall.rewindHeld, rewindHoldKey and rewindHoldToken (packages/core/src/types/state.ts:155-190) still describes the !279 rewind hold: a crashed rewind's held calls, …__rewind-<key> instance ids and takeover tokens. Rewinds now happen inside the entry's run-start commit (StartRun.truncateTo), and no runtime sets these fields any more; only assertRewindable reads rewindHeld and releaseRewindHold strips them. Rewrite the JSDoc to say so (or remove the fields together with FU-S5-CORE-UNUSED-REWIND-CLAIM-HELPERS). Found while rewriting the llm-history changesets as-built for the first release since 2026-09-23 (RELEASE-1).

FU-S5-CFW-HANDLE-BEFORE-RUN-START: CF Workflows timing edges at the run-start window — open (LOW, owner: CF-D1) ​

  • getHandle() can return a reconnect handle before the instance's run-start commit lands ("no live run yet" is reported only when no run record exists at all).
  • A send that arrives while the previous turn's instance finishes waits up to priorInstanceSettleTimeoutMs (30 s) before its 409.

FU-RUN-START-PLAN-DURABLE: the run-start plan lives in memory — open (MED, owner: orchestrator) ​

Temporal, CF Workflows and DBOS keep a committed run-start entry's plan (restore target, entry fields, rewind, discarded calls) in memory, keyed by runId (Temporal GenericActivities.committedEntryPlans, CFW AgentSteps.committedEntryPlans), so a retried activity / step.do after the commit rebuilds the entry instead of re-planning. A retry that lands on ANOTHER worker or invocation, after the commit and before the entry state was written, has no plan: a retry entry then fails non-retryably (framework_internal_error) and its committed run ends failed rather than starting from a wrong state (on CFW the engine retries that error twice more, harmlessly). Fix: record the plan's outputs durably inside the run-start commit (an optional StartRun.entryPlan stored on RunMetadata: the restore target's checkpoint id plus the entry fields), across core and all five stores, and drop the fallback.

FU-S5-APPROVAL-ID-DURABLE-RUNTIMES: Temporal / DBOS / CFW do not record approvalId on pending entries — open (LOW, owner: orchestrator) ​

Only the core step iterator (runtime-js, the CF DO) writes PendingClientToolCall.approvalId. Temporal, DBOS and the CFW approval helper emit approvals themselves and do not, so the snapshot infers a pending SERVER tool to be an approval; a client-executed tool that is also approval-gated renders input-available in the snapshot on those runtimes until the re-sent tool-approval-request arrives. Set approvalId in runtime-temporal, runtime-dbos and approval-gate-workflow-helper.ts.

FU-CFW-TERMINATE-RUN-STREAM-ACTIVE: executor.terminateRun() leaves the stream active — open (LOW, owner: CF-D1) ​

After terminateRun(sessionId) the CF Workflows instance is terminated, but the session stream stays active and the run stays live until the next entry supersedes it. A design decision so far; a reader that attaches in between waits until the next entry or its live-wait fail-safe.

FU-S5-ZOD-ERROR-NAME-CHECKS: name / cast error checks should be Zod schemas — open (LOW, owner: orchestrator) ​

CLAUDE.md principle 5. Manual name and cast checks remain in Temporal activity-failures.ts, asLostRetryRace and isSupersededFailure; CFW StepMessageSchema (a z.custom duck type), the as Partial<AgentWorkflowResult> engine-output casts and the message-regex error classification; and the resync / tracker / use-helix-chat guards in ai-sdk.

FU-TEMPORAL-PERSIST-TERMINAL-STATE-TEST-ONLY: delete the test-only persistTerminalState activity — open (LOW, owner: the Temporal round) ​

Production Temporal ends a run through commitTerminalState + finalizeTerminalRun. persistTerminalState is still exported and used by about 10 test files only. Move those tests to the production activities and delete it.

FU-S5-SIBLING-CHECKPOINT-SCANS: CFW and DBOS still scan every checkpoint of a session — open (LOW, owner: CF-D1 / the DBOS round) ​

planRetry now reads only the failed run's checkpoints through the new listRunCheckpoints (review I10). Two sibling scans still read every checkpoint in full: CFW appendFailedStepRows (runtime-cloudflare/src/steps.ts), which runs inside one Worker invocation against D1's per-invocation query budget, and DBOS run-commit.ts. Switch both to listRunCheckpoints.

FU-DBOS-CLIENT-TOOL-RECV-OUTLIVES-FAILED-TURN: a client-tool recv can outlive a failed DBOS turn — open (LOW, owner: the DBOS round) ​

When a DBOS turn fails while one of its client-tool calls is waiting in DBOS.recv, the waiting step can outlive the settled turn. A late submit for that call then lands on a run that already ended failed. Bound or cancel the pending recv when the turn settles.

FU-S5-ONAGENTFAIL-WRITES-K5-JS-DO: onAgentFail state writes are dropped on runtime-js and the CF DO — open (LOW, owner: orchestrator) ​

On every terminal exit the K5 commitState carries the terminal status and error. On completion it also carries the agent's onAgentComplete state writes on every runtime. onAgentFail writes land in K5 only on Temporal and CF Workflows: runtime-js (and the CF DO, which runs its executor) fires onAgentFail after K5 and drops what the hook writes. DBOS hooks cannot write state. Apply onAgentFail writes to K5 on JS / DO (as onAgentComplete does), or document the hook as read-only everywhere. Add a conformance cell for it.

FU-S5-DBOS-SUPERSEDED-HOOK-PAYLOAD: DBOS onAgentSuperseded gets a placeholder payload — open (LOW, owner: the DBOS round) ​

runTurnMappingStops (runtime-dbos/src/workflows/standard-workflow.ts) fires onAgentSuperseded with { error, finalState: {}, stepCount: 0, durationMs: 0 } and context.stepCount: 0. The other runtimes pass the run's real final state, step count and duration. Load the session's state (as a recorded step) and pass the real values, so Langfuse spans and audit hooks see the same payload on every runtime.

FU-S5-DBOS-ONAGENTRESUMED-THROW-FAILS-RUN: a throwing onAgentResumed fails the run on DBOS only — open (LOW, owner: the DBOS round) ​

On runtime-js, the CF DO, Temporal and CF Workflows an error from onAgentResumed is logged and the resume goes on. On DBOS onAgentResumed is the run's start hook, so a throw fails the run (with one onAgentFail), like a throwing onAgentStart (documented in docs/guide/hooks.md). Pick one rule for every runtime and add a parity cell.

FU-HANDLE-RESUME-LOSER-UNTYPED: a getHandle(...).resume() loser surfaces as a failed result, not a typed error — open (LOW, owner: reliability orchestrator (route)) ​

CP-96 gave resumeLoop (the reconnect handle's resume) a seen run, so a stale handle resume now loses its run-start commit and writes nothing. But the handle factory's safety net reports that loss as an AgentResult with status 'failed' and the message, not as a typed AgentAlreadyRunningError / RunStartRejectedError. This is the existing handle-factory behaviour, not new. Surface the typed loser to the handle's result() consumers.

FU-CP96-DURABLE-LOSER-CLASS: the durable runtimes add two loser classes to the CP-96 race's C3 outcome — open (LOW, owner: reliability orchestrator) ​

CP-96 ruling U3 gives the stale-entry race one caller-facing outcome: the C3 AgentAlreadyRunningError (RunStartConflictError('expected_mismatch')). runtime-js and the CF DO always report it. Temporal, DBOS and CF Workflows report it too when the stale entry's commit lands after the winner started (the run-id CAS or the run-status CAS; proven by the real-engine cp96-entry-interleave probes' "B held at its commit" cases), but their entries re-plan from a fresh session read (U6), so they ADD two classes: a stale entry that plans after the winner finished loses at its plan-time gate with RunStartRejectedError('entry_status_changed'), and one that arrives while the winner's workflow / instance is open is refused before its commit by an owner check (Temporal's executor assertNoLiveOwner: a plain AgentAlreadyRunningError, no conflictCause; DBOS's entry step assertDbosOwnerless and CF Workflows' assertCfwOwnerless: RunStartConflictError('run_active'); CFW's executor awaitPriorInstanceSettled times out with a plain AgentAlreadyRunningError). Documented per runtime in docs/internals/execution-flow.md §Entry gates. Unifying them on C3 needs the executor's seen run status carried into the durable entry (which reverses part of U6: an interrupt-then-resume would then lose), or an equivalent signal that tells the entry's plan-time gate "a newer run happened since the caller looked".

FU-D1-RUN-START-OWNER-GUARD-SPLIT-READ: D1's run-start owner guard decides on an earlier lease / owner read than the one its batch pins — open (LOW, owner: reliability orchestrator) ​

D1StateStore.commitRunStart (store-cloudflare/src/d1-state.ts) resolves the current run in readCommitView, then reads the live runs. applyRunStartRules' owner guard reads view.current.leaseUntil / ownerToken from the FIRST read, while the claim guard's runViewGuardSql pins lease_until / owner_token from the SECOND (CP-96's audit fix pins the current run's status, not its lease or owner). A JS lease renewal that lands between the two reads lets the owner guard pass on the stale (expired) lease while the batch's pin matches the renewed one, so the run-start supersedes a run whose lease was just renewed. The outcome is typed (the renewing run stops with RunSupersededError, superseded), never silent corruption, and the window is two consecutive reads. Fix: in the same claim-guard branch that pins view.current's status, also pin its lease_until IS ? and owner_token IS ? (keeping the bind count constant, e.g. two NULL-safe binds in the no-current-run branch too), and add a d1-guarded-commit race test that renews the lease between the two reads.

FU-REJECT-CAUSES-CORE-HELPER: REJECT_CAUSES / isRejectCause are duplicated in two executors — open (LOW, owner: reliability orchestrator (route)) ​

runtime-temporal and runtime-cloudflare each define the allow-list of RunStartRejectedError causes they revive from a crossed workflow / instance boundary (REJECT_CAUSES, isRejectCause). A new cause must be added in both. Move the list and the guard into core, next to RunStartRejectedError, and import it in both executors (DBOS can use it too).

Released under the MIT License.