Skip to content

Conformance findings catalogue ​

Total: 186. Design: docs/superpowers/specs/2026-09-23-runtime-conformance-suite-design.md

Legend ​

  • Evidence paths are relative to packages/ (e.g. runtime-js/src/run-loop.ts:480 is packages/runtime-js/src/run-loop.ts, line 480). DOAB = runtime-cloudflare/src/do/durable-object-agent-base.ts (the Cloudflare Durable Object agent base; expanded in the evidence column below). Line numbers are as of when the finding was last verified.
  • Evidence prefixes: conformance: names the conformance scenario / assertion ids and backends that witness it; contrast: cites code that does it RIGHT (not a defect site); related: links another finding; live: is a live observation outside the conformance suite — in the consumer app, or a real-workerd / real-provider check run by a peer session (the note says which); inferred: is a step not directly observed; severity explains the chosen severity; consumer is code in the consumer repo.
  • Runtimes / backends: JS = runtime-js; DO = Cloudflare Durable Object runtime (runtime-cloudflare/src/do/, which embeds the JS executor); CF = Cloudflare; CF Workflows / CFW = the Cloudflare Workflows runtime (runtime-cloudflare/src/workflow.ts, executor.ts, steps.ts); D1 = Cloudflare D1 (store-cloudflare); DBOS = runtime-dbos; Temporal = runtime-temporal.
  • Other abbreviations: CAS = compare-and-set (a versioned status transition); HITL = human-in-the-loop (client tools / approvals); SSE = server-sent events; HWM = highWaterMark (stream backpressure threshold); TOCTOU = time-of-check to time-of-use race; EOF = end of stream.
  • IDs: RM = read model & durable state; MA = agent materialization & session context; CP = session control plane; DI = failure-surfacing discipline. Severity: S1 = wrong model behavior, data loss, lost user action, or unbounded spend; S2 = silent failure with a workaround or bounded impact; S3 = latent or observability-only.

Read model & durable state ​

IDSevFindingEvidenceProvenanceVerificationStatus
RM-01S1Every store's getMessages defaults limit to 50; any caller omitting a limit silently gets the oldest 50.store-memory/src/in-memory-state.ts:350-353; store-redis/src/redis-state.ts:1197-1200; store-postgres/src/state/postgres-state-store.ts:828-830; store-cloudflare/src/d1-state.ts:690-692; runtime-cloudflare/src/do/do-state-store.ts:1299-1301; core/src/store/state-store.ts:48-56 (GetMessagesOptions: limit "default: 50") and :241 (the getMessages contract)sourceverifiedopen
RM-02S1DO /snapshot calls getMessages without a limit → oldest 50 raw rows. Measured: 154-row session reloads with 34 of 106 tool calls.runtime-cloudflare/src/do/durable-object-agent-base.ts:3702-3704source, live-reproverifiedopen
RM-03S1DOStateStoreClient.getMessages paginates over the capped /snapshot: total ≤ 50, hasMore false; published chat.getMessages({offset: 50}) returns [].runtime-cloudflare/src/do/do-clients.ts:209-240; ai-sdk/src/cloudflare/chat-handler-factory.ts:162-193sourceverifiedopen
RM-04S1ai-sdk buildSnapshot omits the limit → 50 on every store and runtime.ai-sdk/src/handler/snapshot.ts:169-171sourceverifiedopen
RM-05S2Past the cap, buildSnapshot logic corrupts: isNewTurn compares capped count to startMessageCount; partial merges into an old assistant; UI ids stop at msg-49 while the live stream emits msg-N; content-replay slice wrong.ai-sdk/src/handler/snapshot.ts:239, 338-340, 499sourceverifiedopen
RM-06S2DO appendPartialContentIfActive appends the in-progress assistant after the oldest 50 (wrong position).runtime-cloudflare/src/do/durable-object-agent-base.ts:3892+sourceverifiedopen
RM-07S1One message in the first 50 failing MessageSchema → fetchFromDO returns null → loadState null → "session does not exist". fetchFromDO maps any non-2xx / schema failure / throw to null.runtime-cloudflare/src/do/do-clients.ts:101-155, 157-207sourceverifiedopen
RM-08S2Snapshot is not atomic: buildSnapshot / handleChatStream issue separate /snapshot fetches for state, run, checkpoint, stream info.runtime-cloudflare/src/do/do-clients.ts; ai-sdk/src/handler/snapshot.ts:163-202; runtime-cloudflare/src/do/do-clients.ts:162, 247, 292, 620 (loadState / getLatestCheckpoint / getCurrentRun / getStreamInfo — each a full /snapshot); runtime-cloudflare/src/do/do-clients.ts:212-215 (only getMessages passes excludePartialContent); runtime-cloudflare/src/do/durable-object-agent-base.ts:3891-3945 (appendPartialContentIfActive re-reads chunks; falls back to getAllChunks — the whole multi-turn stream — when the run row is not running, e.g. suspended on a client tool); ai-sdk/src/handler/handle-chat-stream.ts:238, 241, 1974 + runtime-cloudflare/src/do/do-clients.ts:863, 1016 (a DO chat POST costs 4 /snapshots on a fresh turn or tool-result resume, 7 on the auto-resume race branch, more on attach); severity: fix must be ONE atomic, bounded-cost snapshot, not a larger limit (orchestrator B2 contract); ai-sdk/src/handler/snapshot.ts:169 (getMessages) → :192 (getStreamInfo) → :202 (getCurrentRun) → :270, :283 (computeIsRunSettled; a settled run skips the partial-content merge) — three separate reads, so a snapshot taken across run completion can combine messages from BEFORE the final commit with stream / run state from AFTER it; live: shape 1 (active, text missing) — the example-refresh-flake session (S5) instrumented the canonical CF DO getSnapshot() during a first-turn run (VCR replay, local workerd): {status:'active', streamSequence:145, 4 messages ending in a tool message, runStatus:'completed', isRunSettled:true, includePartialContent:true}; a client resuming from that cursor gets no text. 2 of 71 active snapshots with three concurrent pollers; live: shape 2 (ended, text missing) — the same session observed status:'ended', streamSequence:70 with messages lacking the final step (assistant = two tool calls, no text) while a later read showed [tool, tool, text] at the same sequence; a client that resumes only on 'active' (the opennext example ChatClient, the consumer app) never resumes and the final text is missing until another reload. 1 of 40 in isolation; ~8 of 90 comprehensive-refresh runs under six-worker stress (the opennext example refresh flake, comprehensive-refresh.spec.ts:370); related: a planned Phase B witness (not yet written) is a host-tier cf-do cell polling /snapshot across run completion with the invariant: for every snapshot, the last assistant text equals the concatenation of text deltas with id <= streamSequence, whatever the status (both live shapes violate it)source, live-repro, peer-reportedverifiedopen
RM-09S3DO /messages pages correctly (default 100, max 1000) but no SDK client wraps it.runtime-cloudflare/src/do/durable-object-agent-base.ts:3649-3680sourceverifiedopen
RM-10S1Temporal: every LLM call receives only the system prompt — no conversation history and not even the current user message. projectSessionToAgent reads session.messages (absent on SessionState) → []; runLLMStep builds the prompt from it. Also affects retry and child workflows. With a real provider through llm-vercel every call — even a fresh session's first turn — fails outright ("Invalid prompt: messages must not be empty"), because the adapter hoists the system prompt out of messages, leaving them empty.runtime-temporal/src/workflow.ts:84-104 (projectSessionToAgent: messages come from ('messages' in session ? session.messages : undefined) ?? [] at :93-96 — SessionState, core/src/types/session.ts:88, has no messages field, so always []), :901 (the step loop's session = refreshState(...)) and :914 (its projection is the LLM step's state); also :817 (the resume deferred-finishWith projection); runtime-temporal/src/activities.ts:1975-1976 (runLLMStep builds the prompt via buildMessagesForLLM(input.state.messages, ...)), :4724-4730 (refreshState returns stateStore.loadState(sessionId) — a SessionState, never messages); llm-vercel/src/vercel-adapter.ts:158-161, 202 (system messages are hoisted into the AI SDK system: option and removed from messages; with no history the conversation messages are therefore empty); live: B8 real-provider check — on runtime-temporal every real-provider LLM call failed with "Invalid prompt: messages must not be empty" (the AI SDK's empty-messages prompt validation), even a fresh session's first turn; conformance: transcript.* on temporal-* — the turn-2 model input carries none of the tool calls turn 1 persisted; conformance: transcript.finish-with-fails-then-retries on temporal-* — .llm-input-has-turn1-calls: the turn-2 model input carries none of the tool calls turn 1 persisted; conformance: transcript.heal on temporal-* — .llm-input-has-seed-calls-on-{completed,failed,interrupted}: the continuation model input carries none of the seeded tool calls; conformance: history.prior-turns on temporal-memory — the turn-2 model call's input is exactly ["history"] (the system prompt); neither turn's user message nor the turn-1 reply is present; conformance: smoke.client-tool-completes on temporal-* — after the responder answers the client tool, the model's second call carries only the system prompt (no tool result, no history); runtime-temporal/src/workflow.ts:964 (baseMessageCount = agentState.messages.length + assistantAdded) → :1009-1014 → activities.ts:1745 (fireDeferredToolHooks messageIndex = baseMessageCount) and activities.ts:2158-2176 (assistant onMessage, messageIndex = nextState.messages.length): hook messageIndex computed off the empty history — visible in onMessage / afterTool payloadssource, probe, live-repro, peer-reportedverifiedopen
RM-11S1DBOS: the model receives the oldest 10,000 messages.runtime-dbos/src/steps/load-state.ts:99-104; runtime-dbos/src/workflows/shared.ts:1518-1567; runtime-dbos/src/workflows/shared.ts:1833 (hook messageIndex based on the 10k-capped messages.length)sourceverifiedopen
RM-12S2DBOS between-step checkpoint messageCount capped at 10,000 and becomes "latest".runtime-dbos/src/workflows/shared.ts:2907-2911sourceverifiedopen
RM-13S2Corrupt-row skipping adjusts total per page on memory / Postgres / D1 → paging stops early and drops the newest message.store-memory/src/in-memory-state.ts:401-408; store-postgres/src/state/postgres-state-store.ts:876-883; store-cloudflare/src/d1-state.ts:760-768; core/src/store/get-all-messages.ts:30-39sourceverifiedopen
RM-14S2Redis getMessages({limit: 0}) returns every message (LRANGE 0 -1); reachable via limit: checkpoint.messageCount when the count is 0.store-redis/src/redis-state.ts:959-962, 2583; runtime-js/src/js-agent-executor.ts:2108, 2759sourcepartly-inferredopen
RM-15S3D1 cloneSession paging duplicates rows past a corrupt page; infinite loop on ≥ a page of consecutive corrupt rows.store-cloudflare/src/d1-state.ts:2370-2375sourcepartly-inferredopen
RM-16S1Postgres truncateMessages does not decrement states.message_count; later checkpoints are inflated, so truncation misses orphans and retry reloads post-checkpoint messages.store-postgres/src/state/postgres-state-store.ts:918-929, 2186-2190sourcepartly-inferredopen
RM-17S1DO legacy promoteStaging checkpoint counts only valid messages, while truncateMessages deletes by sequence → after one corrupt row, resume / interrupt deletes the newest valid messages.runtime-cloudflare/src/do/do-state-store.ts:2078-2090, 2180sourceverifiedopen
RM-18S3DOStateStoreClient.getLatestCheckpoint.messageCount = capped snapshot length (latent; current consumer ignores it).runtime-cloudflare/src/do/do-clients.ts:282sourceverifiedopen
RM-19S1workspaceRef is persisted only by the memory store; Redis, Postgres, D1, DO drop it silently; persistRef "succeeds".core/src/types/session.ts:306; store-postgres/, store-cloudflare/ and runtime-cloudflare/src/do/do-state-store.ts have no workspaceRef column/field; store-redis/src/redis-state.ts:1035 (cloneSession nulls it, by design: the ref is scoped to the source session)sourceverifiedopen
RM-20S1Local workspaces (local-bash, local-sandbox, docker) are destroyed at the end of every run (closeAll in the run-loop finally); the next resume silently opens a fresh empty workspace.runtime-js/src/run-loop.ts:1190-1211; core/src/workspace/registry.ts:820-832, 1200+ (closeAll → closeEntry); workspace-local-bash/src/workspace.ts:114; workspace-docker/src/workspace.ts:84-86sourceverifiedopen
RM-21S1Lazy workspace open kills the run: buildWorkspaceRegistry closes over the first step's state; persistRef bumps the stored version on that stale object; the loop's live state keeps the old version; the next CAS save throws StaleStateError, reported as executor_superseded ("Another executor has taken over this run").runtime-js/src/run-loop.ts:3738-3810, 908, 3108-3113 (cf. the same bug fixed for sub-agent hooks, :1302-1311)source, live-reproverifiedopen
RM-22S3Redis truncateMessages does not update the hash messageCount (listSessions counts stale).store-redis/src/redis-state.ts:950, 1071-1084sourceverifiedopen
RM-23S2React useAutoResync / useCheckpointSnapshot / useResumableChat replace the thread with the capped snapshot.ai-sdk/src/react/index.ts:247-305, 457-530, 658+sourceverifiedopen
RM-24S3The remote protocol exposes no history at all (snapshot is customState only). Design gap to decide in sub-project 1.core/src/types/remote-protocol.ts:190-207sourceverifieduntested — Design gap (remote protocol exposes no history); decided in sub-project 1.
RM-25S2JS from_checkpoint restores checkpoint messages in memory but truncates the store to the latest checkpoint; LLM context and durable log diverge.runtime-js/src/js-agent-executor.ts:2102-2112, 2248-2266sourcepartly-inferredopen
RM-26S2Memory / Redis / Postgres cloneSession do not check the checkpoint belongs to the source session; D1 silently clones everything when the checkpoint is missing.store-memory/src/in-memory-state.ts:258-263; store-redis/src/redis-state.ts:845-850; store-postgres/src/state/postgres-state-store.ts:573-582; store-cloudflare/src/d1-state.ts:2384-2390sourceverifiedopen
RM-27S3Redis and D1 cloneSession are non-atomic; a half-finished clone is reused by JS initializeSession.store-redis/src/redis-state.ts:899-928; store-cloudflare/src/d1-state.ts:2404-2476; runtime-js/src/js-agent-executor.ts:3334-3341sourcepartly-inferredopen
RM-28S1DBOS retry() and EVERY resume() mode (continue, with_message, with_confirmation, from_checkpoint — all share one workflow-start call) start the workflow with input: undefined; normalizeAgentInput's non-AgentInput fallback turns it into a user message whose content is JSON.stringify(undefined) (undefined), which the workflow body appends unconditionally, the store persists, and paged reads then silently skip as a corrupt row — the session's message history reads short / is corrupted after any retry or resume.runtime-dbos/src/lifecycle/retry.ts:126-128 (retryImpl standardAgentWorkflow call, input: undefined); runtime-dbos/src/lifecycle/resume.ts:512-514 (the single standardAgentWorkflow start shared by every resume mode, input: undefined; with_message only pre-appends its message at :447-450, then falls through to the same start); runtime-dbos/src/workflows/shared.ts:272-291 (normalizeAgentInput fallback: content = JSON.stringify(input), undefined for undefined input), :1099 (called on every entry, resume included), :1167-1172 (messagesToAppend = [...historyPrefix, ...userMessages] appended unconditionally — isResume drops only the history prefix); conformance: lifecycle.retry-history-readable, lifecycle.resume-history-readable (with_message) and lifecycle.continue-history-readable (lifecycle.resume-continue, mode continue) all fail on dbos-postgres with a short read (paged messages < getMessageCount); with_confirmation / from_checkpoint are source-only (same start call)source, probeverifiedopen
RM-29S1A terminal short-circuit leaves sibling tool_use blocks without a tool_result — the persisted transcript is invalid and a continued session's next turn fails at the provider ('tool result missing'), breaking the persisted-history pairing invariant. Four shapes: (a) the 2nd+ successful finishWith call in one step is skipped but stays in the assistant message; (b) any call co-emitted with __finish__ is committed but never executed; (c) a second __finish__ gets no synthetic result (heals use .find); (d) a finishWith call co-emitted with a client tool or a suspending sub-agent is deferred by the phase-1 suspension and never executed on resume (JS / DO / Temporal / CF Workflows; DBOS runs it) — its tool_use is orphaned mid-history and the run goes on to an extra model call. On Temporal even a single finishWith completion persists no tool result. Fixed: every call a terminal step does not execute now persists a typed not-executed tool_result (a/b/c), Temporal persists its finishWith results, a deferred finishWith now runs on resume before any further model call (d), and a new turn on a completed/failed/interrupted session pairs trailing unpaired calls before appending its input.core/src/orchestration/step-iterator.ts:1837-1845 (pre-fix, origin/main f5c8a14bdf) (phase-2 finishWith break; skipped calls not appended); core/src/orchestration/step-processor.ts:301-367 (pre-fix, origin/main f5c8a14bdf) (__finish__ plan keeps every call, returns no pending calls); core/src/orchestration/step-iterator.ts:978-988 (pre-fix, origin/main f5c8a14bdf) (.find heal; comment concedes shape (b)); core/src/orchestration/message-builder.ts:251-264 (pre-fix, origin/main f5c8a14bdf) (legacy findUnpairedFinishCallId heal — __finish__ only); runtime-js/src/js-agent-executor.ts:3387-3392 (pre-fix, origin/main f5c8a14bdf) (__finish__-only continuation repair for the implicit-continuation path), run-loop.ts:4927 (a second, separate __finish__-only continuation repair for the persistent-companion re-consult path); runtime-dbos/src/workflows/shared.ts:2454-2509 (pre-fix, origin/main f5c8a14bdf) (loop break), 1870-1876 (heal); runtime-cloudflare/src/workflow.ts:2190-2219 (pre-fix, origin/main f5c8a14bdf) (finishWith break, same defect shape as core), 630-658 (continuation-time legacy heal via findUnpairedFinishCallId, mirrors runtime-js/runtime-dbos/runtime-temporal); runtime-cloudflare/src/steps.ts:1433-1437 (pre-fix, origin/main f5c8a14bdf) (per-step heal for a same-step dangling __finish__ — CF has both this and the workflow.ts:630-658 continuation heal, but like every runtime neither covers shape (a) or (b)); runtime-temporal/src/activities.ts:2398-2448 (pre-fix, origin/main f5c8a14bdf) (runPhase2FinishWith appends no tool message at all), 2140-2154 (.find heal; its comment finishWith tools append their own result is false for this runtime); core/src/orchestration/step-iterator.ts:1688 (pre-fix, origin/main f5c8a14bdf) (shape (d): suspension defers phase 2 "until resume"), runtime-js/src/run-loop.ts:273, 628 (resume re-enters with isResume and goes straight to a new LLM step — nothing re-derives the deferred finishWith; DO routes through the same JSAgentExecutor); runtime-temporal/src/workflow.ts:1057-1136 (pre-fix, origin/main f5c8a14bdf) (shape (d): suspension branch returns without reading pendingFinishWithCalls; phase 2 exists only on the no-suspension branch, 1157-1160), activities.ts:241-245 (doc comment "deferred-finishWith resumes on the next iteration" is false), activities.ts:3196 (applyResultsAndReload has no finishWith handling); runtime-cloudflare/src/workflow.ts:742-775 (pre-fix, origin/main f5c8a14bdf) (shape (d): resume applies results then re-enters the loop), 1846 (finishWithTools derived only from the CURRENT step pendingToolCalls, so the suspended step deferred call is never dispatched); contrast: (pre-fix, origin/main f5c8a14bdf) runtime-dbos/src/workflows/shared.ts:390, 474 (dispatchClientTool awaits DBOS.recv inside phase 1), ~2450 (phase 2 then runs in the same iteration — shape (d) does not occur on DBOS); live: B8 real-LLM check — turn 1 was SCRIPTED (the incident shape, incl. shape (b): a call co-emitted with __finish__, produced by the scripted model, not by llm-vercel) and only turn 2 went to a real model (Anthropic claude-haiku-4-5). Pre-fix (origin/main 9c1aed51be) turn 2 failed: through the AI SDK with "Tool result is missing for tool call toolu_…", and against the raw Messages API with a 400 "tool_use ids were found without tool_result blocks immediately after"; post-fix (!257, 0fee0b1409) turn 2 completes on runtime-js and runtime-dbos for shapes (a) and (b). This is consistent with RM-35: through llm-vercel a call co-emitted with __finish__ never reaches the runtime, so shape (b) was only reachable via the scripted turn 1; live: incident 2026-09-24 audio-resume — DO SQLite messages row 42 carries two accept_audio calls, row 43 pairs only the first; turn-3 run 9e1e7e16 failed with Tool result is missing for tool call toolu_012Yx…; not yet reachable (conformance cell) — shape (d) fixed: a deferred finishWith now runs on resume before any further model call; proven by the runtime deferred-finishWith tests — core/src/orchestration/__tests__/transcript.test.ts (findDeferredFinishWithCalls), runtime-js/src/__tests__/deferred-finishwith-after-resume.test.ts (incl. the persistent-completion ordering case), runtime-dbos/src/__tests__/integration/deferred-finishwith-after-resume.integ.test.ts, runtime-temporal/src/__tests__/integration/multi-turn-continuation.integ.test.ts + runtime-temporal/src/__tests__/activities-finishWith.test.ts, runtime-cloudflare/src/__tests__/integration/do-multiturn-client-tool-resume.integ.test.ts (DO) and runtime-cloudflare/src/__tests__/integration/client-tool-workflow.integ.test.ts (CFW); the conformance cell is pending the Phase B executor submit-tool-result capability; conformance: transcript.finish-cosiblings .cosibling-not-executed is the direct proof of C2 shape (b) "a call co-emitted with __finish__ is not executed": the co-emitted note tool records an in-process execution witness (its sessionId + toolCallId) whenever its body runs, independent of any state commit, and the turn-1 cosibling is absent from it after both turns; .cosibling-witness-observable is its positive control (the same tool executed in turn 2 is in the witness). Separately, .cosibling-state-not-written (control .cosibling-marker-observable) pins that the cosibling left no customState write. The witness only observes tools run in the vitest process (js-*, dbos-postgres, temporal-* with the in-process worker); an out-of-process backend (cf-do) must surface it as a precondition failure or declared gap, never a pass. Scope: model adapters that surface co-emitted calls (the scripted/mock adapters used by the suite, and non-Vercel adapters); through llm-vercel the co-emitted call is never produced at all (vercel-adapter returns only structured_output — catalogued as RM-35), so this note makes no claim about the Vercel path. These ids passed pre-B8 as well (the cosibling never ran before either; B8 only added its synthetic result), so they pin the guarantee rather than prove the B8 fix — provenBy is unchanged (transcript.finish-cosiblings cells are already listed for .persisted-paired / .skipped-marked / .llm-input-paired); related: CF DO and CF Workflows are proven by runtime-cloudflare do-rm29-transcript.integ.test.ts (real workerd DO x DO-SQLite) and rm29.wf-noiso.test.ts (CFW x D1), pending the Phase B cf-do / cfw conformance cells; RM-29 re-opens if any of those cells fails; conformance: transcript.finish-with-siblings / transcript.finish-cosiblings / transcript.double-finish on js-memory / js-redis / js-postgres / dbos-postgres — proves the fix (shapes a/b/c): pre-fix the persisted turn-1 transcript left the 2nd+ call unpaired (.persisted-paired), no non-winning call carried a superseded_by_terminal_call not-executed result (.skipped-marked), and the turn-2 model input stayed unpaired (.llm-input-paired); post-fix all pass, and .skipped-marked also pins that the winner is not marked; conformance: transcript.heal on js-memory / js-redis / js-postgres — proves the C5 continuation heal: pre-fix a new turn on a completed / failed / interrupted session left the seeded trailing call unpaired (.heal-on-*, .heal-marked-on-*, .llm-input-paired-on-*); post-fix it is paired with an unpaired_on_continuation not-executed result; conformance: .output-is-first-winner (transcript.finish-with-siblings / transcript.finish-cosiblings / transcript.double-finish / transcript.finish-with-fails-then-retries) is pinned, not proof — it passed pre-fix; conformance: not claimed in provenBy — every transcript.* cell on temporal-* (its RM-29 ids, incl. transcript.finish-with-single / transcript.finish-with-fails-then-retries .persisted-paired, .failed-call-keeps-real-error and .failed-call-marked-failed, now pass, but the cell still fails .llm-input-has-* under RM-10); transcript.heal on dbos-postgres (RM-29 ids pass; the interrupted continuation fails under CP-19); transcript.finish-with-single and transcript.finish-with-fails-then-retries on JS / DBOS (their RM-29 ids passed pre-fix; the fails-then-retries .failed-call-marked-failed fix on JS is DI-17's); related: CLAUDE.md "__finish__ history invariant"source, live-repro, peer-reportedverifiedfixed (15 cells)
RM-30S3D1 truncateMessages deletes message rows but never updates states.message_count, so listSessions().messageCount stays inflated after any truncation (the D1 twin of RM-22; checkpoints and getMessageCount use live counts and are unaffected).store-cloudflare/src/d1-state.ts:838-862 (truncateMessages issues the DELETE but never updates states.message_count), :3369, :3446-3457 (listSessions reads the stale denormalized message_count column); contrast: store-cloudflare/src/d1-state.ts:778 (getMessageCount is a live COUNT(*)) and :1392 (the promoteStaging path is also a live COUNT(*)); related: RM-22source, peer-reportedverifiedopen
RM-31S1CF Workflows resume({ mode: 'from_checkpoint', checkpointId }) restores the checkpoint's customState and stepCount but never truncates messages (it truncates to the LATEST checkpoint, then saves the checkpoint state through saveAgentState, which never touches messages), so the model still sees every post-checkpoint message; it also never checks the checkpoint belongs to the session (restoredState.sessionId comes from the checkpoint → a foreign checkpoint can write another session's row) and run_resumed.fromCheckpointId reports the old latest id.runtime-cloudflare/src/executor.ts:695 (truncates to the latest checkpoint, not necessarily the one being resumed to), :714-746 (from_checkpoint restore saves checkpoint.state via saveAgentState but never truncates messages, and spreads restoredState.sessionId from the checkpoint unchecked), :185-240 (saveAgentState never touches messages), :833 (run_resumed.fromCheckpointId reports the stale pre-restore agentState.checkpointId); related: RM-25 (the JS sibling); related: CP-38 (foreign checkpoint)source, peer-reportedverifiedopen
RM-32S3A server tool that returns a plain string is persisted (and shown to the model on later steps) verbatim on JS / CF DO but JSON-quoted on Temporal / DBOS / CF Workflows (hello vs "hello"), so the model's view of tool output and any consumer comparing tool content diverge by runtime. Which form is correct is a contract decision (a conformance cell can assert only one).runtime-js/src/run-loop.ts:2356-2373 (the regular server-tool success path builds the result via createToolResultMessage, then — :2371-2373 — overwrites content with the raw string when typeof toolResult === 'string', so a string tool result is persisted verbatim and everything else JSON.stringified; the comment at :2363-2370 states this preserves pre-DI-17 behavior); core/src/orchestration/message-builder.ts:169 (createToolResultMessage always JSON.stringify's input.result regardless of its type — no string special-case); runtime-dbos/src/workflows/shared.ts:2331 (phase 1 regular-tool result message built via createToolResultMessage) and :2467 (phase 2 finishWith-tool result message, also via createToolResultMessage) — both inherit core's unconditional JSON.stringify, so a string result arrives JSON-quoted; runtime-temporal/src/workflow.ts:998-1000 (the regular per-tool server-tool path — results of executeToolActivity — builds each message via buildServerToolResultMessage, runtime-temporal/src/tool-result-messages.ts:25-43, i.e. core createToolResultMessage: a string result is JSON-quoted, no string special-case); activities.ts:849-857 (executeServerToolWithHooks, now reached from the approval-approve drain at :3822, likewise via createToolResultMessage); runtime-cloudflare/src/steps.ts:2258 (regular server-tool success path builds its result message via createToolResultMessage, same as DBOS — a string result is JSON-quoted); runtime-cloudflare/src/do/durable-object-agent-base.ts:79 (CF DO imports and runs JSAgentExecutor from @helix-agents/runtime-js verbatim, so it inherits the run-loop.ts:2371-2373 string special-case — CF DO persists a string result verbatim like plain JS); inferred: no scenario yet; an executor-tier cell (a tool returning a string, asserting the turn-2 model input content) witnesses it once the contract picks a formsource, peer-reportedverifiedopen
RM-33S2DBOS writes no checkpoint for a turn whose final step makes no tool calls: a single text-only turn leaves no checkpoint of its own (the inter-step checkpoint is only created on the continue path and the terminal return writes none), so getLatestCheckpoint can keep pointing at an earlier turn's checkpoint (or none), and branch / retry / resume from 'the latest checkpoint' can act on stale state.runtime-dbos/src/workflows/shared.ts:2900 (the terminal path returns { completionReason } after endStream — no createCheckpoint) and :2903-2912 (the inter-step dispatcher.createCheckpoint runs only on the continue path); the other checkpoint writers are the forced-completion continue tails (:2665, :2727) and promoteStaging (:2311-2318), which runs only when the step dispatched tool calls; conformance: smoke.branch-completes on dbos-postgres — .source-checkpointed fails: turn 1 (one text-only step) left no checkpoint (latestCheckpointId undefined); .completes and .history-has-turn1 fail only on that precondition (failPreconditions — never branched); inferred: a structured-output (__finish__) completion is NOT claimed here — that step is a tool call and was not verified to take the no-checkpoint terminal path; inferred: retry/from_checkpoint truncating to a stale earlier-turn checkpoint could drop a completed turn's messages — not observed; contrast: runtime-js/src/run-loop.ts:1616-1640 (commitStep → persistViaSaveAndPromote → stateStore.saveStateAndPromoteStaging with the step's checkpointMeta, run for EVERY committed step including the terminal one — :2937-3047) — smoke.branch-completes passes on js-* (and temporal-*); related: RM-12, CP-30, CP-32source, probepartly-inferredopen
RM-34S2Temporal mishandles a phase-2 finishWith tool's customState writes whenever phase 2 produces no output: when the tool RETURNS undefined, Temporal streams its state_patch but never persists the writes, while JS commits them; when the tool THROWS, Temporal streams state_patch chunks for its partial pre-throw writes and its afterTool-on-error writes that it never persists, while JS neither stages nor streams a throwing tool's writes — either way the client sees state the Temporal store never holds.runtime-temporal/src/workflow.ts:1250-1300 (phase-2 finishWith dispatch, no-suspension branch): runPhase2FinishWith (:1252) returns { output, nextState, resultMessages } (activities.ts:2853); since !257 the workflow appends phase2.resultMessages (:1267-1271) but reads nextState.customState and calls persistTerminalState with it ONLY inside the guard at :1273 ("if (phase2.output !== undefined)", annotated "RM-34 ... (catalogued)" at :1272); when output stays undefined the block is skipped and nextState — carrying the finishWith tool's customState writes — is discarded (the loop re-reads state via refreshState at :901) with nothing else persisting it; runtime-temporal/src/activities.ts runPhase2FinishWith (:2447-2857), returned-undefined case: the success path streams the tool's writes as state_patch (:2653-2663) and any afterTool writes (:2698-2708), then leaves output undefined (only set when transformed !== undefined, :2743-2745) and returns { output, nextState, resultMessages } at :2853 — streamed, then dropped by the workflow; runtime-temporal/src/activities.ts runPhase2FinishWith, throw case (catch at :2746): streams state_patch for the partial pre-throw writes (:2769-2779) and for the afterTool-on-error hook's writes (:2813-2823), folds both into nextState.customState (:2783, and after the hook), and falls through to return { output: undefined, nextState, resultMessages } at :2853 — streamed, then dropped by the workflow; contrast: returned undefined — runtime-js runs a phase-2 finishWith through executeServerTool → runServerTool (runtime-js/src/run-loop.ts:1338, :1466, :1935), which on success stages the writes (stageChanges, :2223) and streams them only after staging (:2245-2275); core/src/orchestration/step-iterator.ts:1837 (executePhase2FinishWithCalls, whose loop calls hooks.executeServerTool at :2217) is followed by an UNCONDITIONAL hooks.commitStep (:1947) regardless of whether finishWithOutput (:1872) is defined — so JS persists what it streamed; contrast: throw — runtime-js/src/run-loop.ts:2185 records the throw as toolError and :2201 (if (!toolError && tracker.hasChanges())) skips staging, the in-memory merge AND the state_patch emission — a throwing tool's writes are neither persisted nor streamed on JS (consistent, if lossy); inferred: no scenario yet — observable at the executor tier via PersistedObservation.customState (persistence) but the streamed half needs a stream observation (host/client tier)source, peer-reportedverifiedopen
RM-35S1llm-vercel silently drops every tool call (and the accompanying text) the model co-emits with __finish__: any step whose tool calls include __finish__ is returned as structured_output carrying only the __finish__ input, and core's structured_output branch plans no assistant message and no pending calls — so the co-emitted calls are never executed, never persisted, and never answered, while the agent completes as if they had run. A side-effecting call the model believes it made (and may describe in its structured output) is lost with no error. Through llm-vercel RM-29 shape (b) therefore cannot occur: the co-emitted calls never reach the step processor at all.llm-vercel/src/vercel-adapter.ts:440-456 (toolCalls.find(__finish__) → return { type: 'structured_output', output: finishCall.input, thinking, shouldStop: true, ... } — the other toolCalls and text are discarded; contrast the regular tool_calls return at :470-481, which keeps every call and content: text); core/src/types/runtime.ts:147-158 (StructuredOutputStepResult has no field for text or other tool calls — the drop is structural); core/src/orchestration/step-processor.ts:375-408 (structured_output branch: assistantMessagePlan: null, pendingToolCalls: [], pendingSubAgentCalls: [], status completed — nothing is executed or recorded, not even the __finish__ call itself); llm-vercel/src/vercel-adapter.ts:236-257 (streaming callbacks still forward text deltas and tool-arg-stream chunks live, so a client may have SEEN the dropped text / call arguments stream by; only the persisted transcript and the executed work lose them); live: B8 (peer, real-provider check) did NOT observe a co-emission — the drop itself is inferred from source. What its fixed shape-(b) runs on runtime-js and runtime-dbos did show: turn 2 went to a real model (claude-haiku-4-5) through llm-vercel and no assistant message was persisted for that turn, consistent with the structured_output conversion above; the model's raw tool calls were not captured; related: RM-29 (shape (b) — a call co-emitted with __finish__ — is fixed in core for adapters that surface co-emitted calls; the transcript.finish-cosiblings scenario scripts the co-emission through ScriptedModel, a MockLLMAdapter that replaces llm-vercel — core/src/llm/scripted-model.ts:326 — so it does not cover this path); severity S1 from the verified source mechanism alone (vercel-adapter.ts:440-456 + step-processor.ts:375-408): whenever the model issues other tool calls in a step that also calls __finish__, those calls are discarded with no error, no persisted trace and no tool result, and the run still completes — wrong model behavior per the legend, independent of what the dropped tool does; if that tool has side effects the user action it represents is also silently lost. No live co-emission has been observedsource, peer-reportedverifiedopen
RM-36S1The Durable Object store's saveState writes its checkpoint row without message_count, so the column takes its default of 0: a DO turn whose latest checkpoint comes from saveState (rather than saveStateAndPromoteStaging) leaves getLatestCheckpoint() claiming 0 messages while the session holds more. Retry and from-checkpoint paths that trust checkpoint.messageCount then act on a zero-length history.runtime-cloudflare/src/do/do-state-store.ts:2506-2517 (saveState checkpoint INSERT lists id, session_id, step_count, timestamp, status, state — no message_count, no stream_sequence); runtime-cloudflare/src/do/do-state-store.ts:698 (checkpoints.message_count INTEGER NOT NULL DEFAULT 0); contrast: runtime-cloudflare/src/do/do-state-store.ts:2248-2250 (the promoteStaging checkpoint INSERT records message_count and stream_sequence); the memory / Redis / Postgres stores record the live count; inferred: runtime-js/src/js-agent-executor.ts:2790 (retry sets truncateTo = checkpoint.messageCount with no zero guard) and :2108-2113 (checkpoint restore loads getMessages({ limit: checkpoint.messageCount })) would truncate / restore to zero messages on a DO session whose latest checkpoint came from saveState — a conversation wipe, not observed. The orphan-cleanup truncations (js-agent-executor.ts:2260, :3964; run-loop.ts:3917) guard messageCount > 0 and are unaffected; live: B2-core (RM-33 work), a real workerd DO integration case, read the latest checkpoint as 0 messages against 2 live messages before its fix; related: RM-33 (the DBOS missing-checkpoint defect found in the same work — a separate cause); related: RM-37 (every store's saveState checkpoint records streamSequence 0, the DO by omitting stream_sequence in this same INSERT — a separate cause sharing this fix site)source, peer-reportedpartly-inferredopen
RM-37S3saveState mints its checkpoint with streamSequence 0 on every store that checkpoints there (memory, Redis and Postgres write a literal 0; the Durable Object store omits the column, which defaults to 0). saveState takes no stream position, so it cannot record one. That checkpoint becomes the LATEST one whenever a saveState follows the promote-once sentinel: a turn-start save, a client-tool submit save, or a terminal save. getLatestCheckpoint().streamSequence then reads 0 while the stream has advanced, although 0 is itself a valid replay origin ('replay everything').store-memory/src/in-memory-state.ts:481-489 (the 'staging' sentinel skips only the FIRST saveState after a promote), :491-516 (every later saveState writes a checkpoint with streamSequence: 0 at :509), :806-819 (getLatestCheckpoint returns the last-written checkpoint); store-redis/src/redis-state.ts:1297 (same sentinel), :1422-1431 (saveState checkpoint streamSequence: 0 at :1429); store-postgres/src/state/postgres-state-store.ts:942 (same sentinel), :1093-1109 (saveState checkpoint INSERT passes 0 for stream_sequence at :1106); runtime-cloudflare/src/do/do-state-store.ts:2384 (same sentinel), :2499-2511 (saveState checkpoint INSERT omits stream_sequence), :699 (stream_sequence INTEGER NOT NULL DEFAULT 0); core/src/store/state-store.ts:289-292 (SessionStateStore.saveState takes (sessionId, state) only) vs :578-588 (saveStateAndPromoteStaging takes checkpointMeta.streamSequence); live: probe against store-memory, store-redis and store-postgres dist (this session): after saveStateAndPromoteStaging with streamSequence 42, the first saveState leaves the latest checkpoint at 42 and the second makes it 0 on all three stores; live: probe, JS executor + memory store (this session): during a continued turn's first step the latest checkpoint was the turn-start saveState checkpoint, with streamSequence 0; inferred: agent-server/src/agent-server.ts:779-793 treats checkpoint.streamSequence <= 0 as a fallback and degrades the remote snapshot to snapshot-only (-1) with a warn, so a checkpoint-pinned late-join is unavailable while such a checkpoint is latest (for example a paused session after a client-tool submit). This is loud, not silent; inferred: core/src/recovery/conversation-recovery.ts:250-285 documents resuming a reader from result.checkpoint.streamSequence (:277) after loading messages up to checkpoint.messageCount. A caller following it replays the whole stream from 0 over messages that already hold that content. There is no in-repo caller; inferred: runtime-js/src/js-agent-executor.ts:2445, :2859, runtime-temporal/src/executor.ts:1717 and runtime-cloudflare/src/executor.ts:1040 put checkpoint.streamSequence into stream_resync.fromSequence. No shipped client reads that field (ai-sdk forwards it as data-stream-resync, stream-transformer.ts:1563-1576); contrast: store-cloudflare/src/d1-state.ts:2490 (D1 saveState writes no checkpoint at all); related: RM-36 (the DO saveState checkpoint also omits message_count — a separate cause specific to the DO, but the SAME fix site: runtime-cloudflare/src/do/do-state-store.ts:2499-2511 omits both columns, so a fix there should write both); RM-39 (DBOS records 0 on every checkpoint, a runtime-side cause); DI-21 (the stream-info fallback that can also write 0); severity S3: the only in-repo reader that acts on the value (agent-server) detects 0 and degrades loudly; the silent consequence needs a caller of the public recoverFromCheckpoint pattern, which nothing in the repo issource, probe, peer-reportedpartly-inferredopen
RM-38S1Redis saveState points the session's checkpointId at a checkpoint that does not exist. It generates id A and hands the write to createSessionCheckpoint, which generates its own id B, stores the checkpoint under B, and sets the pointer to B. saveState then overwrites the pointer with A. getLatestCheckpoint finds no checkpoint A and falls back to the index ordered by stepCount, which is not monotonic across turns. It returns an earlier turn's checkpoint, and interrupt / resume / retry truncate to that stale messageCount, deleting whole committed turns. This defeats the pointer fix the CLAUDE.md State Store note describes, for every checkpoint that saveState writes on Redis.store-redis/src/redis-state.ts:1423 (saveState: checkpointId = generateCheckpointId(...), id A), :1426-1431 (createSessionCheckpoint is not passed that id), :1434-1438 (HSET the session pointer to A); store-redis/src/redis-state.ts:2463-2516 (createSessionCheckpoint: :2478 generates its own id B, stores meta and state under B, :2503 ZADDs B scored by stepCount, :2507-2511 sets the pointer to B, which saveState then overwrites); core/src/types/checkpoint.ts:161-166 (generateCheckpointId includes Date.now() and a random suffix, so A !== B); store-redis/src/redis-state.ts:2521-2548 (getLatestCheckpoint: pointer lookup, getCheckpoint(A) is null, so it falls back to ZREVRANGE by stepCount at :2541); live: probe against store-redis dist (this session): after a saveState, getCheckpoint(state.checkpointId) is null. After a new-turn saveState with stepCount 0 and 4 messages, getLatestCheckpoint returned the older stepCount-3 checkpoint (messageCount 3). Memory and Postgres returned the newest checkpoint (messageCount 4); live: probe, JS executor + Redis store (this session): turn 1 (tool step + text; stepCount 2), turn 2 (text), then turn 3 interrupted in its first step. The latest checkpoint read {stepCount 2, messageCount 4}, and the interrupt cleanup truncated the transcript to 4 messages: turn 2's user message and answer were deleted along with turn 3's input. On the memory store the same run kept turn 2 (only turn 3's input was lost, RM-42); runtime-js/src/run-loop.ts:3889-3925 (performOrphanCleanupOnInterrupt truncates to getLatestCheckpoint().messageCount); runtime-js/src/js-agent-executor.ts:2251-2269 (resume) and :2714, :2825-2826 (retry) truncate to the same checkpoint; contrast: store-redis/src/redis-state.ts:3185, :3543 (the Lua promote paths write the pointer to the checkpoint they create); the CLAUDE.md State Store note and core/src/testing/checkpoint-operations.ts:74-107 (the latest-written contract, exercised only through createCheckpoint, never through saveState); related: RM-42 (the turn-start checkpoint excludes the turn input, a separate cause); RM-25 / RM-31 / CP-58 (other truncation-from-a-stale-anchor paths); severity S1: committed turns are silently deleted, observed live through the JS executor on Redissource, probe, peer-reportedverifiedopen
RM-39S3DBOS records streamSequence 0 on EVERY checkpoint it writes. The step promote passes no stream position, so the store defaults it to 0. The inter-step and forced-completion checkpoints hard-code 0, under a comment claiming promoteStaging populates it. DBOS never calls resolveCheckpointStreamSequence, which JS, Temporal and CF Workflows use at the same boundaries. So no DBOS checkpoint ever carries the stream position it was taken at.runtime-dbos/src/workflows/shared.ts:2318 (dispatcher.promoteStaging({ sessionId, stepId }) with no options); runtime-dbos/src/steps/staging.ts:39-44 (runPromoteStaging forwards args.options, here undefined); store-postgres/src/state/postgres-state-store.ts:1891 (promoteStaging streamSequence = options?.streamSequence ?? 0); runtime-dbos/src/workflows/shared.ts:2673-2679, :2735-2741 (forced-completion continue checkpoints, streamSequence: 0) and :2916-2923 (inter-step checkpoint, streamSequence: 0, commented 'populated by promoteStaging; pass 0' — promoteStaging is given nothing); contrast: runtime-js/src/run-loop.ts:1647 (commitStep resolves the live sequence), runtime-temporal/src/activities.ts:3193, :4388, runtime-cloudflare/src/steps.ts:4195, :4777, :4838 (call core resolveCheckpointStreamSequence); the CF Workflows per-step commitStep instead reads getStreamInfo inline at runtime-cloudflare/src/steps.ts:4086-4092 (unguarded) and passes it to promoteStaging; inferred: agent-server/src/agent-server.ts:779-793 degrades every non-terminal DBOS session's remote snapshot to snapshot-only (-1) with a warn, and ai-sdk/src/state/remote-patch-filter.ts:309-325, :379-392 then suppresses that document's patches until a re-snapshot, which is unpinned again. A remote consumer of a DBOS-hosted agent therefore never gets the pinned live join. Not observed; inferred: the same 0 reaches core recoverFromCheckpoint (core/src/recovery/conversation-recovery.ts:374) as for RM-37; related: RM-37 (every store's saveState checkpoint records 0, a store-side cause); RM-33 (DBOS terminal path writes no checkpoint); severity S3: the value is wrong on every DBOS checkpoint, but the only in-repo reader detects 0 and degrades loudly (warn log, re-snapshot signal); no content is lostsource, peer-reportedpartly-inferredopen
RM-40S2DBOS takes a tool step's checkpoint BEFORE it appends that step's tool results. The promote (commit plus checkpoint) runs only when the step dispatched phase-1 calls: regular tools, sub-agents or remote sub-agents. When it does, the phase-1 results are appended after it, and so is the result of any finishWith or companion call in the same step. So the checkpoint's messageCount ends on an assistant message whose tool_use blocks have no results. A continuing step is followed by an inter-step checkpoint that counts them. A TERMINAL step that ran a phase-1 call is not: for example a regular tool co-emitted with the winning finishWith, or a tool step that hits maxSteps. So this short checkpoint stays the LATEST (RM-33), and anything that slices history at it gets a transcript that breaks the RM-29 pairing invariant: branch-from-checkpoint, or recoverFromCheckpoint.runtime-dbos/src/workflows/shared.ts:1834 (the assistant message is appended first), :1897-1908 (phase-1 partition: finishWith, companion, sub-agent and remote calls are excluded from phase1NonSubagentCalls), :2311-2319 (promoteStaging only when phase-1 regular / sub-agent / remote calls exist; the store counts messages at promote time, store-postgres/src/state/postgres-state-store.ts:1893-1897); runtime-dbos/src/workflows/shared.ts:2321-2365 (phase1ToolMessages built and appended AFTER the promote), :2475 (phase-2 finishWith result appended), :2514 (superseded synthetic results), :2558 (companion result); runtime-dbos/src/workflows/shared.ts:2688, :2781-2782 (maxSteps stop), :2906 (terminal return, no checkpoint) vs :2913-2923 (only the continue path re-counts messages into a new checkpoint); runtime-dbos/src/lifecycle/execute.ts:234-238 (branch clones the source session fromCheckpointId); store-postgres/src/state/postgres-state-store.ts:637-645 and store-memory/src/in-memory-state.ts:258-265 (cloneSession copies messages only up to checkpoint.messageCount); inferred: branching a DBOS session from its latest checkpoint after a terminal step that ran a regular tool (tool + finishWith in one step, or a tool step at maxSteps) yields a branch whose last assistant message has no tool results. The branch's next turn either heals them as not-executed (wrong: they ran) or reaches the provider unpaired. Not observed; inferred: a step whose only call is a finishWith or companion call never promotes, so it writes no checkpoint at all — that is RM-33, not this finding; contrast: runtime-js/src/run-loop.ts:2962-3067 (persistViaSaveAndPromote appends the step's messages in the same saveStateAndPromoteStaging call that writes its checkpoint); related: RM-33 (DBOS writes no terminal checkpoint, which is what leaves this one latest); RM-29 (pairing invariant); RM-39 (the same promote records streamSequence 0); severity S2: the checkpoint is wrong on every DBOS phase-1 tool step, but a silent wrong transcript needs a terminal such step and a caller that slices at it (branch from that checkpoint, or recoverFromCheckpoint)source, peer-reportedpartly-inferredopen
RM-41S1CF Workflows checkpoints only when a step staged customState writes. A step's input, assistant and tool-result messages are appended to the store directly, outside any checkpoint. commitStep returns the previous checkpoint when nothing was staged, and the completion path never commits at all. So after any step whose tools write no state, including text-only steps, the latest checkpoint's messageCount lags the committed transcript. When an earlier step did stage state (so a checkpoint with messageCount > 0 exists), the interrupt handler and resume() treat everything past that count as orphans and truncate it. That silently deletes committed assistant messages and tool results, and the turn's own user input. A session that never staged state has no such checkpoint and loses nothing.runtime-cloudflare/src/steps.ts:4074-4084 (commitStep: when hasStagedChanges is false it returns the latest existing checkpoint id and writes none), :4096-4098 (a checkpoint is written only through promoteStaging of staged changes); runtime-cloudflare/src/steps.ts:2193-2222 (a tool stages only when stateTracker.hasChanges()), :2267 (its tool result is appended directly), :1408 (the assistant message is appended directly), :786-792 (initialize: saveAgentState, then the input messages appended); runtime-cloudflare/src/workflow.ts:661, :668, :737 (continuation heal and new-turn input appended directly); :1833-1890 (the isComplete path ends the stream and finalizes the run with no commit-step); the commitStep calls at :1666 (deferred finishWith), :2360-2362 (per step), :2418 and :2440 (finishWith) all no-op when nothing was staged; store-cloudflare/src/d1-state.ts:2490 (D1 saveState writes no checkpoint, so no turn-start checkpoint covers the input either); runtime-cloudflare/src/steps.ts:3962-3985 (handleInterrupt truncates to getLatestCheckpoint().messageCount before marking interrupted, guarded by messageCount > 0 at :3974); runtime-cloudflare/src/executor.ts:679-696 (resume truncates to the same checkpoint, same guard at :694); inferred: a CF Workflows session whose last checkpoint predates one or more stateless steps (for example turn 1 wrote state, turn 2's steps did not), then interrupted or resumed, loses every message committed since that checkpoint. Not observed; CF Workflows has no conformance backend yet; contrast: runtime-js/src/run-loop.ts:1636-1660 (JS commitStep checkpoints every committed step via saveStateAndPromoteStaging with the step's messages, state-writing or not); related: RM-33 (the DBOS twin of the missing terminal checkpoint); RM-42 (JS: the turn input sits outside the turn-start checkpoint); RM-31 (CF Workflows from_checkpoint truncates to the latest checkpoint); severity S1: committed transcript content is silently deleted on an ordinary interrupt / resume (inferred, not observed); severity S1 while partly-inferred: severity rates the impact if the mechanism holds (committed messages deleted is data loss, S1 on the severity scale); verification records confidence separately: each link (the commitStep no-op, the direct appends, the truncate-to-checkpoint in handleInterrupt / resume) is cited from source above, only the end-to-end loss is unobservedsource, peer-reportedpartly-inferredopen
RM-42S1runtime-js appends a continued turn's new user message AFTER the checkpoint that marks the turn's start. So until the turn's first step commits, the latest checkpoint counts only the previous turns. Stopping the run during that first step (a mid-LLM stop, or a durable interrupt flag seen at the top of the loop) runs the orphan cleanup, which truncates to that count and silently deletes the user's own message. The session reads 'interrupted' with the prompt gone, and a later resume continues without it. resume({ mode: 'with_message' }) has the same shape: it truncates, then appends its message outside any checkpoint.runtime-js/src/js-agent-executor.ts:3494-3497 (Step 8: saveAgentState — its saveState writes the turn-start checkpoint, or leaves the previous one latest when the promote sentinel is set), :3499-3526 (Step 9: the new messages are appended AFTER it); runtime-js/src/js-agent-executor.ts:2251-2269 (resume truncates to the latest checkpoint) then :2358-2362 (with_message appends its message; no checkpoint follows until the first step commits); runtime-js/src/run-loop.ts:1013-1030 (soft interrupt after a failed/aborted step → performOrphanCleanupOnInterrupt), :4209-4216 (finalizeDurableInterrupt, the same cleanup), :3889-3925 (truncates to getLatestCheckpoint().messageCount when > 0); live: probe, JS executor + memory store (this session): turn 1 'FIRST user message' completed; turn 2 'SECOND user message' interrupted while its first LLM call was in flight. The latest checkpoint during turn 2 had messageCount 2 (3 messages stored); after the interrupt the stored transcript was [user FIRST, assistant] — the turn-2 message was deleted. The same 3-turn run on Redis lost more (RM-38); live: probe, JS executor + store-postgres (this session; also confirmed by the independent verifier): a 3-turn run whose turn 3 is stopped during its first LLM call keeps turns 1-2 (6 messages; the latest checkpoint during turn 3 was {stepCount 0, messageCount 6}) and deletes turn 3's user message; contrast: a first turn is safe — the fresh session's turn-start checkpoint counts 0 and the cleanup skips messageCount 0. The stop.* conformance scenarios (e2e/src/conformance/scenarios/stop.ts) stop only a first turn, so they do not reach this; live: the resume({ mode: 'with_message' }) path was reproduced on the memory store by the independent catalogue verifier (the resumed message was deleted on re-interrupt). inferred: the companion path (B2-core's B13 review, source-level only, not observed). Entry paths with the same shape: a resume({ mode: 'with_message' }) turn whose message is appended at runtime-js/src/js-agent-executor.ts:2358-2362, and a JS companion resumed via sendMessage, whose message core appends at core/src/orchestration/companion-tool-dispatch.ts:780 before runtime-js/src/run-loop.ts:4604-4616 (resumeChildAgent) re-enters runLoop without a new checkpoint. Either turn, re-interrupted before its first step commits, goes through the same interrupt truncation (runtime-js/src/run-loop.ts:3889-3925), which deletes the delivered message because it was appended outside any checkpoint; inferred: the Cloudflare DO embeds this executor; its saveState checkpoint records message_count 0 (RM-36), which the cleanup skips, so the DO loses the message only when the previous promote checkpoint is still latest (sentinel case). Not observed; contrast: runtime-js/src/run-loop.ts:1636-1660 (once the first step commits, saveStateAndPromoteStaging counts the input and the window closes); related: RM-41 (CF Workflows: messages appended outside any checkpoint, a separate mechanism); RM-38 (Redis resolves the latest checkpoint to an older turn); CP-19 (a stale durable interrupt flag interrupts the next run — combined, the next turn's message is deleted too); severity S1: the user's message is silently deleted from the durable transcript, observed livesource, probe, peer-reportedverifiedopen
RM-43S2runtime-js resume({ mode: 'with_message' }) and CF Workflows' continued turns create the turn's run record BEFORE appending the turn's input. On CF Workflows that covers every execute() on an existing session: a completed, failed or interrupted one (the continuation branch) and a paused one without pending client tools (the resumable branch). So the run's startMessageCount / startUIMessageCount exclude the user message (and any continuation heal) that belong before the run. JS execute() appends first and records them correctly. The ai-sdk handler trusts these boundaries. In content-replay mode the snapshot keeps only messages before startMessageCount, which drops the turn's user message while the run is active (it is not in the replayed stream either). The streamed assistant's derived id msg-<startUIMessageCount> is the user message's own UI id.runtime-js/src/js-agent-executor.ts:2335-2355 (resume: createRun with startMessageCount = state.messages.length and startUIMessageCount from the store), then :2358-2362 (with_message appends the new user message); runtime-cloudflare/src/workflow.ts:585-586 (isResumable only for running / paused, so completed, failed and interrupted sessions take the continuation branch): :624-627 (createResumeRun) before :661 (trailing heal appended) and :668 (the new turn's user messages appended); the resumable branch has the same order, :713-716 (createResumeRun) before :735-739 (newMessages appended), reached by execute() on a paused session without pending client tools (runtime-cloudflare/src/executor.ts:466-500 rejects only a running run or pending client tools). resume() is unaffected: runtime-cloudflare/src/executor.ts:776 appends its message before the workflow starts; runtime-cloudflare/src/steps.ts:5807-5840 (createResumeRun reads getMessageCount and the UI count at call time, :5828-5829); contrast: runtime-js/src/js-agent-executor.ts:3499-3526 then :3558 (execute appends first, then createRun — the '[R3C-H2] ordering' note at :3501-3508 explains why the count must include the input); ai-sdk/src/handler/snapshot.ts:233-250 (content-replay filter keeps messages.slice(0, startMessageCount) while the session is active), :338-340 (isNewTurn from startMessageCount); ai-sdk/src/handler/handle-chat-stream.ts:1884-1887 (deriveExistingMessageId = msg-<startUIMessageCount>), used at :1150, :1308, :1977; inferred: reloading during such a turn with content replay on shows the conversation without the user's latest message until the run ends, and the replayed assistant bubble takes the id the snapshot gives that user message. Not observed; related: RM-44 (Temporal records the boundary racing the workflow, a separate mechanism); RM-45 (DBOS hard-codes it to 0); RM-05 (snapshot logic past the 50-message cap); severity S2: visible content is wrong only while a resumed / continued run is active and only in content-replay modesource, peer-reportedpartly-inferredopen
RM-44S2Temporal execute() on an existing session reads the new run's boundaries AFTER it has started the workflow: startSequence from the stream, and startMessageCount / startUIMessageCount from the store. The workflow is already running in parallel, appending the turn's user message in its initialize activity and then streaming the first step. So the recorded boundary depends on scheduling. It can count the user message or not, and can start after some of the run's own chunks. A content-replay reconnect then drops or duplicates the user message, or starts replay after the turn's first chunks.runtime-temporal/src/executor.ts:430 (client.startWorkflow) precedes :538-548 (continuation run metadata: streamInfo.latestSequence, getMessages().total and the UI count, read at that moment) and :582 (createRun); runtime-temporal/src/activities.ts:6660 (initializeAgentState activity) appends the turn's input at :7187, concurrently with the executor reads above; runtime-temporal/src/executor.ts:496-519 (the in-code note: the executor-side createRun was hoisted out of the claim block; it does not order the reads against the workflow); ai-sdk/src/handler/snapshot.ts:233-250, :338-340 and ai-sdk/src/handler/handle-chat-stream.ts:1143-1152, :1302-1311 (content replay starts at currentRun.startSequence and slices history at startMessageCount / derives msg-<startUIMessageCount>); contrast: the first turn is deterministic — runtime-temporal/src/executor.ts:527-536 computes it from the input, not from the store; inferred: a reconnect during a continued Temporal turn can miss the start of the answer (startSequence past the first chunks) or show the user message twice / not at all. Not observed; related: RM-43 (JS resume / CF Workflows record the boundary deterministically before the append); RM-46 (Temporal retry creates no run at all); severity S2: wrong visible content only on a content-replay reconnect during a continued Temporal turn, and only when the race goes the wrong waysource, peer-reportedpartly-inferredopen
RM-45S2DBOS seeds the run record with hard-coded zero boundaries instead of reading them, on runs that start on an existing history. Every resume() mode records startMessageCount / startUIMessageCount 0, and so do a persistent companion's continuation-restart and sendMessage turns. The continuation restart also records startSequence 0 while streaming to the shared session stream. The ai-sdk handler reads 0 counts as "the run started at the beginning of the session". In content-replay mode it skips the history filter (applied only when startMessageCount > 0), and gives the run's streamed assistant the id msg-0, which is the session's FIRST message's UI id.runtime-dbos/src/lifecycle/resume.ts:508-524 (the single standardAgentWorkflow start shared by every resume mode: runMeta { turn, startSequence (read at :435), startMessageCount: 0, startUIMessageCount: 0 }); runtime-dbos/src/steps/execute-companion-tool.ts:929-945 (companion continuation restart: initialRunMeta { turn: 1, startSequence: 0, startMessageCount: 0, startUIMessageCount: 0 }; the in-code comment calls the counts bookkeeping only); runtime-dbos/src/workflows/persistent-workflow.ts:245-255 (that run is begun on streamId = sessionId, the shared session stream, which initStream does not clear — runtime-dbos/src/lifecycle/execute.ts:595-596); runtime-dbos/src/steps/execute-companion-tool.ts:1141-1152 (sendMessage inbox runMeta: counts 0). Its startSequence 0 is CORRECT: each sendMessage turn streams to its own new stream (streamId = childRunId = <subSessionId>-tc-<toolCallId>, :1127, :1145; persistent-workflow.ts:630, :652; run-lifecycle.ts:158-164 stamps it on the session); runtime-dbos/src/workflows/standard-workflow.ts:485-499 (beginRun writes runMeta verbatim into the run record); ai-sdk/src/handler/snapshot.ts:233-250 (history filter only when startMessageCount > 0), :338-340 (isNewTurn), :499; ai-sdk/src/handler/handle-chat-stream.ts:1884-1887 (msg-<startUIMessageCount>), :1143-1152 (replay from currentRun.startSequence); contrast: runtime-dbos/src/lifecycle/execute.ts:169-211, :692 (execute projects real counts for a fresh or continued turn); inferred: reloading a root DBOS session during a resumed run with content replay shows the run's committed content twice (history plus replay), and the streamed bubble collides with the first message's id. A reconnect to a companion session during a continuation restart would also replay its earlier turns from the shared stream. Not observed; related: CP-30 (DBOS retry hard-codes startMessageCount 0 among ignoring its options — same pattern, catalogued there); RM-43, RM-44 (the boundary recorded at the wrong moment on JS / CF Workflows / Temporal); severity S2: visible duplication / id collision, only on a content-replay reconnect during a DBOS resumed run (root sessions); the companion legs are narrower still, since a UI rarely attaches to a companion session directlysource, peer-reportedpartly-inferredopen
RM-46S2Temporal retry() creates no run record. It reactivates the session and starts a new workflow, but getCurrentRun() keeps returning the failed run, and that record still reads 'running' because Temporal never finalizes it (RM-47). The ai-sdk handler therefore treats the retry as part of the failed run. Content replay starts at the failed run's startSequence, re-delivering the failed attempt's chunks ahead of the retry's. The derived message id and history slice come from the failed run's counts, before retry's truncation and appended messages. Temporal child sessions likewise never get a run record.runtime-temporal/src/executor.ts:1554-1808 (retry: CAS failed → active at :1590, truncation and appended retry messages, stream_resync at :1712-1731, startWorkflow at :1776 — no createRun anywhere in the method); the only createRun calls are :582 (execute) and :1328 (resume); runtime-temporal/src/workflow-dispatch.ts:337, :680 (children start via wf.startChild; nothing creates a run for the child session); ai-sdk/src/handler/snapshot.ts:202, :270, :283-300 (currentRun drives isRunSettled and the partial-content start sequence); ai-sdk/src/handler/handle-chat-stream.ts:1143-1152, :1302-1311 (content replay from currentRun.startSequence, id from its startUIMessageCount); contrast: runtime-js/src/js-agent-executor.ts:2878 (JS retry creates a new run); runtime-cloudflare/src/executor.ts:1074 starts a workflow whose body creates a run through createResumeRun (runtime-cloudflare/src/workflow.ts:624-627 or :713-716); inferred: reconnecting to a Temporal retry with content replay replays the failed attempt's output before the retry's, and the snapshot's partial merge reads chunks from the failed run's start. Not observed; related: RM-47 (Temporal writes no terminal run status); DI-16 (runs never superseded); RM-44; severity S2: visible duplication only on a content-replay reconnect during a Temporal retry; run-listing is wrong in all casessource, peer-reportedpartly-inferredopen
RM-47S3Temporal never writes a terminal run status. A completed, failed or interrupted turn's run record stays 'running' forever, and one that suspended stays 'suspended_step_partial' / 'suspended_client_tool' / 'suspended_awaiting_children' after it resumes and finishes. listRuns() / getCurrentRun() therefore misreport every finished Temporal turn, while JS, DBOS and CF Workflows record 'completed' / 'failed' / 'interrupted'. An in-code comment accepts this as observability-grade.runtime-temporal/src/activities.ts:4575-4600 (commitSuspendedStep's updateRunStatus is the only run-status write in the package, and only for the three suspended statuses: suspended_step_partial, suspended_client_tool, suspended_awaiting_children); runtime-temporal/src/executor.ts:496-509 (the comment: a completing turn's record 'stays at the initial running status'; 'nothing load-bearing reads the terminal status'); contrast: runtime-js/src/run-loop.ts:976, :1042, :3216, :3374, :4226 (JS writes failed / interrupted / completed); runtime-dbos/src/steps/run-lifecycle.ts:331 (finalizeRun); runtime-cloudflare/src/steps.ts:5782-5799 (finalizeRun); inferred: ai-sdk/src/handler/snapshot.ts:56-86 (computeIsRunSettled) and ai-sdk/src/handler/handle-chat-stream.ts:284-306 key on currentRun.status. They only act while the session is active or paused, so a finished turn's stale 'running' matters when it is still the current run of an active session (RM-46: a Temporal retry); related: RM-46 (retry creates no run, so the stale record stays current); DI-16 (runs never superseded); severity S3: run records are misreported on every Temporal turn, but the host layer consults run status only for active or paused sessions; the visible consequence is RM-46'ssource, peer-reportedpartly-inferredopen
RM-48S3@helix-agents/llm-vercel never reports LLM step boundaries. It routes chunks to the step callbacks only from streamText's onChunk, and ai never passes step parts to onChunk. Its own fullStream loop, which does receive start-step / finish-step, ignores them. So onStepStart / onStepEnd never fire. runtime-js (and the CF DO), Temporal and CF Workflows emit step_start / step_end only from those callbacks, so with the shipped adapter no step chunk reaches the Helix stream. includeStepEvents: true is then a silent no-op, and per-step usage and finishReason never appear on the stream.llm-vercel/src/vercel-adapter.ts:222-298 (onChunk → mapVercelChunkToStreamChunk; the step_start / step_end cases at :259-277 are the only calls to onStepStart / onStepEnd); llm-vercel/src/chunk-mapper.ts:133-154, :230-240 (the finish-step → step_end and start-step → step_start mappings, unreachable through onChunk); llm-vercel/src/vercel-adapter.ts:348-373 (the fullStream loop handles text-delta, reasoning-delta, reasoning-end, tool-call, finish and error only — start-step / finish-step fall through unused); live: ai 6.0.49 dist/index.mjs:6059-6061 (eventProcessor calls onChunk only for text-delta, reasoning-delta, source, tool-call, tool-result, tool-input-start, tool-input-delta, raw), while :6639-6643 and :6816-6827 enqueue start-step / finish-step into the stream that fullStream (:7017-7027) exposes unfiltered; live: ai 6.0.91 (the consumer's pin; npm tarball inspected this session) dist/index.mjs:6263-6265 has the identical onChunk guard — step parts never reach onChunk there either; live: probe (this session) — VercelAIAdapter.generateStep over ai/test MockLanguageModelV3 with onTextDelta / onStepStart / onStepEnd callbacks: only onTextDelta fired; the same model's streamText fullStream yielded [start, start-step, text-start, text-delta, text-end, finish-step, finish]; runtime-js/src/run-loop.ts:3637-3669 (step_start / step_end emitted only inside onStepStart / onStepEnd); runtime-temporal/src/activities.ts:2997-3029; runtime-cloudflare/src/steps.ts:1246-1278 (same pattern); ai-sdk/src/transformer/stream-transformer.ts:1344-1381 (includeStepEvents turns step_start / step_end into start-step / finish-step, carrying usage and finishReason; with no step chunks it emits nothing); live: the S6 worker captured the opennext example's chat-route SSE (examples/opennext-cloudflare-do/src/lib/agent-client.ts:50 sets includeStepEvents: true): no start-step / finish-step events; related: DI-20 — its schema rejection reaches only step-reporting adapters BECAUSE of this; fixing this finding on an ai version before 6.0.231 exposes llm-vercel users to DI-20 unless DI-20 is fixed first. RM-49 (DBOS emits a step_start of its own but never a step_end); inferred: content blocks — the transformer's step_end handler also closes open text / reasoning blocks (ai-sdk/src/transformer/stream-transformer.ts:1365-1368), but tool_arg_stream_start and tool_start close them too (:737, :1095; emitted for every tool call, runtime-js/src/run-loop.ts:2001, :2110, :3607), and every non-terminal step ends in a tool call. So a missing step_end changes block boundaries only if a step emits text AFTER its last tool call and the next step emits text as well (the two texts would share one block). Not observed; severity S3: a documented option silently does nothing and per-step usage / finishReason are missing from the stream; token usage itself is still recorded (the adapter returns it in the StepResult), and message content is affected only in the narrow text-after-tool-call case above, not observedsource, probe, peer-reportedverifiedopen
RM-49S3DBOS emits a step_start for every LLM call but never a step_end. callLLMStep writes step_start itself before calling the adapter, and passes the adapter no streaming callbacks, so this happens whatever adapter is used. No code in runtime-dbos writes step_end. With includeStepEvents: true, a DBOS stream therefore carries a start-step per step and never a matching finish-step, with no per-step usage or finishReason. JS, Temporal and CF Workflows emit both, or with llm-vercel neither (RM-48). So DBOS is the only runtime whose step events are unbalanced under every adapter.runtime-dbos/src/steps/call-llm.ts:85-114 (callLLMStep writes step_start before generateStep, unconditionally, with no stepId), :129-132 (the comment: streaming callbacks are not wired — no onStepStart / onStepEnd reaches the adapter); runtime-dbos/src/steps/call-llm.ts:100 is the only step chunk written in runtime-dbos/src; no step_end in non-test source (grep -rn "step_end\|onStepEnd" packages/runtime-dbos --exclude-dir=__tests__ --exclude-dir=node_modules --exclude-dir=dist: no matches; only test mocks such as src/__tests__/integration/helpers/mock-llm.ts:161 call onStepEnd); contrast: runtime-js/src/run-loop.ts:3646-3669, runtime-temporal/src/activities.ts:3006-3029, runtime-cloudflare/src/steps.ts:1255-1278 (step_end with usage and finishReason, from the adapter callback); ai-sdk/src/transformer/stream-transformer.ts:1344-1354 (start-step emitted per step_start), :1365-1381 (finish-step only from step_end, which DBOS never sends); related: not claimed here — the missing stepId is not DBOS-specific — no runtime sets stepId on step_start / step_end (core/src/types/stream.ts:433-448 declares it optional); inferred: a UI consuming step events (a per-step progress indicator or usage readout) sees steps start on DBOS but never finish. Not observed; related: RM-48 (llm-vercel reports no step boundaries, so the other runtimes emit none); DI-20 (the strict start-step schema — the DBOS start-step carries no stepId, so it serializes as { type } and is not rejected); inferred: content blocks — as for RM-48, a missing step_end loses only the transformer's step-boundary block close (ai-sdk/src/transformer/stream-transformer.ts:1365-1368); tool calls close blocks too (:737, :1095), so only text after a step's last tool call followed by text in the next step could merge. Not observed; severity S3: unbalanced observability chunks; message content is affected only in the narrow case above, not observedsource, peer-reportedpartly-inferredopen
RM-50S2Redis createSession is not atomic. It claims the session hash with HSETNX sessionId in one round trip, then writes status and every other field in a separate pipeline. loadState treats any non-empty hash as an existing session, so a read that lands between the two sees a hash holding only sessionId. parseStoredSessionStatus then throws a typed InvalidStoredStateError: a session that is merely mid-creation is reported as CORRUPT, where it should read as not-yet-created (null). The error contract routes that error to corruption remediation (delete and restart), not to retry.store-redis/src/redis-state.ts:560-563 (HSETNX sessionId claim, its own round trip), :577-609 (the status / fields hset queued on a pipeline), :636 (pipeline.exec, a second round trip); store-redis/src/redis-state.ts:1498-1501 (loadState returns null only for an empty hash; a sessionId-only hash is treated as present), :1550 (parseStoredSessionStatus(sessionId, mainFields.status) on the missing status); core/src/store/parse-stored-session-status.ts (throws InvalidStoredStateError on a status outside SessionStatusSchema); core/src/errors/agent-errors.ts:70-87 (the contract: route InvalidStoredStateError to remediation, e.g. delete the corrupt row and restart, rather than retry); store-redis/src/redis-state.ts:950 (secondary: appendMessages ends with an unguarded HINCRBY messageCount on the session hash, which would create a status-less hash if it ran before createSession); contrast: store-cloudflare/src/d1-state.ts:2220-2241 and store-postgres/src/state/postgres-state-store.ts:462-512 (createSession writes the row in a single INSERT, so a reader sees all of it or none) — Redis only; live: probe (this session, store-redis dist against Redis :6390): createSession on one connection while a second connection loops loadState on the same id. 142 of 200 runs threw InvalidStoredStateError, and nothing else was thrown; live: B8 (peer-reported) saw the same error in 153 of 200 runs of the same tight loop against a private Redis. It also hit CI: pipeline 2885420760, job 16754754845 (e2e ai-sdk-redis-temporal non-streaming sub-agent) failed on it and passed on the automatic retry, job 16754894710; related: RM-27 (non-atomic Redis / D1 cloneSession, same store and same class of defect); RM-08 (non-atomic snapshot reads); severity S2: a concurrent reader of a session being created (a snapshot / status poll racing execute, a sub-agent observer) gets a corruption error instead of "not found". It is transient and retry clears it, but the typed contract tells callers not to retry; conformance: session.create-observed-whole on js-redis / temporal-redis proves the fix (MR !269). Each cell creates 48 sessions, and 16 concurrent observe() loops read each one. Every observer must first read "no session" before the run starts, and must never see a half-created one. Pre-fix, measured with the scenario at 4bbf1678aa against main 59d47f1e0f store-redis: .never-partial (and only that id) failed 10/10 isolated runs per lane. Sessions observed half-created per run were 10-35 of 48 on js-redis (runs: 26, 16, 10, 12, 24, then 24, 26, 26, 35, 15) and 23-34 of 48 on temporal-redis (32, 27, 33, 30, 28, then 32, 30, 34, 26, 29). Post-fix: 10/10 pass on both lanes, and the scenario passes on all 7 executor backends; store-redis/src/__tests__/create-session-atomic.integ.test.ts (regression, MR !269): a createSession on one connection against a loadState looped on a second failed 161-173 of 200 iterations pre-fix and 0 of 200 post-fix (0 of 1000 on a longer run). A deterministic variant runs a loadState before every command createSession dispatches: 20 of 20 attempts failed pre-fix, 0 of 20 post-fix. The core contract cases "a loadState racing createSession sees no session or the whole initial session" and "appendMessages on a missing session rejects and creates nothing" (core/src/testing/session-lifecycle.ts, core/src/testing/message-operations.ts) fail on the pre-fix Redis store and pass on memory, Redis, Postgres, D1 and the DO store; store-redis/src/redis-state.ts (fix): createSession is one Lua script (CREATE_SESSION_ATOMIC_SCRIPT) — HEXISTS sessionId claim, clears stale per-session keys, writes custom state and indexes, and writes the full hash last. appendMessages is one Lua script (APPEND_MESSAGES_ATOMIC_SCRIPT) guarded on HEXISTS sessionId. It throws Session not found for a missing sessionsource, live-repro, peer-reportedverifiedfixed (2 cells)
RM-51S1Cloudflare DO: DOStreamManagerClient.createResumableReader (and createReader, which delegates to it) silently drops every SSE frame after the first one in a single network read. next() splits the read into a local lines array, moves only the trailing partial line back into the carried buffer, and then returns from inside the line loop on the first chunk. The rest of lines is discarded, so those frames are never yielded or buffered, and neither is an event: end frame in the same read. workerd was observed to coalesce back-to-back DO stream writes into one read on wrangler 4.141 / workerd 1.20260925 (a burst of text deltas), so the ai-sdk Cloudflare chat handler's live-attach, resume-raced-auto-resume, AlreadyRunning-attach and terminal-replay paths can stream a truncated assistant message and raise no error. A dropped end frame does not hang the reader (the DO closes the SSE body after it, so the reader sees a clean EOF), but a dropped client-tool call never reaches the client, so the tool is never dispatched and the run stays paused. In the opennext example the post-client-tool confirmation stopped at "The page", and only a reload (served from the snapshot) showed the full text.runtime-cloudflare/src/do/do-clients.ts:472-560 (the next() read loop): :482-483 (lines is local to one read; only the trailing partial line is kept in the carried buffer), :548-554 (return from inside for (const line of lines) on the first chunk, which discards the remaining complete lines of that read; an end frame later in the same read is discarded because this return is taken first, so the end handler at :513-515 never sees it); :596-601 (createReader delegates to createResumableReader, so it drops frames the same way: same cause, same site, same fix); runtime-cloudflare/src/__tests__/do-clients.test.ts:501-507 (the defect was already known: the existing test notes that each next() yields at most one chunk per read, calls it "a known pre-existing limitation", and asserts only the first chunk; its stated mechanism, a per-call buffer, is out of date since the buffer moved outside next(), but its conclusion still holds); runtime-cloudflare/src/do/durable-object-agent-base.ts:4304-4337 (sendInDOTerminalFrame writes the end/fail frame and then closes the SSE writer) and core/src/stream/sse-connection-manager.ts:296-323 (broadcastTerminal writes the terminal frame and then closes each live writer): after a dropped end frame the reader's next read() returns done, so the consumer sees a clean EOF rather than a hang; ai-sdk/src/transformer/stream-transformer.ts:733-792 (tool-input-available is emitted from tool_start, or at finalize :544-576 only for a tool call whose arg-stream chunks were seen) and ai-sdk/src/react/use-resume-client-tools.ts:1-45 (the hook dispatches only tool parts in state input-available): when every frame of a client tool call lands after the first frame of its read, the client never runs the tool and the run stays paused with no error, until a reload rehydrates the pending part from the snapshot or the client-tool deadline (core/src/constants.ts:117, default 5 min) expires; ai-sdk/src/cloudflare/chat-handler-factory.ts:168 (createCloudflareChatHandler wires DOStreamManagerClient as the handler stream manager); ai-sdk/src/handler/handle-chat-stream.ts:1178 (streamLiveResponse reads through createResumableReader; reached from :335 active-stream attach, :821 a client-tool resume that raced the runtime auto-resume, and :1000 the AlreadyRunning fallback) and :1351 (replayTerminalStream also reads through createResumableReader); contrast: runtime-cloudflare/src/do/do-clients.ts:1253-1343 (DOExecutionHandleImpl.parseSSEToChunks, used by the fresh-turn handle.stream() path, is an async generator that yields inside the line loop, so it resumes the same lines array and loses nothing — not defective, no separate id); runtime-cloudflare/src/do/do-executor.ts:176-270 (SSEStreamReader, also a generator) and store-cloudflare/src/durable-stream.ts:197-240 (pushes every parsed frame into a buffer) are also correct; live: probe (this verification, unit level on origin/main): a mocked DO /sse body whose single read carries four text_delta frames plus an event: end frame. createResumableReader yielded only ["The"] and currentSequence stayed 1. The same four frames delivered one per read yielded all four. DOFrontendExecutor.execute(...).stream() (parseSSEToChunks) over the same single-read body yielded all four; live: probe (fix round 1, unit level on origin/main, not committed): one read carrying a text_delta, a client-tool tool_start and an event: end with status paused. createResumableReader and createReader each yielded only ["text_delta"] and then ended cleanly when the body closed (as the DO does after an end frame); the tool_start was never yielded. With the body left open the reader waited on the next read, which the DO never leaves open after a terminal frame; live: B8 (peer-reported) — the opennext refresh-exact-messages C case fails 3/3 on wrangler 4.141 before the fix and passes after; the fix (MR !272, head f83d14de93 on main 0e618772ee) queues the split lines across next() calls; runtime-cloudflare src/__tests__/do-clients.test.ts with the pre-fix do-clients.ts: 2 failed | 50 passed ("yields every frame when one read() carries several SSE frames" got 1 of 4 deltas; "yields every delta and honors a trailing end frame from the same read"), with the fix: 52 passed; the old "known pre-existing limitation" test now asserts the full output; runtime-cloudflare/src/do/do-clients.ts (fix: MR !272 (186b0c6a48 on main) — the reader queues the split lines across next() calls; red/green in runtime-cloudflare/src/__tests__/do-clients.test.ts (pre-fix file: 2 failed | 50 passed; fixed: 52 passed, re-run on 186b0c6a48) and the opennext refresh-exact-messages C case (3/3 fail pre-fix on wrangler 4.141, 9/9 pass after, peer-reported by B8); ships in @helix-agents/runtime-cloudflare 5.20.0: the patch changeset .changeset/do-resumable-reader-multi-frame.md is pending beside the transcript-validity minor, and the CHANGELOG still tops at 5.19.0. The do-clients.ts line references above are to the pre-fix file); inferred: the client-tool stall (tool_start dropped, so useResumeClientTools never runs and the run stays paused until a reload or the client-tool deadline) is shown by the reader-level probe plus source only; it was not reproduced end to end; inferred: production is likely exposed independently of the local wrangler pin, because deployed workers run on the current edge workerd; coalescing was observed only on wrangler 4.141 / workerd 1.20260925, and main's resolved wrangler (4.63) passes the opennext case. The terminal-replay path is likely hit hardest: it reads a finished stream whose history arrives in large reads, so a replayed message keeps about one delta per read; related: CP-28 and CP-25 (other DOStreamManagerClient reader defects: errors turned into null/EOF, and clean EOF on eviction); RM-43 (run-boundary counts on the same chat-handler attach/replay paths, a different mechanism); severity S1: besides silently truncating attached or replayed assistant text (display-only, a reload recovers it), the same drop can swallow a client-tool call on the attach paths, which include the resume after a client-tool submit. The client then never dispatches the tool and the run stays paused with no error until the user reloads or the client-tool deadline expires, so the user action is lost unless they happen to reload. A dropped end frame alone is benign: the DO closes the body, so the reader ends cleanlysource, probe, peer-reportedverifiedopen

Agent materialization & session context ​

IDSevFindingEvidenceProvenanceVerificationStatus
MA-01S1DO wake: an agent-factory throw is logged at warn and ensureExecutionContinuesImpl returns; the watchdog retries 5× at 30 s, then fails the stream with auto_resume_exhausted, whose message blames "a step that crashes the instance". The real error is lost. Live: draft factory "started without pending meta … did beforeStart run?" ×5.runtime-cloudflare/src/do/durable-object-agent-base.ts:682-696, 1061-1185, 1204-1256; runtime-cloudflare/src/do/wake-recovery.ts:179-187source, live-reproverifiedopen
MA-02S2Same warn-and-return on wake for LLM-adapter creation failure and missing subAgentNamespace.runtime-cloudflare/src/do/durable-object-agent-base.ts:705-711, 721-733sourceverifiedopen
MA-03S1Factory context is {env, sessionId, userId?, runId}; userId is passed only on /start; persisted userId / tags / metadata are never passed. AgentResolver is synchronous, forcing consumers to stage async data in beforeStart.runtime-cloudflare/src/do/types.ts:91-124; runtime-cloudflare/src/do/durable-object-agent-base.ts:1696-1713, 2096-2116, 2494-2516, 682-696sourceverifiedopen
MA-04S2/resume and /retry report any factory exception as 400 Unknown agent type.runtime-cloudflare/src/do/durable-object-agent-base.ts:2104-2116, 2502-2516sourceverifiedopen
MA-05S2Side lookups swallow factory failures: getToolByName returns undefined so output-schema validation of a submitted result is silently skipped, and passes sessionId: ownerAgentType; retention lookups silently default.runtime-cloudflare/src/do/durable-object-agent-base.ts:600-614, 3019-3032, 3085-3099sourceverifiedopen
MA-06S1Client-tool submit to a paused (HITL) session + factory failure → submit acknowledged 200, run stranded forever (paused sessions have no watchdog).runtime-cloudflare/src/do/durable-object-agent-base.ts:3107-3112, 2999-3003, 949-951sourceverifiedopen
MA-07S2beforeStart runs before the 409 checks, so staged data leaks when a start is rejected (the brief's non-determinism).runtime-cloudflare/src/do/durable-object-agent-base.ts:1644-1690sourceverifiedopen
MA-08S1DO wake path builds JSAgentExecutor without executorOptions(): workspace providers, metrics and limits dropped only on wake; workspace tools then error.runtime-cloudflare/src/do/durable-object-agent-base.ts:768-771 vs :518-532, 1783, 2148, 2548; live: incident 2026-09-24 audio-resume — right after the wake, "Workspace failed to open during eager strategy: No provider registered for providerId cloudflare-dynamic-worker" (ai-run1.out:2994), the log's only occurrencesource, live-reproverifiedopen
MA-09S1DO wake path fires no ServerHooks; onComplete never fires after an auto-wake or auto_resume_exhausted. Live: host row stuck "writing" with a live Stop button forever.runtime-cloudflare/src/do/durable-object-agent-base.ts:662-842; live: incident 2026-09-24 audio-resume — woken run 19935d73 completed but the host never advanced the article (runs row + incident ANALYSIS step 4); no log line shows the hook directly; conformance: materialize.wake-hooks on cf-do (host and client tiers) — evicted mid model call, woken by a real request (the chat snapshot GET → onStart → recoverStrandedExecution('wake')), the run completes and the agent's onAgentComplete fires once, but the host's hooks.onComplete never fires (.host-on-complete-fired-once: no host.onComplete record). Only the fresh-instance 'wake' recovery is claimed; the warm 'watchdog' alarm path is not exercised; conformance: since Miniflare 5 the evicted DO's own pre-eviction alarm fires at its deadline (~30s after the run started) and its 'watchdog' recovery would complete the run too. So materialize.wake-hooks' .recovered-by-the-wake proves the wake ran the recovery: the recovered model call started after the wake request and before the earliest alarm deadline (Node-side ScriptedModel timestamps). Every hook id takes it as a precondition, so a fix MR can name this cell in provenBy only for the 'wake' path. Checked by a temporary perturbation, never committed: the conformance DO's recoverStrandedExecution skips 'wake' and keeps 'watchdog'. (A) With the 15s bound, .run-completes-after-wake does not settle, and the dependent ids fail as its precondition. (B) With the bound widened to 45s, the pre-eviction alarm completes the run inside the window, so .run-completes-after-wake PASSES. But .recovered-by-the-wake fails ('the recovered model call started 3003xms after the run started', the alarm deadline), and the hook ids fail as its precondition instead of passing through the alarmsource, live-reproverifiedopen
MA-10S2Unguarded throws on wake (rewriteSubAgentTools) propagate into onStart, failing every request to that DO until it stops throwing.runtime-cloudflare/src/do/durable-object-agent-base.ts:698, 714, 1461-1464sourcepartly-inferredopen
MA-11S2CF Workflows resolves only statically registered agents; factory agents are impossible.runtime-cloudflare/src/registry.ts:120-128; runtime-cloudflare/src/workflow.ts:519sourceverifiedopen
MA-12S2Temporal's resolver gets no context (no env, sessionId, userId).runtime-temporal/src/activities.ts:123-125; runtime-temporal/README.md:102-103sourceverifiedopen
MA-13S1DBOS materializes agents from process-global registries with two different keys, both last-writer-wins: every tool (incl. __finish__ and sub-agent / remote tools) is registered by tool.name alone — so two DIFFERENT agents that declare a same-named tool overwrite each other process-wide, and each runs whichever implementation registered last — while serialized definitions, the live model, skills and the cache strategy are keyed by agent.name, so per-session configurations of one agent name overwrite each other.runtime-dbos/src/lifecycle/shared-registration.ts:68 (ToolRegistryHolder.registry[tool.name] = tool — tool-name key, shared across agents), :74 (RemoteTransportHolder, same key); runtime-dbos/src/lifecycle/shared-registration.ts:80, 98-107, 116, 125 (AgentDefinitionRegistryHolder / LiveModelRegistryHolder / SkillsRegistryHolder / LiveCacheRegistryHolder keyed by agent.name; only the model overwrite logs a warning); runtime-dbos/src/workflows/standard-workflow.ts:457-467 (getTools resolves the running agent's tools by NAME from ToolRegistryHolder at call time); related: MA-18 (skills map), MA-20 (the tool-name collision bypasses approval gates)sourceverifiedopen
MA-14S2Sub-agent /subagent/:type/start builds a synthetic request dropping headers; companion /start sends only metadata (no userId / tags); wake passes no identity to companions.runtime-cloudflare/src/do/durable-object-agent-base.ts:3163-3185, 714; runtime-cloudflare/src/do/persistent-companion-tools.ts:78-80, 257-264, 371-380sourceverifiedopen
MA-15S3Three unrelated AgentNotFoundError classes; no error code for agent resolution failure.runtime-cloudflare/src/registry.ts:46; runtime-cloudflare/src/do/types.ts:777; runtime-temporal/src/registry.ts:89sourceverifiedopen
MA-16S1Memory tools omit userId / tags / metadata from the entity context (saves go to the wrong entity, searches come back empty); CF Workflows auto-inject / realtime extraction omit identity; completion extraction has no user identity on any runtime.memory/src/tools/save-memory.ts:41-46; memory/src/tools/search-memory.ts:34-39; runtime-cloudflare/src/steps.ts:938-943, 5809-5816; core/src/memory/types.ts:286-298sourceverifiedopen
MA-17S2Semantic memory is a silent no-op on Temporal and DBOS.runtime-temporal/src and runtime-dbos/src wire no MemoryManagersourceverifiedopen
MA-18S2DBOS skills map is last-writer-wins without a warning; load_skill resolver cache never clears.runtime-dbos/src/lifecycle/shared-registration.ts:115-117; runtime-dbos/src/steps/execute-tool.ts:170-183sourceverifiedopen
MA-19S3DO /resume and /retry take agentType from the request body without checking the persisted agentType; an idle DO does not verify it on /start.runtime-cloudflare/src/do/durable-object-agent-base.ts:2024, 2411sourceverifiedopen
MA-20S1DBOS approval gate fails open through the tool-name registry collision (MA-13): when two agents in one process declare a same-named tool — one approval-gated, one not — and the ungated one registered last, the gated agent's call resolves to the OTHER agent's ungated tool, isApprovalGatedTool is false so no gate is evaluated at all, and the other agent's implementation runs without approval. (A gated tool merely MISSING from the registry does not run: getTools throws and executeToolStep answers 'Tool not found'; the descriptor-missing branch needs a hand-built Tool that bypasses defineTool and logs a warning.)runtime-dbos/src/lifecycle/shared-registration.ts:68 (tools keyed by tool.name process-wide — MA-13); runtime-dbos/src/workflows/standard-workflow.ts:457-467 (getTools → ToolRegistryHolder.registry[name]; a missing tool throws); runtime-dbos/src/workflows/shared.ts:1817-1820 (liveToolsByName from getTools), :1978-1994 (gate evaluated only when isApprovalGatedTool(tool) — an ungated same-named tool skips it); runtime-dbos/src/steps/execute-tool.ts:206-213 (a tool missing from the registry → ok:false "Tool not found", never run); runtime-dbos/src/steps/approval-gate.ts:128-155 (in-step registry miss / tool_not_gated → false, logged), :159-176 and runtime-dbos/src/workflows/shared.ts:696-714 (descriptor-missing → false, logged; unreachable via defineTool); severity S1 (kept): the collision needs two agents sharing a tool name with different gating in one DBOS process, but its consequence is an approval-gated side effect executed with no approval and no error; the reverse registration order fails closed (the ungated agent's calls start suspending for approval)sourceverifiedopen
MA-21S2DBOS persistent mode (the long-lived inbox workflow, mode: 'persistent') drops the agent's live outputSchema on ordinary follow-up turns. The persistent workflow passes outputSchema to runAgentLoopOneTurn for the first turn and for a message drained at idle timeout, but not for a normal inbox message (a follow-up execute() routed into the live workflow, or a companion sendMessage to a DBOS persistent child). On those turns planStepProcessing gets no schema and forced completion is disabled. The model still sees __finish__ (it is in the serialized tools), but its arguments are accepted unvalidated and saved as the session output. A turn that ends on plain text, max_tokens or maxSteps completes its run with no output instead of entering forced completion. Nothing reports either case.runtime-dbos/src/workflows/persistent-workflow.ts:294 (the live Zod outputSchema is resolved once, for every turn); :482-497 (initial turn, passes outputSchema at :495) and :580-597 (idle-timeout drained message, :594); :648-663 (the ordinary inbox turn omits it); runtime-dbos/src/workflows/shared.ts:1297 (hasOutputSchema = args.outputSchema !== undefined), :1752-1754 (planStepProcessing gets outputSchema only when args carries it), :2709 (forced completion is entered only when hasOutputSchema), :2791 (saveOutput of whatever output the plan carries); core/src/orchestration/step-processor.ts:317-333 (with no outputSchema, a __finish__ call's arguments become the output as-is: output = repaired as TOutput); runtime-dbos/src/lifecycle/shared-registration.ts:60-80 (buildEffectiveTools injects __finish__ into the serialized tools whenever the agent has an outputSchema, independent of the per-turn argument, so __finish__ is still offered); contrast: the other runtimes read the schema from the resolved AgentConfig on every step rather than from a per-call argument: core/src/orchestration/step-iterator.ts:886-887 (runtime-js / CF DO), runtime-temporal/src/activities.ts:2111-2112, runtime-cloudflare/src/steps.ts:1338-1339 (CF Workflows); the long-lived inbox workflow is DBOS-only (other runtimes ignore mode / defaultMode; their persistent companions run ordinary per-turn loops that read the schema from the AgentConfig); inferred: the observable effect (a completed follow-up turn with missing or schema-invalid output) follows from the code above; not reproduced by a test; related: MA-20 (another DBOS path that misses per-agent materialization); FU-DBOS-PERSISTENT-INBOX-TURN-DROPS-OUTPUTSCHEMA (peer-reported by B8, filed on MR !270's branch; not on main); severity S2: DBOS persistent mode only, and only turns after the first; the output contract ("completed means schema-valid output") silently does not hold there, with no error to act onsource, peer-reportedpartly-inferredopen

Session control plane ​

IDSevFindingEvidenceProvenanceVerificationStatus
CP-01S1handleChatStream path 5 attaches to a live run without reading params.messages: a new user message is dropped with HTTP 200. A unit test pins this as intended.ai-sdk/src/handler/handle-chat-stream.ts:284-354; ai-sdk/src/__tests__/handler/handle-chat-stream.test.ts:414source, live-reproverifiedopen
CP-02S1AlreadyRunning fallback in handleFreshTurnPath drops the message with HTTP 200.ai-sdk/src/handler/handle-chat-stream.ts:963-1008sourceverifiedopen
CP-03S1Generic execute() failure → HTTP 200 SSE containing only data-resume-rejected; no client code reads it.ai-sdk/src/handler/handle-chat-stream.ts:1026-1030, 2153-2164sourceverifiedopen
CP-04S2streamHandle null / throwing after retries → 200 rejected-only, although the run was started and persisted.ai-sdk/src/handler/handle-chat-stream.ts:1933-1966sourceverifiedopen
CP-05S2Abandonment path polls 5 s for paused to clear, then falls into the AlreadyRunning drop.ai-sdk/src/handler/handle-chat-stream.ts:446-480sourceverifiedopen
CP-06S1No idempotency: UIMessage.id is dropped on input; a retried POST after completion re-executes and duplicates the turn.ai-sdk/src/handler/handle-chat-stream.ts:1602-1719sourceverifiedopen
CP-07S1HelixChatTransport never checks /start / /resume responses (e.g. 409 ALREADY_COMPLETED → message dropped), openSSE ignores response.ok and sends no fromSequence, network error in sessionExists → /start, sends only the last user message's text, re-sends the previous user text on a tool-output submit, and cannot send approvals.ai-sdk/src/transport/helix-chat-transport.ts:96-135, 211-268, 322-349sourceverifiedopen
CP-08S2ai@6 Chat.sendMessage never rejects (third-party design); the SDK layers no delivery contract on top.ai@6.0.91 dist/index.mjs:12781-12794, 12498-12512sourceverifiedopen
CP-09S2useResumeClientTools: addToolOutput is fire-and-forget; a tool without a handler is silently left input-available.ai-sdk/src/react/use-resume-client-tools.ts:460-495sourceverifiedopen
CP-10S1Client-tool calls owned by a sub-agent are invisible to handleChatStream (only root pending map read) → rejected as unknown-tool-call with 200.ai-sdk/src/handler/extract-resume-intent.ts:435; ai-sdk/src/handler/handle-chat-stream.ts:238, 371, 446-449; runtime-js/src/run-loop.ts:1768-1845sourcepartly-inferredopen
CP-11S1DO discards approval decisions: the DO forwards an approval-response submit to the low-level JsClientToolResolver.submit, which drops approved/reason (result and error both undefined), so the decision never reaches the approval gate. (The JS runtime's own public JSAgentExecutor.submitToolResult persists { approved, reason } correctly — the resolver primitive is reached only by the DO.)runtime-cloudflare/src/do/durable-object-agent-base.ts:2988-2994, 3040-3046; runtime-js/src/client-tool-resolver.ts:654-655 (submit(): approval-response → result/error undefined); runtime-js/src/js-agent-executor.ts:962-968 (consumer: isApprovalResponseResult(undefined) is false, so the gate never sees a decision); contrast: runtime-js/src/js-agent-executor.ts:3138-3143 (public submitToolResult persists { approved, reason })sourcepartly-inferredopen
CP-12S1ai-sdk exposes no stop wiring anywhere (handler, CF factory, Express adapter, transports, hooks); useChat().stop() only aborts the local fetch while the server run continues to completion. interrupt() exists on handles, agent-server, and the remote protocol but is unreachable from the client stack.ai-sdk/src/{handler,cloudflare,adapters,transport,react,client} (grep: no stop / interrupt wiring); core/src/types/runtime.ts:571-597; agent-server/src/http/handler.ts:124-155sourceverifiedopen
CP-13S1"Last message wins" supersede in beforeStart is unreachable through handleChatStream while a run is live (path 5 attaches before /start). The consumer's research and writer cancel routes rely on it and never stop the run.ai-sdk/src/handler/handle-chat-stream.ts:298-345; runtime-cloudflare/src/do/types.ts:207-210; live: writer cancel route's handleChat('cancel') attached as passive subscriber, message dropped, run continued seq 2303→4452 over 4+ minsource, live-reproverifiedopen
CP-14S1DO /abort cancels nothing: the execution-state abort controller's signal is never passed to the executor; /abort awaits the whole run, then marks it failed.runtime-cloudflare/src/do/durable-object-agent-base.ts:2792-2819, 1768-1810; runtime-cloudflare/src/do/execution-state.ts:130-148sourceverifiedopen
CP-15S2CF Workflows handles call instance.abort(), which the real WorkflowInstance does not have (it has terminate); likely throws in production after persisting failed.runtime-cloudflare/src/executor.ts:1501-1503, 1612-1632; runtime-cloudflare/src/bindings.ts:107sourcepartly-inferredopen
CP-16S1Temporal, DBOS, CF Workflows never cancel an in-flight LLM call (no abortSignal passed); interrupt takes effect at step boundaries.runtime-temporal/src/activities.ts:2025-2040; runtime-dbos/src/workflows/shared.ts:1729-1743 (the only dispatcher.callLLM caller passes no signal; steps/call-llm.ts:44, 122 declare and forward it, so the plumbing exists but is never fed); runtime-cloudflare/src/steps.ts:1152, 484sourceverifiedopen
CP-17S2DBOS interrupt calls cancelWorkflow immediately (graceful path unreachable), emits no run_interrupted, does not end the stream, leaves paused sessions paused with pending calls cleared.runtime-dbos/src/lifecycle/interrupt.ts:67-125sourceverifiedopen
CP-18S2Stop on a paused session: DO returns 400; agent-server waits then 504s. DO /interrupt returns 400 "No agent to interrupt" whenever the in-memory session id is unset (after eviction).runtime-cloudflare/src/do/durable-object-agent-base.ts:2737-2789; agent-server/src/agent-server.ts:886-964sourceverifiedopen
CP-19S1A durable interrupt flag set just before a run ends is not cleared by DO /start / /retry, JS execute, or CF Workflows execute → it interrupts the next run. On DBOS the standard workflow reads the flag without ever clearing it and execute does not clear it either, so the session's next execute is interrupted (resume does clear it, runtime-dbos/src/lifecycle/resume.ts:375).runtime-cloudflare/src/do/durable-object-agent-base.ts:2764 (/interrupt sets the durable flag); contrast: runtime-cloudflare/src/do/durable-object-agent-base.ts:2086 (/resume DOES clear the flag — /start and /retry do not); runtime-js/src/run-loop.ts:4164-4195 (pollDurableInterruptFlag: the flag is cleared only when a running loop observes it); runtime-cloudflare/src/workflow.ts:1696-1701; runtime-dbos/src/workflows/shared.ts:1493-1504 (standard loop: checkInterruptFlag then finalize interrupted, no clear); runtime-dbos/src/lifecycle/interrupt.ts:67 (setInterruptFlag); contrast: runtime-dbos/src/lifecycle/resume.ts:375 (resume clears it — lifecycle/execute.ts never does); store-postgres/src/state/postgres-state-store.ts:2246-2257 (checkInterruptFlag is a plain SELECT, non-consuming); conformance: stop.mid-tool-ignores-abort on dbos-postgres — .continuation-completes: the new message's run ends interrupted; probe: the interrupt_flags row for the session was still present after the continuation ran; conformance: transcript.heal on dbos-postgres — .continuation-completes-on-interrupted: a new message on the stopped session ends interrupted without ever calling the modelsource, probepartly-inferredopen
CP-20S2Client-side handle failures swallowed: DOExecutionHandleImpl.interrupt/abort only warn on non-2xx; DOAgentExecutor handle abort ignores the response.runtime-cloudflare/src/do/do-clients.ts:1346-1376; runtime-cloudflare/src/do/do-executor.ts:747-753sourceverifiedopen
CP-21S2JS ExecuteOptions.abortSignal is ignored by execute / resume / retry (fresh controllers).runtime-js/src/js-agent-executor.ts:1805, 2374, 2894sourceverifiedopen
CP-22S2Temporal abort has no non-cancellable scope; persistTerminalState can be cancelled, leaving status active.runtime-temporal/src/workflow.ts:1459-1507; runtime-temporal/src/executor.ts:791-802sourcepartly-inferredopen
CP-23S1Parent interrupt does not reach blocking companions / children on Temporal, CF Workflows (persistent waits), or the DO park / waitForResult loops.runtime-temporal/src/workflow.ts:242-245; runtime-cloudflare/src/workflow.ts:3117-3170; runtime-cloudflare/src/do/persistent-companion-tools.ts:837-917, 992-1048; runtime-cloudflare/src/do/persistent-companion-tools.ts:983 pollChildToTerminal while(true) /status poll never checks parent abort nor interrupts child (live: stop during companion__spawnAgent took 31,367ms to unwind; companion__waitForResult up to 39s)source, live-reproverifiedopen
CP-24S1Disputed: source shows the DO interrupt aborting the in-flight model call (js-agent-executor.ts:1884-1900 → run-loop.ts:1254 → step-iterator.ts:831 → llm-vercel/src/vercel-adapter.ts:217), but the consumer measured 37–62 s to settle a stop pressed mid-step (possibly pre-dating its switch to /interrupt, or a long tool call, or /interrupt awaiting cleanup + onComplete). Settled by stop.* cells.consumer docs/operations/audio-stream-attach-latency.mdconsumer-documented, live-reproverifiedwithdrawn — Resolved by live repro 2026-09-23: DO in-flight provider abort fires ~1ms after /interrupt (loop unwinds ~+104ms, /interrupt responds ~+220ms incl. onComplete). The 37–62s consumer measurement is CP-23 (companion wait ignores abort).
CP-25S2SSEConnectionManager evicts slow / stale clients with a clean writer.close() (docstring assumes the client reconnects); DOStreamManagerClient maps EOF to done; handleChatStream emits [DONE]; useChat never reconnects → the UI stream silently ends mid-run. The browser protocol has no explicit end / fail frame.core/src/stream/sse-connection-manager.ts:1-12, 172, 202-208; runtime-cloudflare/src/do/do-clients.ts:443-466; live: 0 evict_slow_consumer/evict_stale events across ~45 connections in 5 runs; path real in code, not hit livesourceverifiedopen
CP-26S3DO SSE TransformStream highWaterMark: 256*1024 without a size function counts chunks (262,144), not bytes; backpressure eviction almost never fires.runtime-cloudflare/src/do/durable-object-agent-base.ts:3564-3566; core/src/stream/sse-connection-manager.ts (SSE_READABLE_HIGH_WATER_MARK); live: workerd identity TransformStream keeps writes pending until read; chunk-count HWM effectively mootsourceverifiedopen
CP-27S2No byte on attach: the first byte is the 15 s heartbeat; clients cannot distinguish "connected, waiting" from "connecting".ai-sdk/src/response/sse-builder.ts:66-94source, live-reproverifiedopen
CP-28S2DOStreamManagerClient / DOFrontendExecutor turn failures into null / [] / EOF / sequence 0; getHandle returns null on any failure; a new session's JSON-null snapshot is treated as a schema failure and the preflight silently falls back to startFromSequence: 0.runtime-cloudflare/src/do/do-clients.ts:418-421, 603-766, 860-884, 943-978, 1014-1037; runtime-cloudflare/src/do/durable-object-agent-base.ts:3684-3699source, live-reproverifiedopen
CP-29S2Resume options must be passed identically to prepareHelixChatRequest and prepareHelixReconnectRequest or the stream silently dies; values are frozen at transport construction, so later POSTs carry a stale X-Resume-From-Sequence.ai-sdk/src/client/prepare-helix-chat-request.ts; consumer lib/helix-stream-transport.ts:21-34source, consumer-documentedverifiedopen
CP-51S3handleChatStream GET with no messages on an active session whose current run is not running falls through to replayTerminalStream.ai-sdk/src/handler/handle-chat-stream.ts:486-499, 901-918peer-reportedunverifiedopen
CP-30S1DBOS retry ignores all options (_opts): checkpointId and message dropped; no truncation; startMessageCount hard-coded 0.runtime-dbos/src/lifecycle/retry.ts:52, 125-126sourceverifiedopen
CP-31S1Temporal resume never reads options.mode: with_message is dropped and from_checkpoint becomes continue.runtime-temporal/src/executor.ts:1160-1489; runtime-temporal/src/executor.ts:1343-1349 (the resume workflow input always carries newMessages: []; options are read only for fromSequence, :1210); runtime-temporal/src/executor.ts:1063-1070 (the getHandle() reconnect handle's resume(_options) calls this.resume(agent, sessionId) without its options, so even fromSequence is dropped; forwarding them would not help while resume() ignores mode); executor.ts:859-863 and runtime-temporal/src/handle.ts:221-225 (the execute() handle's and the factory handle's resume() throw) — so the documented handle.resume({ mode: 'with_message' | 'from_checkpoint' }) (runtime-temporal/README.md, Resume Interrupted or Suspended Execution) is honoured by no Temporal handle; not yet reachable by conformance: lifecycle.resume-with-message on Temporal never gets a genuinely interrupted session (CP-56 makes the first run end completed), so resume() takes the completed-session short-circuit (CP-60) before any mode handlingsourceverifiedopen
CP-32S2DBOS resume from_checkpoint silently becomes continue.runtime-dbos/src/lifecycle/resume.ts:447 (the only mode branch — from_checkpoint falls through to the continue path); runtime-dbos/src/lifecycle/resume.ts:512-531 (the workflow restarts from the current state; resumedFromCheckpointId is the session's latest state.checkpointId — opts.checkpointId is never read)sourceverifiedopen
CP-33S2Temporal and CF Workflows branch ignore checkpointId / messageIndex and copy all messages.runtime-temporal/src/activities.ts:6645-6650; runtime-cloudflare/src/steps.ts:676-704sourceverifiedopen
CP-34S2DO /resume accepts checkpointId / modifyState / appendMessages in its schema and returns 200 while ignoring them.runtime-cloudflare/src/do/durable-object-agent-base.ts:351-356, 2194-2197sourceunverifiedopen
CP-35S1CF Workflows resume() (and the handle.resume() path) truncates messages and cleans up the stream before a status CAS that passes no expectedVersion and accepts 'active' as a source state (TOCTOU). A losing concurrent resume can delete the winner's messages. Because 'active' is accepted, a resume against a live run, or two concurrent resumes, all pass the CAS: the live turn's uncheckpointed messages are deleted and a second workflow instance starts beside the running one. Separately, retry() calls cleanupToStep without startSequence.runtime-cloudflare/src/executor.ts:676 (cleanupOrphanedStaging), :695 (truncateMessages to latestCheckpoint.messageCount) and :707-711 (cleanupToStep) all run BEFORE :756-760 (compareAndSetStatus(["paused","interrupted","active"] → "active") with no expectedVersion); the comment at :748-755 claims a version pin that is never passed; runtime-cloudflare/src/executor.ts:1340-1386 (private resumeWorkflow — the handle.resume() path — has the same truncate-then-unpinned-CAS order); runtime-cloudflare/src/executor.ts:978 (retry() calls cleanupToStep(streamId, checkpoint.stepCount) with no startSequence; its CAS at :882 runs before the truncation, so retry() has no TOCTOU); inferred: because the CAS accepts 'active' as a source state, a resume() against a LIVE run (not only a losing concurrent resume) passes the CAS: it truncates the live turn's uncheckpointed messages and starts a second workflow instance next to the running one; two concurrent resumes both win; live: probe (batch-5 verification, unit level on origin/main 3560d6c72a: CloudflareAgentExecutor + a real InMemoryStateStore + a mock workflow binding): session active with [user TURN-1, assistant] under a messageCount-2 checkpoint and turn 2 input user TURN-2 appended (first step in flight); resume({mode:"continue"}) succeeded, deleted TURN-2 ([TURN-1, answer-1, TURN-2] → [TURN-1, answer-1]) and created agent__a__s1__resume__1 beside the live agent__a__s1. Two concurrent resume() calls on a paused session both fulfilled and created __resume__1 and __resume__2; related: the original title's "cleanupToStep without startSequence" now applies only to retry() (runtime-cloudflare/src/executor.ts:978); resume() passes resumingRun?.startSequence at :707-711 and :1369-1373. runtime-temporal/src/executor.ts:1245-1268 carries the same version-pin comment, but there the CAS runs before any truncation; live: peer-reported (batch-5 candidate #1, B2)source, probe, peer-reportedverifiedopen
CP-36S3CF Workflows retry flattens a multimodal retry message with JSON.stringify.runtime-cloudflare/src/executor.ts:986-989sourceverifiedopen
CP-37S2JS retry ignores abortSignal and usageStore (the DO passes one for nothing).runtime-js/src/js-agent-executor.ts:2893-2963sourceverifiedopen
CP-38S2JS / Temporal / CF retry accept a checkpointId belonging to another session.runtime-js/src/js-agent-executor.ts:2709-2718; runtime-temporal/src/executor.ts:1613-1622; runtime-cloudflare/src/executor.ts:905-914; runtime-cloudflare/src/executor.ts:714-746 (CFW from_checkpoint writes restoredState.sessionId from the checkpoint — a foreign checkpoint targets another session's row)sourceverifiedopen
CP-39S1DO: non-blocking companions never notify the parent (persistentAgents: undefined in the DO rewrite disables the notifier).runtime-cloudflare/src/do/durable-object-agent-base.ts:1557; runtime-js/src/run-loop.ts:480, 819sourceverifiedopen
CP-40S1DO: a companion suspended on a client tool / approval cannot be resumed (root ownership write fails silently), and separately the companion wait / pollChildToTerminal loops never end for a child that is paused OR interrupted: while (true) with no deadline and no parent-abort check, and the failure window resets on every successful poll, so only continuous fetch failure exits.runtime-js/src/run-loop.ts:1796-1802; runtime-cloudflare/src/do/persistent-companion-tools.ts:1038-1086, 861-866; runtime-cloudflare/src/do/persistent-companion-tools.ts:983-1050 (poll loop); runtime-cloudflare/src/do/persistent-companion-tools.ts:1068-1087 (mapRemoteStatusToRefStatus: interrupted maps to interrupted, paused falls through to the caller's running fallback); runtime-cloudflare/src/do/persistent-companion-tools.ts:105-131 (failure window; recordSuccess at 1007); core/src/tools/companion/helpers.ts:9 (TERMINAL_STATUSES excludes interrupted — intentionally: interrupted is resumable; the defect is the unbounded wait, not the set)sourcepartly-inferredopen
CP-41S1Temporal / CF Workflows: sendMessage to a running companion sets parent_message, which those runtimes treat as an interrupt; the message is never processed, returns delivered: true, and waitForResult can hang.core/src/orchestration/companion-tool-dispatch.ts:769-801 (handleSendMessage appends the message, sets the 'parent_message' interrupt flag on a running child at :782-784, and returns delivered: true); runtime-temporal/src/workflow.ts:903-911 (the per-step durable interrupt check calls finishInterrupted for ANY flag reason) and :827-838 (on resume, a set flag of any reason makes the deferred-finishWith path end the run interrupted); runtime-cloudflare/src/workflow.ts:1696-1705 (the per-iteration check hands ANY flag to handleInterruptFlag, :1536-1591, which ends the run interrupted); a flag seen inside the step instead surfaces as state.aborted (runtime-cloudflare/src/steps.ts:877-884) and is handled at runtime-cloudflare/src/workflow.ts:1782-1830 (discard staged, failAgentStream, finalizeRun 'failed', result 'cancelled') — neither branches on the reason; runtime-cloudflare/src/steps.ts:396-397, :435-436 (loadAgentState turns ANY interrupt flag into aborted: true / abortReason with no reason branch; 'parent_message' appears in runtime-cloudflare/src/steps.ts only in comments, :5185 and :5403); runtime-temporal/src/activities.ts:5827-5850 (consumeInterruptFlag returns { interrupted, reason } and clears the flag, with no reason branch); runtime-temporal/src/workflow.ts:903-911 (the step-boundary check ends the run via finishInterrupted for any reason, parent_message included); contrast: runtime-js/src/run-loop.ts:673-699 (on reason 'parent_message' the loop reloads messages via getAllMessages, clears the flag and continues); runtime-dbos/src/steps/execute-companion-tool.ts:1156-1163 (not a correct contrast, and not claimed in this title: DBOS sends the message to the child's inbox FIRST, then sets parent_message for a running child); runtime-dbos/src/workflows/shared.ts:1492-1502 (the running turn exits 'interrupted' on any flag), runtime-dbos/src/workflows/persistent-workflow.ts:546-550, :614-618 (the recv loop clears the flag and runs the inbox message as the next turn). So DBOS does deliver the message, but on every parent_message it cuts the in-flight turn short as 'interrupted' — a narrower partial defect of the same kind, unlike JS, which continues the turn. CP-42 covers DBOS after a continuation, a different situation; related: likely fix direction (B2-core's B13 review, peer-reported): on reason 'parent_message', clear the flag, reload messages and continue the loop, as JS doessource, peer-reportedpartly-inferredopen
CP-42S1DBOS: after a continuation, sendMessage targets the dead initial workflow while interrupting the live continuation child; returns delivered: true.runtime-dbos/src/steps/execute-companion-tool.ts:922, 1104, ~1165sourcepartly-inferredopen
CP-43S1JS / DO: sendMessage to an interrupted companion resumes it with no history (state built from loadState, which has no messages).runtime-js/src/run-loop.ts:4604 (resumeChildAgent), :4650 (the child state is loaded with stateStore.loadState — a SessionState, which carries no messages: core/src/types/session.ts:88), :4675-4678 (it is cast straight to AgentState with no getAllMessages hydration) and :4685-4691 (runSubAgentLoop runs the child on it); call site :5217; contrast: runtime-js/src/run-loop.ts:4937 (the COMPLETED-child continuation path reads the preserved history with getAllMessages before running the child); inferred: messages is undefined (not []) on the loaded state, so the resumed loop may throw on a state.messages access (e.g. core/src/orchestration/step-iterator.ts:651) rather than run with an empty history; the resume catch (run-loop.ts:4698-4719) would then mark the ref failed. Either way the child never sees its history — the exact failure shape is not traced end to end; related: CP-68 — on JS this path is currently unreachable: a parent-interrupted companion's session and parent ref are both recorded 'failed' (CP-68), so handleSendMessage throws "not active (status: failed)" (core/src/orchestration/companion-tool-dispatch.ts:706-727) and never returns resumeMetadata; reaching resumeChildAgent (companion-tool-dispatch.ts:770, :791-799) also needs a ref status of 'interrupted', which no runtime-js updateSubSessionRef call writes — so a CP-68 fix must record 'interrupted' for CP-43 to become observablesourcepartly-inferredopen
CP-44S2JS: a companion suspended on a client tool is recorded completed with no output.runtime-js/src/run-loop.ts:4496-4508, 4812-4818sourceverifiedopen
CP-45S2Remote child interrupted / paused with any (stale) output is treated as success.core/src/orchestration/remote-sub-agent-dispatch.ts:607-618; runtime-js/src/execution/remote-subagent-executor.ts:643-656, 762-772; runtime-dbos/src/steps/execute-remote-subagent.ts:558-566sourceverifiedopen
CP-46S2Temporal: startChild failure for companion spawn / resume / continue is only logged; the tool reports success; the ref stays running.runtime-temporal/src/workflow-dispatch.ts:770, 801, 902sourceverifiedopen
CP-47S1CF Workflows: children notify agent__{type}__{sessionId}, but a resumed parent runs as …__resume__N, so completions go to a finished instance and the parent times out.runtime-cloudflare/src/workflow.ts:2546, 3320, 3602, 3735 (parentWorkflowId: agent__AGENTTYPE__SESSIONID at each child-dispatch site); runtime-cloudflare/src/executor.ts:786sourcepartly-inferredopen
CP-48S2DBOS: a missing child definition throws inside dispatch and fails the whole parent instead of returning a tool error.runtime-dbos/src/workflows/standard-workflow.ts:473-481; runtime-dbos/src/workflows/shared.ts:2157sourceverifiedopen
CP-49S2DBOS: a submit whose owning workflow cannot be resolved is acknowledged already_completed without delivery.runtime-dbos/src/dbos-agent-executor.ts:534-536; runtime-dbos/src/lifecycle/resolve-owner-workflow.ts:27-40sourceverifiedopen
CP-50S2Temporal terminatePersistentChild is a no-op; DBOS terminate is flag-only (child finalizes as completed after idle timeout).runtime-temporal/src/activities.ts (~6044); runtime-dbos/src/steps/execute-companion-tool.ts:263-273; runtime-dbos/src/workflows/persistent-workflow.ts:680-694sourceverifiedopen
CP-52S1Client-tool submit landing during a paused run's teardown is dropped: ensureExecutionContinues returns early on isExecuting=true ('the live run will observe the state change') but the run already decided to suspend; handleResumePath's /resume 409s and the fallback returns data-resume-rejected; session stranded 'paused' with the result submitted, never resumed. Deterministic on a fresh session's first client-tool round (3/3 on two articles); delaying the POST 1.5s → 3/3 complete.runtime-cloudflare/src/do/durable-object-agent-base.ts:644-653; ai-sdk/src/handler/handle-chat-stream.ts ~748-800; core/src/orchestration/step-iterator.ts:1745 then :1765 (onAgentSuspended is awaited after persistSuspension, so the session reads paused for the whole hook — a slow hook widens the window); runtime-cloudflare/src/do/durable-object-agent-base.ts:1061-1073 (recoverStrandedExecution only recovers 'active', so the stranded 'paused' session is never re-driven); live: live-repro (batch-6, origin/main e37d676945): executor tier, cf-do-d1 in the e2e workerd pool — lifecycle-hooks-parity scenario 2 with a 1 s onAgentSuspended delay fails 4/4 (expected 'suspended_client_tool' to be 'completed'), session paused with tc-1 submitted, still paused 5 s later; 3/3 pass without the delay. CI: jobs 16758514869, 16759542448, 16760323616, 16764206430; main pipelines 2887177079 / 2886530232 (peer-reported, B8); conformance: submit.race-a on cf-do (host tier) — with the DO held inside onComplete after the client-tool pause (.window-open: a caller inside the hook AND isExecuting true), the auto-send's /resume 409s ('Agent already running') and the caller gets only data-resume-rejected (.submit-accepted, .continuation-reaches-caller), and the submitted answer never resumes the run: .run-completes does not settle within 15000ms, .not-stranded-paused sees status paused with the answered call still pending, .model-sees-result-once never gets a second model call; contrast: submit.no-race on cf-do (the same round answered 1500ms after the pause, no hold; also 3/3 at 0ms) — the run completes server-side (.run-completes, .model-sees-result-once and .not-stranded-paused pass): only the held window strands it; contrast: submit.no-race passes on js-chat-host (host tier); submit.race-a is a declared gap there (no host hook to hold, no execution flag); related: CP-53 (the no-hold control submit.no-race fails the same two caller-side ids against an IDLE DO — a different mechanism, ledgered there); conformance: race A's caller-side ids (.submit-accepted, .continuation-reaches-caller) are CP-52's while the run is held; after CP-52 is fixed they may fail under CP-53 until CP-53 is fixed — re-attribute from the observed run, never pre-assign (the same rule for the -client variant); conformance: submit.race-a-client on cf-do (client tier, the real useHelixChat / DefaultChatTransport auto-send) — with DI-20 fixed (!268) the Chat now answers the tool: .submitted and .window-open pass, the auto-send's continuation is only data-resume-rejected (.submit-accepted, .continuation-reaches-caller), and the run is stranded paused with the answer submitted (.run-completes does not settle within 15000ms, .not-stranded-paused, .model-sees-result-once) — race A's host shape, identical 5/5. Before !268 these ids were masked by DI-20 (the Chat errored on finish-step before the tool part). A CP-52 fix MR has the same caller-side caveat here as for the host cell; conformance: for CP-52's fix MR — the caller-side 409 → data-resume-rejected in race A comes from the same /resume fallback CP-53 describes, so a CP-52 fix alone will likely not turn submit.race-a's .submit-accepted / .continuation-reaches-caller green. The fix MR cannot name submit.race-a|cf-do|host in provenBy until those two ids are re-attributed from its observed run (likely to CP-53); moving them is part of that MR, not a sign the fix failed; conformance: race A's hold bracket releases on the auto-send's response headers (.submit-returned). A fix that withholds that response until the DO is idle would deadlock against the hold and show as a .submit-returned timeout — a harness bracket to change (release on the request's arrival at the host), not a regression (FU-CONF-RACE-A-RELEASE-ON-HEADERS)source, live-repro, peer-reportedverifiedopen
CP-53S1Resume-path fallback reads state once: when the submit auto-resumes an idle DO (producer claim held), the explicit /resume 409s and the single loadState still shows pendingClientToolCalls → rejectedOnlyResponse, while the resumed run continues with zero SSE listeners and emits the next client-tool call nobody receives. Polling the gate (100ms up to 5s) cleared in ~105ms and fixed it (3/3).ai-sdk/src/handler/handle-chat-stream.ts ~748; ai-sdk/src/handler/handle-chat-stream.ts ~774-800; conformance: submit.no-race on cf-do (host tier; no hold, answered 1500ms after the pause) — the auto-send's /resume 409s ('Agent already running') and the caller's continuation is only data-resume-rejected (.submit-accepted, .continuation-reaches-caller), while the woken run completes server-side (.run-completes passes); conformance: submit.race-a on cf-do (host tier) fails the same two caller-side ids under the held window, where the DO is still executing — that is CP-52's mechanism and is ledgered there, not here; contrast: submit.no-race passes on js-chat-host (host tier); related: CP-52; conformance: race A's caller-side ids (.submit-accepted, .continuation-reaches-caller) are CP-52's while the run is held; after CP-52 is fixed they may fail under CP-53 until CP-53 is fixed — re-attribute from the observed run, never pre-assign (the same rule for the -client variants); conformance: submit.no-race-client on cf-do (client tier, the real useHelixChat / DefaultChatTransport auto-send, no hold) — with DI-20 fixed (!268) the Chat answers and auto-sends (.submitted passes), the continuation is only data-resume-rejected (.submit-accepted, .continuation-reaches-caller), and the woken run completes server-side (.run-completes, .model-sees-result-once, .not-stranded-paused pass) — the host control's shape, identical 5/5. Before !268 these ids were masked by DI-20; contrast: submit.no-race-client passes on js-chat-host (client tier)source, live-reproverifiedopen
CP-54S1Abandonment path auto-resumes the OLD run on the DO: abandonPendingTools submits client_tool_abandoned, the submit poke resumes the old run which continues its plan and suspends on a new client tool; execute(new message) then fails 'session is already running with unresolved client-tool calls' — new user messages dropped (1 of 4 persisted live).ai-sdk/src/handler/handle-chat-stream.ts:446-480; runtime-cloudflare/src/do/durable-object-agent-base.ts handleSubmitToolResult → ensureExecutionContinueslive-repro, sourcepartly-inferredopen
CP-55S1Client hydrate re-executes already-applied client tools: after reload the snapshot shows the session 'paused' with input-available parts; useResumeClientTools (per-mount optimistic-done set, empty after reload) re-runs the handlers and auto-sends submit-message POSTs, although the DO already recorded those results (submitted before reload, stranded by CP-52). The snapshot does not expose submitted-but-unconsumed results, so the client cannot distinguish 'pending, never executed' from 'executed, submitted, run stranded'. Live: insertAfter applied twice.ai-sdk/src/react/use-resume-client-tools.ts (hydrate dispatch, per-mount done set); runtime-cloudflare/src/do/durable-object-agent-base.ts /snapshot (pendingClientToolCalls payload); ai-sdk/src/handler/handle-chat-stream.ts orphan recovery (inherited_success) + unknown-tool-call on second POSTlive-repro, sourceverifiedopen
CP-56S1Temporal drops a pending interrupt when the in-flight step is also the terminal step: the durable interrupt flag is only sampled at the TOP of the main loop, before starting the next step (workflow.ts:903-907); a step that completes the run via shouldStop returns straight from the completion branch without ever re-checking the flag (workflow.ts:1559-1567). Combined with CP-16 (no abortSignal reaches the in-flight LLM call), a stop pressed while the last step is running is silently discarded and the run persists 'completed', not 'interrupted'. Reproduced by conformance stop.mid-llm-call on temporal-memory/temporal-redis/temporal-postgres: outcome() and observe().persisted both report 'completed' after release.runtime-temporal/src/workflow.ts:903-907; runtime-temporal/src/workflow.ts:1559-1567; conformance: stop.mid-llm-call on temporal-memory/temporal-redis/temporal-postgres — stop.outcome-interrupted and stop.persisted-interrupted both observe status completed; inferred: the lost interrupt flag is never cleared on Temporal (only consumeInterruptFlag clears it, runtime-temporal/src/activities.ts:5843) -> likely interrupts the next run on the session (CP-19 pattern); conformance: lifecycle.resume-with-message on temporal-memory/temporal-redis/temporal-postgres — lifecycle.resume-setup-interrupted observes status completed instead of interrupted (same held-terminal-step pattern via stop())source, probeverifiedopen
CP-57S1DBOS ActiveHandle.result() hardcodes status 'failed' whenever dbosHandle.getResult() throws, without checking the session's persisted terminal status. interruptImpl cancels the workflow via cancelWorkflow whenever it holds a workflowId (making getResult() reject); the session row then reads 'interrupted' only because a forced CAS fallback (active → interrupted, after a fixed delay) sets it — not via any graceful interrupt path (there is none, see CP-17). So the persisted row says 'interrupted' but the caller-facing AgentResult from the same result() call reports status: 'failed'. Reproduced by conformance stop.mid-llm-call on dbos-postgres: outcome() reports 'failed' while observe().persisted reports 'interrupted'.runtime-dbos/src/handles/active-handle.ts:162-177; runtime-dbos/src/lifecycle/interrupt.ts:65-124 (inside if (workflowId): 83-92 cancelWorkflow, 108-123 forced CAS fallback active → interrupted); conformance: stop.mid-llm-call on dbos-postgres — stop.outcome-interrupted observes status failed while stop.persisted-interrupted passes; related: CP-17 (DBOS interrupt calls cancelWorkflow immediately); conformance: lifecycle.resume-with-message on dbos-postgres — lifecycle.resume-setup-interrupted observes status failed instead of interrupted; conformance: stop.mid-tool-ignores-abort on dbos-postgres — .outcome-interrupted observes status failed ("Workflow ... has been cancelled") while .persisted-interrupted passessource, probeverifiedopen
CP-58S2runtime-js retry() without a message on a session that failed at its first LLM call throws 'No user message found to retry with.': the checkpoint written on failure has messageCount >= 1, so the candidate slice is empty, contrary to the code's stated intent to preserve the original triggering message.runtime-js/src/js-agent-executor.ts retry() (message recovery via checkpoint.messageCount truncation, 'preserve the original triggering message' fallback); runtime-js/src/run-loop.ts persistTerminalState (invoked from onFailed) mints the failure checkpoint whose messageCount already includes the triggering message; conformance: canary drop-retry-message (lifecycle.retry-message on js-memory) — retry({}) throws instead of completing, and the strip-retry-marker-from-model-input canary on the same scenario shows the fault alone (without this defect in the way) only fails lifecycle.retry-message-reaches-model; live: probe (batch-6 verification, origin/main e37d676945, runtime-js + InMemoryStateStore): also reproduces on a continued turn — turn 1 completes, turn 2's first LLM call fails, and retry() without a message throws because the failure checkpoint's messageCount (3) already includes TURN-2; on a store whose saveState mints no checkpoint (D1-like) the same retry() completes and the model receives TURN-2, so on the two stores probed the defect needs a store whose failure path mints a checkpoint (memory does; the D1-like store does not). Postgres is inferred to behave like memory from its checkpoint-minting saveState, not probed; Redis is not established either way, because its saveState pointer names a checkpoint that was never written (RM-38)source, probeverifiedopen
CP-59S2core step-iterator commits 'completed' on the terminal (no-tool) branch without re-checking the abort signal: when the adapter returns normally after abort, a stop that reported ok is silently lost and the run ends completed (runtime-js / Cloudflare DO). JS/DO counterpart of CP-56.core/src/orchestration/step-iterator.ts terminal-without-tool-execution branch (structured_output / __finish__ / end-of-turn text) — commits and returns completed without checking deps.abortSignal.aborted; the abort re-check exists only on the tool-execution/commit branch; conformance: canary ignore-abort-signal (stop.mid-llm-call on js-memory) — probe: stop() reports {"ok":true}, outcome() reports completed, observe().persisted reports completed; related: CP-56 (same class on Temporal); severity S2 (not S1): the committed-completed outcome needs the adapter to return a NORMAL (non-error) result after the abort. The shipped adapter, @helix-agents/llm-vercel, forwards abortSignal (llm-vercel/src/vercel-adapter.ts:217) and on an aborted call returns an error StepResult {type: error, shouldStop: true, stopReason: error} (:530-535, cancellation classified framework_cancelled at :508-523); step-processor.ts:410-428 turns that into statusUpdate failed + isTerminal, step-iterator.ts:1130-1136 returns kind failed before the completed commit (:1138+) is reached, and runtime-js run-loop.ts:993-1058 remaps failed + aborted + interrupted to interrupted. So the stop is honoured with llm-vercel; the defect bites only with an adapter that ignores the signal or when the provider completion races the abortsource, probeverifiedopen
CP-60S1Temporal resume() on a completed (or failed) session silently returns the previous result for EVERY resume mode — with_message drops the message, continue never reaches the model, and with_confirmation / from_checkpoint are equally ignored: resume() short-circuits to attachToCompletedWorkflow whenever the persisted status is completed/failed, before options.mode (or any mode-specific option) is read, so no workflow starts and the model is never called — while JS, DBOS, CF Workflows and CF DO all throw (JS/CF: 'Cannot resume: agent already completed').runtime-temporal/src/executor.ts:1230-1243 (terminal short-circuit: status completed/failed → attachToCompletedWorkflow; the only option read before it is the stream cursor fromSequence at :1210 — options.mode is never inspected anywhere in resume(), see CP-31; failed branch: source-only); runtime-temporal/src/executor.ts:1500-1537 (attachToCompletedWorkflow resolves the prior status/output from persisted state); runtime-js/src/js-agent-executor.ts:2129-2134 (throws 'Cannot resume: agent already completed' / 'agent has failed'); runtime-dbos/src/lifecycle/resume.ts:133-311 (CAS ['interrupted','paused'] → active; a completed session throws 'not in a resumable state'); runtime-cloudflare/src/executor.ts:649-654 and runtime-cloudflare/src/do/do-executor.ts:466-471 (CF Workflows / CF DO throw 'Cannot resume: Agent already completed' / 'Agent has failed'); conformance: lifecycle.resume-with-message (mode with_message) on temporal-memory/temporal-redis/temporal-postgres — the first run ends completed (CP-56), resume() then reports completed with no new model call (lifecycle.resume-message-reaches-model: "no new model call for the session (had 1 before, 1 after)"; lifecycle.resume-completes passes); conformance: lifecycle.resume-continue (mode continue) on the same three backends — lifecycle.continue-reaches-model fails the same way ("no new model call (1 before, 1 after)"), confirming the short-circuit is mode-agnosticsource, probeverifiedopen
CP-61S2agent-server's chat routes never pass the HTTP request to the chat handler, so Last-Event-ID / X-Resume-From-Sequence / X-Existing-Message-Id never reach handleChatStream and its request-size gate never runs: every Node reconnect replays from the run start under a new message id.agent-server/src/http/handler.ts:955-961 (POST /chat passes only sessionId/messages/userId/metadata); agent-server/src/http/handler.ts:966-990 (GET /chat/{id}/stream passes messages: [] only); agent-server/src/types.ts:31-36 (the chatHandler type has no request); ai-sdk/src/handler/handle-chat-stream.ts:551-582 (resolveResumeHints reads headers only from request); ai-sdk/src/handler/handle-chat-stream.ts:198-217 (size gate needs request); related: CP-29 (client-side resume values frozen)sourceverifiedopen
CP-62S1Stop while a server tool is running: Temporal and DBOS never fire the tool's abortSignal (each tool gets a fresh, never-aborted AbortController, and interrupt only sets a flag / cancels the workflow at the next step boundary), so the run cannot end until the tool returns on its own. A cooperative tool waiting on its signal hangs the run indefinitely (outcome never settles). On DBOS, once the tool does return, the cancelled workflow never commits the step's tool results, so persisted history keeps the assistant tool_use calls with no tool_result.runtime-temporal/src/activities.ts:1254, 1329 (executeToolActivity: new AbortController() handed to tool.execute as abortSignal, never aborted; the activity cancellationSignal is not bridged); runtime-temporal/src/workflow.ts:967-988 (regular tools awaited via Promise.all of executeToolActivity; no cancellation scope tied to interrupt); runtime-temporal/src/workflow.ts:245-248 (the interrupt signal only sets inProcessInterrupt), 903-907 (sampled at the top of the next loop iteration); runtime-temporal/src/executor.ts:804-808 (interrupt() only signals the workflow); runtime-dbos/src/workflows/shared.ts:2109-2117 (dispatcher.executeTool called with no signal); runtime-dbos/src/steps/execute-tool.ts:261, 278 (args.signal ?? new AbortController().signal handed to the tool); runtime-dbos/src/lifecycle/interrupt.ts:83-92 (interrupt cancels the workflow; a running step is not preempted), 94-100 (every @DBOS.step() after cancellation throws); runtime-dbos/src/workflows/shared.ts:1836 (assistant message with the tool calls appended BEFORE the tools run), 2148-2149 (tools awaited), 2320 and 2367 (promoteStaging / tool-result appendMessages — steps that throw once the workflow is cancelled); conformance: stop.mid-tool-honors-abort on temporal-memory/temporal-redis/temporal-postgres/dbos-postgres — .outcome-settles does not settle within 30000ms after stop() returned ok (dependents fail their precondition); the lane drains only because the scenario releases the tool in its finally; conformance: stop.mid-tool-ignores-abort on dbos-postgres — .transcript-paired: persisted history unpaired [slow, note] (the note call had completed); conformance: stop.mid-tool-honors-abort's stop lands after slow is running AND the co-emitted note has run (.tool-reached). Before that (FU-CONF-CP62-DBOS-SETTLE-TIMING, CI job 16756806141) a DBOS note step not yet started at the cancel threw 'cancelled' through its three retries, failing the run in ~3s with a different id set; re-checked 5/5 identical on dbos-postgres and the three temporal backends (this fragment unchanged); contrast: stop.mid-tool-honors-abort and stop.mid-tool-ignores-abort both pass on js-memory/js-redis/js-postgres (runtime-js aborts the running tool); conformance: stop.mid-tool-ignores-abort passes on temporal-memory/temporal-redis/temporal-postgres (once the tool returns, the flag is honored at the next step boundary); related: CP-16 (the same missing abort for an in-flight LLM call), CP-17 (DBOS interrupt cancels the workflow immediately), CP-57 (DBOS outcome misreported failed), CP-19 (DBOS interrupt flag outlives the stopped run)source, probeverifiedopen
CP-63S3CF Workflows runs a step's finishWith calls (phase 2) BEFORE dispatching the sub-agents co-emitted in the same step, while core (JS / DO) runs sub-agents in phase 1 and finishWith strictly after, so on CF Workflows a finishWith tool observes state from before its sibling sub-agents ran. Completion is still deferred until the sub-agents return, so the transcript stays paired; only the ordering (and what the finishWith tool can see) diverges.runtime-cloudflare/src/workflow.ts:2377-2443 (phase 2: finishWith tools executed); runtime-cloudflare/src/workflow.ts:2446 (sub-agent dispatch only AFTER phase 2); contrast: core/src/orchestration/step-iterator.ts:1214-1215 (phase 1 = regular server tools + sub-agents in parallel; phase 2 = finishWith, deferred to step 9); inferred: the user-visible effect needs a co-emitted sub-agent whose completion changes state the finishWith tool reads; not yet observed; related: RM-29 (terminal-step ordering and pairing)source, peer-reportedpartly-inferredopen
CP-64S2DBOS: submitToolResult wakes the waiting workflow itself (DBOS.send → recv), so a caller's documented follow-up resume() races it: once the workflow has cleared its pending entry, resume() throws AgentAlreadyRunningError (or 'not in a resumable state' once it completes), and the executor offers no way to attach to the continued run (the execute handle's result is cached as suspended; getHandle's result is a one-shot status snapshot). The continuation's outcome is unobservable to the caller.runtime-dbos/src/dbos-agent-executor.ts:700 (submitToolResult → DBOS.send(wfId, payload, submit.toolCallId, idempotencyKey) into the workflow still blocked on DBOS.recv — the submit itself wakes the run); runtime-dbos/src/lifecycle/resume.ts:171-179 (the accepted-race comment: once clearPendingClientToolCallStep ran, resume() sees no pending calls and the caller 'should retry or fall back to awaiting the original handle.result()'), :311 (throw new AgentAlreadyRunningError(sessionId, 'active') when the session is active with no pending calls), :314-318 ('not in a resumable state' once the woken run completed); runtime-dbos/src/handles/active-handle.ts:126-133 (result() memoises _resultPromise — the execute handle that reported suspended_client_tool keeps returning it, so the documented fallback of awaiting it never sees the continuation); runtime-dbos/src/handles/reconnect-handle.ts:77-117 (getHandle's ReconnectHandle.result() is a one-shot loadState snapshot: an 'active' session reads as status 'failed' — no way to await the continued run); conformance: smoke.client-tool-completes on dbos-postgres — the responder submits, waits for the woken run to clear the pending entry (clientToolSubmitWakesRun: the legal order a late caller hits), then resume()s: .completes fails with outcome threw 'Agent session … is already running (status: active)'; .llm-input-has-result and .llm-input-paired fail only as its preconditions (failPreconditions). Without the wait the same cell failed intermittently (1 of 6 full DBOS matrix runs) when resume() lost the race.; inferred: while CP-64 is open the DBOS continuation's correctness is unobserved — the smoke's model-input ids fail only as preconditions of .completes; inferred: the harness waits only for the pending entries to clear, so resume() may land while the woken run is still active OR after it completed — a fix must make resume() attach (or return the continued outcome) in BOTH cases, otherwise soak verdicts become mixed; related: CP-53 (the DO host-level race B), CP-49source, probeverifiedopen
CP-65S2Temporal resume() on a session that is still suspended with nothing new to drain never settles: the workflow re-suspends via the hasMoreSuspended early return without emitting a new suspension_marker, while the resume handle only observes chunks after the stream’s latest sequence, so handle.result() waits forever.runtime-temporal/src/workflow.ts:640-648 (resume-mode path — input.mode === 'resume' guards this block from :268 — after applyResultsAndReload's drain, the drainResult.hasMoreSuspended branch re-suspends by deriving the suspension status/payload straight from durable state via deriveSuspensionFromState and returns, without ever calling activities.emitSuspensionMarker); runtime-temporal/src/workflow.ts:1090-1095 (contrast: the normal suspend branch calls activities.emitSuspensionMarker(streamId, suspensionKind, ...) BEFORE the durable commit — the only place in the workflow a suspension_marker stream chunk is written; the hasMoreSuspended early return at :640-648 skips this call entirely); runtime-temporal/src/executor.ts:1196-1224 (C-3: when the caller omits fromSequence, resume() probes the stream and defaults fromSequence to its CURRENT latestSequence rather than 0 — 'pins the new handle to observe ONLY chunks emitted by the resumed workflow' — so a resumed run that emits no new chunk leaves the handle nothing to observe); runtime-temporal/src/executor.ts:1394-1399 (resume() returns createTemporalAgentHandleV7({ ..., fromSequence, ... }) — the stream-observed v7 handle, not the native workflowHandle.result() that createActiveHandle uses for execute()); runtime-temporal/src/handle.ts:115-146 (observeAndProject: for await (const chunk of reader) over the reader created from fromSequence; the returned AgentResult resolves only when a terminal output/error/run_interrupted/suspension_marker chunk is observed on the stream, or the reader drains and fallbackStatusQuery runs); runtime-temporal/src/handle.ts:185-188 (result() returns the cached projection or awaits observeAndProject()’s promise — the only resolution path for this handle; with the stream still open and no chunk emitted after fromSequence, the for-await loop at :126 never yields and result() never resolves); live: observed by the B2-core / B5 bucket implementer while writing Temporal from_checkpoint integration tests (resume with no newly submitted client-tool results hangs handle.result()); inferred: no conformance cell yet — scenario: suspend on a client-executed tool, resume({mode: continue}) without submitting, expect terminal suspended under a named settle (now expressible with the responder / start opts)source, peer-reportedverifiedopen
CP-66S2A denied approval-gated tool call is persisted in two different shapes depending on the runtime: failure-shaped ({"error":...}, helixToolFailed true) on CF Workflows and DBOS, but a success-shaped string (helixToolFailed false) on JS / CF DO and Temporal — so the same user decision renders differently in the UI and reads differently to the model. Which shape is correct is a contract decision.runtime-cloudflare/src/steps.ts:4475-4480 (deny path builds the tool_result via createToolResultMessage({ success: false, error: 'Tool call was not approved by the user' }); core createToolResultMessage stamps metadata.helixToolFailed = !success and content JSON.stringify({error}) — failure-shaped); runtime-dbos/src/workflows/shared.ts:934-935 (dispatchApprovalGatedTool deny branch returns { toolCallId, ok: false, error: denyMsg }; dispatched as a phase-1 call at :2094, and the workflow body's phase-1 append at :2331-2338 threads success: res.ok into createToolResultMessage — failure-shaped, same as CF Workflows); runtime-js/src/js-agent-executor.ts:142 (APPROVAL_DENIED_TOOL_RESULT = 'Tool call was not approved by the user'), :973 (orphan-drain deny branch sets result = APPROVAL_DENIED_TOOL_RESULT while outcome.ok is true), :1166-1175 (createToolResultMessage({ result, success: outcome.ok, ... }) — success: true, content is a JSON-stringified plain string, no error wrapper — helixToolFailed false); runtime-temporal/src/activities.ts:3475-3485 (deny path — since !257 / DI-17 built via createToolResultMessage({ success: true, result: 'Tool call was not approved by the user' }): content JSON.stringify of the plain string and metadata.helixToolFailed explicitly false, the same success-shaped string as JS; the comment there states the parity with runtime-js is deliberate); runtime-cloudflare/src/do/durable-object-agent-base.ts:79 (durable-object-agent-base.ts imports JSAgentExecutor from @helix-agents/runtime-js — CF DO embeds the JS runtime's deny path, so it inherits the success-shaped string, not the CF Workflows failure shape); related: CP-11 (the DO discards approval decisions), MA-20 (DBOS approval gate fails open); inferred: no scenario yet; once the contract picks a shape, an executor-tier scenario (approval-gated tool, deny via the responder/submit-approval) asserting the persisted result shape witnesses it on every backendsource, peer-reportedverifiedopen
CP-67S1CF Workflows: resume() after an interrupt never reactivates the root stream the interrupt ENDED, so the resumed run cannot write to it ('Stream is ended, cannot write'). executor.resume() throws to its caller from its own run_resumed write — AFTER it has already started the resume workflow instance. That instance (and every handle.resume(), which skips the executor write) hits the same rejection inside its first LLM step. The step reports it as a recoverable error with no tool calls, so the workflow takes the 'stopped' branch: it marks the session failed with the misleading error "Agent stopped without completing (no structured output)", and its fail-stream step then fails again on the same ended stream (retried under the default step.do policy). The outer catch re-marks the session failed and its fail-stream fails a third time, and neither finalize-run step runs. The resumed turn (and any with_message input) is lost. The loss is silent on the stream, which never shows the resumed turn or its failure. It is surfaced only to a direct executor.resume() caller, as that call's rejection; a handle.resume() caller learns of it only when the handle's result() reports the run failed.runtime-cloudflare/src/workflow.ts:1561-1564 (end-stream-interrupted: an interrupted root run ENDS its stream); runtime-cloudflare/src/executor.ts:649-654 (resume() rejects completed / failed sessions, so interrupted → resume is the supported follow-up path) and :822 (initStream only); resumeWorkflow (handle.resume) :1379-1383 (CAS from paused / interrupted / active) and :1434 (initStream only); :1013-1023 (retry() is the ONLY reactivateStream call in runtime-cloudflare/src outside tests); store-cloudflare/src/stream-durable-object.ts:364-367 (handleWrite rejects every write to a non-active stream with HTTP 400 "Stream is ${status}, cannot write"); ended → active only via handleReactivate (:885-896) or handleReset (:1128-1139, which also deletes all chunks — runtime-cloudflare never calls resetStream); handleInit :420-429 only inserts the init_claimed marker; store-cloudflare/src/durable-stream.ts:557-563 (the writer turns the 400 into a thrown StreamConnectionError), :510-524 (initStream = POST /init — does not reactivate); runtime-cloudflare/src/executor.ts:786-806 (resume() creates the __resume__<n> workflow instance) then :825-840 (writes run_resumed via withStreamWriter to the still-ended stream, so resume() rejects after the instance was already created); runtime-cloudflare/src/steps.ts:988-993 (executeAgentStep's writer on state.streamId) and :1652-1663 (its catch turns ANY throw inside the step body, try at :1006, into { error, recoverable: true, toolCallCount: 0 } — the workflow's llmStepRetry is not exercised); runtime-cloudflare/src/workflow.ts:1972-2027 (not-continuing, zero-tool-call branch: mark-stopped records "Agent stopped without completing (no structured output)", then fail-stream-stopped at :1988 → steps.ts:3722-3742 failAgentStream writes an error chunk to the ended stream before failStream, so the step throws after its default retries and finalize-run-stopped (:2012) never runs); runtime-cloudflare/src/workflow.ts:3927-3968 (outer catch: error-mark-failed, then error-fail-stream throws again on the same write; the enclosing try swallows it and skips finalize-run-catch :3957); contrast: runtime-js/src/js-agent-executor.ts:2417-2427 (runtime-js resume() resumes a paused stream and reactivates an ended / failed one before writing); related: CP-76 (the Temporal twin: resume() likewise calls only initStream, runtime-temporal/src/executor.ts:1386, while its interrupted terminal path ends the stream, runtime-temporal/src/activities.ts:5095-5098; B2-core reproduced it outside the suite. A conformance stop there can end the run completed instead, CP-56, and resume() then short-circuits, CP-60); related: runtime-dbos/src/lifecycle/resume.ts:459 (DBOS resume() calls only initStream, which does not reactivate); whether the DBOS stream is ended at that point depends on its interrupt path (shared.ts:1493-1502 ends it on a graceful flag check; interruptImpl cancels the workflow — CP-17 / CP-57); not assessed; live: B2-core (peer, real workerd via vitest-pool-workers) observed the resumed run failing on the ended stream for resume with continue and with_message; inferred: which write fails first inside the step depends on the LLM adapter (llm-vercel converts a callback throw into a type 'error' step, llm-vercel/src/vercel-adapter.ts:507-535; MockLLMAdapter does not await onTextDelta, core/src/llm/mock-adapter.ts:340-346); any throw that escapes to the steps.ts:1652 catch takes the path above. The stall length is Cloudflare's default step.do retry policy (not defined in this repo)source, peer-reportedpartly-inferredopen
CP-68S2runtime-js: a persistent companion whose parent is interrupted via handle.interrupt() is recorded 'failed', not 'interrupted', in both its own session and the parent's SubSessionRef. The in-process child loop is not given the parent's interrupt state, so the interrupt reaches it only as a hard abort through the parent's AbortController. The companion therefore can never be resumed or continued: a later companion__sendMessage returns a tool error ("not active (status: failed)") instead of resuming it, and spawnAgent under the same name deletes its session and starts fresh, losing the companion's memory.runtime-js/src/js-agent-executor.ts:1885-1894 (handle.interrupt sets interruptState.interrupted AND aborts the run's AbortController, :1924 passes that controller's signal as the run's abortSignal); runtime-js/src/run-loop.ts:4796-4804 (spawn) and :5069-5076 (continuation): the persistent child gets its own AbortController, aborted when the parent abortSignal aborts — the only link from a parent interrupt to the child; runtime-js/src/run-loop.ts:1375-1461 (runChildLoop calls runLoop with abortSignal: signal, :1435, but no interruptState — unlike the root run); runtime-js/src/run-loop.ts:1080-1087 (an aborted signal with no interruptState.interrupted is a HARD abort, abortInfo.aborted false) → onFailed :3328-3350 (not a soft interrupt, so state.status = 'failed', persisted by persistTerminalState :3365-3367); the run's RunOutcome is kind 'failed'; runtime-js/src/run-loop.ts:4515-4519 (runSubAgentLoop throws on a 'failed' outcome) → the spawn wrapper's catch writes the parent ref 'failed' (:4838-4850); the continuation wrapper does the same (:5107-5120). Even a soft interrupt reaching the child (finalizeDurableInterrupt :4209-4252, session 'interrupted') returns kind 'failed', so the ref is still written 'failed'; core/src/orchestration/companion-tool-dispatch.ts:706-727 (handleSendMessage: a ref in TERMINAL_STATUSES — which includes 'failed', core/src/tools/companion/helpers.ts:9-13 — can only be continued when 'completed'; otherwise it throws "Child agent '<name>' is not active (status: failed)"); :502-512 (spawnAgent over a failed ref deletes the child session and re-spawns fresh); contrast: runtime-js/src/run-loop.ts:3335-3339 (a ROOT run given interruptState records a soft interrupt as 'interrupted' and stays resumable) — the child path simply never receives interruptState; related: CP-43 (the interrupted-companion resume path, resumeChildAgent, is unreachable on JS because the ref is never left non-terminal — see there); distinct from CP-23 (parent interrupt not reaching children on Temporal / CF Workflows / the DO park loops) and CP-50 (terminate); not yet reachable by this chain — a DURABLE-flag interrupt of the parent (stateStore.setInterruptFlag, polled at runtime-js/src/run-loop.ts:673 → :700-708) finalizes the parent via finalizeDurableInterrupt (:4209-4252), which sets interruptState without aborting the run AbortController (:704-707) and never touches childAbortControllers (only terminateChild aborts one, :5145-5148), so the in-process child keeps running and is not marked failed. That path is not claimed here and needs its own look (an orphaned running child after a parent interrupt); inferred: CF DO is not claimed — its persistent children are sibling DOs driven by persistent-companion-tools.ts, not the in-process runChildLoop; DBOS / Temporal / CF Workflows were not assessed for this mechanism; severity S2: the misreport is loud rather than silent — the LLM gets a tool error, not a false delivered:true — and re-spawning is a workaround, but it discards the companion sessionsource, peer-reportedverifiedopen
CP-69S2DBOS: a BLOCKING persistent-companion spawn returns only when the child's persistent workflow idles out (persistentIdleTimeoutMs, 24h by default), not when the child's first turn completes. The parent's step sits in a durable poll of the child's SESSION status, which stays 'active' across turns. JS, Temporal, CF Workflows and CF DO return the blocking spawn after the child's turn. Nothing errors: the parent turn simply stalls until the idle timeout plus a 30s margin.runtime-dbos/src/steps/execute-companion-tool.ts:561 (idleTimeoutMs = persistentIdleTimeoutMs ?? 86_400_000 — 24h default), :619 (the spawned child runs as a persistent workflow via ops.startPersistentWorkflow); runtime-dbos/src/steps/execute-companion-tool.ts:712-731 (blocking branch: pollChildSessionTerminal(childSessionId, deps, idleTimeoutMs + BLOCKING_SPAWN_MAX_WAIT_MARGIN_MS); the in-code SEMANTICS comment concedes it blocks until the child's idle timeout, NOT its first turn), :764-766 (250ms poll, 30s margin), :782-806 (pollChildSessionTerminal returns only on a completed / failed / interrupted SESSION status, else 'running' once the bound elapses); runtime-dbos/src/workflows/persistent-workflow.ts:478-496 (each turn runs with skipSessionStatusUpdate: true — the session stays 'active' between turns; also :595, :661), runtime-dbos/src/workflows/shared.ts:1034 (finalizeSessionId undefined in that mode), persistent-workflow.ts:527-534 (only an idle recv timeout or shutdown exits the loop) and :680-697 (finalizeLoop is the only session-status write, 'completed'); contrast: runtime-js/src/run-loop.ts:4776-4862 (spawnPersistentChild runs the child loop in-process) and :4865-4880 (the blocking waitForChildTerminal awaits that runPromise, which resolves when the child run ends); contrast: runtime-temporal/src/workflow-dispatch.ts:680-722 (blocking spawn: wf.startChild of the standard agentWorkflowFn, then awaits childHandle.result() — one child run); contrast: runtime-cloudflare/src/workflow.ts:3326-3360 (blocking spawn waits for the child workflow sub-agent-complete event, which the child emits via notifyParentCompletion on completing its run, e.g. :1882); contrast: runtime-cloudflare/src/do/persistent-companion-tools.ts:444-450, 983 (the DO blocking park polls the child DO /status until terminal); inferred: the child DO runs runtime-js, whose session reaches completed at the end of the turn; related: docs/dev/follow-ups.md FU-DBOS-BLOCKING-SPAWN-SEMANTICS (the tracked follow-up; the user-facing caveat is in docs/guide/sub-agents.md and docs/runtimes/dbos.md, which recommend non-blocking spawn + companion__waitForResult on DBOS); conformance: (pending — the scenario is in B2-core's unmerged MR) history.companion-continue on dbos-postgres fails on named ids; severity S2 (not S1): the stall is bounded (idle timeout + 30s, then 'running' with waitForResult as the recovery path), the limitation is documented to users with a workaround (non-blocking spawn + waitForResult, or a short persistentIdleTimeoutMs), and no data is lost; at the 24h default it is, however, a de facto hang of the parent turn; inferred: the follow-up's second consequence (a failed child turn reported 'completed' after idling out) is only partly mitigated by runtime-dbos/src/steps/run-lifecycle.ts:238-240, 260-320 (shouldPreservePriorRunStatus preserves a failed / interrupted status on the clean-exit finalize only when the most recent turn failed or was interrupted; an earlier failed turn followed by a completed turn can still be reported completed) — not claimed heresource, peer-reportedverifiedopen
CP-70S2DO wake silently strands an already-'paused' session when executor.resume() rejects before claiming the run. ensureExecutionContinuesImpl surfaces a resume failure only when the session reads 'active' afterwards. It re-arms only when this same call reset the session from 'active' to 'paused'. A session that was ALREADY paused matches neither, so the error is logged at warn and dropped. That includes the normal idle-DO wake after a client-tool submit or timeout. The session stays 'paused' forever, with no errorDetail, no failed stream chunk and no onAgentFail, and the client that submitted waits indefinitely.runtime-cloudflare/src/do/durable-object-agent-base.ts:789-835 (the resume() .catch: AgentAlreadyRunningError is benign; otherwise it fails the session ONLY if stateAfter.status === 'active' (:808-815), and re-arms ONLY if didResetFromActive && status === 'paused' (:827-834); an already-paused session falls through with just a warn log); runtime-cloudflare/src/do/durable-object-agent-base.ts:674 (resumableStatuses includes 'paused', so an idle DO whose session is paused is exactly the wake-after-submit path); live: B2-core, via a DO integ fixture whose seeded assistant row was invalid (content: []) — DOStateStore.getMessages skips the row while getMessageCount counts it, B2's history loader rejects resume() with a typed state_history_incomplete before claiming, and the session stayed 'paused' with no failure recorded (do-hibernation-wake.integ.test.ts on B2's branch before its fixture fix); inferred: any other pre-claim resume() rejection (a transient store error, an agent-validation failure) takes the same path; not observed; related: MA-01 (agent-factory throw on wake, also only warn-logged, but that path is retried by the watchdog); CP-52 / CP-53 (other ways a client-tool submit to the DO is lost); RM-13 (corrupt rows skipped while counted)source, peer-reportedpartly-inferredopen
CP-71S1runtime-js strands a session 'active' forever when anything inside the step loop THROWS instead of returning a failed outcome. runLoop and runWithNewRunLoop have try/finally but no catch, so the throw reaches the handle factory's safety net, which its comments call unreachable. That net reports { status: 'failed' } to the caller but writes nothing. The session stays 'active', the run 'running', and the stream 'active' with no terminal chunk, and onAgentFail never fires. Every later execute() throws AgentAlreadyRunningError, retry() refuses ('status is running') and resume() refuses ('already running'), so the API cannot recover the session. Triggers include a systemPrompt function that throws, a schema-invalid __finish__ from an adapter that reports it as a tool call, and an adapter StepResult that breaks the contract.runtime-js/src/run-loop.ts:925 (await runStepIteration with no catch), :568 and :1210 (the runLoop body try has only a finally); runtime-js/src/js-agent-executor.ts:3787-3892 (runWithNewRunLoop: statelessRunLoop at :3860 inside try / finally, no catch), :1936-1947 (comment: the runLoop "never rejects under normal operation"); runtime-js/src/execution/handle-factory.ts:176-197 (the .catch safety net, commented "unreachable in normal operation": it caches { status: failed, error, errorDetail } and persists nothing — no session status, no run status, no stream end, no hook); core/src/orchestration/step-iterator.ts:664-666 (buildMessagesForLLM evaluates a systemPrompt function, unguarded), :884-888 (planStepProcessing, unguarded; the in-code note: ordinary __finish__ completions still parse and throw); core/src/orchestration/step-processor.ts:218-220 (coerceSubAgentCallList calls .map on an absent subAgentCalls); contrast: runtime-js/src/run-loop.ts:960-992 (a StaleStateError surfaced as a failed OUTCOME is persisted: status failed, run failed, stream finalized); core/src/orchestration/step-iterator.ts:850 (an adapter throw is caught and becomes a failed outcome); live: probe (runtime-js + store-memory dist, this session): an agent whose systemPrompt function throws. The result is { status: 'failed', error: 'systemPrompt boom' }. One second later the session is 'active', the run 'running' and the stream 'active', and an onAgentFail hook ran 0 times. A second execute() throws AgentAlreadyRunningError, retry() throws "agent status is 'running'", resume() throws 'agent is already running'; live: probe (this session) — an adapter returning tool_calls without subAgentCalls strands the session identically (TypeError "reading 'map'"); the independent verifier reproduced the same with a __finish__ tool call carrying schema-invalid args (ZodError) and located the throw with an inspector; runtime-cloudflare/src/do/durable-object-agent-base.ts:1826 (the DO /start awaits handle.result(), which RESOLVES to status 'failed', so the .catch arm at :1907 that fails the session and stream never runs), :1855 (the stream is ended only on 'completed') — the DO /start path does not rescue the session; inferred: the DO wake path treats 'active + not executing' as a stranded run and auto-resumes it (runtime-cloudflare/src/do/durable-object-agent-base.ts:668-674); a deterministic throw would recur on each resume until the watchdog cap fails the stream (MA-01, auto_resume_exhausted), with the real error lost. Not verified; Temporal, DBOS and CF Workflows not checked; inferred: llm-vercel is not exposed to the __finish__ trigger — it returns __finish__ as structured_output (llm-vercel/src/vercel-adapter.ts:440-456), which does not parse there; severity S1: the caller is told 'failed' while the durable session is permanently wedged 'active', with no failure hook, no terminal stream chunk and no API path to recover it; a user systemPrompt function is an ordinary triggersource, probe, peer-reportedverifiedopen
CP-72S3Temporal: two companion__sendMessage calls to the same INTERRUPTED persistent child in one parent step compute the same child workflow id. The id is …__persistent__<childSessionId>__resume__<stepCount>, with no tool-call id in it. Companion calls run concurrently, so both activities can see the child ref as 'interrupted', append their message, report delivered: true, and return resumeMetadata. The child is started with WORKFLOW_ID_REUSE_POLICY_ALLOW_DUPLICATE, so the second wf.startChild fails only while the first child workflow is still open. When it does fail, the error is only logged. If the first child run had already loaded its history, the second message stays in the child's log unprocessed, although the parent was told it was delivered.runtime-temporal/src/workflow-dispatch.ts:786-814 (sendMessage-to-interrupted branch: childWorkflowId built at :790 from parentWorkflowId, childSessionId and stepCount only; wf.startChild at :792; the catch at :807-813 logs persistent_agent.resume_failed and continues); runtime-temporal/src/workflow-dispatch.ts:616-630 (dispatchCompanionTools runs every companion call of the step concurrently via Promise.all); core/src/orchestration/companion-tool-dispatch.ts:769-799 (handleSendMessage: wasInterrupted from the ref read, the message is appended, the ref is set to 'running', and delivered: true with resumeMetadata is returned) — two concurrent calls can both read 'interrupted' before either update lands; contrast: runtime-cloudflare/src/workflow.ts:3571 (CF Workflows suffixes its companion continuation instance id with the companion call id — the B2 fix direction: append the toolCallId); inferred: when the calls do not overlap, the second call reads 'running' and takes the parent_message interrupt path instead (CP-41), so the collision needs the two activities to interleave. It also needs the first child still running when the second start is attempted: under ALLOW_DUPLICATE (runtime-temporal/src/workflow-dispatch.ts:806-807) a closed first child lets the second start succeed, and that run reloads the whole log, so nothing is lost. The same stepCount-only suffix is used for the continuation id (runtime-temporal/src/workflow-dispatch.ts:833); a second continuation there is expected to fail its completed-to-active CAS first, so it is not claimed. Not observed; related: CP-41 (sendMessage to a running companion on Temporal is treated as an interrupt); CP-46 (Temporal startChild failures for companion spawn / resume / continue are only logged); severity S3: needs two sendMessage calls to the same interrupted child in one step, interleaved, with the first child still running at the second start; the loss is one message, which stays in the durable logsource, peer-reportedpartly-inferredopen
CP-73S2Temporal never records an interrupted persistent companion as interrupted on its parent. When the child workflow ends 'interrupted', the parent-side SubSessionRef keeps a stale 'running': the completion bookkeeping maps only 'completed' / 'failed', and core's ref sync deliberately leaves an interrupted child's ref at 'running'. Core's handleWaitForResult returns only on a terminal ref. Its orphan detection runs only where the runtime supplies a liveness probe, and only runtime-js does. So an untimed companion__waitForResult on that child on Temporal never returns 'interrupted'. It keeps polling until the companion activity itself times out, and the parent's turn stalls, then gets a generic failure.runtime-temporal/src/activities.ts:5704-5745 (markPersistentChildCompleted: maps only a 'completed' / 'failed' child session, returns without writing for any other status, :5717-5719); the only other ref writes are the continuation reset to 'running' (:6174-6178) and the unused helper runtime-temporal/src/state-helpers.ts:179-192 — nothing in runtime-temporal writes 'interrupted' to a SubSessionRef; core/src/orchestration/companion-tool-dispatch.ts:833-866 (resolveAndSyncSubSessionRefs maps only completed / failed; the comment at :855 keeps an interrupted child's ref at 'running'); core/src/tools/companion/helpers.ts:9-13 (TERMINAL_STATUSES = completed, failed, terminated); core/src/orchestration/companion-tool-dispatch.ts:948-1079 (handleWaitForResult: without a timeout the loop exits only on a terminal ref, isCancelled(), or orphan detection; orphan detection at :1038-1041 requires deps.isChildExecutionLive, which only runtime-js supplies, runtime-js/src/run-loop.ts:5157); runtime-temporal/src/activities.ts:6193-6200 (isCancelled reads the activity cancellation signal); runtime-temporal/src/workflow.ts:224-235 (the activity proxy: startToCloseTimeout 5 minutes, maximumAttempts 3); runtime-temporal/src/workflow-dispatch.ts:636-651 (a failed companion activity becomes a generic failed tool result); inferred: the untimed wait is therefore bounded only by the activity timeout and retries (roughly three 5-minute attempts), after which the tool reports a generic activity failure, not 'interrupted'. Whether the timed-out attempts' polling loops are cancelled on the worker (no heartbeat is sent) was not checked. Not observed; contrast: runtime-cloudflare/src/workflow.ts:1566-1575 (a CF Workflows child ending on an interrupt notifies its parent with status 'terminated'), and :3434-3530 (waitForResult waits for that event, bounded by subAgentEventTimeout, and records it); runtime-js returns promptly because its child is recorded 'failed' (CP-68); related: CP-41 (a Temporal sendMessage to a running companion is treated as an interrupt, so the child ends interrupted; its 'waitForResult can hang' is this finding's consequence); CP-23 (a parent interrupt does not reach Temporal children, a different path); CP-68 (JS records the interrupted child 'failed'); related: likely fix (B2-core's B13 review, peer-reported): Temporal writes the ref 'interrupted' when the child workflow ends interrupted, and waitForResult treats an interrupted ref as a result; severity S2: needs an interrupted companion plus an untimed waitForResult; the parent turn stalls for the activity timeout and then gets a misleading failure instead of 'interrupted'; inferred: the fix has two halves — Temporal must write an interrupted child's ref, AND core must treat an interrupted ref as terminal for an untimed waitForResult (TERMINAL_STATUSES at core/src/tools/companion/helpers.ts:9-13 excludes interrupted), or a Temporal-only fix still waits; DBOS not checkedsource, peer-reportedpartly-inferredopen
CP-74S1DBOS persistent agent: a user message sent just as the persistent workflow exits on its idle timeout is silently lost. The workflow drains its inbox, then finalizeLoop marks the session 'completed', but the workflow stays DBOS-PENDING until its finally cleanup step returns. An execute() in that window routes on the DBOS status alone, so it takes the SEND path (DBOS.send to 'inbox') into a workflow that never calls recv again. execute() returns a handle with no error, the message stays unconsumed in dbos.notifications, the session stays 'completed', and no new workflow starts.runtime-dbos/src/workflows/persistent-workflow.ts:527-534 (idle timeout: one non-blocking inbox drain, then finalizeLoop) and :680-696 (finalizeLoop writes session status completed via finalizeRun); runtime-dbos/src/workflows/persistent-workflow.ts:300-352 (the finally block still runs HookStep.cleanupPerCallHookManagerStep after finalizeLoop, so the workflow is PENDING after the session reads completed); runtime-dbos/src/persistent/routing-decision.ts:34 (ALIVE_STATUSES = PENDING, ENQUEUED) and :53 (send when alive) — routing reads only the DBOS workflow status, never the session status, and nothing verifies the send is consumed; live: B8 reproduced locally with no product change (a 1 s idle TTL; execute, wait for 'completed', execute again immediately): 5 of 6 runs lost the second message, the workflow was PENDING at 'completed' with finalizeRunStep as its last step, and the inbox row stayed in dbos.notifications; live: CI pipeline 2885443694 job 16754864348 — multi-message-execute-dbos-persistent.integ.test.ts "idle-timeout exit then new message starts a fresh persistent workflow" failed at :436 (expected 2 user messages, got 1); it passed on the automatic retry (job 16755337462); related: CP-69 (DBOS blocking persistent spawn waits for the idle timeout); other runtimes' persistent routing not checked; related: FU-DBOS-PERSISTENT-SEND-INTO-EXITING-WORKFLOW (pending in MR !259 / !266, not yet on main)source, live-repro, peer-reportedverifiedopen
CP-75S3companion__sendMessage to a persistent child whose session does not exist yet is not treated as 'not started'. Core's handleSendMessage loads the child session, and a null result passes every guard (the C-4 pending-client-tool check and the I1 terminal check both read it through optional chaining), so it goes on to append the message. The child's ref already reads 'running' in that window: a re-spawn deletes the child session and sets the ref 'running' before the child is recreated, and on Temporal and CF Workflows the child session is created only later, by the child workflow. Every store's append now throws for a missing session (memory, Postgres and Redis 'Session not found', D1 'State not found for sessionId'), so the model gets a tool error for a child that is starting normally, and has to retry the send. Since RM-50's fix (MR !269) Redis no longer writes a status-less hash, sets a stale parent_message flag or reports delivered: true here. The DBOS and CF DO runtimes use their own sendMessage handlers and are not affected.core/src/orchestration/companion-tool-dispatch.ts:733 (childSessionState = loadState(child)), :745-757 (C-4 guard, optional-chained, so a null state passes), :765 (I1 terminal guard, same), :780 (appendMessages to the child — throws for a missing session, so :783 setInterruptFlag and :801 delivered: true are not reached); the dispatcher's catch at :416-417 turns the throw into a failed tool result; core/src/orchestration/companion-tool-dispatch.ts:504-526 (re-spawn of a failed / terminated child: deleteSession at :512, then the ref is set 'running' at :520-526, before spawnPersistentChild at :563 recreates the child) and :530-543 (first spawn adds a 'running' ref before the child exists); runtime-temporal/src/activities.ts:6051-6054 (the spawn activity's spawnPersistentChild is a no-op), runtime-temporal/src/workflow-dispatch.ts:680 (the workflow starts the child workflow after the companion activity returns), runtime-temporal/src/activities.ts:6982 (the child session is created only in the child workflow's initializeAgentState) — so the window outlasts the spawn call; runtime-temporal/src/workflow-dispatch.ts:630 (a step's companion calls run concurrently via Promise.all, so a spawn and a sendMessage in the same step race directly); runtime-cloudflare/src/steps.ts:5205-5208 (CF Workflows: spawnPersistentChild is likewise a no-op; the workflow body starts the child afterwards); store-memory/src/in-memory-state.ts:336-341 and store-postgres/src/state/postgres-state-store.ts:772-785 (appendMessages throws "Session not found"); store-cloudflare/src/d1-state.ts:256 with store-cloudflare/src/errors.ts:142 (StateNotFoundError, "State not found for sessionId: …"); store-redis/src/redis-state.ts:1165-1185 (appendMessages is one Lua script guarded on HEXISTS sessionId and throws "Session not found" — the RM-50 fix, MR !269); live: probe (batch-6 verification, store-redis dist on origin/main e37d676945 against Redis :6390): appendMessages to a never-created session threw 'Session not found'; afterwards sessionExists returned false, no key for the id existed, and loadState returned null. The pre-!269 probe (appendMessages did not throw, left a status-less hash that loadState rejected as InvalidStoredStateError, and a stale parent_message flag) no longer reproduces; inferred: runtime-js has the same window — the child session is created inside the detached child IIFE after its first await (runtime-js/src/run-loop.ts:4814-4822) — but it is short; not observed; contrast: runtime-dbos/src/steps/execute-companion-tool.ts:556 / :576 also adds the ref before createSession, but DBOS runs companion calls sequentially in the workflow body (runtime-dbos/src/workflows/shared.ts:2521-2536) and routes sendMessage through its own handler (execute-companion-tool.ts:967), so no DBOS sendMessage can run inside that window; the CF DO runtime also has its own sendMessage (runtime-cloudflare/src/do/persistent-companion-tools.ts), which reads the child DO's /status instead — neither is this defect; related: RM-50 (fixed by MR !269): its atomic appendMessages removed the Redis delivered:true / status-less-hash / stale-flag branch this record used to carry, which is why the title and severity were narrowed; runtime-js/src/run-loop.ts:4591-4599 (a failed append of the child's initial message is only warn-logged); severity S3 (narrowed from S2 in batch 6): on every store the model now sees a transient tool error for a child that is starting normally and can retry the send; nothing is corrupted or reported as delivered; related: peer-reported by B8 from its caller audit; the source and the Redis probe were re-checked in batch 6source, probe, peer-reportedpartly-inferredopen
CP-76S1Temporal: resume() after a stop never reactivates the root stream that the stop ENDED, so the resumed run cannot write to it. execute() and retry() both reactivate an ended or failed stream before starting their workflow; resume() only calls initStream, which does not. The resumed workflow's first stream write is rejected ("Cannot write to stream … in 'ended' state"), and — observed only on a peer branch — the run fails with the executor's synthetic "Stream closed without terminal observation". The resumed turn does not complete, and the stream never shows it. (That a with_message resume also drops its message is CP-31, independent of this.) This is the Temporal twin of CP-67 (CF Workflows), with its own fix site.runtime-temporal/src/executor.ts:1301-1302 (resume() reads getStreamInfo only for latestSequence), :1364-1375 (starts the __resume-N workflow), :1386 (then initStream only; there is no reactivateStream call anywhere in resume()); contrast: runtime-temporal/src/executor.ts:417-426 (execute() reactivates an ended / failed stream BEFORE startWorkflow, with a comment naming this exact throw) and :1702-1708 (retry() does the same); runtime-temporal/src/activities.ts:5095-5098 (an interrupted root run ends its stream: endStream(streamId, undefined)); store-memory/src/in-memory-stream.ts:236-240 and store-redis/src/redis-stream.ts:653 (writers throw "Cannot write to stream ${streamId} in '${status}' state" for a non-active stream); runtime-temporal/src/executor.ts:1438-1443 (the handle's fallback when the observed stream closes with the session still active or paused: status 'failed', error 'Stream closed without terminal observation'); live: B2-core (peer) reproduced it on origin/main and on its branch reliability-b2-core-llm-history (head 3ec6301a51): packages/e2e/src/__tests__/cross-runtime-resume-matrix.integ.test.ts:162-163 (TEMPORAL_ENDED_STREAM_SKIP) skips the completion legs 'B2 C3/C4 (completion): resume(with_message) after a stop completes…' (:910-913) and 'B2 C7 (interrupted route, completion): resume(from_checkpoint) after a stop completes…' (:995-997) on all three Temporal combos. That file is not on main; the skip exists only on that branch. That branch also fixes CP-31 (its executor.resume() handles with_message), so its '(input)' twins passing does not carry over to main, where CP-31 drops the message; B2 also reports a repro on origin/main, not re-checked here (no Temporal server in this session); inferred: which the caller sees first depends on timing — the handle observes from the ended stream, so result() may resolve with the synthetic failure while the workflow is still retrying its LLM activity against the ended stream; the session status left behind was not established; related: CP-31 (Temporal resume() ignores options.mode and message: runtime-temporal/src/executor.ts:1343-1349 always sends newMessages: []) — the with_message / from_checkpoint losses are that finding, not this one; CP-67 (the CF Workflows twin; its own related note flagged this gap as latent from source); CP-56 (a stop on the terminal step can end the run completed instead, and then resume() short-circuits, CP-60) — this gap is reached when the stop lands mid-run and leaves the session interrupted; severity S1: a lost user action — resume after an interrupt is the supported follow-up, and every such resume on Temporal fails to complete its turn (HITL paused resumes are not affected: they do not end the stream)source, live-repro, peer-reportedpartly-inferredopen
CP-77S2Cloudflare DO: in the remote-agents-cloudflare-do example, the producer's planner DO intermittently fails to continue after its two parallel sub-agent children finish. The planner calls its sub-agent tool twice in one step; the DO rewrites that tool onto DOStubTransport, so each child runs in its own DO. Both children make their final __finish__ LLM call, but the planner's next LLM call (the one carrying both children's results) is never made within the e2e test's 90 s window, so the consumer's snapshot never gains state.research and the test times out. Root cause unknown.runtime-cloudflare/src/do/durable-object-agent-base.ts:1441-1475 (rewriteSubAgentTools: with subAgentNamespace set, every local sub-agent tool becomes createRemoteSubAgentTool over a DOStubTransport, one child DO per call, timeoutMs default 10 min); runtime-cloudflare/src/do/do-stub-transport.ts:88 (DOStubTransport, the parent-to-child-DO transport those calls use); live: examples/remote-agents-cloudflare-do/producer/src/agents/planner.ts:41-47 (threadTool = createSubAgentTool(ThreadResearcherAgent)); examples/remote-agents-cloudflare-do/producer/src/my-agent-server.ts:22 (subAgentNamespace set); the recorded planner response 878ab27f8ce90470 carries two subagent__thread-researcher calls in one step; live: examples/remote-agents-cloudflare-do/tests-e2e/cross-service.spec.ts:40-51 (polls /api/snapshot/<id> for state.research, timeout 90 s); playwright.config.ts sets retries: 0; live: replay fixtures (examples/remote-agents-cloudflare-do/test-utils/openai-fixtures/): d503428640967220 is a thread researcher's final call (its response is __finish__), and 6676bfdb88d8563e is the planner's call whose input ends with both thread outputs as tool results. Peer-reported failing runs' replay logs stop after the 9th LLM call (d503428640967220) and never request 6676bfdb88d8563e; live: peer-reported CI sightings: jobs 16755700166 (MR !262), 16756092320 (MR !268), 16756390704 (MR !259); not re-checked here (no GitLab access in this session); inferred: candidate causes, none confirmed: (a) a missed child-completion or stream-end wake in core/src/orchestration/remote-sub-agent-dispatch.ts when two DOStubTransport children end close together; (b) a CP-70-like DO wake strand of the planner DO; (c) interference from a concurrent /api/child-snapshot read; inferred: the rewritten tool times out after 10 minutes by default, so the 90 s test window cannot tell a permanent hang from a delay longer than 90 s; related: CP-70 (DO wake silently strands a paused session); FU-REMOTE-06 (fold read timeouts bound the parent, not the request); severity S2 (provisional): a parent that never continues after its children complete is a silent stall with no error; S2 rather than S1 until a repro shows it is unbounded and not example-specificsource, peer-reportedunverifiedopen
CP-78S2Remote sub-agents: a 409 ALREADY_COMPLETED from /start reaches the parent as a plain Error, and the parent FAILS a delegation that succeeded. HttpRemoteAgentTransport.start() retries its POST on network errors and 5xx responses. If the first POST started the child but its response was lost, and the child finished before the retry (1 s, then 2 s, then 4 s later by default), the retry gets ALREADY_COMPLETED. Both transports map only ALREADY_RUNNING, NOT_FOUND and FEATURE_UNAVAILABLE to typed errors, and every remote dispatch site attaches only on RemoteAgentAlreadyRunningError and rethrows anything else. The sub-agent call therefore fails with "Remote agent start failed (409): …" (failureReason 'transport-error') although the child completed, and its output never reaches the parent.core/src/transport/http-transport.ts:134-148 (start() → fetchWithRetry, then toProtocolError on !ok), :341-380 (fetchWithRetry: a thrown fetch error is retried at :371-375, a 5xx at :364-368, a 4xx is returned at :359-362; maxRetries 3 and retryBaseDelayMs 1000 by default, :126-127), :317-329 (toProtocolError maps only ALREADY_RUNNING / NOT_FOUND / FEATURE_UNAVAILABLE; ALREADY_COMPLETED falls through to a generic Error); runtime-cloudflare/src/do/do-stub-transport.ts:297-307 (the same three-code mapping; DOStubTransport.start at :110 does not retry, so its gap is latent); agent-server/src/agent-server.ts:369-383 (startAgent throws ALREADY_COMPLETED for a session that is 'completed' or 'failed' — so a child that FAILED quickly is reported as a transport-error with the 409 body instead of remote-failed with the child's own error); agent-server/src/http/handler.ts:862-863 (mapped to HTTP 409); core/src/orchestration/remote-sub-agent-dispatch.ts:385-396 (start catch: attach only on RemoteAgentAlreadyRunningError, else rethrow) and :729-736 (the outer catch returns kind failed, failureReason transport-error, with the error message); runtime-js/src/execution/remote-subagent-executor.ts:373-384, runtime-dbos/src/steps/execute-remote-subagent.ts:438-449, runtime-cloudflare/src/steps.ts:3133-3145 (the same attach-or-rethrow start catch in each runtime-specific dispatch); live: probe (this session, core dist): HttpRemoteAgentTransport.start with a fetch that throws 'fetch failed' on the first call and returns 409 {code:'ALREADY_COMPLETED'} on the retry made 2 calls and threw a plain Error (not RemoteAgentAlreadyRunningError): 'Remote agent start failed (409): {"error":"Session s1 already completed","code":"ALREADY_COMPLETED"}'; inferred: the end-to-end effect (the parent's sub-agent tool result reports a failure for a child that completed) follows from the dispatch code above; not reproduced against a live producer. The pre-start getStatus check (remote-sub-agent-dispatch.ts:311-338) handles a child that is already completed on re-entry, but not one that completes during start()'s own retries; contrast: a Cloudflare DO producer does not answer ALREADY_COMPLETED at all (the only error code runtime-cloudflare/src/do/durable-object-agent-base.ts emits is ALREADY_RUNNING, :1590, :1676, :1687); what a retried /start does to a completed DO session was not assessed; related: CP-07 (HelixChatTransport ignores /start responses including ALREADY_COMPLETED — a different component); docs/guide/cross-service-remote-agents.md "At-least-once POSTs" (documents the ALREADY_RUNNING → attach rule, which does not cover this case); severity S2: needs a lost /start response plus a child that finishes within the retry backoff, but the result is a wrong failure reported to the model for a completed delegation, and the parent may repeat the worksource, probe, peer-reportedpartly-inferredopen
CP-79S1DBOS standard mode: an interrupt() that lands before the workflow's in-body appendMessages step silently drops that turn's user input. DBOS execute() passes the input only as a workflow argument, and the workflow appends it several durable steps in. interrupt() on the handle execute() returned calls DBOS.cancelWorkflow, so the next pre-append step throws and the append never runs. The interrupt path then CASes the session active → interrupted. The stored session ends interrupted with no user message and no error / errorDetail (the same handle's result() reports failed, CP-57, which does not name the lost input), and nothing re-supplies the input: resume() starts the workflow with input: undefined, so the LLM answers with only the system prompt (plus prior turns) and the run reads completed. Fresh sessions and continued turns are both affected. Only a handle that holds the workflow id cancels: a getHandle() / ReconnectHandle interrupt only sets the flag, so the append runs and the input survives. JS appends the input inside execute() before it returns the handle. Temporal and CF Workflows interrupt through a flag or signal that never stops the input append.runtime-dbos/src/lifecycle/execute.ts:426-432 (createSession) / :456-460 (continuation CAS to active) and :681-695 (DBOS.startWorkflow receives input; execute() persists no user message before it returns the handle); runtime-dbos/src/workflows/standard-workflow.ts:490 (beginRun step) and :528 (runAgentLoopOneTurn); runtime-dbos/src/workflows/shared.ts:1047 (cleanupOrphanedStaging), :1054 (loadState), :1078 (initStream), :1145 (getMessageCount) and only then :1169-1172 (appendMessages of messagesToAppend, the history prefix + user input, built at :1167) — five-plus durable steps before the input is durable; runtime-dbos/src/lifecycle/interrupt.ts:67 (setInterruptFlag), :86 (DBOS.cancelWorkflow: every later @DBOS.step throws, so the append is never reached), :115-116 (after 500 ms, CAS active → interrupted from outside the workflow). The graceful in-loop flag check at runtime-dbos/src/workflows/shared.ts:1494 sits AFTER the append and is never reached; runtime-dbos/src/handles/active-handle.ts:587-591 (the execute() handle passes its workflowId to interruptImpl, so it takes the cancel path); runtime-dbos/src/lifecycle/get-handle.ts:88 and runtime-dbos/src/handles/reconnect-handle.ts:135-140 (a getHandle() handle passes null, the flag-only path at interrupt.ts:125-134, so the append runs and the input survives); agent-server/src/agent-server.ts:906-918 (interruptAgent writes the flag, then calls the same-replica activeHandles handle, so a UI stop hits this only when it lands on the replica that started the run); runtime-dbos/src/lifecycle/resume.ts:447-450 (only with_message appends anything) and :512-529 (the resumed workflow gets input: undefined, isResume: true, so the dropped input is never re-appended); live: live-repro (this verification, RUNTIME_DBOS_INTEG=1, private Postgres DB on :5433, runtime-dbos integ context from origin/main, no product change): (a) with no hold, execute() then handle.interrupt() straight away: 10 of 10 runs ended interrupted with 0 messages stored; (b) with a hold on getMessageCount, a fresh session ended interrupted with an empty transcript and 0 LLM calls. resume() then completed and the LLM input was only [system]. The transcript became a lone assistant message with no user message; (c) with a hold, turn 2 of a completed session ended interrupted with only turn 1 [user TURN-1, assistant] stored, so the turn-2 input was gone. An independent review re-run (same setup, private DB) reproduced it: fresh unheld 5/5, continued turns unheld 3/3, and resume() answering [system] only and reporting completed; contrast: runtime-js/src/js-agent-executor.ts:3500-3530 (execute() awaits saveAgentState + appendMessages before the handle exists, so no interrupt can precede the append; a first turn keeps its input); contrast: runtime-temporal/src/executor.ts:804-807 (interrupt is a workflow signal, not a cancel) and runtime-temporal/src/workflow.ts:245-248 (the handler only sets a flag), so the initializeAgentState activity still appends the input (runtime-temporal/src/activities.ts:7185-7187) before the loop observes the interrupt; contrast: runtime-cloudflare/src/executor.ts:1299-1319 (CF Workflows interrupt = durable flag + best-effort event, never a terminate), so the workflow init step still appends the input; the CF DO embeds the JS executor (append before the handle); related: CP-17 (DBOS interrupt calls cancelWorkflow immediately, so the graceful path is unreachable — the root mechanism; a graceful interrupt would let the append run before the flag check at shared.ts:1494, so a CP-17 fix likely fixes this too, but the input loss is a distinct S1 observable, as with CP-57); CP-57 (the same cancel makes ActiveHandle.result() report failed while the stored session reads interrupted); RM-42 (runtime-js deletes a continued turn's input on an interrupt during the first step — a different mechanism, the orphan-cleanup truncation); CP-19 (stale durable interrupt flag); RM-43 / RM-41 (weaker: CF Workflows appends the turn input outside the turn-start boundaries); live: first reported by B2 (peer-reported). The debug report the brief cites (debug-interrupt-resume-dbos-report.md) could not be found on disk, so this verification re-derived the claim from source and reproduced it independently (the live-repro item above); severity S1: the user's own message is silently lost and the stored session carries no error, which a stop through the execute() handle right after send reliably triggers (10/10 unheld; through agent-server only when the stop lands on the executing replica). A later resume() cannot recover it, and it produces an answer to no question that reads completedsource, live-repro, peer-reportedverifiedopen
CP-80S2Temporal: the handle resume() returns settles result() on a terminal LLM error before the session is persisted as failed. The v7 resume handle projects stream chunks, not the workflow result, and treats the first error chunk as terminal. runLLMStep's onError callback writes that chunk inside the LLM activity. The workflow runs commitStep and then persistTerminalState(failed) only after the activity returns. So result() reports failed while the stored session still reads 'active'. A caller that acts on the result straight away gets a stale read. The documented recovery, retry(), then throws "Cannot retry: agent status is 'running'" because its ['failed'] CAS fails. The window is inherent: it was observed with no injected delay. The completed path is not affected, because the output chunk is written by persistTerminalState after its saveState. execute() and retry() handles are not affected either, because they await the workflow result.runtime-temporal/src/executor.ts:1394-1400 (resume() returns createTemporalAgentHandleV7, the stream-projecting handle), contrasted with :691-703 (execute() and retry() use createActiveHandle, which awaits workflowHandle.result(), so the workflow has already persisted); runtime-temporal/src/handle.ts:126-135 (result() resolves on the first chunk projectChunk accepts) and :266-277 (any 'error' chunk projects to status 'failed'); runtime-temporal/src/activities.ts:2951-2963 (runLLMStep's onError writes an 'error' chunk, recoverable: false, from inside the LLM activity); runtime-temporal/src/workflow.ts:1234 (commitStep runs after the activity returns) and :1439-1471 (only then does the terminal-failed branch call persistTerminalState(failed) with skipErrorEmit: true, because the chunk was already written); runtime-temporal/src/activities.ts:5013 (persistTerminalState saveState is the write that makes the session read failed); runtime-temporal/src/executor.ts:1590-1601 (retry() CASes ['failed'] → 'active' and throws while the session still reads active); contrast: runtime-temporal/src/activities.ts:5013 then :5044-5063 (on completion persistTerminalState saves status completed BEFORE it writes the output chunk, so a completed resume() result is never ahead of the store; the peer claim that runPhase2FinishWith emits the output chunk early is refuted: activities.ts:5052 is the only output-chunk writer in runtime-temporal); live: live-repro (this verification, HELIX_TEMPORAL_BACKEND=test-server, runtime-temporal integ harness with InMemoryStateStore, origin/main 3560d6c72a, no product change): client-tool suspend → submitToolResult → resume(), the resumed LLM step returns a non-recoverable error. result() returned 'failed' while loadState() read straight after returned status 'active' (2 of 2 runs, with no delay injected and with an 800 ms hold on the terminal saveState). The same session read 'failed' 1.5 s later. The completed variant (resumed step calls __finish__) read 'completed' at resolve time in both runs. An immediate executor.retry() after the failed result threw "Cannot retry: agent status is 'running'"; inferred: an adapter that returns an error step with shouldStop: false (a recoverable error; the core MockLLMAdapter can, llm-vercel always returns shouldStop: true) also writes the error chunk, so the resume handle would report 'failed' for a run that continues. Not reproduced; related: CP-57 (DBOS result() reports failed while the store reads interrupted — the same result/store disagreement class, different mechanism); CP-65 (the same Temporal resume handle never settles when nothing new is drained); DI-24 (Temporal keeps the classified error detail only on the stream chunk, not in the persisted row); the batch-5 peer candidate #3 prompted the check, but its mechanism is refuted (see contrast) and this failed-path variant was found and reproduced independently, so peer-reported is not claimed; severity S2: the terminal status is right and converges within the persist latency, but a caller that reads the session or calls the documented retry() as soon as result() says failed gets a stale 'active' read or a thrown retry, with no signal to waitsource, live-reproverifiedopen
CP-81S1Duplicate of CP-52. CF DO: a client-tool submit that arrives while the suspending run is still inside onAgentSuspended is dropped. persistSuspension has already written the session paused, so /status and the harness report the suspension while the hook is still running. /submit-tool-result stores the result and calls ensureExecutionContinues, which returns early because isExecuting is still true. The run then exits suspended, the scenario's explicit resume() right after the submit did not rescue it in the repro, and nothing wakes the session later: the watchdog recovers only active sessions. The session stays paused with the result stored and never consumed.core/src/orchestration/step-iterator.ts:1745 (persistSuspension) then :1765 (onAgentSuspended is awaited after it, so the session reads paused for the whole hook); runtime-cloudflare/src/do/durable-object-agent-base.ts:3107-3109 (handleSubmitToolResult → ensureExecutionContinues) and :644-653 (returns early on _executionState.isExecuting: 'the live run will observe the state change', but the run has already decided to suspend); runtime-cloudflare/src/do/durable-object-agent-base.ts:1061-1073 (recoverStrandedExecution returns unless the session is 'active', so a stranded 'paused' session with a submitted result is never recovered); live: live-repro (batch-6 verification, origin/main e37d676945, no product change): the lifecycle-hooks-parity scenario 'tracingContext set by onAgentSuspended persists through resume → completion' copied into a temp cf-do-d1 test (e2e workerd pool, Miniflare) with a 1 s delay added to onAgentSuspended failed 4 of 4 runs with the CI signature (expected 'suspended_client_tool' to be 'completed'); the session read paused with tc-1 submittedAt set and the result stored, and still read paused with tc-1 pending 5 s later. Without the delay it passed 3 of 3; contrast: runtime-js (probe, batch-6 verification): submit + resume({ mode: 'continue' }) issued while a 1 s onAgentSuspended hook was still running were accepted and the run completed with the real result, because JS continuation is caller-driven and the resume CAS takes the already-paused session. Temporal (caller-driven resume whose CAS accepts paused, interrupted and active) and DBOS (DBOS.send is queued durably to the waiting call's recv topic) have no such window, from source; related: CP-52 (the same drop, same mechanism and fix site: ensureExecutionContinues returns early on isExecuting while the run is tearing down a suspension; CP-52 observed it through the ai-sdk host on a fresh session's first client-tool round); peer-reported by B8 (2026-09-27) from CI jobs 16758514869, 16759542448, 16760323616, 16764206430 and main pipelines 2887177079 / 2886530232 (lifecycle-hooks-parity.cf.test.ts, cf-do-d1); severity S1: lost user action (the submitted client-tool result is never consumed and the session is stranded paused)source, live-repro, peer-reportedverifiedwithdrawn — Duplicate of CP-52: the same submit-during-suspension-teardown drop, with the same mechanism (ensureExecutionContinues returns early on isExecuting after the run has decided to suspend) and fix site. The onAgentSuspended hook only widens CP-52's teardown window. This record's executor-tier repro, the watchdog gap and the CI signature were moved to CP-52's evidence; track and fix the defect there.
CP-82S1A partial client-tool batch continues the paused run, and the unanswered calls get a fabricated 'runtime restarted' failure. When one step emits two client-tool calls (c1, c2) and only c1 is submitted, the CF Durable Object's /submit-tool-result wakes the run itself (ensureExecutionContinues → executor.resume({ mode: 'continue' })), and runtime-js's resume() does the same when a caller resumes after a partial submit. resume() has no 'every pending call is answered' gate. Its orphan drain, written for calls stranded by a crashed process, treats every pending entry with no submission as lost and synthesizes runtime_restarted for c2. The model then gets c1's real answer plus an error telling it the runtime restarted, and the run completes. The client's real c2 answer, when it arrives, is refused as already_completed, which the ai-sdk handler treats as success, so the user's action is dropped with no error anywhere. The ai-sdk handler resumes after every tool-result POST without checking for other pending calls, so js-chat-host is exposed the same way. Temporal and DBOS keep waiting for c2 and use its real answer.runtime-cloudflare/src/do/durable-object-agent-base.ts:3107-3109 (handleSubmitToolResult wakes the DO after any locally-routed accepted submit) and :662-790 (ensureExecutionContinuesImpl: only a resumable-status check, then executor.resume(agent, sessionId, { mode: "continue" }) at :779; nothing checks whether other pending calls are still unanswered); runtime-js/src/js-agent-executor.ts:2502 (resume() drains orphan client-tool calls before the loop) and runtime-js/src/client-tool-resolver.ts:898-903 (drainOrphans' last branch: an entry with neither submittedResult nor submittedError yields outcome { ok: false, error: 'runtime_restarted' }, whatever its age or deadline); ai-sdk/src/handler/handle-chat-stream.ts:659-719 (the tool-result path submits each intent in the POST; 'accepted' and 'already_completed' are both success at :712-713) then :748 (callExecutorResume mode continue, with no check for other pending calls); live: live-repro (batch-6 verification, origin/main e37d676945, no product change) on the CF DO: a temp cf-do-d1 test in the e2e workerd pool (Miniflare, no Docker). One step emits c1 and c2 (client tools); submitToolResult(c1) returned accepted; 4 s later the session read completed with no pending calls after 2 LLM calls, and the second model input was [system, user, assistant(c1, c2), tool c1 {"a":"ONE"}, tool c2 {"error":"The 'ask' could not be completed because the runtime restarted while it was pending. …"}]. A late submitToolResult(c2) returned already_completed; live: probe (batch-6 verification, unit level, runtime-js JSAgentExecutor + InMemoryStateStore): the same batch, submitToolResult(c1), then resume({ mode: 'continue' }): result completed, pending cleared, persisted [user, assistant(c1, c2), tool c1, tool c2 runtime-restarted error, assistant final]; a late submitToolResult(c2) returned already_completed; contrast: live-repro (batch-6 verification) on Temporal (HELIX_TEMPORAL_BACKEND=test-server, runtime-temporal integ harness): after submit(c1) + resume(), the session stayed waiting with c2 pending and 1 LLM call; submit(c2) + resume() then completed with both real answers in the transcript (runtime-temporal/src/activities.ts:3585, an unsubmitted entry stays pending; the resumed handle itself never settled while c2 was pending, the CP-65 mechanism). On DBOS (RUNTIME_DBOS_INTEG=1, private Postgres DB on :5433, Redis db 14 on :6390) the workflow waits on a per-call DBOS.recv (runtime-dbos/src/client-tool-resolver.ts:74): after submit(c1) it stayed active with c2 pending, and submit(c2) completed the run with both real answers; inferred: (unverified) even a full batch from the stock client may race on the DO. The Helix chat transport sends one /submit-tool-result request per tool output, concurrently (ai-sdk/src/transport/helix-chat-transport.ts:228-236, Promise.all), each accepted submit wakes the DO (runtime-cloudflare/src/do/durable-object-agent-base.ts:3107-3109), and the wake runs executor.resume detached (:777-790), so c1's wake-and-orphan-drain can run before c2's submit lands. A stock useChat only auto-sends once every tool call of the message has a result, so a one-by-one client (or a direct /submit-tool-result caller) hits it deterministically; whether the concurrent full-batch path loses the race in practice has not been reproduced; related: CP-52 (a submit landing in the DO suspension teardown is dropped: the opposite failure, the run never wakes); CP-65 (the Temporal resume handle does not settle while a call is still pending); peer-reported by S5 (2026-09-27, with a reviewer probe on its branch); the DO and JS behaviour were reproduced independently on main here; the fix is planned in S5's snapshot MR (ruling R42), cell submit.partial-batch; severity S1: lost user action. The client's real answer to c2 is discarded, the model is told the runtime restarted, and the run completes as a success; the late submit is acknowledged as already_completed, so nothing surfaces the losssource, live-repro, peer-reportedverifiedopen
CP-83S2resume() against a live run is accepted on Temporal (every mode) and on runtime-js (from_checkpoint), and it takes the session away from that run. Temporal: the resume CAS accepts active as a source state, and nothing distinguishes an active session that is awaiting client-tool submissions from one whose run is still executing. Accepting active is deliberate and must stay: in the v7 stateless model a session awaiting a client tool can read active (the SF-D2 comment at executor.ts:1253-1264), and a Temporal session re-suspended on a partial client-tool batch reads active while the remaining calls are pending, so a later resume() on it is the only way to finish the batch (CP-82's Temporal contrast, CP-65). The defect is that an active session with no pending client-tool calls, i.e. a live run, is accepted too: resume() starts a __resume-N workflow beside the live one. In the probe the live execute() run ended failed, the resumed run completed, and the live run's answer never reached the transcript. runtime-js: from_checkpoint skips the running-session pre-check the other modes have, so the same happens there: the live run is superseded and fails, the resumed run fails on the ended stream, and the session ends failed. The reference behaviour is DBOS's guard, which accepts active only when client-tool calls are pending and otherwise throws AgentAlreadyRunningError; the CF DO also refuses (409). CF Workflows has the same acceptance, recorded as CP-35.runtime-temporal/src/executor.ts:1253-1264 ([SF-D2]: the CAS accepts 'active' on purpose, because a session awaiting a client-tool submission can read 'active') and :1265-1275 (compareAndSetStatus(['paused', 'interrupted', 'active'] → 'active'): no check for pending client-tool calls or a live workflow, so an 'active' session with none pending, i.e. a running one, is accepted too) and :1364 (a new __resume-N workflow is started; options.mode is never read, CP-31, so every mode behaves the same); contrast: reference behaviour, runtime-dbos/src/lifecycle/resume.ts:156-311 (the DBOS guard: when the CAS finds 'active', it loads the session and accepts only if pendingClientToolCalls is non-empty; 'Active session with no pending client tools — genuinely running' throws AgentAlreadyRunningError at :311). A fix that drops 'active' from the Temporal CAS wholesale would strand the partial-batch continuation (CP-82 contrast, CP-65); the discriminator is pending client-tool calls (or a live workflow), not the status alone; runtime-js/src/js-agent-executor.ts:2098-2126 (from_checkpoint loads the checkpoint with no status check; the running/completed/failed pre-checks at :2128-2142 apply only to the other modes) and :2304-2325 (the from_checkpoint save sets state.version = the version just loaded, so it is no guard against a live run: a version pin only arbitrates between concurrent resumers, since a resumer reads the live run's current version); live: live-repro (batch-6 verification, origin/main e37d676945, no product change; HELIX_TEMPORAL_BACKEND=test-server, runtime-temporal integ harness, InMemoryStateStore): an agent whose LLM call is held, so the run is live and no client-tool call is pending; while the execute() run is active, resume({ mode: 'continue' }) and resume({ mode: 'with_message' }) both returned a handle without throwing. After the hold was released the original execute() result was failed, the resumed handle completed, 2 LLM calls were made, and the persisted transcript was [user hi, assistant b]: the live run's answer is not in it; live: probe (batch-6 verification, unit level: runtime-js JSAgentExecutor + InMemoryStateStore): with the execute() run's LLM call held, resume({ mode: 'from_checkpoint', checkpointId: <the latest checkpoint> }) returned a handle. The live run failed with "Executor superseded for session … - another executor has taken over", the resumed run failed with "Cannot write to stream … in 'ended' state", and the session ended failed with only [user hi] persisted; contrast: runtime-js continue / with_message / with_confirmation on a running session throw (untyped, DI-31); DBOS (RUNTIME_DBOS_INTEG=1 probe, batch-6 verification) threw AgentAlreadyRunningError for continue, with_message and from_checkpoint on a running session with no pending client tools (runtime-dbos/src/lifecycle/resume.ts:311); the CF DO /resume returns 409 while isExecuting (runtime-cloudflare/src/do/durable-object-agent-base.ts:2050-2051); related: CP-35 (CF Workflows resume() also accepts a live run and starts a second instance, and truncates first); CP-31 (Temporal resume ignores options.mode); CP-60 (Temporal resume on a completed session returns the old result); CP-65 and CP-82 (the Temporal partial-batch path that legitimately resumes an active session); DI-31 (the refusals that do happen on runtime-js are untyped); found while checking the B2 report on the JS untyped error (batch-6 item 13), not itself peer-reported; related: the open MR !279 (B2-core, not yet on main) addresses this record on both affected runtimes (Temporal and runtime-js; the CF Workflows variant is CP-35), with red-first tests: its from_checkpoint guard (a typed AgentAlreadyRunningError on runtime-js, CF Workflows and Temporal) and, from its fix round 2de33b56b9, Temporal and CF Workflows resume() in continue / with_message / with_confirmation now reject an active session with no pending client-tool calls (a live run) before any write, as runtime-js and DBOS already did, while an active session with pending calls stays resumable. This record stays open, describing main, until the lifecycle.resume-while-running conformance cell proves the fix; severity S2: a stray or duplicated resume() (a reconnecting client, a retrying caller) kills the user's in-flight turn: its caller gets failed and its answer is lost. It needs a resume against a running session, and the turn can be resentsource, probe, live-reproverifiedopen

Failure-surfacing discipline ​

IDSevFindingEvidenceProvenanceVerificationStatus
DI-01S2auto_resume_exhausted reaches only the stream; compareAndSetStatus cannot persist errorDetail, so /status and snapshots lack the code; the code is outside ErrorCode.runtime-cloudflare/src/do/durable-object-agent-base.ts:1219-1233, 1250-1254; core/src/store/state-store.ts:349-357; core/src/errors/helix-error.ts:11-40sourceverifiedopen
DI-02S2DO wake-path resume failure and start-path uncaught failure set failed via updateStatus with no error or detail.runtime-cloudflare/src/do/durable-object-agent-base.ts:809-815, 1925-1931sourceverifiedopen
DI-03S2Hard abort carries NO error code: on JS (and the CF DO, which runs the JS executor) the step iterator returns a plain Error('Agent execution aborted'), and resolveErrorDetail → toErrorDetail projects a plain Error to { message } only — so the persisted errorDetail, the failed stream chunk, AgentResult.errorDetail and onAgentFail all lack framework_cancelled (an abort is indistinguishable from any uncoded failure). CF Workflows' abort branch likewise returns errorDetail: { message } with no code. (classifyError's framework_internal_error fallback is NOT on this path — it applies only to throws reaching CF Workflows' top-level catch.)core/src/orchestration/step-iterator.ts:485, 1922-1927 (abort before the LLM call / before commit → { kind: failed, error: new Error('Agent execution aborted') }, no errorCode / errorDetail); runtime-js/src/run-loop.ts:1080-1087 (hard-abort branch keeps kind failed), :1143 (resolveErrorDetail(finalOutcome.error, finalOutcome.errorDetail): the abort outcome carries no errorDetail); core/src/errors/error-detail.ts toErrorDetail (plain Error → { message, cause } — no code, no category); runtime-cloudflare/src/workflow.ts:1783-1831 (CF Workflows abort branch: status cancelled, errorDetail { message: abortReason ?? 'Aborted' }, no code); contrast: runtime-cloudflare/src/workflow.ts:3928 (classifyError → framework_internal_error only for errors thrown into the top-level catch); core/src/errors/classify-error.ts:16-31 (only a DOMException / Error named AbortError maps to framework_cancelled)sourceverifiedopen
DI-04S2Temporal top-level workflow catch persists errorDetail: { message }, dropping the classified code.runtime-temporal/src/workflow.ts:1570-1618 (the top-level catch: persistTerminalState with errorDetail: { message: errMsg } at :1592, and the returned result carries the same message-only detail at :1618); related: the citation was :1476-1497 until batch 5 (fix round 1) re-pointed it; that range now holds the stop-condition checksourceverifiedopen
DI-05S1DBOS never fires onAgentFail; failures fire onAgentComplete with output: undefined, so hooks and Langfuse record failures as successes.runtime-dbos/src/workflows/shared.ts:1374-1386, 2806-2816sourceverifiedopen
DI-06S2DBOS after a crash: per-workflow hook registry empty → silent fallback to the constructor hook manager; agent.hooks / options.hooks lost.runtime-dbos/src/steps/hooks.ts:199sourceverifiedopen
DI-07S2DefaultHookManager stops at the first throwing handler; suspend / resume hook failures are swallowed, so later handlers (e.g. Langfuse) silently never run.core/src/hooks/hook-manager.ts:39-55; core/src/orchestration/step-iterator.ts:1769-1810sourceverifiedopen
DI-08S2Skills catalog / preload failures go to console.warn and yield an empty catalog; skill-fs silently skips a missing root; load_skill "not found" is a successful string result.core/src/skills/resolve.ts:40-42, 65-74; skill-fs/src/filesystem-provider.ts:84-88; core/src/skills/skill-tools.ts:58-60, 97-99sourceverifiedopen
DI-09S2Langfuse spans orphan into runId-keyed traces whenever the in-memory run record is missing (restart, other worker, missing stateStore); CF Workflows fires suspend / resume only to agent.hooks; Checkpoint.tracingContext.lastActiveSpanId is never written.tracing-langfuse/src/stateless-span-emitter.ts:303, 411, 723; tracing-langfuse/src/langfuse-hooks.ts:430-436, 476-494; runtime-cloudflare/src/workflow.ts:136-143; core/src/types/checkpoint.ts:97sourceverifiedopen
DI-10S2Tokens of failed LLM calls are never recorded (debug log only); DBOS drops reasoningTokens; Temporal and CF Workflows drop cacheWriteTokens; recordTokens failures after a successful call fail the step (retry may double-count).llm-vercel/src/vercel-adapter.ts:531-536; core/src/usage/usage-hooks.ts:55-56; runtime-dbos/src/workflows/shared.ts:1784-1808; runtime-temporal/src/activities.ts:2054-2064; runtime-cloudflare/src/steps.ts:1368-1378sourceverifiedopen
DI-11S2DO usageStore factory failure silently falls back to the internal store (one warning); cross-session DOUsageStore writes misattribute; remote /usage unavailability silently omits child usage from rollups.runtime-cloudflare/src/do/durable-object-agent-base.ts resolveUsageStore; runtime-cloudflare/src/do/do-usage-store.ts:171-176, 305-308; core/src/orchestration/remote-sub-agent-completion.ts:240-250sourceverifiedopen
DI-12S2Memory: extraction LLM error → [] and the batch is never retried; retrieval, dedup and router errors swallowed; one suspension drain permanently disables background embedding for the process.memory/src/extraction/extractor.ts:65-67; memory/src/memory-manager.ts:69-78, 180-182; memory/src/execution/background-embedding-executor.ts:34-37, 66; memory/src/retrieval/retriever.ts:41-45sourceverifiedopen
DI-13S3Temporal persistTerminalState: agent-resolution failure silently skips terminal hooks; onAgentFail receives new Error(message) (code lost).runtime-temporal/src/activities.ts:4976-4981, 5017-5019sourceverifiedopen
DI-14S3Memory dedup UPDATE / DELETE ids are not validated against the candidate list or entity (adjacent; possible cross-entity mutation).memory/src/extraction/dedup.ts:258-300; memory/src/memory-manager.ts:460-482sourceverifiedopen
DI-15S2Capability rejections are untyped: Temporal and DBOS reject a workspace-declaring agent with a plain Error, not a typed framework_not_supported, so callers cannot branch on it. CF Workflows is worse: it does not reject at the call at all — the workflow body catches the same plain Error and returns a late status: failed result whose errorDetail is { message } only.core/src/workspace/runtime-integration.ts:34-46 (assertRuntimeSupportsWorkspaces throws new Error(...)); runtime-temporal/src/executor.ts:306, 1172, 1563 (execute / resume / retry call sites); runtime-dbos/src/dbos-agent-executor.ts:300, 373, 411 (execute / resume / retry call sites); runtime-cloudflare/src/workflow.ts:525-545 (call site inside the workflow body; caught and returned as { status: failed, errorDetail: { message } }); conformance: materialize.workspace-run on temporal-memory/temporal-redis/temporal-postgres/dbos-postgres — materialize.workspace-run.rejected-typedsource, probeverifiedopen
DI-16S3A resumed/woken session opens a new run without closing the stranded one: nothing ever writes run status 'superseded', so a permanent running run row stays in listRuns; if a wake bails before createRun (MA-01/MA-02) the ghost stays the current run — its runId/turn always leak into the snapshot, and its startSequence leaks too because the row is still status running.runtime-js/src/js-agent-executor.ts:2331-2352 (resume creates a run; previousRun read only for turn); runtime-js/src/js-agent-executor.ts:2863-2880 (retry — same); core/src/types/session.ts:353 (superseded defined, never written by any runtime); ai-sdk/src/handler/snapshot.ts:66 (only read site for superseded, deny-list check); runtime-dbos/src/steps/run-lifecycle.ts:190 (the switch case: superseded maps to null, no session-level transition); runtime-cloudflare/src/do/do-state-store.ts:2922-3039 (getCurrentRun = highest turn; listRuns returns every row by default); runtime-cloudflare/src/do/durable-object-agent-base.ts:3709-3712 (startSequence gated on status running), 3765-3766 (runId/turn exposed unconditionally for any currentRun); live: incident 2026-09-24 audio-resume — run 8c49120f (turn 1) stuck running, step_count 0, no completed_at, under later turns; inferred: the ghost becoming the CURRENT run needs a wake that bails before createRun (MA-01/MA-02); not observedsource, live-repropartly-inferredopen
DI-17S2runtime-js (and so CF DO, whose DO base embeds runtime-js's JSAgentExecutor) and Temporal's regular server-tool path build a failed server tool's persisted tool_result by hand (JS: abort, invalid input, thrown; Temporal: a failed per-tool activity result, a thrown tool) without metadata.helixToolFailed, while DBOS and CF Workflows build it through core createToolResultMessage; the ai-sdk converter decides error-vs-success ONLY from that flag, so a reloaded JS / CF DO / Temporal server-tool failure renders as a successful tool call.runtime-js/src/run-loop.ts:2010-2030 (pre-fix, origin/main 9c1aed51be) (abort + invalid-input branches build ToolResultMessage by hand, no metadata) and :2332-2338 (thrown-tool branch, same); runtime-cloudflare/src/do/durable-object-agent-base.ts:79 (the DO base imports and runs runtime-js JSAgentExecutor — CF DO inherits the runtime-js tool-result build); runtime-temporal/src/workflow.ts:892-906 (pre-fix, origin/main 9c1aed51be) (the per-tool executeToolActivity results map — failures included — builds each ToolResultMessage by hand: content JSON.stringify({ error: r.error }), no metadata); runtime-temporal/src/activities.ts:965-972 (pre-fix, origin/main 9c1aed51be) (executeServerToolWithHooks catch: a thrown server tool builds its ToolResultMessage by hand, no metadata); runtime-temporal/src/workflow-dispatch.ts:642-660 (pre-fix, origin/main 9c1aed51be) (companion tool dispatch: both the thrown-activity catch and the failed compResult build the ToolResultMessage by hand, no metadata); ai-sdk/src/converter/helix-to-aisdk-converter.ts:141-142 (isToolResultError reads only metadata[COMMON_METADATA_KEYS.TOOL_FAILED]); contrast: core/src/orchestration/message-builder.ts:189, 220 (createToolResultMessage / createSubAgentResultMessage stamp the flag); runtime-dbos/src/workflows/shared.ts:2330 and runtime-cloudflare/src/steps.ts:2107 (DBOS and CF Workflows build server-tool results — failures included — through createToolResultMessage); conformance: discipline.failed-tool-marked on js-memory / js-redis / js-postgres / temporal-memory / temporal-redis / temporal-postgres — proves the fix for a failing REGULAR server tool: pre-fix .thrown-marked-failed and .invalid-input-marked-failed failed there (the result was present, single and carried the error, but lacked helixToolFailed); post-fix every id passes. dbos-postgres is contrast, not proof: it passed pre-fix; conformance: transcript.finish-with-fails-then-retries on js-memory / js-redis / js-postgres — proves the runtime-js fix for a failing finishWith: pre-fix .failed-call-marked-failed failed there; post-fix it passes (it passed on DBOS and Temporal pre-fix); not yet reachable (conformance cell) — CF DO (inherits the runtime-js fix via JSAgentExecutor), CF Workflows, the runtime-js abort branch and Temporal companion tool dispatch (runtime-temporal/src/workflow-dispatch.ts ~642-660) are covered by runtime unit tests only (runtime-js/src/__tests__/run-loop-tool-error-metadata.test.ts, runtime-temporal/src/__tests__/tool-result-helix-tool-failed.test.ts), pending the Phase B cf-do / cfw conformance cells; DI-17 re-opens if any of those cells failssource, peer-reportedverifiedfixed (9 cells)
DI-18S2CF Workflows handle.result() stops polling on ANY workflow instance status other than 'running' / 'queued' and memoizes whatever that maps to. A non-terminal status therefore becomes a permanent { status: 'failed', error: 'Unknown error' } for that handle while the run goes on to complete. Cloudflare's own types declare such statuses — 'waiting' ("hibernating and waiting for sleep or event"), 'waitingForPause', 'unknown' — but the runtime's WorkflowStatus type omits them. Observed on a retry / resume handle under vitest-pool-workers local-dev Workflows. Production is not observed, but the agent workflow sleeps and waits for events on ordinary paths (sub-agent and companion waits, respawn backoff), which is exactly when Cloudflare documents 'waiting'.runtime-cloudflare/src/executor.ts:1457-1499 (createResumedHandle — returned by retry() :1087, resume() :842 and handle.resume() :1437): result() memoizes at :1471 / :1473; the poll at :1477-1486 treats every status except 'running' / 'queued' as done (:1480-1483); runtime-cloudflare/src/executor.ts:1566-1609 (createActiveHandle — the execute() handle): the same memo (:1583 / :1585) and the same done-on-anything-else poll; runtime-cloudflare/src/executor.ts:400-443 (mapWorkflowStatusToAgentResult: unless status is 'complete' AND the persisted session is 'completed', it returns failed with error status.error?.message ?? state?.error ?? 'Unknown error' — a live run has neither, hence 'Unknown error'); runtime-cloudflare/src/bindings.ts:165-181 (WorkflowStatus / WorkflowStatusSchema list only queued | running | complete | errored | terminated | paused); @cloudflare/workers-types (latest/index.d.ts, InstanceStatus) additionally declares 'waiting' — commented "instance is hibernating and waiting for sleep or event to finish" — 'waitingForPause' and 'unknown'; runtime-cloudflare/src/workflow.ts:2620-2665 (ephemeral sub-agent wait: step.waitForEvent), :3326-3360 and :3440-3460 (persistent companion waits: step.waitForEvent), :970 and :1002 (respawn-poll backoff: step.sleep) — ordinary agent-workflow paths during which Cloudflare documents the instance as waiting; live: peer-reported (vitest-pool-workers local-dev Workflows) — a retry / resume handle resolved result() to { failed, Unknown error } while the workflow run completed; inferred: that the production Workflows API returns 'waiting' from instance.status() during these waits is taken from the workers-types comment, not observed; which transient status local dev reported was not identified; severity S2: the caller gets a wrong terminal outcome, but it is bounded to that handle's memoized result — the run itself completes and its persisted session state is correct, so re-reading session state is a workaround; it would be S1 if confirmed in production on sub-agent / companion runs, where result() would report failure for runs that succeedsource, peer-reportedpartly-inferredopen
DI-19S1DBOS reports a FAILED ephemeral sub-agent to its parent as a success with undefined output: executeSubAgentInWorkflow awaits the child workflow's getResult() and builds { ok: true, result: childOutput.state.output } without reading the child's terminal status. A child that ends 'failed' through a failure-finalize path (for example forced-completion exhaustion) returns normally from its workflow rather than throwing, so the parent's tool result and subagent_end say the delegation succeeded, and the parent model continues on a success that never happened.runtime-dbos/src/workflows/standard-workflow.ts:267-273 (const childOutput = await childHandle.getResult(); stepResult = { ok: true, result: childOutput.state.output, … } — no check of childOutput.state.status); only a THROWN error reaches the ok:false branch (:274-281); runtime-dbos/src/workflows/standard-workflow.ts:558-576 (the workflow returns { state, completionReason } normally after the loop, whatever the session's final status, including 'failed'); runtime-dbos/src/workflows/shared.ts:1356 (finalizeFailure) and :1414-1424 (finalizeForcedFailure → finalizeFailure) — failure-finalize paths RETURN a RunOneTurnResult; they do not throw; contrast: runtime-js/src/run-loop.ts:4516-4519 (runSubAgentLoop: a child outcome of kind failed is THROWN, so the parent sees the delegation fail); inferred: the parent-facing effect (tool result ok:true with no output, subagent_end reported as success) follows from the code; not reproduced by a test; related: CP-69 / FU-DBOS-BLOCKING-SPAWN-SEMANTICS (the persistent-child variant); the remote-child variant is already handled (agent-server synthesizeEndEvent); related: reported by B8 from a read-only triage of MR !247, whose commit 10662e6f8d fixes this (!247 is being closed as superseded)source, peer-reportedpartly-inferredopen
DI-20S2With includeStepEvents: true (the opennext example's and the consumer app's setting) AND a model adapter that reports step boundaries (MockLLMAdapter, the conformance ScriptedModel, custom adapters — NOT @helix-agents/llm-vercel, which never reports them), the ai-sdk StreamTransformer emits step-boundary UI chunks carrying fields that the ai package's UI-chunk schema forbids: finish-step carries stepId / usage / finishReason, and start-step carries stepId whenever the step has one. ai 6.0.0 through 6.0.230 declare both chunk types as strictObject({ type }) (6.0.231+ switched to looseObject and accept them), and DefaultChatTransport throws on the first chunk that fails to parse. So a stock useChat / Chat client on an affected ai version, including 6.0.49 (the workspace) and 6.0.91 (the consumer's pin), ends every such response with status 'error' ('Type validation failed'). In a client-tool round it errors before the tool part arrives, so no handler runs and no result is ever auto-sent. The healthy stream is reported to the user as a failure, and any client tool stalls.ai-sdk/src/transformer/stream-transformer.ts:1344-1354 (pre-fix, origin/main e0aae15185) (transformStepStart emits { type: start-step, stepId }); ai-sdk/src/transformer/stream-transformer.ts:1365-1381 (pre-fix, origin/main e0aae15185) (transformStepEnd emits { type: finish-step, stepId, usage, finishReason }); live: ai 6.0.49 dist/index.mjs ~4745-4750 declares start-step and finish-step as z.strictObject({ type }); DefaultChatTransport.processResponseStream (~11948-11961) throws the first failed parse. The reliability orchestrator confirmed the same strict schema and throw in ai 6.0.91, the consumer's pin (index.mjs ~4912-4916, ~12421-12430); related: the opennext example enables it — examples/opennext-cloudflare-do/src/lib/agent-client.ts:50 and coordinator-agent-client.ts:24, :33 (transformerOptions: { includeStepEvents: true }); live: pre-fix, the Phase B cf-do host-driver integ (a real DO chat route parsed through ai's real schema) observes a 'Type validation failed' error naming the finish-step chunk, and the client-tier driver running the real useHelixChat / DefaultChatTransport observes Chat status 'error' on both the cf-do and js-chat-host lanes. start-step passed there only because those steps carried no stepId; live: the owning worker (reliability-ui-chunk-schema) reproduced it with ai's real Chat + DefaultChatTransport across ai 6.0.0 through 6.0.230 (strict schema); ai 6.0.231+ switched these chunk types to looseObject, so the break depends on the consumer's pinned ai version; inferred: the same worker reports further schema-invalid emissions of the same class (a tool-approval-response chunk type no ai 6 version accepts; file.filename; tool-approval-request.isAutomatic; replayed tool-output-denied carrying message / dynamic) — not individually verified here; the fix MR covers them under this finding; contrast: llm-vercel/src/vercel-adapter.ts:222 maps streamText's onChunk, and ai calls onChunk only for text-delta / reasoning-delta / source / tool-call / tool-result / tool-input-start / tool-input-delta / raw (ai 6.0.49 dist/index.mjs:6059), so the adapter never calls onStepStart / onStepEnd and runtime-js emits step_start / step_end only from those callbacks (runtime-js/src/run-loop.ts:3637-3660) — with llm-vercel no step chunk reaches the UI stream, so this defect does not fire. DBOS emits step_start itself (runtime-dbos/src/steps/call-llm.ts:100) but without stepId, which passes the strict schema; inferred: the opennext example e2e passes because its real llm-vercel adapter emits no step chunks at all (the owning worker captured its chat-route SSE: no start-step / finish-step); related: an earlier hypothesis linked this to the consumer app "stops updating live" report; with llm-vercel (no step chunks) it cannot be the cause there unless the app uses a step-reporting adapter — not confirmed; related: fixed by !268 (0e618772ee, "fix(ai-sdk): emit only ai-schema-valid UI chunks (DI-20)") — ai-sdk/src/transformer/stream-transformer.ts emits bare start-step / finish-step, no tool-approval-response, no isAutomatic, file.filename in providerMetadata.helix; ai-sdk/src/transformer/replay-events.ts emits bare tool-output-denied; ai-sdk/src/types.ts narrows every AISDK*Event to the ai 6.0.0 chunk shape; live: ai-sdk/src/__tests__/ui-chunk-schema-contract.test.ts validates the full emission surface (every StreamChunk type × option matrix, replay for every ReplayToolState, handler paths incl. resume-with-rejections and reconnect-skip) as wire JSON against pinned ai 6.0.0 / 6.0.49 / 6.0.91 / 6.0.230 / 6.0.281, plus a stock AbstractChat + DefaultChatTransport round-trip and a client-tool auto-send (C4) per version; red on the pre-fix code (34 failing, including each latent emission individually: tool-approval-response on all versions, file.filename / isAutomatic / replayed tool-output-denied on 6.0.0–6.0.230), green after; compile-time exact-keys guard vs ai 6.0.0 UIMessageChunk; live: e2e/src/helpers/stock-ai-client.ts drives the stock ai client in every ai-sdk-* family (js, redis-js, temporal, redis-temporal, dbos, d1-js integ, d1 DO-streams cf) and e2e/src/__tests__/ai-sdk-do-chat-route.cf.test.ts drives createCloudflareChatHandler → a real DO in workerd; 7/8 red on the pre-fix build ("Unrecognized key: finishReason"), all green after (DBOS passed pre-fix by omission — it never emits step_end, RM-49; pinned there by an it.fails finish-step non-vacuity test); conformance: stream.ui-chunk-schema (host tier) parses every exchange of a text run and a paused client-tool round through ai 6.0.49's real UI-chunk pipeline (parseJsonEventStream + the strict uiMessageChunkSchema into readUIMessageStream). Pre-fix, .client-accepts-stream failed with 'Type validation failed' naming the finish-step chunk and .step-events-bare saw finish-step carry finishReason, on cf-do and js-chat-host, 5/5 per cell. Post-fix (!268) both ids pass, 5/5 per cell. Non-vacuous: .step-events-bare requires at least one start-step and one finish-step frame, so a stream with no step events cannot pass; conformance: stream.ui-chunk-schema-client (client tier, the real useHelixChat / useResumeClientTools hooks over ai 6.0.49's DefaultChatTransport): pre-fix every error the Chat reported across its text, answered-tool, paused and reloaded sessions was that same finish-step rejection, and in the tool round the Chat errored before the tool part arrived. Post-fix every id passes on cf-do and js-chat-host, 5/5 per cell. Non-vacuous: .client-step-events-present requires the Chat to have read start-step and finish-step and to hold a step-start part; conformance: corroborating, NOT proof (so not in provenBy — they assert nothing about step frames, and would keep passing if the hosts stopped emitting step events): the same rejection masked OTHER scenarios' client cells pre-fix, and they now pass: smoke.hello-completes (smoke.completes: pre-fix outcome chat:error) on cf-do and js-chat-host, and submit.no-race-client on js-chat-host (pre-fix .submitted: the Chat errored on finish-step{finishReason:'tool-calls'} before the tool part, so nothing was auto-sent). They run the same hosts (includeStepEvents: true) and the same step-reporting ScriptedModel; the harness self-tests (harness/cf-do/__tests__/client-driver.integ.test.ts) pin that the Chat reads finish-step in the same text run and client-tool round and still completes; conformance: unmasked by the fix, submit.race-a-client and submit.no-race-client on cf-do now reach the host and fail under the catalogued control-plane findings, re-attributed from the observed run (identical 5/5): race A's five post-.submitted ids to CP-52 (the held window, stranded paused) and the control's .submit-accepted / .continuation-reaches-caller to CP-53 (data-resume-rejected while the woken run completes). Those cells are not DI-20 proofsource, live-repro, peer-reportedverifiedfixed (4 cells)
DI-21S3resolveCheckpointStreamSequence swallows any stream-info read failure with a bare catch {} and returns its fallback, with no log. Every runtime caller passes 0 as the fallback, so a failed read silently records streamSequence 0 on the checkpoint. agent-server's remote snapshot then logs a degrade as 'zero_fallback_stream_sequence', and the original error is gone. The same 0 is written when the stream is merely absent, which the docstring intends.core/src/store/stream-manager.ts:545-578 (docstring: best-effort, never abort a commit; :572-577 try { getStreamInfo } catch { return fallback } — no logger, no rethrow); runtime-js/src/run-loop.ts:1647-1651 (fallback = checkpointMeta.streamSequence, which the step-iterator sets to 0), :1786-1790 (fallback 0); runtime-temporal/src/activities.ts:3193-3197, :4388-4392; runtime-cloudflare/src/steps.ts:4195-4199, :4777-4781, :4838-4842 (fallback 0); inferred: agent-server/src/agent-server.ts:719-721, :779-793 (the reader documents 0 as 'the step-iterator's fallback when the stream-info read failed' and degrades to -1 with a warn that names only the symptom); related: RM-37, RM-39 (checkpoints that record 0 without any read at all); severity S3: observability only — the degrade itself is logged downstream; what is lost is the cause of the failed readsource, peer-reportedpartly-inferredopen
DI-22S3CF Workflows' terminal-failure seam does not follow the resolveErrorDetail contract. Its outer catch runs classifyError + toErrorDetail instead, so it (a) ASSERTS framework_internal_error with retryable: false for any error it does not recognise, where runtime-js and Temporal's executor leave code and retryable unset for the same error; (b) drops an ErrorDetail attached to the error (latent today); and (c) never looks at error.cause, so a classified error wrapped by another error loses its code and gets framework_internal_error in its place (resolveErrorDetail does not read cause either, so on the other runtimes the wrapped code is also lost, but left unset rather than replaced). It persists that detail as the session's errorDetail, sends it on the failed stream chunk and returns it. Example: a D1StateError wrapping a transient fetch failed is recorded as a definite, non-retryable internal error. (No runtime classifies such a store error as retryable, and every runtime sends an unclassified error with recoverable: false; what is CFW-specific is the asserted classification and the lost detail.)runtime-cloudflare/src/workflow.ts:3927-3928 (outer catch: helixError = classifyError(error)), :3937-3943 (error-mark-failed persists errorDetail: toErrorDetail(helixError)), :3946-3953 (error-fail-stream sends the same detail on the failed chunk), :3984-3988 (the workflow result carries it too); core/src/errors/classify-error.ts:13-104 (branches only for HelixError, AbortError and seven framework error classes; any other Error, and any non-Error, becomes framework_internal_error with retryable: false, :90-104); it never inspects error.cause and never calls getErrorDetail; store-cloudflare/src/errors.ts:43-80 (D1StateError extends Error and keeps the underlying error only as cause) and store-cloudflare/src/d1-state.ts:341, :439, :510, :616 (D1 operations wrap every failure in it); runtime-cloudflare/src/steps.ts:3587-3598 (markAgentFailed stores the given errorDetail on the session as-is); live: probe (this session, core + store-cloudflare dist): toErrorDetail(classifyError(new D1StateError('saveState', 's1', new TypeError('fetch failed')))) = { code: 'framework_internal_error', category: 'framework', retryable: false, … }, while resolveErrorDetail of the same error has no code and no retryable; for an Error built by fromErrorDetail({ code: 'provider_rate_limited', retryable: true }), classifyError returned framework_internal_error / retryable: false while resolveErrorDetail kept provider_rate_limited / retryable: true; contrast: (partial) runtime-js/src/run-loop.ts:1201 and runtime-temporal/src/executor.ts:762 derive the detail with resolveErrorDetail, which keeps an attached ErrorDetail and, for a plain Error, sets no code and no retryable flag rather than asserting false; Temporal's workflow-level catch does not use it either (DI-04); contrast: the wire behaviour is the same everywhere — every runtime's LLM onError sends an unclassified error with recoverable: false and no code (docs/guide/error-handling.md, Runtime Error Handling; runtime-cloudflare/src/steps.ts:1198-1211), so no runtime tells a consumer that a transient store error is retryable; inferred: the D1 trigger — a transient fetch failure that outlasts the step.do retries and escapes to the outer catch — was peer-reported and was not reproduced here; whether the error keeps its D1StateError class across the step boundary does not matter, because classifyError treats every unrecognised Error the same way. The attached-ErrorDetail loss is latent: nothing in runtime-cloudflare throws an error built by fromErrorDetail today; related: DI-04 (Temporal's top-level workflow catch persists a detail that does not come from resolveErrorDetail either); DI-03 (a hard abort carries no error code); DI-24 (Temporal / DBOS lose the classified detail at their terminal seam); CLAUDE.md "Structured terminal-failure detail" (states the resolveErrorDetail precedence as the cross-runtime contract and records that CF Workflows never calls it); severity S3: the failure itself is reported; the persisted detail asserts a classification the error does not support and can drop a real one, so a consumer keying on code / retryable is misled on CF Workflows onlysource, probe, peer-reportedpartly-inferredopen
DI-23S2LLM adapters call the async callbacks.onError fire-and-forget, so a failing error emit becomes an unhandled promise rejection that can crash the process. llm-vercel's catch path calls callbacks.onError(helixError) without awaiting it (its stream onError path does await it). The core MockLLMAdapter does the same on both of its error paths. Every runtime that passes callbacks (JS, Temporal, CF Workflows) implements onError as an awaited write of an error chunk to the run's stream. When that stream is already ended or failed, the write rejects, nothing awaits the promise, and under Node's default --unhandled-rejections=throw the process exits — on Temporal that is the worker. On llm-vercel the dead-stream case feeds itself: a rejected onTextDelta write is what sends the adapter into the catch path in the first place.llm-vercel/src/vercel-adapter.ts:525-529 (catch path: callbacks.onError(helixError) with no await) — contrast :310-315 (the streamText onError path awaits callbacks.onError) and :222 / :236 (onTextDelta is awaited inside the onChunk callback passed to streamText; a rejected write surfaces to the adapter outer catch at :507, as the probe below shows); core/src/llm/mock-adapter.ts:301-303 (graceful-exhaust path) and :412-414 (error response path): input.callbacks.onError(error) with no await; runtime-temporal/src/activities.ts:2921-2923 (emit = await writer.write) and :2951-2964 (onError awaits emit of an error chunk); runtime-js/src/run-loop.ts:3548-3550 and :3579 (the same shape, also used by the CF DO, which embeds the JS executor); runtime-cloudflare/src/steps.ts:1198-1211 (CF Workflows onError awaits writer.write); store-memory/src/in-memory-stream.ts:236-240 and store-redis/src/redis-stream.ts:653 (a write to an ended / failed stream throws); no package registers a process 'unhandledRejection' handler (grep over core, runtime-*, agent-server, llm-vercel src); contrast: runtime-dbos/src/steps/call-llm.ts:115-132 (DBOS passes no stream callbacks to generateStep; the workflow body emits chunks after the step), so DBOS is not affected; live: probe (this session, core + store-memory dist): MockLLMAdapter error response with an onError that writes to an ended InMemoryStream — generateStep returned type 'error' normally, then node exited with code 1 on the unhandled "Cannot write to stream s1 in 'ended' state"; live: probe (this session, llm-vercel + store-memory dist, ai/test MockLanguageModelV3): with the run's stream ended, onTextDelta's write rejected, the adapter returned type 'error', and node then exited with code 1 on the unawaited onError write. When the provider call itself fails (doStream rejects, invalid prompt, gateway auth), the error goes through the awaited stream onError path instead and the process survived; live: B2-core (peer) conformance-temporal run on its branch reliability-b2-b11 (~/hb2b11/.superpowers/sdd/2026-09-25-llm-history-completeness/conf.log, read here): vitest reported 65 'Uncaught Exception' entries, all 'Cannot write to stream cf-history.companion-continue-temporal-*-s' (44 in 'failed' state, 21 in 'ended'), each with the stack Object.write → emit → Object.onError (runtime-temporal activities.ts on that branch) ← ScriptedModel.generateStep (core/src/llm/mock-adapter.ts:302) ← runLLMStep. All 202 tests still passed: vitest caught the rejections that would have exited a plain Node process; inferred: a production Temporal worker (llm-vercel, real stream) crashing on this path follows from the probes and the Node default; not observed in a deployed worker. A Temporal worker could also run with a non-default unhandled-rejection mode, which this repo does not set; related: CP-23 (a child whose parent was interrupted keeps making LLM calls after its stream ended — the peer-reported source of the dead-stream writes in the log above); CP-67 / CP-76 (resume() writing to an ended root stream is another way to reach it); severity S2: bounded to an LLM error or rejected callback write while the stream is already ended or failed, but the effect is a process crash that takes down every other run on that worker, with the real cause only in the crash logsource, probe, live-repro, peer-reportedpartly-inferredopen
DI-24S2Temporal and DBOS drop the classified ErrorDetail of an ordinary terminal LLM/provider error at their terminal-failure seam, so the workflow result and the persisted SessionState.errorDetail carry only { message }. On Temporal, runLLMStep has the full detail in memory but returns only the error string in its terminal result; the workflow then builds { message } and persists and returns that (the stream's error chunk, written inside the activity, keeps the code). On DBOS, the LLM call is a DBOS step whose result is serialized and deserialized even on its first run, so the HelixError comes back as a plain Error without code, and the detail derived from it is { message } — for the persisted state and, from source, for DBOS's own failed chunk too. A reconnect / getStatus consumer therefore cannot see the code or retryable flag of the failure. This breaks the documented errorDetail-propagation invariant (CLAUDE.md). runtime-js keeps the detail only when the adapter throws; an error-type step result loses it there too (DI-30).runtime-temporal/src/activities.ts:2394-2398 (runLLMStep's terminal result for a failed plan carries only error: plan.statusUpdate.error — the plan's errorDetail is not returned) and :291-295 (the terminal type has no errorDetail field); runtime-temporal/src/workflow.ts:1459-1473 (non-forced LLM error-stop: failErrorDetail = { message: failError }, passed to persistTerminalState and returned); runtime-temporal/src/activities.ts:4977-4994 (persistTerminalState only backfills an ABSENT detail, so the { message } is what persists); live: probe (this session): a copy of e2e/src/__tests__/empty-error-detail-temporal.integ.test.ts with HELIX_TEMPORAL_BACKEND=test-server and a HelixError{ provider_rate_limited, retryable: true } error response — stream chunk errorDetail { message, code: 'provider_rate_limited', category: 'provider', retryable: true }; workflow result errorDetail {"message":"Rate limited"}; persisted SessionState.errorDetail {"message":"Rate limited"} (status failed). The repo test asserts only the chunk, so it passes. Also reproduced by the conformance verifier (peer); live: probe (batch-6 verification, origin/main e37d676945, HELIX_TEMPORAL_BACKEND=test-server, runtime-temporal integ harness, a HelixError { code: provider_rate_limited, retryable: true } error step): still reproduces on main — workflow result errorDetail {"message":"Rate limited"} and persisted SessionState.errorDetail {"message":"Rate limited"}; runtime-dbos/src/steps/call-llm.ts:85-86 (callLLMStep is a @DBOS.step); @dbos-inc/dbos-sdk 4.17.6 dist/src/dbos-executor.js:571-578 (a step returns funcResult.deserialized — its serialized-then-parsed result — even on first execution) with dist/src/config.js:147 (default serializer DBOSJSON); core/src/orchestration/step-processor.ts:410-418 (an error step's statusUpdate.errorDetail = toErrorDetail(stepResult.error), which is { message } for a plain Error); runtime-dbos/src/workflows/shared.ts:2797-2813 (that detail goes to finalizeRun, else { message }), :2841-2848 and :2868-2871 (the same detail on DBOS's error chunk and failStream); runtime-dbos/src/steps/run-lifecycle.ts:357 (persisted onto SessionState); live: probe (this session, DBOS serializeFunctionInputOutput with DBOSJSON): { type: 'error', error: HelixError(provider_rate_limited), shouldStop: true } round-trips to a plain Error with code undefined (HelixError.isInstance false); planStepProcessing then yields statusUpdate.errorDetail {"message":"Rate limited"} where the in-memory original yields { message, code: 'provider_rate_limited', category: 'provider', retryable: true }; inferred: the DBOS end-to-end effect (persisted errorDetail and the failed chunk without the code) follows from the step round-trip above; not run on a live dbos-postgres backend; contrast: runtime-js/src/run-loop.ts:1143 keeps the classification only when the adapter THROWS; an error-type step result (llm-vercel's path) loses it at core/src/orchestration/step-iterator.ts:1123-1127 — DI-30; related: DI-30 (runtime-js and the CF DO lose the same detail at the step iterator when the adapter returns an error step); DI-04 (Temporal top-level workflow catch persists { message }: a different seam); DI-22 (CF Workflows outer catch asserts a classification instead of resolving it); severity S2: the failure itself is reported, but its machine-readable code and retry flag are silently lost for durable readers (reconnect, getStatus, AgentResult) on two runtimessource, probe, live-repro, peer-reportedpartly-inferredopen
DI-25S2The runtimes disagree on how a run that exhausts maxSteps is reported for an agent without outputSchema, and no contract says which is right. Exhaustion: JS, the CF Durable Object (which embeds the JS loop), Temporal and DBOS end a run that runs out of maxSteps mid-tool-loop completed with undefined output, usually right after a tool result. CF Workflows ends it failed with "Max steps (N) reached", so retry() and a parent's sub-agent result behave differently there. completed is the majority practice, but the forced-completion design called it "a silent semantic downgrade", the CF Workflows test keeps failed on purpose, and docs/internals/subagent-execution.md uses a sub-agent failing with "Max steps exceeded" as its example. Signal: on the runtimes that complete, nothing marks the stop as a truncation. Only DBOS writes RunMetadata.completionReason ('max_steps'), and DBOS's value is unreliable in general (DI-27). A caller can only guess it: stepCount === maxSteps is a clue but not proof (a complete answer can also land on the last allowed step), and a run that ends with no output right after a tool result is a stronger hint. The fix must decide the contract and document it. (Whether an unset maxSteps is capped at all is a separate cause: DI-28.)core/src/orchestration/step-iterator.ts:2096-2154 (maxSteps stop; for a non-outputSchema agent: "keep the legacy completed + undefined behavior", :2122, then { kind: 'completed', output: undefined }, :2153-2154; runtime-js and the CF DO run this iterator); runtime-temporal/src/workflow.ts:1476-1567 (shouldStopExecution at maxSteps; a non-outputSchema agent falls through to persistTerminalState status 'completed' at :1557-1567); runtime-dbos/src/workflows/shared.ts:2688-2808 (maxStepsExceeded → stopTag 'max_steps' → finalizeRun status 'completed' with completionReason 'max_steps', the only runtime that writes the field); runtime-cloudflare/src/workflow.ts:546 (maxSteps defaults to 50), :1693 (loop guard) and :3855-3925 (a non-outputSchema agent that exits the loop at maxSteps is marked failed with "Max steps (N) reached"); runtime-cloudflare/src/__tests__/workflow.test.ts:1185-1208 pins that failure under the title "should fail with Max steps reached for a non-outputSchema agent (unchanged)", while core/src/orchestration/__tests__/step-iterator-forced.test.ts:513 pins the opposite ("a non-outputSchema agent at maxSteps still completes with undefined output"); core/src/types/session.ts:48-56 (CompletionReason includes 'max_steps' alongside, not as, 'failed') and :429-434 (RunMetadata.completionReason: "Populated when the run reaches a terminal state"); runtime-js, runtime-temporal and runtime-cloudflare never write completionReason (only runtime-dbos/src/workflows/shared.ts:2788-2806 does); contrast: evidence for completed — docs/guide/agents.md:176 ("Without outputSchema, the agent runs until maxSteps or a stopWhen condition"), the core iterator comment above, and four of five runtimes; contrast: evidence for failed — docs/superpowers/specs/2026-07-20-deterministic-forced-completion-design.md:34-35 ("normalized to completed / output: undefined / error: null on most runtimes — a silent semantic downgrade"), fixed only for outputSchema agents, with non-outputSchema behavior left "unchanged" at :396 (each runtime kept what it did; no contract was chosen); the CF Workflows test title above ("(unchanged)"); docs/internals/subagent-execution.md:529-551 (the canonical sub-agent failure example is error: "Max steps exceeded"); live: live-repro (batch-5 verification, origin/main 3560d6c72a, no product change; an agent without outputSchema, maxSteps 3, the LLM calls a server tool every step): runtime-js (JSAgentExecutor + InMemoryStateStore) result completed, session completed, run record completed with no completionReason, last message role tool; Temporal (HELIX_TEMPORAL_BACKEND=test-server, runtime-temporal integ harness) result completed, session completed, run record 'running' (RM-47) with no completionReason; DBOS (RUNTIME_DBOS_INTEG=1, private Postgres DB on :5433, Redis db 13 on :6390) result completed, run completed with completionReason 'max_steps'. CF Workflows: the existing workflow.test.ts case above passes on origin/main (status failed, error contains 'Max steps'); related: RM-47 (Temporal never writes a terminal run status at all); DI-26 (the same runtimes also disagree on whether an error stop reason fails a non-outputSchema run); DI-27 (DBOS completionReason is failed for every ordinary text completion); DI-28 (the documented maxSteps default of 50 is not applied on JS, the DO or Temporal); the peer report (B2) framed JS/Temporal as the defect; this record states no contract and leaves the choice to the fix; severity S2: no data is lost, but the same agent is 'failed' on CF Workflows and 'completed' elsewhere, so retry() and a parent's sub-agent result diverge, and the completing runtimes give no machine-readable truncation signal at all (stepCount === maxSteps is only a clue; DBOS's completionReason aside, DI-27)source, live-repro, peer-reportedverifiedopen
DI-26S1Temporal and DBOS report completed when an agent without outputSchema stops on an error stop reason. The affected reasons are content_filter, refusal, max_tokens, error and unknown on a text step. The model's answer was blocked, refused or truncated, yet the stored session reads completed with no error. JS, the CF Durable Object and CF Workflows fail the same run: JS says "Content blocked by safety filters" / "Response truncated: token limit reached", CF Workflows "Agent stopped with error: REASON". Core's determineFinalStatus is the shared classifier, and isErrorStopReason defines the set. planStepProcessing gives a text step no status update, so a runtime must call that classifier itself. Temporal and DBOS never do. Their only failure signal is plan.statusUpdate failed, which only an error-type step sets. DBOS's run completionReason is no signal either: it reads 'failed' for every text completion, ordinary ones included (DI-27). llm-vercel returns exactly this shape (text, empty or not, shouldStop: true, stopReason mapped from the provider's finish reason) for max_tokens, content_filter, error and unknown, so this is the production path for a filtered or truncated reply. It never produces refusal; only a custom adapter can. outputSchema agents are not affected, because forced completion classifies these reasons on every runtime, except on DBOS persistent-mode follow-up turns, which run without the schema (MA-21).core/src/orchestration/stop-checker.ts:86-100 (determineFinalStatus: an error stop reason → failed) and core/src/types/runtime.ts:75-83 (isErrorStopReason: max_tokens, content_filter, refusal, error, unknown); core/src/orchestration/step-processor.ts:233-246 (a text step plans statusUpdate: null, so the failure must come from determineFinalStatus); runtime-temporal/src/activities.ts:2388-2399 (runLLMStep sets terminal only from plan.statusUpdate, so a text step with an error stop reason has none) and runtime-temporal/src/workflow.ts:1476-1567 (shouldStopExecution is true for shouldStop text; a non-outputSchema agent goes straight to persistTerminalState status 'completed'); runtime-dbos/src/workflows/shared.ts:2797-2808 (isFailed = plan.statusUpdate?.status === 'failed' only, so finalizeRun writes status 'completed'; the completionReason 'failed' written beside it is written for every text stop reason, end_turn included, so it does not mark this case: DI-27); llm-vercel/src/vercel-adapter.ts:438, :485-495 (a reply with text returns shouldStop: true with the mapped stopReason) and :498-506 (a reply with no content, the usual shape of a content-filtered finish, also returns type text, shouldStop: true and the mapped stopReason); core/src/llm/stop-reason.ts:25-40 (mapVercelFinishReason: 'length' → max_tokens, 'content-filter' → content_filter, 'error' → error, 'other' / unrecognised → unknown; there is no refusal case, so llm-vercel never yields refusal); contrast: core/src/orchestration/step-iterator.ts:1194-1204 and :2106-2115 (runtime-js, and the CF DO that embeds it, call determineFinalStatus after the commit and fail the run); runtime-cloudflare/src/steps.ts:1545-1595 (CF Workflows completes only when !isErrorStopReason, else fails with "Agent stopped with error: REASON"); live: live-repro (this verification, origin/main 3560d6c72a, no product change; an agent without outputSchema whose single LLM step returns text with shouldStop: true): runtime-js (JSAgentExecutor + InMemoryStateStore) content_filter → result failed "Content blocked by safety filters", session failed; Temporal (HELIX_TEMPORAL_BACKEND=test-server, runtime-temporal integ harness) content_filter → result completed, session completed, no error, and max_tokens → completed; DBOS (RUNTIME_DBOS_INTEG=1, private Postgres DB on :5433, Redis db 13 on :6390) content_filter → result completed, session completed (the run completionReason read 'failed', which DBOS writes for every text completion: DI-27); live: probe/live-repro (batch-6 verification, origin/main e37d676945): every runtime persists the content_filter assistant row. runtime-js (JSAgentExecutor + InMemoryStateStore, and a store whose saveState mints no checkpoint) fails the run ("Content blocked by safety filters") with [user, assistant FILTERED-PARTIAL] persisted under a messageCount-2 checkpoint, and retry() without a message then throws "No user message found to retry with" (CP-58); Temporal (test-server) and DBOS (private Postgres) store the same two rows with the session completed, and retry() refuses the completed session ("Cannot retry: agent status is 'completed'" / "not in 'failed' state"), so the filtered reply cannot be retried at all there. The batch-6 peer claim that JS never stores the filtered row is refuted; related: MA-21 (DBOS persistent-mode follow-up turns drop outputSchema, so an outputSchema agent is exposed there too: the forced-completion gate at runtime-dbos/src/workflows/shared.ts:2711 is skipped); DI-27 (DBOS completionReason is failed for every text completion); DI-25 (the same runtimes disagree on the maxSteps terminal status); DI-19 (DBOS reports a failed sub-agent to its parent as a success); found while verifying the batch-5 peer candidate on Temporal retry rollback (refuted), not itself peer-reported; severity S1: a run whose answer was blocked by the provider's safety filter, refused, or cut off at the token limit is reported and stored as a success with no error, so neither the caller nor the UI knows the reply is missing or truncated, and retry() refuses a 'completed' sessionsource, probe, live-reproverifiedopen
DI-27S3DBOS records completionReason: 'failed' on the run of every ordinary text completion. When a run ends on a text step, the standard loop sets its stop tag to the step's provider stop reason (end_turn, stop_sequence, content_filter, …). mapStopReasonToCompletionReason knows only the framework tags (finish_tool, max_steps, stop_when, idle_timeout, shutdown, interrupted, aborted) and maps everything else to 'failed'. So a normal end_turn answer is stored with status completed and completionReason: 'failed'. The value reaches the run record (getCurrentRun() / listRuns() / getRun() on the stores that persist it: postgres, redis and memory) and the standard workflow's return value. No hook, handle result, stream chunk or tracing integration reads it, and the status field is correct, so the harm is to callers and dashboards that trust completionReason. A stopWhen stop cannot produce 'stop_when' either, since that stop carries the tool_use reason. The forced-completion design noted the mismatch for outputSchema agents and fixed it only there.runtime-dbos/src/workflows/shared.ts:2774-2788 (stopTag: finish_tool for finishWith / __finish__ output, else plan.stopReason when plan.isTerminal (:2781-2782), else max_steps, else plan.stopReason; then mapStopReasonToCompletionReason(stopTag)) and :2797-2806 (finalizeRun gets status 'completed' with that completionReason); runtime-dbos/src/completion-reason.ts:21-43 (mapStopReasonToCompletionReason: only finish / finish_tool, max_steps, stop_when, idle_timeout, shutdown, interrupted and aborted are recognised; the default branch returns 'failed', and it is reached by end_turn, stop_sequence and every other provider stop reason); runtime-dbos/src/steps/run-lifecycle.ts:324-331 (finalizeRun passes completionReason to stateStore.updateRunStatus); store-postgres/src/state/postgres-state-store.ts:1398-1400, store-redis/src/redis-state.ts:2505-2506 and store-memory/src/in-memory-state.ts:712-713 persist it on the run record; runtime-dbos/src/workflows/standard-workflow.ts:528-576 returns it as the workflow result; core/src/types/session.ts:48-56 (CompletionReason) and :429-434 (RunMetadata.completionReason: "Populated when the run reaches a terminal state"); nothing in core, runtime-dbos handles, agent-server, ai-sdk or tracing-langfuse reads completionReason (source search on origin/main 186b0c6a48), so the record and the DBOS workflow result are its only surfaces; contrast: runtime-dbos/src/workflows/shared.ts:1356-1504 (the forced-completion finalizers write 'failed' / 'finish_tool' / 'interrupted' deliberately) and docs/superpowers/specs/2026-07-20-deterministic-forced-completion-design.md:24 (the design table already recorded DBOS's mismatched completionReason: 'failed' for an outputSchema agent's terminal text; forced completion fixed that case only); live: probe (batch-5 fix round 1, unit level on origin/main 186b0c6a48, not committed): runAgentLoopOneTurn with the mock StepDispatcher from runtime-dbos/src/__tests__/run-agent-loop-one-turn.test.ts and a non-outputSchema agent whose single LLM step is text with shouldStop: true. finalizeRun received status completed and completionReason failed for end_turn, stop_sequence and content_filter alike, and the returned completionReason was failed each time; the batch-5 reviewer's independent probe got the same for end_turn, stop_sequence, content_filter and max_tokens; inferred: a stopWhen stop is planned with stopReason tool_use, which also maps to 'failed', so DBOS never writes 'stop_when' for it (and DBOS does not pass stopWhen to shouldStopExecution at all, shared.ts:2694-2696 pass only maxSteps). Not probed; related: DI-25 (only DBOS writes completionReason, so this is the only runtime where the field exists and it is wrong for the common case); DI-26 (DBOS completes an error-stop-reason run; its 'failed' completionReason is not a signal of that, because of this finding); severity S3: observability-only. The run and session status are correct and no framework code branches on the value, but a caller or dashboard that reads completionReason sees every ordinary DBOS text answer as a failuresource, probeverifiedopen
DI-28S1maxSteps is documented as defaulting to 50 (the AgentConfig.maxSteps JSDoc, docs/guide/agents.md and the budget table in docs/guide/finishing-agents.md), but only CF Workflows and DBOS apply that default. JS, the CF Durable Object (which runs the JS loop) and Temporal pass an unset maxSteps through as undefined, and the shared stop check skips the step limit when it is undefined. An agent that sets no maxSteps therefore has no step cap on those three runtimes: a model that keeps calling tools keeps calling the LLM until it stops on its own, with no error and no signal. The fix belongs where the config default is applied, not in the terminal-status code (that is DI-25).core/src/types/agent.ts:255-262 (AgentConfig.maxSteps: "Maximum number of steps before agent is forced to stop (default: 50)"); core/src/orchestration/stop-checker.ts:63 (the step limit is checked only when config.maxSteps !== undefined) and core/src/orchestration/step-iterator.ts:2103 (maxSteps: deps.agent.maxSteps, passed through with no default; runtime-js and the CF DO run this iterator); runtime-temporal/src/executor.ts:391 and :1351 (maxSteps: agent.maxSteps into the workflow input, no default) and runtime-temporal/src/workflow.ts:1478-1491 (input.maxSteps is used as-is); runtime-dbos/src/serialize.ts:87 (maxSteps: agent.maxSteps ?? 50) and runtime-cloudflare/src/workflow.ts:546 (const maxSteps = agentConfig.maxSteps ?? 50): the two runtimes that apply the documented default; contrast: docs/guide/agents.md:293-295 ("Maximum number of LLM calls before the agent stops. Default is 50.") and docs/guide/finishing-agents.md:243-246 (budget table: maxSteps default 50); live: probe (batch-5 fix round 1, re-run on the fix-round-1 head f4e47117d5 over origin/main 186b0c6a48, unit level, not committed): runtime-js JSAgentExecutor + InMemoryStateStore + MockLLMAdapter; defineAgent without maxSteps (agent.maxSteps is undefined), the LLM calls a tool 70 times and then answers. The run completed at stepCount 71 after 71 LLM calls, past the documented default of 50, with no error. Temporal and the CF DO are covered from source only; related: DI-25 (how a run that does exhaust maxSteps is reported: failed on CF Workflows, completed elsewhere; a separate cause and fix site); severity S1: unbounded spend. An agent that relies on the documented default has no step budget on JS, the DO or Temporal, so a model stuck in a tool loop keeps calling the LLM (probe: 71 steps where 50 were promised); only the per-turn LLM behaviour ends itsource, probeverifiedopen
DI-29S1AgentConfig.stopWhen is honoured only on runtime-js and the CF Durable Object (which runs the JS loop). Temporal and DBOS call core's shouldStopExecution with maxSteps only, so the predicate is never evaluated. CF Workflows never calls shouldStopExecution at all: its step decides whether to continue from the tool-call count and step type alone. On those three runtimes an agent whose stopWhen fires keeps calling the LLM and running tools until the model stops on its own or maxSteps is reached (on CF Workflows a non-outputSchema agent then fails with "Max steps (N) reached"). Nothing signals that the stop condition was skipped. The docs present stopWhen as a stop condition with no runtime caveat.core/src/types/agent.ts:305-309 (AgentConfig.stopWhen: "Optional predicate to determine when agent should stop. Called after each step with the result.") and core/src/orchestration/stop-checker.ts:67-70 (shouldStopExecution evaluates config.stopWhen only when the caller passes it); core/src/orchestration/step-iterator.ts:2104 (the iterator passes stopWhen: deps.agent.stopWhen; runtime-js and the CF DO run this iterator); runtime-temporal/src/workflow.ts:1478-1480 (shouldStopExecution(stepResp.stepResult, stepCountForCommit, { maxSteps: input.maxSteps }): no stopWhen; the comment at :1476 says "maxSteps, stopWhen"); runtime-dbos/src/workflows/shared.ts:2694-2696 (shouldStopExecution(stepResult, currentStepCount, { maxSteps: agentSerialized.maxSteps }): no stopWhen); runtime-cloudflare/src/steps.ts:1641 (CF Workflows executeAgentStep: shouldContinue = toolCallCount > 0 || stepResult.type === 'text'; runtime-cloudflare/src never calls shouldStopExecution or reads stopWhen) and runtime-cloudflare/src/workflow.ts:1972-1977 (a comment says this branch "happens when stopWhen condition is met", which no code path produces); runtime-temporal/src/workflow.ts:1557-1558 is a similarly stale comment ("a stop not covered by forced completion (e.g. stopWhen)"); contrast: docs/guide/agents.md:176 ("Without outputSchema, the agent runs until maxSteps or a stopWhen condition") and the ### stopWhen section (:372-393 after batch 6; :371-392 on origin/main e37d676945), which on main carried no runtime caveat; docs/internals/execution-flow.md:30 on main listed stopWhen among the conditions shouldStopExecution() checks with no runtime caveat. Batch 6 adds the DI-29 caveats to all three places, so they now qualify the promise rather than make it; live: probe (batch-6 verification, origin/main e37d676945, no product change; an agent without outputSchema, maxSteps 5, stopWhen: () => true, the scripted LLM calls a server tool twice and then answers). runtime-js (JSAgentExecutor + InMemoryStateStore): completed after 1 LLM call, stepCount 1. Temporal (HELIX_TEMPORAL_BACKEND=test-server, runtime-temporal integ harness): completed after 3 LLM calls, stepCount 3. DBOS (RUNTIME_DBOS_INTEG=1, private Postgres DB on :5433, Redis db 14 on :6390): completed after 3 LLM calls. CF Workflows (runAgentWorkflow unit harness from runtime-cloudflare/src/__tests__/workflow.test.ts, the LLM returning non-stopping text every step): 5 LLM calls, then failed "Max steps (5) reached". The CF DO is covered from source only (it runs the iterator above); related: DI-25 (the same runtimes disagree on how a maxSteps stop is reported); DI-27 (a DBOS stopWhen stop could never record completionReason 'stop_when' anyway); DI-28 (Temporal applies no default maxSteps, so there an ignored stopWhen leaves the loop bounded only by the model); closes FU-CONF-STOPWHEN-PARITY (docs/dev/follow-ups.md), filed from the batch-5 reviewer's F9 source lead; severity S1: wrong model behaviour and extra spend. An agent that relies on stopWhen to end its loop (the docs' examples stop on a final_answer tool or a [DONE] marker) keeps running tools and calling the LLM after the stop it asked for, on three of five runtimes, with no errorsource, probe, live-reproverifiedopen
DI-30S2runtime-js, and the CF Durable Object that runs its loop, drop the classified ErrorDetail of an ordinary terminal LLM error. The adapter reports the error as an error-type step result, which is how llm-vercel reports every provider failure. planStepProcessing puts the classified detail on plan.statusUpdate.errorDetail. The step iterator ignores it and returns { kind: 'failed', error: new Error(plan.statusUpdate.error) }, a plain Error carrying only the message. resolveErrorDetail then projects that to { message }, so AgentResult.errorDetail, the persisted SessionState.errorDetail and the detail passed to failStream carry no code, category or retryable. Only an adapter that throws keeps the detail, because the thrown error object passes through. This is the runtime-js counterpart of DI-24 (Temporal and DBOS), at a different seam, and it contradicts DI-24's contrast line that runtime-js keeps the detail.core/src/orchestration/step-processor.ts:410-418 (an error step with shouldStop: true plans statusUpdate { status: 'failed', error: message, errorDetail: toErrorDetail(stepResult.error) }; the StatusUpdatePlan doc at :46-51 says the detail exists so the durable runtimes can match the JS runtime); core/src/orchestration/step-iterator.ts:1123-1127 (if plan.statusUpdate?.status === 'failed' → return { kind: 'failed', error: new Error(plan.statusUpdate.error ?? 'Agent failed') }: neither the original error nor plan.statusUpdate.errorDetail is carried); runtime-js/src/run-loop.ts:1143 (resolveErrorDetail(finalOutcome.error, finalOutcome.errorDetail): no errorDetail on the outcome and a plain Error, so toErrorDetail gives { message }) and :3345-3349 (onFailed persists that detail onto the session); contrast: core/src/orchestration/step-iterator.ts:850-877 (an adapter that THROWS: the thrown error is returned as-is via ensureError, so a thrown HelixError keeps its classification); llm-vercel/src/vercel-adapter.ts:532-537 (llm-vercel's catch path returns { type: 'error', error: helixError, shouldStop: true }: the production path is the error-step path, not the throw path); live: probe (batch-6 verification, origin/main e37d676945, unit level: runtime-js JSAgentExecutor + InMemoryStateStore, a HelixError { code: provider_rate_limited, retryable: true }). Returned as an error step (MockLLMAdapter errorInstance, recoverable: false): AgentResult.errorDetail {"message":"Rate limited"} and persisted SessionState.errorDetail {"message":"Rate limited"}. Thrown from generateStep: both carried { message, code: 'provider_rate_limited', category: 'provider', retryable: true }. The CF DO is covered from source only (it runs the same iterator); related: DI-24 (Temporal and DBOS lose the same detail at their own terminal seams; its contrast line citing runtime-js/src/run-loop.ts:1143 as keeping the detail holds only for thrown errors); DI-03 (hard abort carries no code, the same plain-Error pattern on the abort path); peer-reported by B2 (2026-09-27, a runtime-js step-seam test saw errorDetail.code undefined for a classified HelixError), mechanism located and reproduced here; severity S2: the failure is reported, but its machine-readable code and retry flag are lost for every reader on runtime-js and the CF DO (result, getStatus / reconnect, and from source the failStream detail), so a caller cannot tell a retryable rate limit from a permanent failuresource, probe, peer-reportedverifiedopen
DI-31S3resume() refusals are untyped on runtime-js and partly on DBOS, although the interrupt/resume guide documents AgentAlreadyRunningError for a running session and AgentNotResumableError for a terminal one. runtime-js's pre-checks for continue, with_message and with_confirmation throw plain Errors: "Cannot resume: agent is already running", "… already completed" and "… has failed". The same running-session case is typed only when a concurrent caller loses the CAS a moment later, so one condition yields two error types. DBOS types the running case (AgentAlreadyRunningError) in every mode but throws a plain Error for a completed or failed session. No runtime ever throws AgentNotResumableError: only classifyError references it. A caller therefore cannot tell 'already running, wait' from 'terminal, use retry or a new turn' by type or code on these runtimes. (Temporal accepts a resume on a running session instead of refusing it, CP-83, and returns the previous result for a completed one, CP-60.)runtime-js/src/js-agent-executor.ts:2136-2144 (the non-checkpoint pre-checks throw new Error('Cannot resume: agent already completed' / 'agent has failed' / 'agent is already running')) and :2233-2239 (the CAS loser for the same running session throws the typed AgentAlreadyRunningError); runtime-dbos/src/lifecycle/resume.ts:311 (typed AgentAlreadyRunningError for an active session with no pending client tools) and :314-318 (a plain Error 'Session … is not in a resumable state (current: completed)' for a terminal session); core/src/errors/agent-errors.ts:17-25 (AgentNotResumableError) is exported but thrown nowhere in any runtime; core/src/errors/classify-error.ts:65-72 maps it to code state_not_resumable, so a terminal-session refusal on these runtimes reaches classifyError as a plain Error instead; contrast: docs/guide/interrupt-resume.md, Error Handling: ### AgentAlreadyRunningError (:629-643 after batch 6; :612-625 on origin/main e37d676945: "Thrown when trying to resume an agent that's already running", with an instanceof example) and ### AgentNotResumableError (:645-659 after batch 6; :627-640 on main: "Thrown when trying to resume an agent in a terminal state"). Batch 6 adds the DI-31 caveat after them; live: probe (batch-6 verification, origin/main e37d676945, unit level: runtime-js JSAgentExecutor + InMemoryStateStore, the run's LLM call held): resume() in continue, with_message and with_confirmation on the running session each threw name 'Error', "Cannot resume: agent is already running"; after completion each threw name 'Error', "Cannot resume: agent already completed". DBOS (RUNTIME_DBOS_INTEG=1, private Postgres DB on :5433, Redis db 14 on :6390): continue, with_message and from_checkpoint on the running session threw AgentAlreadyRunningError; on the completed session each threw a plain Error "… is not in a resumable state (current: completed)"; related: CP-83 (Temporal resume() in every mode, and runtime-js from_checkpoint, accept a running session instead of refusing it); CP-60 (Temporal resume on a completed session returns the old result); CP-71 (a stranded-active JS session makes resume() refuse with this untyped 'already running'); peer-reported by B2 (2026-09-27, the JS continue case), extended here to every runtime and mode; related: the open MR !279 (B2-core) adds a typed AgentAlreadyRunningError guard for from_checkpoint only (on runtime-js, CF Workflows and Temporal). Whichever of !279 and this batch lands second narrows this record; the runtime-js continue / with_message / with_confirmation pre-checks, the DBOS terminal-session refusal and the never-thrown AgentNotResumableError are not covered by it; severity S3: the refusal itself is loud and correct, but its type contradicts the documented contract, so a caller branching on the error class (as the guide shows) misclassifies itsource, probe, live-repro, peer-reportedverifiedopen

Released under the MIT License.