Consistent snapshots and checkpoint truth
This release makes a chat snapshot one consistent, checkpoint-anchored read, and makes checkpoints and runs trustworthy on every runtime × store. A page refresh at any moment no longer loses, duplicates or permanently hides assistant content (RM-08). Design: docs/superpowers/specs/2026-09-26-consistent-snapshot-checkpoint-truth-design.md (read its "As built" changelog first).
It affects:
- every deployment (schema migrations, in-flight workflows; see the deploy checklist);
- custom
SessionStateStore/StreamManagerimplementations (a new store contract); - code that switches over
AgentResult.status, or relies onretry()/ failure checkpoints; - code that calls
buildSnapshot/handleChatStream/createCloudflareChatHandler, reads the DO/snapshotwire, usesDOStateStoreClient, or uses the React hooks; - Temporal custom clients (
TemporalClient.describeWorkflow) and CF Workflows bindings (WorkflowInstance.terminate).
Versions (docs/dev/release-process.md): the 0.x packages (core, the stores, sdk, tracing-langfuse) ship this as a minor release; the packages already past 1.0 ship a major: @helix-agents/ai-sdk 36, runtime-js 21, runtime-dbos 17, runtime-cloudflare 6, runtime-temporal 3 and agent-server 2. Read the changesets.
Deploy checklist (do these in order)
Run the schema migrations BEFORE deploying the new code. Every new column is nullable, and old code ignores them.
Store Migrations Adds Postgres V14, V15 V14: checkpoints kind,run_id,keep_count; runslease_until,start_digest,owner_token,owner_unknown_since; statescurrent_run_id,latest_truncate_checkpoint_id. V15: runsentry_step_count,input_row_count,parent_session_id,parent_tool_call_id; the__agents_abandoned_run_startstableD1 V17, V18 the same columns (V17), then the run bookkeeping and the abandoned_run_startstable (V18)DO-SQLite (CF DO) 11, 12 the same, applied by the Durable Object on first use Redis none new hash fields; see the Redis index note below Memory none Pre-release builds of this branch extended V14 / V17 / 11 in place; V15 / V18 / 12 add their columns only when missing, so every database converges on one schema.
Redis index note. The old
saveStatecould change a session's status without moving its[<keyPrefix>:]idx:sessions:status:<status>index entry, so a session can stay listed under an old status inlistSessions({ status }).RedisStateStore.rebuildIndexes()only ADDS entries, it never removes one, so running it alone does not fix that. To clean up, delete the[<keyPrefix>:]idx:sessions:status:*sorted sets, then runrebuildIndexes()once: it re-adds every session under its current status.Drain, or terminate, in-flight Temporal, DBOS and CF Workflows workflows. This is REQUIRED. Workflow command and step order, step names, workflow inputs and workflow ids changed, and none of it is version-gated (no Temporal
patched(), noDBOS.patch). A workflow started by the previous release cannot replay on this release's code, and the runtime cannot detect or settle that itself. The reverse holds too, so a rollback needs the same drain.Temporal. Affected: every agent workflow RUNNING at deploy time, including child and companion-continuation workflows. A Temporal agent workflow exits at every HITL boundary, so a session paused on a client tool, an approval or a child has no running workflow of its own and is safe (its children may not be). What changed: a suspension commits before its marker (
commitSuspendedStep→ marker →pauseSuspendedStream), a resume's first command isstartRunCommit, the terminal tail iscommitTerminalState+finalizeTerminalRun, a companion continuation id is__continue__<toolCallId>(was__continue__<stepCount>), and a resume workflow id is…__resume-<runId>. Choose one:Drain: stop new entries, then wait until this is empty, and only then deploy the new workers:
bashtemporal workflow list --query 'ExecutionStatus="Running" AND WorkflowType="<your agent workflow type>"'Terminate the in-flight ones:
bashtemporal workflow terminate --query 'ExecutionStatus="Running" AND WorkflowType="<your agent workflow type>"' \ --reason "helix snapshot-checkpoint-truth upgrade"The sessions recover on the new release: the next
execute()/resume()/retry()finds the owner workflow closed (or a pre-upgrade run with no owner token) and supersedes the stalerunningrun (run.stale-running-supersedable).Worker Versioning (Build IDs or worker deployments), so the previous release's workers keep serving their pinned executions until those drain.
Symptom if you deployed without draining: the old workflows show repeated
WorkflowTaskFailedevents with causeNON_DETERMINISTIC_ERRORand stayRUNNING, and every laterexecute()on their sessions getsAgentAlreadyRunningError. Recover by terminating them as above.DBOS. The workflow input changed (
runMeta→entry), steps were added and reordered, and workflow ids changed (…__resume__<runId>). A paused DBOS session is in-flight too: it is a PENDING workflow blocked inrecv, waiting for a client-tool result or an approval. Before deploying, let pending workflows finish (answer their pending calls), or cancel them (DBOS.cancelWorkflow(workflowId), or the DBOS admin / Conductor), then deploy. A cancelled workflow is closed, so the session's next entry supersedes its run; the pending interaction is lost and the user re-sends it.CF Workflows. Every
execute()/resume()/retry()runs in its own instance with a per-call key (see the instance-id deploy note). Let in-flight instances finish, or terminate them.runtime-js / CF DO. Nothing to drain beyond the legacy-mode note: a run that is mid-turn during the deploy is superseded by the next entry (a dead process stops renewing its lease; a DO that is not executing it is ownerless).
Deploy the Cloudflare
StreamServerworker before (or together with) the workers that run agents — the claim compare-and-set only protects a newer run once the stream Durable Objects run the new code (see the StreamServer note).Deploy the rest together. Producers and consumers of the DO
/snapshotwire, the chat routes and the React client change together (the wire shape changed; see below).Pre-upgrade sessions keep working in legacy mode until their first new commit.
Breaking changes (write side)
The store contract (custom SessionStateStore implementations)
Every message-log or snapshot-visible mutation is now a commit: it writes a checkpoint {id, kind, runId, messageCount, streamSequence, stepCount, state, keepCount} and moves the session's latest pointer in the same atomic operation.
| Method | Change |
|---|---|
appendMessages(sessionId, messages, commit) | commit: MessageCommit ({ streamSequence, runId, startRun? }) is required. With startRun it is the guarded run-start commit (kind: 'input', rows may be empty); without, an append commit |
truncateMessages(sessionId, keepCount, commit) | commit required; a truncate checkpoint with keepCount, and the session's latestTruncateCheckpointId |
saveStateAndPromoteStaging / promoteStaging | the step commit. PromoteStagingOptions.streamSequence / .runId and checkpointMeta.runId are required (a superseded runId throws RunSupersededError and writes nothing) |
commitState(sessionId, state, commit, opts) (new) | the state commit (K5, the terminal state commit) |
cloneSession | writes a clone checkpoint |
saveState(sessionId, state, { fenceRunId } | { fenceIdle }) | writes no checkpoint any more. Runtimes pass a fence; a write for a superseded run throws RunSupersededError and writes nothing |
createCheckpoint | never moves the latest pointer |
getLatestCheckpoint | the checkpoint the pointer names (never the highest stepCount) |
renewRunLease(sessionId, runId, ttlMs) (new) | a CAS on the store's own clock; rejects a run that is not current and running with RunSupersededError |
setRunOwnerUnknownSince(...) (new) | set-if-null / clear, for the CF Workflows stuck-instance escape |
abandonRunStart(sessionId, runId) (new) | 'committed' when the run exists, else a tombstone ('abandoned'): a later run-start of that runId throws RunStartRejectedError('entry_abandoned') |
listRunCheckpoints(sessionId, runId, { offset, limit }) (new) | the run's own checkpoints (Checkpoint.runId === runId), full, ordered by messageCount, from an INDEXED read by run (never a scan of the session); legacy runId: null rows are never returned. planRetry uses it. Postgres V15, D1 V18 and DO-SQLite 12 add the index |
updateRunStatus | a terminal status (incl. superseded) is sticky: changing a terminal run throws RunSupersededError |
getRun | read-your-writes for committed runs |
createRun / updateRunStatus / getCurrentRun / listRuns / getRun | required: runtimes no longer feature-detect them, and an executor built on a store without them is rejected at construction |
Checkpoint gains the required fields kind: CommitKind | null, runId: string | null and keepCount: number | null (null on legacy rows; a stored kind that is not a commit kind throws UnknownCommitKindError instead of reading as legacy); RunMetadata gains ownerToken, leaseUntil, ownerUnknownSince, startDigest, entryStepCount, inputRowCount and parent.
Also changed for every store:
SessionState.checkpointSourceis removed (saveStateno longer has a staging-checkpoint mode).isSessionStateStore()checks for the new methods too, so an old store no longer passes it.getMessagesthrowsCorruptMessageRowErrorfor a row it cannot parse, where it used to skip the row (snapshot and LLM history reads turn it intostate_history_incompletewith causecorrupt_row).- D1:
cloneSessionwith a checkpoint that does not exist now throwsstate_checkpoint_not_found. - Redis:
getLatestCheckpointwith a pointer naming a missing checkpoint throws a typedstate_checkpoint_not_found; it used to fall back to the higheststepCount. - Removed exports:
calculateUIMessageCountandresolveCheckpointStreamSequence(core; userunBoundaryFrom/responseIndexForandresolveCommitSequence),ContentReplayOptions(ai-sdk).
The run-start commit must, atomically: (1) be idempotent by runId (same startDigest → write nothing, return the run; a different digest → RunStartRejectedError('idempotency_mismatch')); (2) CAS on startRun.expectCurrentRunId (a live current run → RunStartConflictError('expected_mismatch'), an idle one → RunStartRejectedError('expected_mismatch_idle')); (3) guard the owner (the same ownerToken; ownerless: 'lease' with the current run's lease absent or expired on the store clock; or ownerless: 'asserted'), else RunStartConflictError('run_active'); (4) mark every other live run superseded; (5) apply truncateTo as a truncate checkpoint by the new run; (6) create the run, the rows and the input checkpoint. Every runId-carrying commit is a CAS on "this run is current and live". A store that implements the methods without these guards passes the type check but breaks run.single-live. Run the shared contract suites from @helix-agents/core/testing (commitCheckpointOperationTests, runStartOperationTests, fenceOperationTests, the checkpoint operations, and streamManagerContractTests for a stream manager); the pure rules are exported (applyRunStartRules, planRunStartRewind, buildCommitCheckpoint, assertRunIsCurrentLive, assertRunStatusTransition, …).
Store-specific: Redis and D1 throw RedisCommitContentionError / D1CommitContentionError when a commit keeps losing its optimistic retry; the Redis commit is one Lua script with every key passed in KEYS. Redis Cluster is not supported (CROSSSLOT; see the store README).
Custom stream managers
See Custom stream managers below: claimStream is a compare-and-set, endStream / failStream / pauseStream take a run fence, an own-session chunk stamped with another runId is dropped once a newer run claimed the stream (chunkFenceRunId), resumable readers expose catchUpHead / caughtUp, and ResumableReaderOptions.waitForLive is new.
AgentResult.status: 'superseded' and onAgentSuperseded
A run that a newer run took over (a concurrent entry, a lost JS lease, an owner-gone takeover) stops with status: 'superseded' and errorDetail.code: 'state_run_superseded' (non-retryable) on every runtime — never failed. Exhaustive switches over AgentResult.status / RunStatus must handle it.
onAgentCompleteandonAgentFaildo not fire for it; the new optionalonAgentSupersededhook does.- tracing-langfuse closes its span with
level: 'WARNING',statusMessage: 'superseded'. handleChatStreamdoes not treat it as a stream terminal: the superseding run streams on the same session stream.- A late client-tool submit to a superseded run answers 409
{code: 'state_run_superseded'};DOFrontendExecutor.submitToolResult/createHelixChatTransport().submitToolResultthrowRunSupersededError.
New errors and codes
| Error / code | When |
|---|---|
RunStartConflictError (name: 'AgentAlreadyRunningError') | a live run holds the session (run_active, or expected_mismatch against a live run). Existing AgentAlreadyRunningError handling still matches it |
RunStartRejectedError (state_run_start_conflict) | idempotency_mismatch, expected_mismatch_idle (retried once by the runtimes and the chat handler), entry_abandoned. NOT matched as "already running" |
RunSupersededError (state_run_superseded) | a write by a superseded run; non-retryable |
RunStartAbandonedError / RunStartOutcomeUnknownError (core, transport_error) | a Temporal / DBOS / CFW entry that reached no run-start outcome within its fail-safe (below) |
StreamClaimStaleReadError (retryable) | the stream claim read a lagging owner |
state_checkpoint_sequence_unavailable | a commit could not read the live stream head (no silent 0 any more, DI-21); the run fails after native retries |
state_step_messages_uncommitted | the DBOS / CFW step-atomicity assertion |
state_history_incomplete cause | now short_read / corrupt_row / store_error / snapshot_unstable |
RunStartAbandonedError and RunStartOutcomeUnknownError are exported from @helix-agents/core only; Temporal, DBOS and CF Workflows all throw them. The run-start fail-safe is new: its bounds are Temporal entryOutcomeTimeoutMs, DBOS runStartFailSafeMs and CF Workflows entryOutcomeFailSafeMs (entryStartTimeoutMs overrides it), each 120 000 ms by default. On RunStartOutcomeUnknownError, read the session before retrying: the entry may still commit.
Failure checkpoints and retry() (user ruling D1)
- Every terminal exit — completion, failure and stop — now writes the K5
statecheckpoint (with the terminal status and the error; on completion also theonAgentCompletewrites, whileonAgentFailwrites are kept only on Temporal and CF Workflows) before the terminal run status and the stream end. Before, a failure or stop wrote no checkpoint (R28 / R40), so code that took "the latest checkpoint of a failed session" as its last good step now gets K5. retry()never targets K5. CoreplanRetrypicks the failed run's lateststepcommit (or a laterappendcommit of that run with no assistant row); if the run failed on its first step, its run-start commit — the rows rewind to just before the run's own triggering user message, which the retried run re-appends, followed byoptions.messagewhen given. A resumed run's drained results (a client-tool or approval answer appended before its first assistant row) are part of its entry: a retry keeps them, and the target is the latest suchappendcommit, so an approved tool's state writes are kept.customState,stepCountandforcedCompletioncome from that target. An explicitcheckpointIdis restored as given. A session that failed BEFORE its last entry's run-start commit keeps its whole log and needsoptions.message(elsevalidation_error). Retry outcomes match the old model.- runtime-js
retry()with no message re-runs the failed run's trigger after a FIRST-step failure (CP-58). It used to throw. After a run that committed steps, a retry withoutoptions.messagestill throwsvalidation_error. - DBOS
retry()is a run-start commit that honourscheckpointId(an atomictruncateTo) andmessage(CP-30); it used to ignoreRetryOptions.
Runs on every entry; ownerToken; the JS run lease
Every entry (execute, resume, retry, client-tool / approval continuations, children, companions) starts a new run in a run-start commit that supersedes the previous live run. The losing caller of a race writes nothing (no input row, no resumeCount, no stream change). resumeCount is now counted by the WINNER of a resume / retry entry on Temporal, DBOS and CF Workflows; runtime-js and the CF DO never count it.
runtime-js keeps a run lease: JS_RUN_LEASE_MS (30 000 ms), configurable with the executor option runLease: { ttlMs }. It is renewed every ttl / 3 on a timer, through long model and tool calls. A JS event loop blocked for longer than about ttl − ttl/3 (20 s by default) — a synchronous CPU-bound tool — may let a concurrent entry supersede the run; the old run then stops superseded. With no concurrent entry the run renews and continues: only a renewal the store rejects stops it. Raise ttlMs if your tools block that long.
Client-tool batches: resume before submit is refused (R44)
resume() (continue and with_message) on a session whose client-tool batch is partially submitted now throws a retryable state_not_resumable and writes nothing — on runtime-js, Temporal, DBOS, CF Workflows and the CF DO (/resume → 409). Before, DBOS in particular resumed while a call was still awaited and the late result was lost. Instead:
- submit every result of the batch (
submitToolResult; once the batch is complete the run continues), or let the call's deadline lapse (a lapsed deadline counts as its timeout result), then resume if needed; - to abandon the batch,
interrupt()(a pending durable interrupt bypasses the gate) or submit error results; from_checkpointis not refused: a rewind that discards the unanswered calls proceeds, and one whose restored batch is still partial re-suspends at once on the unanswered calls (R47) and waits for the rest of the batch.
Runtime-specific breaking changes
- runtime-js: an executor built on a store without the run methods throws at construction. New executor option
runLease: { ttlMs }and exportJS_RUN_LEASE_MS; new exportscommitRunStartandfinishCommittedRun(re-exported from core) for custom loops;recover()(internal, used by the CF DO's wake recovery) re-enters arunningrun under its ownrunId. - runtime-temporal:
TemporalClient.describeWorkflow(workflowId, runId)is required (the owner-liveness probe): a custom client that implements only the old methods no longer typechecks. Peers:@temporalio/*^1.13.0(was^1.11) andzodis now a peer. New executor optionentryOutcomeTimeoutMs; new activities optionsownerProbe,workflowProbe,ownerWaitandllmAttemptFenceIntervalMs(defaults suit production; tests shorten them). The workflow result gains thesupersededstatus. The workflow input fieldterminalFinalizeTimeoutMs(default 3 min) bounds the post-K5 finalization (run status, stream end); past it the workflow closes and reader-path recovery finishes the run. A companion continuation recordsmetadata.__helix.companionContinueon the child session, whichlistSessionsmetadata filters can see. A detached child's live chunks written after its parent's turn ended are dropped (its content stays in its own session). - runtime-cloudflare (CF Workflows):
WorkflowInstance.terminate()is required on the binding, andWorkflowStatus.statusgainswaiting,waitingForPauseandunknown. The executor rejects a store without the run methods (assertCfwStoreSupportsRuns).execute()/resume()/retry()now wait for their entry's outcome (the run-start commit, or the instance ending without one) before returning.AgentSteps.initializeAgentStateandAgentSteps.createResumeRunare removed (the run-start entry step replaces them) andrunIdis required on the step inputs. New adminexecutor.terminateRun(sessionId)for an operator-pausedinstance;cfwUnknownTerminateAfterMs(default 600 000 ms) terminates an instance stuckunknownand writes arun_terminatedchunk (UIdata-run-terminated). - runtime-cloudflare (DO):
DurableObjectAgentBase/DOStateStoreown the run-start commit (runStartCommitSync, onetransactionSync);DOStoreNotInitializedError. - runtime-dbos:
delivered: falsecompanion results carrychild_busy/child_exited/child_start_unconfirmed; any step that runs out of retries settles the turnfailed.
UI message ids
A run's response message id (msg-<startUIMessageCount>) is now the first-row index of the UI message that holds the response (core responseIndexFor): a continuation (a client-tool or approval round) streams into the open assistant group's id, and system rows count like any other indexed row. Snapshot and live ids now agree (ids.snapshot-matches-live). Ids are non-contiguous, and data keyed by a message id may shift after the upgrade.
snapshot.state
FrontendSnapshot.state is the custom state AT the snapshot's cursor (the checkpoint's state), not the latest saveState. Later changes arrive as state_patch frames after the cursor.
Read side: snapshots are checkpoint-anchored; contentReplay is removed
This part affects code that calls buildSnapshot / handleChatStream / createCloudflareChatHandler, reads the DO /snapshot wire, or uses DOStateStoreClient directly.
Why
A snapshot used to be assembled from independent reads (messages, stream info, run, chunks) plus a server-side merge of in-flight "partial" content. Commits landing between those reads tore it: status: 'active' with the stream head as cursor but the final text missing, or status: 'ended' with the last step missing. The 50-row cap corrupted long sessions.
A snapshot is now taken from one checkpoint:
| Field | Meaning |
|---|---|
messages | the committed rows [0, checkpoint.messageCount), UI-converted; never partial, never capped |
streamSequence | the checkpoint's cursor: everything committed up to it is in messages, nothing after it |
state | the custom state at the cursor |
status | derived from the current run: a live run is active / paused |
resumeMessageId | when live: the id of the UI message the run's response streams into |
The canonical client
const snap = await fetch(`/api/chat/${sessionId}/snapshot`).then((r) => r.json());
const opts = snapshotResumeOptions(snap);
// Seed WITHOUT the resume message; only when live, resume — the server replays
// that message's committed prefix, then streams the tail after the cursor.
const chat = useHelixChat({
id: sessionId,
messages: opts.messages,
resume: opts.resume,
transport: new DefaultChatTransport({
api,
prepareSendMessagesRequest: prepareHelixChatRequest({ api }),
prepareReconnectToStreamRequest: prepareHelixReconnectRequest({
api,
resumeFromSequence: opts.resumeFromSequence,
existingMessageId: opts.existingMessageId,
replayPrefix: true,
}),
}),
resync: { snapshotUrl: `/api/chat/${sessionId}/snapshot` },
});Why the resume message is removed (C-1). ai 6.0.281 starts a resumed response from an EMPTY message that REPLACES the last one; ≤ 6.0.230 continue the last message. Seeding the whole snapshot and resuming only the tail therefore drops the in-flight message's committed steps on 6.0.281. With the message removed and X-Helix-Replay-Prefix: 1 (replayPrefix: true / resumeHeaders({ …, replayPrefix: true })), the server reads ONE snapshot (with HandleChatStreamDeps.snapshotOptions — pass the options your snapshot route uses), replays that message's committed parts in order (only ai 6.0.0 chunk types, between transient data-helix-replay{phase: 'begin' | 'end'} markers) and streams the tail after that snapshot's cursor — identical on every version, also for a message that spans runs (a HITL continuation) and when early chunks were trimmed. useHelixChat never hands a replayed tool call to onToolCall and never re-sends a replayed tool's committed result; a replayed tool that is still PENDING (awaiting the browser) is dispatched once after the replay and its output is sent. The message's committed metadata rides on start; an answered approval replays as its answer (never re-asked). handleChatStream takes the same flag as replayPrefix: true or the request header, and honours it only for a CLIENT-supplied existingMessageId; an id the snapshot does not hold answers a typed 409 state_not_resumable (cause: 'resume_message_unknown'), never the tail alone. The flag and the seed always travel together: snapshotResumeOptions() / useResumableChat() return messages + replayPrefix; a client that cannot suppress replayed tools (a plain or pre-built Chat) must seed the whole snapshot and resume WITHOUT the flag (snapshotResumeOptions(snap, { replayPrefix: false })) — re-anchors do this automatically (with a dev warning).
Migrating
contentReplay / contentReplayEnabled (removed)
Delete contentReplay: true from handleChatStream / handleChat calls and contentReplayEnabled from buildSnapshot deps. Resume from the snapshot cursor (above). A resume without a cursor still replays the current run from its start under the run's response id (msg-<startUIMessageCount>) — what contentReplay was used for on a cold load.
FrontendSnapshot
- Removed:
timestamp,startSequence,startMessageCount. UsestreamSequence(the cursor) andresumeMessageId. SnapshotOptions.includePartialContentis removed; there is no partial content to include. In-flight content arrives through the resumed stream.
status follows the run
status is active / paused whenever the current run is live — including the short window after a run's run-start commit, before its stream is claimed, when the stream still shows the previous turn's ended. Resume on active and paused; a resume in that window waits for the run to go live. If it never does within HandleChatStreamDeps.liveWaitTimeoutMs (default 30 000 ms), the handler answers a retryable 503{code: 'transport_stream_drop', cause: 'run_not_live'} — retry the snapshot.
Run-start conflicts
- A send (a request whose trailing message is a new user message) refused with
RunStartConflictError(namedAgentAlreadyRunningError) answers 409{error, code: 'state_already_running', retryable: true}. It no longer attaches to the running run's stream, which rendered that run's text as the reply and dropped the message. A stockuseChat/useHelixChatends the send with status'error'; read the body withparseHelixChatError(chat.error)and retry withchat.regenerate(). The handler also reads the stream head BEFOREexecute()and streams only after it, so a new turn never replays earlier turns. - Requests with no new user content (a reconnect, a passive attach) and tool-result resumes keep attaching to the winner when a run is already live.
RunStartRejectedError(expected_mismatch_idle): retried once by the handler.- Any other
RunStartRejectedError(e.g.idempotency_mismatch): a typed 409{error, code: 'state_run_start_conflict', cause}instead of a silent attach.
Cloudflare DO
GET /snapshotreturns theSnapshotViewwire:sessionId, status, head, streamSequence, legacy, checkpointId, stepCount, checkpoint {id, kind, runId, keepCount, stepCount, messageCount, streamSequence}, messages, state, run {runId, status, startUIMessageCount, startSequence, turn}, plus the session fields (streamId, sessionStatus, createdAt, updatedAt, agentType, version, error, errorDetail, interruptContext, the client-tool maps, root / parent ids,streamError,streamFinalOutput).timestampand the top-levelstartSequence/startMessageCount/startUIMessageCount/runId/turnare removed.- A missing session is 404
{error: 'Session not found', code: 'state_session_not_found'}(was 200null); a read failure is 500{error, code, cause?}(e.g.state_history_incomplete/corrupt_row). DOStateStoreClient:getMessagespagesGET /messages(limit≤ 1000);getLatestCheckpoint/getCurrentRunreturn the real checkpoint and run (nullwhen there is none) — the run's realstartSequenceandturnincluded, so the chat handler's stale-cursor clamp and mid-block run floor hold on the DO; a non-404 failure throws a typedDOPeerErrorinstead of returningnull.createCloudflareChatHandler().getSnapshot({ includeReasoning })—includeThinkingwas silently ignored and is now mapped (deprecated).
React hooks and client helpers
The snapshot is UI-shaped (
buildSnapshotoutput); the hooks no longer re-convert it (RM-23). If you hand-rolled aServerSnapshotwith rawMessage[], switch toFrontendSnapshot.useAutoResync/useResumableChat/useCheckpointSnapshotretry a 5xx snapshot (fetchSnapshotWithRetry: 4 fetches, 250 / 500 / 1000 ms) and surface a typedSnapshotUnavailableErrorafter exhaustion — render an explicit error state with a retry control, never an empty conversation.A resync re-anchors, safely by default. On a resync trigger the hooks stop the stream, fetch the LATEST snapshot, set its messages and — when live — resume from its cursor under
resumeMessageId. The simplest wiring isuseHelixChatalone:tsconst chat = useHelixChat({ id: sessionId, transport, messages: snapshot.messages, resume: snapshot.status === 'active' || snapshot.status === 'paused', resync: { snapshotUrl: `/api/chat/${sessionId}/snapshot` }, }); if (chat.resyncError) return <RetryBanner error={chat.resyncError} />;useAutoResync/useResumableChatresolve their stream control fromstop+resume, elsechat(e.g. theuseHelixChatresult), else the one mounteduseHelixChat. With none (or several chats mounted) a resync now surfaces a typedConfigurationError(missing: ['stop', 'resume']) viaresyncError/onErrorand leaves the messages alone — it no longer replaces messages under a stream that is still writing. Feed the hookscollectResyncTriggers(messages).Stream control is per chat.
withHelixStreamControl/createHelixChatTransportkey each connection bychatId: two chats sharing a transport never abort each other;abortActive(chatId?),hasActiveConnection(chatId?)(now a function) andactiveKind(chatId). A reconnect the transport aborts before its response resolvesnull(ai6.0.91–6.0.230 would report a thrown abort as a chat error).controlChat(chat, transport, { getStatus })(framework-agnostic;useHelixChatuses it) returns an abort-aware, boundedstop, aresume, and asendMessagethat stops an open resume first so Chat's state stays right.Resync control comes from one source.
stop+resumetogether, elsechat, else the mounteduseHelixChatwhose id ischatIdor whosesetMessagesis the hook'ssetMessages(never "the one mounted chat"). A lonestoporresumeis a typedConfigurationError; astopthat does not abort the transport (e.g. a barechat.stop) logs a development warning.stop()on a resumed stream.ai6.0.0–6.0.230 (incl. the locked 6.0.49) callreconnectToStreamwithout anabortSignal, sochat.stop()cannot stop a resumed stream.useHelixChatnow wraps its transport withwithHelixStreamControl(each connection has its ownAbortController; a new connection aborts the previous one) and returns astopthat aborts the connection and resolves once the chat is idle. With plainuseChat, wrap the transport yourself and calltransport.abortActive(chatId)beforechat.stop().createHelixChatTransportowns its connections the same way (abortActive(chatId)— scoped to one chat).controlChat'sstoprejects with a typedStreamStopTimeoutErrorif the chat does not go idle within its bound, and a send waits for a re-anchor in flight (and vice versa).Resume headers.
prepareHelixReconnectRequestnow lets per-request headers win over its configured getters, sochat.resumeStream({ headers: resumeHeaders(target) })resumes at a re-anchor's cursor.resumeHeaderssendsX-Helix-Replay-Prefixonly whentarget.replayPrefixis set — set it exactly when the chat was seeded WITHOUT the resume message (snapshotResumeOptions(snap)gives both); with a whole-snapshot seed leave it unset, or the replayed prefix duplicates the message onai≤ 6.0.230.Retried attempts. A Temporal activity retry or a CF Workflows
step.doretry re-streams a step under a newattemptId. A resume / attach / terminal replay drops the superseded attempt's content (the committed step keepsstep_committed.committedAttemptIdwhen that attempt streamed in the step — a commit naming none ('') or an id that never streamed changes nothing; an in-flight step the attempt that STARTED last, by the stream position of each attempt'sstep_start, which every runtime writes as the attempt's first chunk before calling the model — never by a clock), including its step markers and files. A live attempt switch — a newer attempt, or astep_committednaming another attempt — emits adata-attempt-supersededevent (a resync trigger), and so does a resume whose cursor is past superseded content the client was already sent.Types.
onResyncreceivesAISDKResyncTriggerEvent(AISDKStreamResyncEvent | AISDKAttemptSupersededEvent; narrow onevent.typebefore readingcheckpointId).data-stream-resynccarriesrunId;ResyncTrackerdedups a resync by(runId, checkpointId)and rescans an array that shrank or was replaced.setMessagesin the hook options is typed(messages: UIMessage[]) => void(it was(messages: unknown[]) => void):useChat'ssetMessagesfits directly for the defaultUIMessage; with a custom message type, adapt it ((m) => setMessages(m as MyUIMessage[])). A setter declared overunknown[]still type-checks.fetchSnapshotWithRetry(url, { signal })aborts (never retried); a 429 is retried like a 5xx; a 200 whose body is not JSON / not a snapshot throwsSnapshotUnavailableError(reason: 'invalid_json'/'invalid_body'). The hooks abort a fetch in flight on unmount or asnapshotUrlchange, and only the latest re-anchor lands.Reasoning renders once: a
thinkingchunk withisComplete: trueis the whole block, and the transformer adds only what the streamed deltas did not deliver.snapshotResumeOptions(snapshot)→{ resume, resumeFromSequence?, existingMessageId?, messages }(messages: what to seed — without the resume message when resuming);useResumableChatexposesstatus,streamSequence,resumeMessageIdandshouldResume.
useHelixChat().retry() and failed resumes
UseHelixChatResultgainsretry(). After a failed RESUME it re-anchors on the latest snapshot (which restores the in-flight message the C-1 seed left out) and resumes from its cursor. After a failed SEND it re-sends (regenerate()) only when the latest snapshot does not hold the user turn yet; otherwise it re-attaches, so a turn is never persisted twice. Withoutresyncit isregenerate(). Wire your Retry button tochat.retry(), notchat.regenerate().useHelixChatwithresyncre-anchors on its own when the mount-time resume fails (409 / 503 / a network error), so the hidden in-flight message comes back.
Pending approvals survive a reload
A snapshot renders a pending approval as an approval-requested part with its approval id (from PendingClientToolCall.approvalId, which the core approval path writes on runtime-js and the CF DO; on Temporal, DBOS and CF Workflows a pending server tool is treated as an approval), and the resumed stream re-sends a skipped, unanswered tool-approval-request once. sendAutomaticallyWhen / extractResumeIntent accept an approval-responded part with providerExecuted: true. The CF DO now records an approval answer as { approved, reason }.
Wire errors (every host)
| Case | Status | Body |
|---|---|---|
| snapshot of a missing session | 404 | {error: 'Session not found', code: 'state_session_not_found'} |
| snapshot read failure | 500 | {error, code, cause?} — e.g. state_history_incomplete / corrupt_row |
| late submit to a superseded run | 409 | {error, code: 'state_run_superseded'} |
| a resume / attach of a run whose stream FAILED (incl. before its first commit) | 409 | {error: 'stream_failed', code, retryable, sessionId, streamId, message, errorDetail?}: code / retryable come from the run's errorDetail (parseHelixChatError reads them) |
| a send (new user message) refused because a run is live | 409 | {error, code: 'state_already_running', retryable: true} (parseHelixChatError) |
an entry the store rejected (RunStartRejectedError, not idle-retried) | 409 | {error, code: 'state_run_start_conflict', cause} |
| a replay-prefix resume for a message id the snapshot does not hold | 409 | {error, code: 'state_not_resumable', cause: 'resume_message_unknown', retryable: true} |
a resume while a client-tool batch is partially submitted (CF DO /resume, R44) | 409 | {error, code: 'state_not_resumable'} |
a resume in the run-start window whose run never goes live within liveWaitTimeoutMs | 503 | {error, code: 'transport_stream_drop', cause: 'run_not_live'} (Retry-After: 1) |
| a failed committed-prefix read for a replay-prefix resume | 503 | {error, code, cause?} (retryable) |
@helix-agents/ai-sdk/cloudflare exports snapshotNotFoundResponse(), snapshotErrorResponse(error) and submitToolResultErrorResponse(error) for route handlers. A snapshot route should answer:
try {
const snapshot = await chat.getSnapshot({ sessionId });
return snapshot ? Response.json(snapshot) : snapshotNotFoundResponse();
} catch (error) {
return snapshotErrorResponse(error);
}On the client, DOFrontendExecutor.submitToolResult and createHelixChatTransport().submitToolResult throw a typed RunSupersededError for the 409. agent-server's GET /status gains runId / runStatus (a superseded run is never reported as a failure); its 500 envelope keeps code: 'INTERNAL_ERROR' and adds errorCode / cause for a typed failure. RemoteStatusResponseSchema parses a runStatus it does not know (a newer peer's) as 'unknown' (RemoteRunStatus) instead of failing the response.
Pre-upgrade sessions (legacy mode)
A session whose latest checkpoint predates this release (kind: null) is served in legacy mode: the complete live log, with the stream head as the cursor, whatever its status. It applies until the session's first new commit (its next run-start commit re-anchors it). Legacy checkpoints can claim too few rows (the old DO message_count = 0, DBOS promotes, CF Workflows appends), which is why the live log is served. This covers terminal sessions and paused sessions that cannot be drained (suspended on a client tool or an approval for days) on runtime-js, the CF DO, Temporal and CF Workflows: they are never replayed from sequence 0. On DBOS a paused session is an in-flight workflow (see the deploy checklist). A legacy session's snapshot status comes from the session and the stream (not from a pre-upgrade run record, which may never have been ended: the previous release wrote no Temporal terminal run status, RM-47), so a session that finished before the upgrade reads ended / failed, not active. A RUNNING legacy session is not covered — drain running sessions before upgrading (see the deploy checklist). A pre-upgrade running run record with no lease or owner becomes supersedable under the new owner rules (no lease on runtime-js, a closed workflow on Temporal, !isExecuting on the CF DO), so it never bricks a session.
Custom stream managers
ResumableReaderOptions.waitForLive (new, optional): until the reader yields its first chunk past fromSequence, a terminal status — or a stream that does not exist yet — means "not live yet": return a reader (never null) and park on your normal wake-up path instead of ending; close() must release a parked next(). A manager that ignores it keeps the old behavior (a resume in the run-start window then ends early on that backend). Once the reader has seen the stream go live — active at attach, a transition to active after it, a chunk past the cursor, or a missing stream created — normal terminal semantics apply, so a run that goes live and ends before its first chunk ends the reader promptly; a terminal status left from the previous turn never does. The shared contract cases are WL.* in @helix-agents/core/testing's streamManagerContractTests.
StreamManager.claimStream(streamId, runId, expect?) is a compare-and-set (CP-85): with expect: { owner } it sets the stream's owner run only while the current owner (null = none) is expect.owner, and resolves true when runId owns the stream afterwards, false otherwise (nothing written; a missing stream is never created). Runtimes call it only through core's claimStreamForCommittedRun. A manager whose claimStream still resolves undefined is treated as having claimed (the pre-CAS, unconditional behaviour). Contract case: RF.claim.cas.
Deploy note: the Cloudflare StreamServer claim (rolling deploys)
The /claim endpoint of the Cloudflare StreamServer Durable Object (store-cloudflare) gained the compare-and-set (expectOwner). During a rolling deploy, a NEW client talking to a not-yet-updated StreamServer gets the old, unconditional claim (the old object ignores expectOwner and does not report claimed, which the client reads as claimed). The claim then cannot protect a newer run's claim that lands in that window. Deploy the StreamServer (the worker that hosts the stream Durable Objects) before, or together with, the workers that run agents, and drain in-flight runs as for the rest of this release.
Deploy note: Cloudflare Workflows instance ids change (drain or terminate first)
On Cloudflare Workflows every execute(), resume() and retry() now runs in its own instance keyed on a key minted by that call (agent__<name>__<sessionId>__turn__<key>, …__resume__<key>, …__retry__<key>; ids that Cloudflare would reject are encoded as h_<hash>_<tail> by buildWorkflowInstanceId). A re-spawned sub-agent is …__respawn__<parent-run suffix>, and a companion continuation / resume is …__continue__<step>-<callId> / …__resume__<step>-<callId>. getHandle() finds the live instance from the session's run record, never by rebuilding an id. An instance started by the previous release has an old-shape id and runs the old workflow body, which this release's body cannot replay. Before you deploy, let in-flight CF Workflows instances finish, or terminate them (a terminated run's session is then failed or superseded by the next entry). Code that rebuilt agent__<name>__<sessionId> to look an instance up must call executor.getHandle(agent, sessionId) instead.