Skip to content

Consistent snapshots and checkpoint truth ​

This release makes a chat snapshot one consistent, checkpoint-anchored read, and makes checkpoints and runs trustworthy on every runtime × store. A page refresh at any moment no longer loses, duplicates or permanently hides assistant content (RM-08). Design: docs/superpowers/specs/2026-09-26-consistent-snapshot-checkpoint-truth-design.md (read its "As built" changelog first).

It affects:

  • every deployment (schema migrations, in-flight workflows; see the deploy checklist);
  • custom SessionStateStore / StreamManager implementations (a new store contract);
  • code that switches over AgentResult.status, or relies on retry() / failure checkpoints;
  • code that calls buildSnapshot / handleChatStream / createCloudflareChatHandler, reads the DO /snapshot wire, uses DOStateStoreClient, or uses the React hooks;
  • Temporal custom clients (TemporalClient.describeWorkflow) and CF Workflows bindings (WorkflowInstance.terminate).

Versions (docs/dev/release-process.md): the 0.x packages (core, the stores, sdk, tracing-langfuse) ship this as a minor release; the packages already past 1.0 ship a major: @helix-agents/ai-sdk 36, runtime-js 21, runtime-dbos 17, runtime-cloudflare 6, runtime-temporal 3 and agent-server 2. Read the changesets.

Deploy checklist (do these in order) ​

  1. Run the schema migrations BEFORE deploying the new code. Every new column is nullable, and old code ignores them.

    StoreMigrationsAdds
    PostgresV14, V15V14: checkpoints kind, run_id, keep_count; runs lease_until, start_digest, owner_token, owner_unknown_since; states current_run_id, latest_truncate_checkpoint_id. V15: runs entry_step_count, input_row_count, parent_session_id, parent_tool_call_id; the __agents_abandoned_run_starts table
    D1V17, V18the same columns (V17), then the run bookkeeping and the abandoned_run_starts table (V18)
    DO-SQLite (CF DO)11, 12the same, applied by the Durable Object on first use
    Redisnonenew hash fields; see the Redis index note below
    Memorynone

    Pre-release builds of this branch extended V14 / V17 / 11 in place; V15 / V18 / 12 add their columns only when missing, so every database converges on one schema.

    Redis index note. The old saveState could change a session's status without moving its [<keyPrefix>:]idx:sessions:status:<status> index entry, so a session can stay listed under an old status in listSessions({ status }). RedisStateStore.rebuildIndexes() only ADDS entries, it never removes one, so running it alone does not fix that. To clean up, delete the [<keyPrefix>:]idx:sessions:status:* sorted sets, then run rebuildIndexes() once: it re-adds every session under its current status.

  2. Drain, or terminate, in-flight Temporal, DBOS and CF Workflows workflows. This is REQUIRED. Workflow command and step order, step names, workflow inputs and workflow ids changed, and none of it is version-gated (no Temporal patched(), no DBOS.patch). A workflow started by the previous release cannot replay on this release's code, and the runtime cannot detect or settle that itself. The reverse holds too, so a rollback needs the same drain.

    Temporal. Affected: every agent workflow RUNNING at deploy time, including child and companion-continuation workflows. A Temporal agent workflow exits at every HITL boundary, so a session paused on a client tool, an approval or a child has no running workflow of its own and is safe (its children may not be). What changed: a suspension commits before its marker (commitSuspendedStep → marker → pauseSuspendedStream), a resume's first command is startRunCommit, the terminal tail is commitTerminalState + finalizeTerminalRun, a companion continuation id is __continue__<toolCallId> (was __continue__<stepCount>), and a resume workflow id is …__resume-<runId>. Choose one:

    • Drain: stop new entries, then wait until this is empty, and only then deploy the new workers:

      bash
      temporal workflow list --query 'ExecutionStatus="Running" AND WorkflowType="<your agent workflow type>"'
    • Terminate the in-flight ones:

      bash
      temporal workflow terminate --query 'ExecutionStatus="Running" AND WorkflowType="<your agent workflow type>"' \
        --reason "helix snapshot-checkpoint-truth upgrade"

      The sessions recover on the new release: the next execute() / resume() / retry() finds the owner workflow closed (or a pre-upgrade run with no owner token) and supersedes the stale running run (run.stale-running-supersedable).

    • Worker Versioning (Build IDs or worker deployments), so the previous release's workers keep serving their pinned executions until those drain.

    Symptom if you deployed without draining: the old workflows show repeated WorkflowTaskFailed events with cause NON_DETERMINISTIC_ERROR and stay RUNNING, and every later execute() on their sessions gets AgentAlreadyRunningError. Recover by terminating them as above.

    DBOS. The workflow input changed (runMeta → entry), steps were added and reordered, and workflow ids changed (…__resume__<runId>). A paused DBOS session is in-flight too: it is a PENDING workflow blocked in recv, waiting for a client-tool result or an approval. Before deploying, let pending workflows finish (answer their pending calls), or cancel them (DBOS.cancelWorkflow(workflowId), or the DBOS admin / Conductor), then deploy. A cancelled workflow is closed, so the session's next entry supersedes its run; the pending interaction is lost and the user re-sends it.

    CF Workflows. Every execute() / resume() / retry() runs in its own instance with a per-call key (see the instance-id deploy note). Let in-flight instances finish, or terminate them.

    runtime-js / CF DO. Nothing to drain beyond the legacy-mode note: a run that is mid-turn during the deploy is superseded by the next entry (a dead process stops renewing its lease; a DO that is not executing it is ownerless).

  3. Deploy the Cloudflare StreamServer worker before (or together with) the workers that run agents — the claim compare-and-set only protects a newer run once the stream Durable Objects run the new code (see the StreamServer note).

  4. Deploy the rest together. Producers and consumers of the DO /snapshot wire, the chat routes and the React client change together (the wire shape changed; see below).

  5. Pre-upgrade sessions keep working in legacy mode until their first new commit.

Breaking changes (write side) ​

The store contract (custom SessionStateStore implementations) ​

Every message-log or snapshot-visible mutation is now a commit: it writes a checkpoint {id, kind, runId, messageCount, streamSequence, stepCount, state, keepCount} and moves the session's latest pointer in the same atomic operation.

MethodChange
appendMessages(sessionId, messages, commit)commit: MessageCommit ({ streamSequence, runId, startRun? }) is required. With startRun it is the guarded run-start commit (kind: 'input', rows may be empty); without, an append commit
truncateMessages(sessionId, keepCount, commit)commit required; a truncate checkpoint with keepCount, and the session's latestTruncateCheckpointId
saveStateAndPromoteStaging / promoteStagingthe step commit. PromoteStagingOptions.streamSequence / .runId and checkpointMeta.runId are required (a superseded runId throws RunSupersededError and writes nothing)
commitState(sessionId, state, commit, opts) (new)the state commit (K5, the terminal state commit)
cloneSessionwrites a clone checkpoint
saveState(sessionId, state, { fenceRunId } | { fenceIdle })writes no checkpoint any more. Runtimes pass a fence; a write for a superseded run throws RunSupersededError and writes nothing
createCheckpointnever moves the latest pointer
getLatestCheckpointthe checkpoint the pointer names (never the highest stepCount)
renewRunLease(sessionId, runId, ttlMs) (new)a CAS on the store's own clock; rejects a run that is not current and running with RunSupersededError
setRunOwnerUnknownSince(...) (new)set-if-null / clear, for the CF Workflows stuck-instance escape
abandonRunStart(sessionId, runId) (new)'committed' when the run exists, else a tombstone ('abandoned'): a later run-start of that runId throws RunStartRejectedError('entry_abandoned')
listRunCheckpoints(sessionId, runId, { offset, limit }) (new)the run's own checkpoints (Checkpoint.runId === runId), full, ordered by messageCount, from an INDEXED read by run (never a scan of the session); legacy runId: null rows are never returned. planRetry uses it. Postgres V15, D1 V18 and DO-SQLite 12 add the index
updateRunStatusa terminal status (incl. superseded) is sticky: changing a terminal run throws RunSupersededError
getRunread-your-writes for committed runs
createRun / updateRunStatus / getCurrentRun / listRuns / getRunrequired: runtimes no longer feature-detect them, and an executor built on a store without them is rejected at construction

Checkpoint gains the required fields kind: CommitKind | null, runId: string | null and keepCount: number | null (null on legacy rows; a stored kind that is not a commit kind throws UnknownCommitKindError instead of reading as legacy); RunMetadata gains ownerToken, leaseUntil, ownerUnknownSince, startDigest, entryStepCount, inputRowCount and parent.

Also changed for every store:

  • SessionState.checkpointSource is removed (saveState no longer has a staging-checkpoint mode).
  • isSessionStateStore() checks for the new methods too, so an old store no longer passes it.
  • getMessages throws CorruptMessageRowError for a row it cannot parse, where it used to skip the row (snapshot and LLM history reads turn it into state_history_incomplete with cause corrupt_row).
  • D1: cloneSession with a checkpoint that does not exist now throws state_checkpoint_not_found.
  • Redis: getLatestCheckpoint with a pointer naming a missing checkpoint throws a typed state_checkpoint_not_found; it used to fall back to the highest stepCount.
  • Removed exports: calculateUIMessageCount and resolveCheckpointStreamSequence (core; use runBoundaryFrom / responseIndexFor and resolveCommitSequence), ContentReplayOptions (ai-sdk).

The run-start commit must, atomically: (1) be idempotent by runId (same startDigest → write nothing, return the run; a different digest → RunStartRejectedError('idempotency_mismatch')); (2) CAS on startRun.expectCurrentRunId (a live current run → RunStartConflictError('expected_mismatch'), an idle one → RunStartRejectedError('expected_mismatch_idle')); (3) guard the owner (the same ownerToken; ownerless: 'lease' with the current run's lease absent or expired on the store clock; or ownerless: 'asserted'), else RunStartConflictError('run_active'); (4) mark every other live run superseded; (5) apply truncateTo as a truncate checkpoint by the new run; (6) create the run, the rows and the input checkpoint. Every runId-carrying commit is a CAS on "this run is current and live". A store that implements the methods without these guards passes the type check but breaks run.single-live. Run the shared contract suites from @helix-agents/core/testing (commitCheckpointOperationTests, runStartOperationTests, fenceOperationTests, the checkpoint operations, and streamManagerContractTests for a stream manager); the pure rules are exported (applyRunStartRules, planRunStartRewind, buildCommitCheckpoint, assertRunIsCurrentLive, assertRunStatusTransition, …).

Store-specific: Redis and D1 throw RedisCommitContentionError / D1CommitContentionError when a commit keeps losing its optimistic retry; the Redis commit is one Lua script with every key passed in KEYS. Redis Cluster is not supported (CROSSSLOT; see the store README).

Custom stream managers ​

See Custom stream managers below: claimStream is a compare-and-set, endStream / failStream / pauseStream take a run fence, an own-session chunk stamped with another runId is dropped once a newer run claimed the stream (chunkFenceRunId), resumable readers expose catchUpHead / caughtUp, and ResumableReaderOptions.waitForLive is new.

AgentResult.status: 'superseded' and onAgentSuperseded ​

A run that a newer run took over (a concurrent entry, a lost JS lease, an owner-gone takeover) stops with status: 'superseded' and errorDetail.code: 'state_run_superseded' (non-retryable) on every runtime — never failed. Exhaustive switches over AgentResult.status / RunStatus must handle it.

  • onAgentComplete and onAgentFail do not fire for it; the new optional onAgentSuperseded hook does.
  • tracing-langfuse closes its span with level: 'WARNING', statusMessage: 'superseded'.
  • handleChatStream does not treat it as a stream terminal: the superseding run streams on the same session stream.
  • A late client-tool submit to a superseded run answers 409 {code: 'state_run_superseded'}; DOFrontendExecutor.submitToolResult / createHelixChatTransport().submitToolResult throw RunSupersededError.

New errors and codes ​

Error / codeWhen
RunStartConflictError (name: 'AgentAlreadyRunningError')a live run holds the session (run_active, or expected_mismatch against a live run). Existing AgentAlreadyRunningError handling still matches it
RunStartRejectedError (state_run_start_conflict)idempotency_mismatch, expected_mismatch_idle (retried once by the runtimes and the chat handler), entry_abandoned. NOT matched as "already running"
RunSupersededError (state_run_superseded)a write by a superseded run; non-retryable
RunStartAbandonedError / RunStartOutcomeUnknownError (core, transport_error)a Temporal / DBOS / CFW entry that reached no run-start outcome within its fail-safe (below)
StreamClaimStaleReadError (retryable)the stream claim read a lagging owner
state_checkpoint_sequence_unavailablea commit could not read the live stream head (no silent 0 any more, DI-21); the run fails after native retries
state_step_messages_uncommittedthe DBOS / CFW step-atomicity assertion
state_history_incomplete causenow short_read / corrupt_row / store_error / snapshot_unstable

RunStartAbandonedError and RunStartOutcomeUnknownError are exported from @helix-agents/core only; Temporal, DBOS and CF Workflows all throw them. The run-start fail-safe is new: its bounds are Temporal entryOutcomeTimeoutMs, DBOS runStartFailSafeMs and CF Workflows entryOutcomeFailSafeMs (entryStartTimeoutMs overrides it), each 120 000 ms by default. On RunStartOutcomeUnknownError, read the session before retrying: the entry may still commit.

Failure checkpoints and retry() (user ruling D1) ​

  • Every terminal exit — completion, failure and stop — now writes the K5 state checkpoint (with the terminal status and the error; on completion also the onAgentComplete writes, while onAgentFail writes are kept only on Temporal and CF Workflows) before the terminal run status and the stream end. Before, a failure or stop wrote no checkpoint (R28 / R40), so code that took "the latest checkpoint of a failed session" as its last good step now gets K5.
  • retry() never targets K5. Core planRetry picks the failed run's latest step commit (or a later append commit of that run with no assistant row); if the run failed on its first step, its run-start commit — the rows rewind to just before the run's own triggering user message, which the retried run re-appends, followed by options.message when given. A resumed run's drained results (a client-tool or approval answer appended before its first assistant row) are part of its entry: a retry keeps them, and the target is the latest such append commit, so an approved tool's state writes are kept. customState, stepCount and forcedCompletion come from that target. An explicit checkpointId is restored as given. A session that failed BEFORE its last entry's run-start commit keeps its whole log and needs options.message (else validation_error). Retry outcomes match the old model.
  • runtime-js retry() with no message re-runs the failed run's trigger after a FIRST-step failure (CP-58). It used to throw. After a run that committed steps, a retry without options.message still throws validation_error.
  • DBOS retry() is a run-start commit that honours checkpointId (an atomic truncateTo) and message (CP-30); it used to ignore RetryOptions.

Runs on every entry; ownerToken; the JS run lease ​

Every entry (execute, resume, retry, client-tool / approval continuations, children, companions) starts a new run in a run-start commit that supersedes the previous live run. The losing caller of a race writes nothing (no input row, no resumeCount, no stream change). resumeCount is now counted by the WINNER of a resume / retry entry on Temporal, DBOS and CF Workflows; runtime-js and the CF DO never count it.

runtime-js keeps a run lease: JS_RUN_LEASE_MS (30 000 ms), configurable with the executor option runLease: { ttlMs }. It is renewed every ttl / 3 on a timer, through long model and tool calls. A JS event loop blocked for longer than about ttl − ttl/3 (20 s by default) — a synchronous CPU-bound tool — may let a concurrent entry supersede the run; the old run then stops superseded. With no concurrent entry the run renews and continues: only a renewal the store rejects stops it. Raise ttlMs if your tools block that long.

Client-tool batches: resume before submit is refused (R44) ​

resume() (continue and with_message) on a session whose client-tool batch is partially submitted now throws a retryable state_not_resumable and writes nothing — on runtime-js, Temporal, DBOS, CF Workflows and the CF DO (/resume → 409). Before, DBOS in particular resumed while a call was still awaited and the late result was lost. Instead:

  • submit every result of the batch (submitToolResult; once the batch is complete the run continues), or let the call's deadline lapse (a lapsed deadline counts as its timeout result), then resume if needed;
  • to abandon the batch, interrupt() (a pending durable interrupt bypasses the gate) or submit error results;
  • from_checkpoint is not refused: a rewind that discards the unanswered calls proceeds, and one whose restored batch is still partial re-suspends at once on the unanswered calls (R47) and waits for the rest of the batch.

Runtime-specific breaking changes ​

  • runtime-js: an executor built on a store without the run methods throws at construction. New executor option runLease: { ttlMs } and export JS_RUN_LEASE_MS; new exports commitRunStart and finishCommittedRun (re-exported from core) for custom loops; recover() (internal, used by the CF DO's wake recovery) re-enters a running run under its own runId.
  • runtime-temporal: TemporalClient.describeWorkflow(workflowId, runId) is required (the owner-liveness probe): a custom client that implements only the old methods no longer typechecks. Peers: @temporalio/* ^1.13.0 (was ^1.11) and zod is now a peer. New executor option entryOutcomeTimeoutMs; new activities options ownerProbe, workflowProbe, ownerWait and llmAttemptFenceIntervalMs (defaults suit production; tests shorten them). The workflow result gains the superseded status. The workflow input field terminalFinalizeTimeoutMs (default 3 min) bounds the post-K5 finalization (run status, stream end); past it the workflow closes and reader-path recovery finishes the run. A companion continuation records metadata.__helix.companionContinue on the child session, which listSessions metadata filters can see. A detached child's live chunks written after its parent's turn ended are dropped (its content stays in its own session).
  • runtime-cloudflare (CF Workflows): WorkflowInstance.terminate() is required on the binding, and WorkflowStatus.status gains waiting, waitingForPause and unknown. The executor rejects a store without the run methods (assertCfwStoreSupportsRuns). execute() / resume() / retry() now wait for their entry's outcome (the run-start commit, or the instance ending without one) before returning. AgentSteps.initializeAgentState and AgentSteps.createResumeRun are removed (the run-start entry step replaces them) and runId is required on the step inputs. New admin executor.terminateRun(sessionId) for an operator-paused instance; cfwUnknownTerminateAfterMs (default 600 000 ms) terminates an instance stuck unknown and writes a run_terminated chunk (UI data-run-terminated).
  • runtime-cloudflare (DO): DurableObjectAgentBase / DOStateStore own the run-start commit (runStartCommitSync, one transactionSync); DOStoreNotInitializedError.
  • runtime-dbos: delivered: false companion results carry child_busy / child_exited / child_start_unconfirmed; any step that runs out of retries settles the turn failed.

UI message ids ​

A run's response message id (msg-<startUIMessageCount>) is now the first-row index of the UI message that holds the response (core responseIndexFor): a continuation (a client-tool or approval round) streams into the open assistant group's id, and system rows count like any other indexed row. Snapshot and live ids now agree (ids.snapshot-matches-live). Ids are non-contiguous, and data keyed by a message id may shift after the upgrade.

snapshot.state ​

FrontendSnapshot.state is the custom state AT the snapshot's cursor (the checkpoint's state), not the latest saveState. Later changes arrive as state_patch frames after the cursor.

Read side: snapshots are checkpoint-anchored; contentReplay is removed ​

This part affects code that calls buildSnapshot / handleChatStream / createCloudflareChatHandler, reads the DO /snapshot wire, or uses DOStateStoreClient directly.

Why ​

A snapshot used to be assembled from independent reads (messages, stream info, run, chunks) plus a server-side merge of in-flight "partial" content. Commits landing between those reads tore it: status: 'active' with the stream head as cursor but the final text missing, or status: 'ended' with the last step missing. The 50-row cap corrupted long sessions.

A snapshot is now taken from one checkpoint:

FieldMeaning
messagesthe committed rows [0, checkpoint.messageCount), UI-converted; never partial, never capped
streamSequencethe checkpoint's cursor: everything committed up to it is in messages, nothing after it
statethe custom state at the cursor
statusderived from the current run: a live run is active / paused
resumeMessageIdwhen live: the id of the UI message the run's response streams into

The canonical client ​

ts
const snap = await fetch(`/api/chat/${sessionId}/snapshot`).then((r) => r.json());
const opts = snapshotResumeOptions(snap);
// Seed WITHOUT the resume message; only when live, resume — the server replays
// that message's committed prefix, then streams the tail after the cursor.
const chat = useHelixChat({
  id: sessionId,
  messages: opts.messages,
  resume: opts.resume,
  transport: new DefaultChatTransport({
    api,
    prepareSendMessagesRequest: prepareHelixChatRequest({ api }),
    prepareReconnectToStreamRequest: prepareHelixReconnectRequest({
      api,
      resumeFromSequence: opts.resumeFromSequence,
      existingMessageId: opts.existingMessageId,
      replayPrefix: true,
    }),
  }),
  resync: { snapshotUrl: `/api/chat/${sessionId}/snapshot` },
});

Why the resume message is removed (C-1). ai 6.0.281 starts a resumed response from an EMPTY message that REPLACES the last one; ≤ 6.0.230 continue the last message. Seeding the whole snapshot and resuming only the tail therefore drops the in-flight message's committed steps on 6.0.281. With the message removed and X-Helix-Replay-Prefix: 1 (replayPrefix: true / resumeHeaders({ …, replayPrefix: true })), the server reads ONE snapshot (with HandleChatStreamDeps.snapshotOptions — pass the options your snapshot route uses), replays that message's committed parts in order (only ai 6.0.0 chunk types, between transient data-helix-replay{phase: 'begin' | 'end'} markers) and streams the tail after that snapshot's cursor — identical on every version, also for a message that spans runs (a HITL continuation) and when early chunks were trimmed. useHelixChat never hands a replayed tool call to onToolCall and never re-sends a replayed tool's committed result; a replayed tool that is still PENDING (awaiting the browser) is dispatched once after the replay and its output is sent. The message's committed metadata rides on start; an answered approval replays as its answer (never re-asked). handleChatStream takes the same flag as replayPrefix: true or the request header, and honours it only for a CLIENT-supplied existingMessageId; an id the snapshot does not hold answers a typed 409 state_not_resumable (cause: 'resume_message_unknown'), never the tail alone. The flag and the seed always travel together: snapshotResumeOptions() / useResumableChat() return messages + replayPrefix; a client that cannot suppress replayed tools (a plain or pre-built Chat) must seed the whole snapshot and resume WITHOUT the flag (snapshotResumeOptions(snap, { replayPrefix: false })) — re-anchors do this automatically (with a dev warning).

Migrating ​

contentReplay / contentReplayEnabled (removed) ​

Delete contentReplay: true from handleChatStream / handleChat calls and contentReplayEnabled from buildSnapshot deps. Resume from the snapshot cursor (above). A resume without a cursor still replays the current run from its start under the run's response id (msg-<startUIMessageCount>) — what contentReplay was used for on a cold load.

FrontendSnapshot ​

  • Removed: timestamp, startSequence, startMessageCount. Use streamSequence (the cursor) and resumeMessageId.
  • SnapshotOptions.includePartialContent is removed; there is no partial content to include. In-flight content arrives through the resumed stream.

status follows the run ​

status is active / paused whenever the current run is live — including the short window after a run's run-start commit, before its stream is claimed, when the stream still shows the previous turn's ended. Resume on active and paused; a resume in that window waits for the run to go live. If it never does within HandleChatStreamDeps.liveWaitTimeoutMs (default 30 000 ms), the handler answers a retryable 503{code: 'transport_stream_drop', cause: 'run_not_live'} — retry the snapshot.

Run-start conflicts ​

  • A send (a request whose trailing message is a new user message) refused with RunStartConflictError (named AgentAlreadyRunningError) answers 409 {error, code: 'state_already_running', retryable: true}. It no longer attaches to the running run's stream, which rendered that run's text as the reply and dropped the message. A stock useChat / useHelixChat ends the send with status 'error'; read the body with parseHelixChatError(chat.error) and retry with chat.regenerate(). The handler also reads the stream head BEFORE execute() and streams only after it, so a new turn never replays earlier turns.
  • Requests with no new user content (a reconnect, a passive attach) and tool-result resumes keep attaching to the winner when a run is already live.
  • RunStartRejectedError(expected_mismatch_idle): retried once by the handler.
  • Any other RunStartRejectedError (e.g. idempotency_mismatch): a typed 409{error, code: 'state_run_start_conflict', cause} instead of a silent attach.

Cloudflare DO ​

  • GET /snapshot returns the SnapshotView wire: sessionId, status, head, streamSequence, legacy, checkpointId, stepCount, checkpoint {id, kind, runId, keepCount, stepCount, messageCount, streamSequence}, messages, state, run {runId, status, startUIMessageCount, startSequence, turn}, plus the session fields (streamId, sessionStatus, createdAt, updatedAt, agentType, version, error, errorDetail, interruptContext, the client-tool maps, root / parent ids, streamError, streamFinalOutput). timestamp and the top-level startSequence / startMessageCount / startUIMessageCount / runId / turn are removed.
  • A missing session is 404 {error: 'Session not found', code: 'state_session_not_found'} (was 200 null); a read failure is 500 {error, code, cause?} (e.g. state_history_incomplete / corrupt_row).
  • DOStateStoreClient: getMessages pages GET /messages (limit ≤ 1000); getLatestCheckpoint / getCurrentRun return the real checkpoint and run (null when there is none) — the run's real startSequence and turn included, so the chat handler's stale-cursor clamp and mid-block run floor hold on the DO; a non-404 failure throws a typed DOPeerError instead of returning null.
  • createCloudflareChatHandler().getSnapshot({ includeReasoning }) — includeThinking was silently ignored and is now mapped (deprecated).

React hooks and client helpers ​

  • The snapshot is UI-shaped (buildSnapshot output); the hooks no longer re-convert it (RM-23). If you hand-rolled a ServerSnapshot with raw Message[], switch to FrontendSnapshot.

  • useAutoResync / useResumableChat / useCheckpointSnapshot retry a 5xx snapshot (fetchSnapshotWithRetry: 4 fetches, 250 / 500 / 1000 ms) and surface a typed SnapshotUnavailableError after exhaustion — render an explicit error state with a retry control, never an empty conversation.

  • A resync re-anchors, safely by default. On a resync trigger the hooks stop the stream, fetch the LATEST snapshot, set its messages and — when live — resume from its cursor under resumeMessageId. The simplest wiring is useHelixChat alone:

    ts
    const chat = useHelixChat({
      id: sessionId,
      transport,
      messages: snapshot.messages,
      resume: snapshot.status === 'active' || snapshot.status === 'paused',
      resync: { snapshotUrl: `/api/chat/${sessionId}/snapshot` },
    });
    if (chat.resyncError) return <RetryBanner error={chat.resyncError} />;

    useAutoResync / useResumableChat resolve their stream control from stop + resume, else chat (e.g. the useHelixChat result), else the one mounted useHelixChat. With none (or several chats mounted) a resync now surfaces a typed ConfigurationError (missing: ['stop', 'resume']) via resyncError / onError and leaves the messages alone — it no longer replaces messages under a stream that is still writing. Feed the hooks collectResyncTriggers(messages).

  • Stream control is per chat. withHelixStreamControl / createHelixChatTransport key each connection by chatId: two chats sharing a transport never abort each other; abortActive(chatId?), hasActiveConnection(chatId?) (now a function) and activeKind(chatId). A reconnect the transport aborts before its response resolves null (ai 6.0.91–6.0.230 would report a thrown abort as a chat error). controlChat(chat, transport, { getStatus }) (framework-agnostic; useHelixChat uses it) returns an abort-aware, bounded stop, a resume, and a sendMessage that stops an open resume first so Chat's state stays right.

  • Resync control comes from one source. stop + resume together, else chat, else the mounted useHelixChat whose id is chatId or whose setMessages is the hook's setMessages (never "the one mounted chat"). A lone stop or resume is a typed ConfigurationError; a stop that does not abort the transport (e.g. a bare chat.stop) logs a development warning.

  • stop() on a resumed stream. ai 6.0.0–6.0.230 (incl. the locked 6.0.49) call reconnectToStream without an abortSignal, so chat.stop() cannot stop a resumed stream. useHelixChat now wraps its transport with withHelixStreamControl (each connection has its own AbortController; a new connection aborts the previous one) and returns a stop that aborts the connection and resolves once the chat is idle. With plain useChat, wrap the transport yourself and call transport.abortActive(chatId) before chat.stop(). createHelixChatTransport owns its connections the same way (abortActive(chatId) — scoped to one chat). controlChat's stop rejects with a typed StreamStopTimeoutError if the chat does not go idle within its bound, and a send waits for a re-anchor in flight (and vice versa).

  • Resume headers. prepareHelixReconnectRequest now lets per-request headers win over its configured getters, so chat.resumeStream({ headers: resumeHeaders(target) }) resumes at a re-anchor's cursor. resumeHeaders sends X-Helix-Replay-Prefix only when target.replayPrefix is set — set it exactly when the chat was seeded WITHOUT the resume message (snapshotResumeOptions(snap) gives both); with a whole-snapshot seed leave it unset, or the replayed prefix duplicates the message on ai ≤ 6.0.230.

  • Retried attempts. A Temporal activity retry or a CF Workflows step.do retry re-streams a step under a new attemptId. A resume / attach / terminal replay drops the superseded attempt's content (the committed step keeps step_committed.committedAttemptId when that attempt streamed in the step — a commit naming none ('') or an id that never streamed changes nothing; an in-flight step the attempt that STARTED last, by the stream position of each attempt's step_start, which every runtime writes as the attempt's first chunk before calling the model — never by a clock), including its step markers and files. A live attempt switch — a newer attempt, or a step_committed naming another attempt — emits a data-attempt-superseded event (a resync trigger), and so does a resume whose cursor is past superseded content the client was already sent.

  • Types. onResync receives AISDKResyncTriggerEvent (AISDKStreamResyncEvent | AISDKAttemptSupersededEvent; narrow on event.type before reading checkpointId). data-stream-resync carries runId; ResyncTracker dedups a resync by (runId, checkpointId) and rescans an array that shrank or was replaced. setMessages in the hook options is typed (messages: UIMessage[]) => void (it was (messages: unknown[]) => void): useChat's setMessages fits directly for the default UIMessage; with a custom message type, adapt it ((m) => setMessages(m as MyUIMessage[])). A setter declared over unknown[] still type-checks.

  • fetchSnapshotWithRetry(url, { signal }) aborts (never retried); a 429 is retried like a 5xx; a 200 whose body is not JSON / not a snapshot throws SnapshotUnavailableError (reason: 'invalid_json' / 'invalid_body'). The hooks abort a fetch in flight on unmount or a snapshotUrl change, and only the latest re-anchor lands.

  • Reasoning renders once: a thinking chunk with isComplete: true is the whole block, and the transformer adds only what the streamed deltas did not deliver.

  • snapshotResumeOptions(snapshot) → { resume, resumeFromSequence?, existingMessageId?, messages } (messages: what to seed — without the resume message when resuming); useResumableChat exposes status, streamSequence, resumeMessageId and shouldResume.

useHelixChat().retry() and failed resumes ​

  • UseHelixChatResult gains retry(). After a failed RESUME it re-anchors on the latest snapshot (which restores the in-flight message the C-1 seed left out) and resumes from its cursor. After a failed SEND it re-sends (regenerate()) only when the latest snapshot does not hold the user turn yet; otherwise it re-attaches, so a turn is never persisted twice. Without resync it is regenerate(). Wire your Retry button to chat.retry(), not chat.regenerate().
  • useHelixChat with resync re-anchors on its own when the mount-time resume fails (409 / 503 / a network error), so the hidden in-flight message comes back.

Pending approvals survive a reload ​

A snapshot renders a pending approval as an approval-requested part with its approval id (from PendingClientToolCall.approvalId, which the core approval path writes on runtime-js and the CF DO; on Temporal, DBOS and CF Workflows a pending server tool is treated as an approval), and the resumed stream re-sends a skipped, unanswered tool-approval-request once. sendAutomaticallyWhen / extractResumeIntent accept an approval-responded part with providerExecuted: true. The CF DO now records an approval answer as { approved, reason }.

Wire errors (every host) ​

CaseStatusBody
snapshot of a missing session404{error: 'Session not found', code: 'state_session_not_found'}
snapshot read failure500{error, code, cause?} — e.g. state_history_incomplete / corrupt_row
late submit to a superseded run409{error, code: 'state_run_superseded'}
a resume / attach of a run whose stream FAILED (incl. before its first commit)409{error: 'stream_failed', code, retryable, sessionId, streamId, message, errorDetail?}: code / retryable come from the run's errorDetail (parseHelixChatError reads them)
a send (new user message) refused because a run is live409{error, code: 'state_already_running', retryable: true} (parseHelixChatError)
an entry the store rejected (RunStartRejectedError, not idle-retried)409{error, code: 'state_run_start_conflict', cause}
a replay-prefix resume for a message id the snapshot does not hold409{error, code: 'state_not_resumable', cause: 'resume_message_unknown', retryable: true}
a resume while a client-tool batch is partially submitted (CF DO /resume, R44)409{error, code: 'state_not_resumable'}
a resume in the run-start window whose run never goes live within liveWaitTimeoutMs503{error, code: 'transport_stream_drop', cause: 'run_not_live'} (Retry-After: 1)
a failed committed-prefix read for a replay-prefix resume503{error, code, cause?} (retryable)

@helix-agents/ai-sdk/cloudflare exports snapshotNotFoundResponse(), snapshotErrorResponse(error) and submitToolResultErrorResponse(error) for route handlers. A snapshot route should answer:

ts
try {
  const snapshot = await chat.getSnapshot({ sessionId });
  return snapshot ? Response.json(snapshot) : snapshotNotFoundResponse();
} catch (error) {
  return snapshotErrorResponse(error);
}

On the client, DOFrontendExecutor.submitToolResult and createHelixChatTransport().submitToolResult throw a typed RunSupersededError for the 409. agent-server's GET /status gains runId / runStatus (a superseded run is never reported as a failure); its 500 envelope keeps code: 'INTERNAL_ERROR' and adds errorCode / cause for a typed failure. RemoteStatusResponseSchema parses a runStatus it does not know (a newer peer's) as 'unknown' (RemoteRunStatus) instead of failing the response.

Pre-upgrade sessions (legacy mode) ​

A session whose latest checkpoint predates this release (kind: null) is served in legacy mode: the complete live log, with the stream head as the cursor, whatever its status. It applies until the session's first new commit (its next run-start commit re-anchors it). Legacy checkpoints can claim too few rows (the old DO message_count = 0, DBOS promotes, CF Workflows appends), which is why the live log is served. This covers terminal sessions and paused sessions that cannot be drained (suspended on a client tool or an approval for days) on runtime-js, the CF DO, Temporal and CF Workflows: they are never replayed from sequence 0. On DBOS a paused session is an in-flight workflow (see the deploy checklist). A legacy session's snapshot status comes from the session and the stream (not from a pre-upgrade run record, which may never have been ended: the previous release wrote no Temporal terminal run status, RM-47), so a session that finished before the upgrade reads ended / failed, not active. A RUNNING legacy session is not covered — drain running sessions before upgrading (see the deploy checklist). A pre-upgrade running run record with no lease or owner becomes supersedable under the new owner rules (no lease on runtime-js, a closed workflow on Temporal, !isExecuting on the CF DO), so it never bricks a session.

Custom stream managers ​

ResumableReaderOptions.waitForLive (new, optional): until the reader yields its first chunk past fromSequence, a terminal status — or a stream that does not exist yet — means "not live yet": return a reader (never null) and park on your normal wake-up path instead of ending; close() must release a parked next(). A manager that ignores it keeps the old behavior (a resume in the run-start window then ends early on that backend). Once the reader has seen the stream go live — active at attach, a transition to active after it, a chunk past the cursor, or a missing stream created — normal terminal semantics apply, so a run that goes live and ends before its first chunk ends the reader promptly; a terminal status left from the previous turn never does. The shared contract cases are WL.* in @helix-agents/core/testing's streamManagerContractTests.

StreamManager.claimStream(streamId, runId, expect?) is a compare-and-set (CP-85): with expect: { owner } it sets the stream's owner run only while the current owner (null = none) is expect.owner, and resolves true when runId owns the stream afterwards, false otherwise (nothing written; a missing stream is never created). Runtimes call it only through core's claimStreamForCommittedRun. A manager whose claimStream still resolves undefined is treated as having claimed (the pre-CAS, unconditional behaviour). Contract case: RF.claim.cas.

Deploy note: the Cloudflare StreamServer claim (rolling deploys) ​

The /claim endpoint of the Cloudflare StreamServer Durable Object (store-cloudflare) gained the compare-and-set (expectOwner). During a rolling deploy, a NEW client talking to a not-yet-updated StreamServer gets the old, unconditional claim (the old object ignores expectOwner and does not report claimed, which the client reads as claimed). The claim then cannot protect a newer run's claim that lands in that window. Deploy the StreamServer (the worker that hosts the stream Durable Objects) before, or together with, the workers that run agents, and drain in-flight runs as for the rest of this release.

Deploy note: Cloudflare Workflows instance ids change (drain or terminate first) ​

On Cloudflare Workflows every execute(), resume() and retry() now runs in its own instance keyed on a key minted by that call (agent__<name>__<sessionId>__turn__<key>, …__resume__<key>, …__retry__<key>; ids that Cloudflare would reject are encoded as h_<hash>_<tail> by buildWorkflowInstanceId). A re-spawned sub-agent is …__respawn__<parent-run suffix>, and a companion continuation / resume is …__continue__<step>-<callId> / …__resume__<step>-<callId>. getHandle() finds the live instance from the session's run record, never by rebuilding an id. An instance started by the previous release has an old-shape id and runs the old workflow body, which this release's body cannot replay. Before you deploy, let in-flight CF Workflows instances finish, or terminate them (a terminated run's session is then failed or superseded by the next entry). Code that rebuilt agent__<name>__<sessionId> to look an instance up must call executor.getHandle(agent, sessionId) instead.

Released under the MIT License.