The run-start commit gates on the session and the run (CP-96)
This guide covers the release in which every store's run-start commit refuses an entry whose session or current run changed since the entry planned (CP-96). It is a major release of @helix-agents/runtime-js, @helix-agents/runtime-temporal, @helix-agents/runtime-dbos, @helix-agents/runtime-cloudflare and @helix-agents/ai-sdk, and a minor release of @helix-agents/core (a 0.x package: minor marks an observable behaviour change) and of the store-memory, store-redis, store-postgres and store-cloudflare packages.
No deploy drain is needed. No workflow input, workflow argument or instance param changed, so an in-flight run re-enters through the new code. No migration either: neither new field is persisted.
What was wrong
A resume() or retry() that read the session just before another entry's run started and finished could still win its run-start commit: the commit compared only the current run's id, and the session-status check ran at plan time. The stale entry then ran a second turn on a session the winner had already finished. A stale retry() rewound the winner's completed turn and lost its result.
The new rule
An entry commits only if the session and the current run are unchanged since the reads its plan was made from. The store checks two new required StartRun fields inside the commit's atomic step; the rule order is in Execution Flow §Entry gates.
| Field | Meaning | Refused with |
|---|---|---|
expectCurrentRunStatus | the status of the current run the entry saw (null iff no run seen) | RunStartRejectedError('expected_run_status_changed') |
entryStatuses | the session status(es) the plan read (non-empty); the session gate | RunStartRejectedError('entry_status_changed') |
What your callers see
- A stale entry loses typed and writes nothing. A run-status mismatch is mapped to the existing
AgentAlreadyRunningError(RunStartConflictError('expected_mismatch')): "a newer run started" and "the run I saw moved on" are one race, so handle both as "already running" (attach, or wait).expected_run_status_changednever reaches your code as itself. - A session-status loss is
RunStartRejectedError('entry_status_changed'). It means the session changed under the entry (for example an/abort). Reload, then plan again. ThroughhandleChatStreamit is a 409state_run_start_conflict. It is now enforced by every store; before, only the CF Durable Object did. - CF Workflows. A session-gate loser now reaches
execute()/resume()/retry()asRunStartRejectedError('entry_status_changed'). Before, every lost Workflows entry surfaced as the C3 conflict. - runtime-js and the CF DO: a concurrent
execute()that read a live run loses (AgentAlreadyRunningError, a retryable 409 on a send), even when that run finished before its commit. Before, it ran with a stale UI boundary. Temporal, DBOS and CF Workflows re-plan inside their builder, so a later entry there runs as an ordinary next turn. - On Temporal, DBOS and CF Workflows the exact CP-96 interleave was already refused (by an owner check — Temporal: the executor's
assertNoLiveOwner; DBOS: the entry step'sassertDbosOwnerless— or the run-id CAS); the gates now also cover the window between plan and commit. - The loser's class in the CP-96 race differs per runtime. On runtime-js and the CF DO it is always the C3
AgentAlreadyRunningError(expected_mismatch). On Temporal, DBOS and CF Workflows the entry re-plans from a fresh read, so a stale entry gets one of three: the C3 conflict when its commit lands after the winner started; aRunStartRejectedError('entry_status_changed')when it plans after the winner finished; or, while the winner's workflow / instance is open, an owner-check refusal before its commit (Temporal: a plainAgentAlreadyRunningErrorwith noconflictCause; DBOS and CF Workflows:RunStartConflictError('run_active')). The C3 and owner-check errors classify asstate_already_running,entry_status_changedasstate_run_start_conflict. Handle all of them as "another entry won" and reload. Details: Execution Flow §Entry gates (FU-CP96-DURABLE-LOSER-CLASS).
Custom SessionStateStore implementers
- Read the session status atomically with the commit and pass it to core's
applyRunStartRulesas the new requiredsessionStatusinput (aSessionStatus, mapped from your raw status if you store another form). It must be the value in the same transaction / script /transactionSyncthat creates the run, not a read before it. - Read the current run's status in that same step.
applyRunStartRulestakes it fromcurrent.statusand applies the run-status CAS after the run-id CAS. - If your commit can retry on a stale view (a CAS or script guard miss), pin the session status and the current run's status in the guard, not only the version: a status update does not always bump it.
- Run
runStartOperationTestsfrom@helix-agents/core/testing. Its sevenrun.entry-status-gate:scenarios cover both gates, "refusal writes nothing", the replay-is-a-noop rule and the validation.
applyRunStartRules runs assertValidEntryGate first: an empty entryStatuses, or an expectCurrentRunStatus that is null exactly when expectCurrentRunId is not, throws a typed validation_error.
Code that builds StartRun
Set both fields. Take them from the reads the plan used:
const current = await store.getCurrentRun(sessionId); // read BEFORE the session read the plan uses
const session = await store.loadState(sessionId);
if (!session) throw new Error(`Session ${sessionId} not found`); // plan nothing for a missing session
const startRun: StartRun = {
turn,
startUIMessageCount,
expectCurrentRunId: current?.runId ?? null,
expectCurrentRunStatus: current?.status ?? null,
entryStatuses: [session.status], // already a SessionStatus
ownerToken,
ownerless,
};session.status is already a SessionStatus. Code that holds an AgentState instead converts its agent status with core's agentStatusToSessionStatus(...).
Neither field is part of startDigest, so an idempotent replay is decided before either gate. Use core's matchesSeenRun(current, seenRunId, seenRunStatus?) and shouldRetryRunStart({ seenRunStatus, entryStatuses, ... }) if you build your own retry-once-on-stale-read loop.
Packages
@helix-agents/ai-sdk:RunStartRejectedShapeaccepts theexpected_run_status_changedcause. It never reaches the chat handler unmapped.- Stores: Redis's Lua session-status guard is now unconditional; D1 pins the raw session status in its claim guard (the current run's status is covered by its live-run view guard: a terminal run's status never changes); DO-SQLite reads the status once for the gate and the finish facts, and the Durable Object base no longer checks statuses itself (the store does, in the same
transactionSync).