Skip to content

The run-start commit gates on the session and the run (CP-96) ​

This guide covers the release in which every store's run-start commit refuses an entry whose session or current run changed since the entry planned (CP-96). It is a major release of @helix-agents/runtime-js, @helix-agents/runtime-temporal, @helix-agents/runtime-dbos, @helix-agents/runtime-cloudflare and @helix-agents/ai-sdk, and a minor release of @helix-agents/core (a 0.x package: minor marks an observable behaviour change) and of the store-memory, store-redis, store-postgres and store-cloudflare packages.

No deploy drain is needed. No workflow input, workflow argument or instance param changed, so an in-flight run re-enters through the new code. No migration either: neither new field is persisted.

What was wrong ​

A resume() or retry() that read the session just before another entry's run started and finished could still win its run-start commit: the commit compared only the current run's id, and the session-status check ran at plan time. The stale entry then ran a second turn on a session the winner had already finished. A stale retry() rewound the winner's completed turn and lost its result.

The new rule ​

An entry commits only if the session and the current run are unchanged since the reads its plan was made from. The store checks two new required StartRun fields inside the commit's atomic step; the rule order is in Execution Flow §Entry gates.

FieldMeaningRefused with
expectCurrentRunStatusthe status of the current run the entry saw (null iff no run seen)RunStartRejectedError('expected_run_status_changed')
entryStatusesthe session status(es) the plan read (non-empty); the session gateRunStartRejectedError('entry_status_changed')

What your callers see ​

  • A stale entry loses typed and writes nothing. A run-status mismatch is mapped to the existing AgentAlreadyRunningError (RunStartConflictError('expected_mismatch')): "a newer run started" and "the run I saw moved on" are one race, so handle both as "already running" (attach, or wait). expected_run_status_changed never reaches your code as itself.
  • A session-status loss is RunStartRejectedError('entry_status_changed'). It means the session changed under the entry (for example an /abort). Reload, then plan again. Through handleChatStream it is a 409 state_run_start_conflict. It is now enforced by every store; before, only the CF Durable Object did.
  • CF Workflows. A session-gate loser now reaches execute() / resume() / retry() as RunStartRejectedError('entry_status_changed'). Before, every lost Workflows entry surfaced as the C3 conflict.
  • runtime-js and the CF DO: a concurrent execute() that read a live run loses (AgentAlreadyRunningError, a retryable 409 on a send), even when that run finished before its commit. Before, it ran with a stale UI boundary. Temporal, DBOS and CF Workflows re-plan inside their builder, so a later entry there runs as an ordinary next turn.
  • On Temporal, DBOS and CF Workflows the exact CP-96 interleave was already refused (by an owner check — Temporal: the executor's assertNoLiveOwner; DBOS: the entry step's assertDbosOwnerless — or the run-id CAS); the gates now also cover the window between plan and commit.
  • The loser's class in the CP-96 race differs per runtime. On runtime-js and the CF DO it is always the C3 AgentAlreadyRunningError (expected_mismatch). On Temporal, DBOS and CF Workflows the entry re-plans from a fresh read, so a stale entry gets one of three: the C3 conflict when its commit lands after the winner started; a RunStartRejectedError('entry_status_changed') when it plans after the winner finished; or, while the winner's workflow / instance is open, an owner-check refusal before its commit (Temporal: a plain AgentAlreadyRunningError with no conflictCause; DBOS and CF Workflows: RunStartConflictError('run_active')). The C3 and owner-check errors classify as state_already_running, entry_status_changed as state_run_start_conflict. Handle all of them as "another entry won" and reload. Details: Execution Flow §Entry gates (FU-CP96-DURABLE-LOSER-CLASS).

Custom SessionStateStore implementers ​

  1. Read the session status atomically with the commit and pass it to core's applyRunStartRules as the new required sessionStatus input (a SessionStatus, mapped from your raw status if you store another form). It must be the value in the same transaction / script / transactionSync that creates the run, not a read before it.
  2. Read the current run's status in that same step. applyRunStartRules takes it from current.status and applies the run-status CAS after the run-id CAS.
  3. If your commit can retry on a stale view (a CAS or script guard miss), pin the session status and the current run's status in the guard, not only the version: a status update does not always bump it.
  4. Run runStartOperationTests from @helix-agents/core/testing. Its seven run.entry-status-gate: scenarios cover both gates, "refusal writes nothing", the replay-is-a-noop rule and the validation.

applyRunStartRules runs assertValidEntryGate first: an empty entryStatuses, or an expectCurrentRunStatus that is null exactly when expectCurrentRunId is not, throws a typed validation_error.

Code that builds StartRun ​

Set both fields. Take them from the reads the plan used:

typescript
const current = await store.getCurrentRun(sessionId); // read BEFORE the session read the plan uses
const session = await store.loadState(sessionId);
if (!session) throw new Error(`Session ${sessionId} not found`); // plan nothing for a missing session

const startRun: StartRun = {
  turn,
  startUIMessageCount,
  expectCurrentRunId: current?.runId ?? null,
  expectCurrentRunStatus: current?.status ?? null,
  entryStatuses: [session.status], // already a SessionStatus
  ownerToken,
  ownerless,
};

session.status is already a SessionStatus. Code that holds an AgentState instead converts its agent status with core's agentStatusToSessionStatus(...).

Neither field is part of startDigest, so an idempotent replay is decided before either gate. Use core's matchesSeenRun(current, seenRunId, seenRunStatus?) and shouldRetryRunStart({ seenRunStatus, entryStatuses, ... }) if you build your own retry-once-on-stale-read loop.

Packages ​

  • @helix-agents/ai-sdk: RunStartRejectedShape accepts the expected_run_status_changed cause. It never reaches the chat handler unmapped.
  • Stores: Redis's Lua session-status guard is now unconditional; D1 pins the raw session status in its claim guard (the current run's status is covered by its live-run view guard: a terminal run's status never changes); DO-SQLite reads the status once for the gate and the finish facts, and the Durable Object base no longer checks statuses itself (the store does, in the same transactionSync).

Released under the MIT License.