Skip to content

CF Workflows terminal settle and pre-commit entry failures ​

This guide covers the release in which Cloudflare Workflows records a truthful outcome for every run whose instance stops, and in which an entry that fails before its run-start commit no longer overwrites a session that already has a run. It is a major release of @helix-agents/runtime-cloudflare and @helix-agents/runtime-js, and a minor release of @helix-agents/core and @helix-agents/store-cloudflare (0.x packages: minor marks an observable behaviour change).

No deploy drain and no migration. No workflow param, instance id or schema changed. Deploy order only affects one optional field: an old stream worker (StreamServer) or Durable Object returns no errorDetail in a failed stream's info until it is redeployed (see Stream info).

abort() on Workflows writes the terminal commit ​

handle.abort(reason) still terminates the owner instance and ends the run failed with a typed framework_cancelled. What changed:

  • The failure is now the terminal state commit (session failed → run failed → stream failed), like every other terminal exit. Before, it was a plain state save with no checkpoint. retry() never targets that commit.
  • handle.result() and getHandle(...).result() report the persisted detail ({ message: reason, code: 'framework_cancelled', category: 'framework', retryable: false }). Before, they reported { message } only (DI-33). The aborted session's output is cleared.
  • result() reports only its own run. Every handle (the reconnect handle from getHandle() included) reports the outcome of the run it was built for, never a newer run's: it can return superseded, or completed with no output when a newer run has committed since (a failed run then reports its own record's error as { message }). The reconnect handle's result() and canResume() read the session's state at call time, not at getHandle().
  • abort() can now reject. If the terminal commit cannot be written (a transient store error), nothing is written, the run stays live, and abort() rejects with a HelixError transport_error, retryable: true, with the store error as cause. Before, it resolved and left the session active behind a failed run.
typescript
try {
  await handle.abort('User cancelled');
} catch (err) {
  if (HelixError.isInstance(err) && err.code === 'transport_error' && err.retryable) {
    await handle.abort('User cancelled'); // nothing was written; try again
  } else {
    throw err;
  }
}

On Miniflare 4 (older @cloudflare/vitest-pool-workers), terminate() is not implemented: the commit still lands, but result() waits until the instance ends on its own. The real engine and Miniflare 5 stop it.

Instances closed outside the product are settled on read ​

A run whose instance was terminated from the Cloudflare dashboard or wrangler, errored, or disappeared before its terminal commit used to stay active / running with a live-looking stream, and result() reported 'Unknown error' (DI-34). Now the next executor reader settles it: each handle's result(), getHandle(), canResume() and abort(). The session, run and stream end failed with a detail that says how the instance closed (abort() settles with its own cancellation instead, framework_cancelled carrying the caller's reason):

InstanceerrorDetail
terminated{ message, code: 'framework_cancelled', category: 'framework', retryable: false }
erroredthe engine's error message, no code
complete without a terminal commit / gone{ message } naming the instance; no code
complete, its output its own failurethat output's validated errorDetail (its error path's cleanup failed)

No lifecycle hook (onAgentFail) fires for a settled run. A run whose terminal commit did land is finished with that outcome instead. A live, operator-paused or unreadable instance is never settled.

What to check: code that waited for canResume() to stay true, or a session to stay active, after an external terminate. Code that alerts on executor log events: cfw.handle.finish_committed_run_failed and cfw.reconnect.finish_committed_run_failed are now cfw.handle.settle_failed and cfw.reconnect.settle_failed.

A host that only serves the chat route (handleChatStream) calls none of these readers, so there the session stays active until its next send, whose entry supersedes the stranded run.

The entry fail-safe throws RunStartStalledError ​

execute() / resume() / retry() on Workflows wait for their entry for at most entryOutcomeFailSafeMs (120 000 ms). A run that committed but has not started by then (no run_started in its own stream, or no entry state) gets the same bound again until its instance closes or is paused, measured from when the call first saw the commit (a healthy instance under load can commit just before the bound and start just after it; a transient owner read or abandon failure does not end it). A real stall can therefore take up to about twice the bound to surface. A run that still has not started used to get a handle whose stream did not exist. Now it never gets a handle: the executor terminates the instance and throws core's RunStartStalledError (transport_error).

It is never retryable (retryable: false for every outcome): the run-start commit already appended the call's input, so repeating the same call (execute() again) would append it a second time. The recovery is executor.retry(agent, sessionId), which re-runs the failed run from its own entry. Every outcome's message names it the same way: executor.retry(agent, sessionId) (pass options.message if the run committed a step or its entry carried no user message of its own). Without a message such a run's retry() is refused validation_error: a run whose start is unknown may have committed a step before it was stopped, and a resume() in continue / from_checkpoint mode with no message commits only heal rows, so its run has no trigger to re-run. What settle says:

  • ownerStopped: true, settle: 'failed': the instance stopped, and the run was settled failed with the error's own detail. Call executor.retry(agent, sessionId).
  • ownerStopped: true, settle: 'not_committed': the instance stopped, but the run's failed write did not land, so nothing was written. A reader settles the run first (getHandle(), canResume() or a handle's result()), then call executor.retry(agent, sessionId).
  • ownerStopped: true, settle: 'not_found': the session or its run could not be read, so nothing was settled. Read the session first.
  • ownerStopped: true, settle: 'superseded': a newer run had already taken the session over, so this call's input belongs to a run that lost. Read the session before doing anything else.
  • ownerStopped: false: the instance could not be stopped and nothing was settled; the run may still proceed. Read the session (or getHandle()) before doing anything else.
  • ownerStopped: false, startUnknown: true: the reads that prove a start (the run, its own stream, the session) kept failing, so the executor could not tell a started run from a stalled one. It never terminates an instance on a read it could not make: a failed read is re-read, an unreadable start gets one more grace pass (the same bound again), and if it is still unreadable the call throws this without stopping or settling anything. This case can take up to about three times the bound. Read the session (or getHandle()) before doing anything else. The same values come back when the abandon failed and the terminate that follows it failed too.
  • ownerStopped: true, startUnknown: true: only when the abandon itself failed. A failed abandon leaves no tombstone, so the call terminates the instance (the only fence against a late commit) once the run re-reads find no run, then re-reads the run. A run that committed before the terminate landed and whose start cannot be read is settled like any stall (a stopped instance can never finish it; settle says what that did), and the message says its start could not be confirmed, never that it did not start. The run may have committed a step before it was stopped, so its retry() may need options.message (above).

The message says the same as settle and names the recovery (an outcome that says to read the session first then adds "if its run failed, recover with" the same retry()). It names entryOutcomeFailSafeMs as the bound, not as the time the call waited: with the start grace the wait runs past it. A run that ended on its own before the fail-safe could fail it is no stall: the call returns its handle.

typescript
import { isRunStartStalledError } from '@helix-agents/core';

try {
  const handle = await executor.execute(agent, input, { sessionId });
} catch (err) {
  if (!isRunStartStalledError(err)) throw err;
  if (err.settle === 'failed' || err.settle === 'not_committed') {
    // Never repeat execute(): its input is already committed.
    await executor.getHandle(agent, sessionId); // a reader settles a not_committed run
    // Pass { message } if the run committed a step or its entry carried no user message of its own.
    const handle = await executor.retry(agent, sessionId);
  } else {
    const handle = await executor.getHandle(agent, sessionId); // the run may still be going
  }
}

An explicit entryStartTimeoutMs keeps its meaning (the call returns the handle when it lapses). Temporal and DBOS still return a handle for a committed run at their fail-safe.

A pre-commit entry failure leaves a session that already has a run untouched ​

An entry that fails before its run-start commit (for example a history read that comes up short, state_history_incomplete) never started a run. Affected: runtime-js and the Durable Object execute() and retry() (runtime-js resume() has no such pre-commit read and is unchanged); Workflows execute(), resume() and retry(). The caller still sees the typed failure (runtime-js and the DO: execute()'s first history read returns a handle whose result() is failed, any other read rejects; Workflows: the call returns a handle whose result() is failed with the typed errorDetail), but:

  • runtime-js and the Cloudflare Durable Object used to mark the session failed with that errorDetail, fail its stream and fire onAgentFail, even when the session had finished earlier runs. Now a session that already has a run (live or ended) is untouched: same status, error, errorDetail, stream and chunks, and no agent lifecycle hook (onAgentFail, …) fires. execute()'s first history read still returns a handle whose result() is failed with the code. The Durable Object host's own hooks.onComplete (the DurableObjectAgentBase config hook, not an agent lifecycle hook) still fires with status: 'failed': it is the DO host's only typed signal of the failed call.
  • Cloudflare Workflows used to record the session failed when no other run was live. It now follows the same rule.
  • A session that never had a run is still recorded failed with the detail.
  • Temporal already behaved this way. DBOS still records the session failed whenever no run is live.

What to check: code that read the session's errorDetail (or waited for onAgentFail) to learn about such a failed call on an existing session. Use the call's own error or its handle's result() instead.

Over HTTP: a typed error, never started ​

Because such an entry leaves the session as it was, its stream still ends with the previous turn and /status still shows that turn's output. A caller told the call had started read the previous turn as this one's. So every HTTP surface now answers the failure typed, as exactly { error, code } with the status of core httpStatusForHelixError (400 for a completion-support refusal or validation_error, 503 for a retryable error, 409 for a non-retryable state_* error, else 500):

  • The Durable Object's /start and /retry answer once their entry's outcome is known (the executor call returned, which is after the run-start commit). An entry that did not start answers the typed error, for example 409 { "error": "History for session … is incomplete …", "code": "state_history_incomplete" }, instead of 200 { "status": "started" } / "resumed". A lost run start answers the DO's existing 409{ code: 'ALREADY_RUNNING' } envelope. /resume is unchanged. Every typed body the DO answers, including the completion-support refusal's 400 on /start, /resume, /retry and /submit-tool-result, now has its message scrubbed by core sanitizeErrorMessage: control characters stripped, < and > escaped, known secrets redacted and the message truncated to 256 characters, as @helix-agents/agent-server already did. The status and code are unchanged.
  • DOFrontendExecutor (execute(), resume(), submitToolResult()) throws such a body as a typed DOPeerError (a HelixError) carrying the code, retryable only on 503, instead of a plain Error. A 409 without a framework code, or with state_already_running, is still AgentAlreadyRunningError.
  • handleChatStream (the DO chat route through createCloudflareChatHandler, the Workflows chat route, any host): a SEND whose execute() rejects with a HelixError, or returns a handle with preCommitFailure (runtime-js's first history read, Workflows), answers that typed error (message scrubbed) instead of a 200 stream. A plain Error keeps its data-resume-rejected internal-error answer.
  • A handle whose entry failed before its run-start commit carries the new optional AgentExecutionHandle.preCommitFailure (the detail its result() reports).
  • Never 503 once the input is committed. The DO's /start / /retry (the entry's run exists) and handleChatStream (the session's current run changed during execute()) answer a retryable failure that came after the run-start commit with its non-retryable row (409 for state_*, else 500): the input is already in the log, so re-sending would append it twice. Call retry() instead. A failed run read counts as committed. All three hosts build their answer with core typedErrorResponse(error, { inputCommitted }).

What to check: a client of the DO's /start / /retry that treated any 200 as "running" now gets a 4xx / 5xx with { error, code } for such a call; a chat client receives it as an HTTP error (parseHelixChatError reads the code). Persisted state is unchanged by this: the session stays as it was.

A failed stream's info carries its errorDetail ​

getStreamInfo(streamId) of a failed stream returns metadata.errorDetail (the structured detail failStream stored) next to metadata.error on every Cloudflare stream manager: store-cloudflare's DurableObjectStreamManager (its StreamServer /info now returns it), the in-process DOStreamManager, and the Worker-side DOStreamManagerClient (the DO /snapshot wire gains streamErrorDetail). Memory and Redis already returned it. A stored detail that does not parse as an ErrorDetail is omitted (and logged at warn). The StreamServer's /fail validates the detail once when it arrives and stores and broadcasts only the validated value (extra keys removed), so live subscribers and later readers see the same detail. Custom stream managers are checked by the new contract case G5.fail-detail-on-info in streamManagerContractTests.

Multi-turn outputSchema agents on Workflows ​

A new turn's entry now clears the previous turn's output. Before, an outputSchema agent's turn 2 kept turn 1's output, so:

  • a maxSteps cut-off on turn 2 failed framework_max_steps_exhausted instead of entering forced completion;
  • a stopWhen stop on turn 2 completed silently with turn 1's output;
  • a failed turn 2 kept turn 1's output in the session row.

Each turn now enters forced completion exactly like turn 1. No code change is needed.

Core API additions ​

  • RunStartStalledError / isRunStartStalledError (see core reference).
  • failCommittedRun({ ..., requireSessionCommit: true }): when the session's failed write fails for any reason other than the run's fence (only a RunSupersededError; an exhausted StaleStateError / D1CommitContentionError is contention and counts), nothing else is written and the result is outcome: 'not_committed' with the write's error as cause (see Committed run finish).

Released under the MIT License.