CF Workflows terminal settle and pre-commit entry failures
This guide covers the release in which Cloudflare Workflows records a truthful outcome for every run whose instance stops, and in which an entry that fails before its run-start commit no longer overwrites a session that already has a run. It is a major release of @helix-agents/runtime-cloudflare and @helix-agents/runtime-js, and a minor release of @helix-agents/core and @helix-agents/store-cloudflare (0.x packages: minor marks an observable behaviour change).
No deploy drain and no migration. No workflow param, instance id or schema changed. Deploy order only affects one optional field: an old stream worker (StreamServer) or Durable Object returns no errorDetail in a failed stream's info until it is redeployed (see Stream info).
abort() on Workflows writes the terminal commit
handle.abort(reason) still terminates the owner instance and ends the run failed with a typed framework_cancelled. What changed:
- The failure is now the terminal
statecommit (sessionfailed→ runfailed→ stream failed), like every other terminal exit. Before, it was a plain state save with no checkpoint.retry()never targets that commit. handle.result()andgetHandle(...).result()report the persisted detail ({ message: reason, code: 'framework_cancelled', category: 'framework', retryable: false }). Before, they reported{ message }only (DI-33). The aborted session'soutputis cleared.result()reports only its own run. Every handle (the reconnect handle fromgetHandle()included) reports the outcome of the run it was built for, never a newer run's: it can returnsuperseded, orcompletedwith nooutputwhen a newer run has committed since (a failed run then reports its own record's error as{ message }). The reconnect handle'sresult()andcanResume()read the session's state at call time, not atgetHandle().abort()can now reject. If the terminal commit cannot be written (a transient store error), nothing is written, the run stays live, andabort()rejects with aHelixErrortransport_error,retryable: true, with the store error ascause. Before, it resolved and left the sessionactivebehind afailedrun.
try {
await handle.abort('User cancelled');
} catch (err) {
if (HelixError.isInstance(err) && err.code === 'transport_error' && err.retryable) {
await handle.abort('User cancelled'); // nothing was written; try again
} else {
throw err;
}
}On Miniflare 4 (older @cloudflare/vitest-pool-workers), terminate() is not implemented: the commit still lands, but result() waits until the instance ends on its own. The real engine and Miniflare 5 stop it.
Instances closed outside the product are settled on read
A run whose instance was terminated from the Cloudflare dashboard or wrangler, errored, or disappeared before its terminal commit used to stay active / running with a live-looking stream, and result() reported 'Unknown error' (DI-34). Now the next executor reader settles it: each handle's result(), getHandle(), canResume() and abort(). The session, run and stream end failed with a detail that says how the instance closed (abort() settles with its own cancellation instead, framework_cancelled carrying the caller's reason):
| Instance | errorDetail |
|---|---|
terminated | { message, code: 'framework_cancelled', category: 'framework', retryable: false } |
errored | the engine's error message, no code |
complete without a terminal commit / gone | { message } naming the instance; no code |
complete, its output its own failure | that output's validated errorDetail (its error path's cleanup failed) |
No lifecycle hook (onAgentFail) fires for a settled run. A run whose terminal commit did land is finished with that outcome instead. A live, operator-paused or unreadable instance is never settled.
What to check: code that waited for canResume() to stay true, or a session to stay active, after an external terminate. Code that alerts on executor log events: cfw.handle.finish_committed_run_failed and cfw.reconnect.finish_committed_run_failed are now cfw.handle.settle_failed and cfw.reconnect.settle_failed.
A host that only serves the chat route (handleChatStream) calls none of these readers, so there the session stays active until its next send, whose entry supersedes the stranded run.
The entry fail-safe throws RunStartStalledError
execute() / resume() / retry() on Workflows wait for their entry for at most entryOutcomeFailSafeMs (120 000 ms). A run that committed but has not started by then (no run_started in its own stream, or no entry state) gets the same bound again until its instance closes or is paused, measured from when the call first saw the commit (a healthy instance under load can commit just before the bound and start just after it; a transient owner read or abandon failure does not end it). A real stall can therefore take up to about twice the bound to surface. A run that still has not started used to get a handle whose stream did not exist. Now it never gets a handle: the executor terminates the instance and throws core's RunStartStalledError (transport_error).
It is never retryable (retryable: false for every outcome): the run-start commit already appended the call's input, so repeating the same call (execute() again) would append it a second time. The recovery is executor.retry(agent, sessionId), which re-runs the failed run from its own entry. Every outcome's message names it the same way: executor.retry(agent, sessionId) (pass options.message if the run committed a step or its entry carried no user message of its own). Without a message such a run's retry() is refused validation_error: a run whose start is unknown may have committed a step before it was stopped, and a resume() in continue / from_checkpoint mode with no message commits only heal rows, so its run has no trigger to re-run. What settle says:
ownerStopped: true,settle: 'failed': the instance stopped, and the run was settledfailedwith the error's own detail. Callexecutor.retry(agent, sessionId).ownerStopped: true,settle: 'not_committed': the instance stopped, but the run'sfailedwrite did not land, so nothing was written. A reader settles the run first (getHandle(),canResume()or a handle'sresult()), then callexecutor.retry(agent, sessionId).ownerStopped: true,settle: 'not_found': the session or its run could not be read, so nothing was settled. Read the session first.ownerStopped: true,settle: 'superseded': a newer run had already taken the session over, so this call's input belongs to a run that lost. Read the session before doing anything else.ownerStopped: false: the instance could not be stopped and nothing was settled; the run may still proceed. Read the session (orgetHandle()) before doing anything else.ownerStopped: false,startUnknown: true: the reads that prove a start (the run, its own stream, the session) kept failing, so the executor could not tell a started run from a stalled one. It never terminates an instance on a read it could not make: a failed read is re-read, an unreadable start gets one more grace pass (the same bound again), and if it is still unreadable the call throws this without stopping or settling anything. This case can take up to about three times the bound. Read the session (orgetHandle()) before doing anything else. The same values come back when the abandon failed and the terminate that follows it failed too.ownerStopped: true,startUnknown: true: only when the abandon itself failed. A failed abandon leaves no tombstone, so the call terminates the instance (the only fence against a late commit) once the run re-reads find no run, then re-reads the run. A run that committed before the terminate landed and whose start cannot be read is settled like any stall (a stopped instance can never finish it;settlesays what that did), and the message says its start could not be confirmed, never that it did not start. The run may have committed a step before it was stopped, so itsretry()may needoptions.message(above).
The message says the same as settle and names the recovery (an outcome that says to read the session first then adds "if its run failed, recover with" the same retry()). It names entryOutcomeFailSafeMs as the bound, not as the time the call waited: with the start grace the wait runs past it. A run that ended on its own before the fail-safe could fail it is no stall: the call returns its handle.
import { isRunStartStalledError } from '@helix-agents/core';
try {
const handle = await executor.execute(agent, input, { sessionId });
} catch (err) {
if (!isRunStartStalledError(err)) throw err;
if (err.settle === 'failed' || err.settle === 'not_committed') {
// Never repeat execute(): its input is already committed.
await executor.getHandle(agent, sessionId); // a reader settles a not_committed run
// Pass { message } if the run committed a step or its entry carried no user message of its own.
const handle = await executor.retry(agent, sessionId);
} else {
const handle = await executor.getHandle(agent, sessionId); // the run may still be going
}
}An explicit entryStartTimeoutMs keeps its meaning (the call returns the handle when it lapses). Temporal and DBOS still return a handle for a committed run at their fail-safe.
A pre-commit entry failure leaves a session that already has a run untouched
An entry that fails before its run-start commit (for example a history read that comes up short, state_history_incomplete) never started a run. Affected: runtime-js and the Durable Object execute() and retry() (runtime-js resume() has no such pre-commit read and is unchanged); Workflows execute(), resume() and retry(). The caller still sees the typed failure (runtime-js and the DO: execute()'s first history read returns a handle whose result() is failed, any other read rejects; Workflows: the call returns a handle whose result() is failed with the typed errorDetail), but:
- runtime-js and the Cloudflare Durable Object used to mark the session
failedwith thaterrorDetail, fail its stream and fireonAgentFail, even when the session had finished earlier runs. Now a session that already has a run (live or ended) is untouched: same status,error,errorDetail, stream and chunks, and no agent lifecycle hook (onAgentFail, …) fires.execute()'s first history read still returns a handle whoseresult()isfailedwith the code. The Durable Object host's ownhooks.onComplete(theDurableObjectAgentBaseconfig hook, not an agent lifecycle hook) still fires withstatus: 'failed': it is the DO host's only typed signal of the failed call. - Cloudflare Workflows used to record the session
failedwhen no other run was live. It now follows the same rule. - A session that never had a run is still recorded
failedwith the detail. - Temporal already behaved this way. DBOS still records the session
failedwhenever no run is live.
What to check: code that read the session's errorDetail (or waited for onAgentFail) to learn about such a failed call on an existing session. Use the call's own error or its handle's result() instead.
Over HTTP: a typed error, never started
Because such an entry leaves the session as it was, its stream still ends with the previous turn and /status still shows that turn's output. A caller told the call had started read the previous turn as this one's. So every HTTP surface now answers the failure typed, as exactly { error, code } with the status of core httpStatusForHelixError (400 for a completion-support refusal or validation_error, 503 for a retryable error, 409 for a non-retryable state_* error, else 500):
- The Durable Object's
/startand/retryanswer once their entry's outcome is known (the executor call returned, which is after the run-start commit). An entry that did not start answers the typed error, for example409 { "error": "History for session … is incomplete …", "code": "state_history_incomplete" }, instead of200 { "status": "started" }/"resumed". A lost run start answers the DO's existing409{ code: 'ALREADY_RUNNING' }envelope./resumeis unchanged. Every typed body the DO answers, including the completion-support refusal's400on/start,/resume,/retryand/submit-tool-result, now has its message scrubbed by coresanitizeErrorMessage: control characters stripped,<and>escaped, known secrets redacted and the message truncated to 256 characters, as@helix-agents/agent-serveralready did. The status andcodeare unchanged. DOFrontendExecutor(execute(),resume(),submitToolResult()) throws such a body as a typedDOPeerError(aHelixError) carrying the code, retryable only on503, instead of a plainError. A409without a framework code, or withstate_already_running, is stillAgentAlreadyRunningError.handleChatStream(the DO chat route throughcreateCloudflareChatHandler, the Workflows chat route, any host): a SEND whoseexecute()rejects with aHelixError, or returns a handle withpreCommitFailure(runtime-js's first history read, Workflows), answers that typed error (message scrubbed) instead of a200stream. A plainErrorkeeps itsdata-resume-rejectedinternal-erroranswer.- A handle whose entry failed before its run-start commit carries the new optional
AgentExecutionHandle.preCommitFailure(the detail itsresult()reports). - Never
503once the input is committed. The DO's/start//retry(the entry's run exists) andhandleChatStream(the session's current run changed duringexecute()) answer aretryablefailure that came after the run-start commit with its non-retryable row (409forstate_*, else500): the input is already in the log, so re-sending would append it twice. Callretry()instead. A failed run read counts as committed. All three hosts build their answer with coretypedErrorResponse(error, { inputCommitted }).
What to check: a client of the DO's /start / /retry that treated any 200 as "running" now gets a 4xx / 5xx with { error, code } for such a call; a chat client receives it as an HTTP error (parseHelixChatError reads the code). Persisted state is unchanged by this: the session stays as it was.
A failed stream's info carries its errorDetail
getStreamInfo(streamId) of a failed stream returns metadata.errorDetail (the structured detail failStream stored) next to metadata.error on every Cloudflare stream manager: store-cloudflare's DurableObjectStreamManager (its StreamServer /info now returns it), the in-process DOStreamManager, and the Worker-side DOStreamManagerClient (the DO /snapshot wire gains streamErrorDetail). Memory and Redis already returned it. A stored detail that does not parse as an ErrorDetail is omitted (and logged at warn). The StreamServer's /fail validates the detail once when it arrives and stores and broadcasts only the validated value (extra keys removed), so live subscribers and later readers see the same detail. Custom stream managers are checked by the new contract case G5.fail-detail-on-info in streamManagerContractTests.
Multi-turn outputSchema agents on Workflows
A new turn's entry now clears the previous turn's output. Before, an outputSchema agent's turn 2 kept turn 1's output, so:
- a
maxStepscut-off on turn 2 failedframework_max_steps_exhaustedinstead of entering forced completion; - a
stopWhenstop on turn 2 completed silently with turn 1's output; - a failed turn 2 kept turn 1's output in the session row.
Each turn now enters forced completion exactly like turn 1. No code change is needed.
Core API additions
RunStartStalledError/isRunStartStalledError(see core reference).failCommittedRun({ ..., requireSessionCommit: true }): when the session'sfailedwrite fails for any reason other than the run's fence (only aRunSupersededError; an exhaustedStaleStateError/D1CommitContentionErroris contention and counts), nothing else is written and the result isoutcome: 'not_committed'with the write's error ascause(see Committed run finish).