Migrating to Truthful Terminal Failures
Overview
This release makes a failed run report the real error on runtime-js, the Cloudflare Durable Object and Cloudflare Workflows, and makes CF Workflows retry a transient LLM failure instead of ending the run.
- No stranded sessions (CP-71). A throw inside the runtime-js run loop now ends the run
failed. The CF DO inherits this. - The classified error survives (DI-30, DI-22). A provider failure keeps its
code,categoryandretryableinAgentResult.errorDetail, the persisted session and the failed stream chunk.resolveErrorDetail()readserror.cause. - Typed aborts and refusals (DI-03, DI-15). A hard abort is
framework_cancelled; a workspace refusal isframework_not_supported. - CF Workflows:
- transient LLM errors retry under
llmStepRetry(CFW mid-stream error); handle.result()waits for a real outcome (DI-18);finishWithruns after the step's sub-agents (CP-63).
- transient LLM errors retry under
Most of this needs no code change. The items below change behaviour you might depend on.
1. An unrecognised error has no code
resolveErrorDetail() used to return the error's message only for an error nothing recognised, except on CF Workflows, which recorded framework_internal_error with retryable: false. Now every runtime records { message }, with no code and no retryable, when nothing on the error or its cause chain is classified.
Action: code that branched on errorDetail.code === 'framework_internal_error' to mean "unknown" should treat a missing code the same way.
To classify an error yourself without the fallback, use the new classifyKnownError(error). It returns a HelixError for a recognised error and undefined otherwise. classifyError() is unchanged.
2. A wrapper keeps the wrapped error's classification
When the classification comes from a cause, the resolved detail keeps the wrapper's message (the failure that surfaced) and takes the wrapped error's code, category and retryable. The wrapped error's message becomes errorDetail.cause.
const inner = new HelixError({
message: 'rate limited',
code: 'provider_rate_limited',
retryable: true,
});
resolveErrorDetail(new Error('D1 write failed', { cause: inner }));
// { message: 'D1 write failed', cause: 'rate limited',
// code: 'provider_rate_limited', category: 'provider', retryable: true }3. Aborts and refusals are typed
- Hard abort:
errorDetailis{ code: 'framework_cancelled', category: 'framework', retryable: false }on runtime-js, the CF DO and CF Workflows. A soft interrupt still endsinterruptedwith no error. - Workspace refusal:
assertRuntimeSupportsWorkspaces()throws aHelixErrorwith codeframework_not_supportedandretryable: false(Temporal, DBOS and CF Workflows). - CF Workflows refuses early.
execute(),resume()andretry()reject a workspace-declaring agent before anything is written. Before, the workflow started and returned a latefailedresult.
Action: match on error.code (or HelixError.isInstance(error)) instead of the message.
4. CF Workflows: transient LLM errors retry
A retryable or unclassified LLM error, thrown or returned by the adapter, is rethrown so step.do retries the LLM step under your llmStepRetry (default 3 retries). The retried attempt's partial output is superseded by its attempt id, so a client that re-anchors on data-attempt-superseded (the useHelixChat resync) shows only the retried text.
- When retries run out, the run fails with the provider's own message and
errorDetail. It used to fail with "Agent stopped without completing (no structured output)". - A classified non-retryable error (for example
provider_auth_error) fails at once. - The
errorstream chunk is written only for a terminal error: a client stops reading at one, so a per-attempt error chunk would hide the retry.
Action: budget for up to llmStepRetry.limit + 1 model calls per step on a transient provider failure. Tune llmStepRetry in WorkflowOptions if that is too many.
5. CF Workflows: handle.result() waits
result() no longer gives up after five minutes, and waits while an instance is operator-paused. A failed status read is not cached. Give result() your own deadline if you need one:
const result = await Promise.race([
handle.result(),
new Promise<never>((_, reject) => setTimeout(() => reject(new Error('timed out')), 600_000)),
]);6. CF Workflows: finishWith runs after sub-agents
In a step that calls both a sub-agent and a finishWith tool, the finishWith tool now runs after the sub-agent, as on runtime-js and the CF DO, so it sees the sub-agent's effects.
Deploy note: the finishWith step names changed (…-phase2-finishwith-…). A workflow instance in flight across the deploy that had already run its finishWith may run it again. Drain in-flight instances first if your finishWith tool is not idempotent.
7. CF Workflows: WorkflowInstance has no abort()
The WorkflowInstance binding type drops abort(), which real Workflows instances never had. The handles terminate the owner instance. Remove abort from any hand-written binding mocks.
Reference
- Design spec, as built:
docs/superpowers/specs/2026-10-01-cfw-terminal-truth-mr1b-design.md. - Error Handling
- Cloudflare Workflows runtime