Skip to content

Migrating to Truthful Terminal Failures ​

Overview ​

This release makes a failed run report the real error on runtime-js, the Cloudflare Durable Object and Cloudflare Workflows, and makes CF Workflows retry a transient LLM failure instead of ending the run.

  • No stranded sessions (CP-71). A throw inside the runtime-js run loop now ends the run failed. The CF DO inherits this.
  • The classified error survives (DI-30, DI-22). A provider failure keeps its code, category and retryable in AgentResult.errorDetail, the persisted session and the failed stream chunk. resolveErrorDetail() reads error.cause.
  • Typed aborts and refusals (DI-03, DI-15). A hard abort is framework_cancelled; a workspace refusal is framework_not_supported.
  • CF Workflows:
    • transient LLM errors retry under llmStepRetry (CFW mid-stream error);
    • handle.result() waits for a real outcome (DI-18);
    • finishWith runs after the step's sub-agents (CP-63).

Most of this needs no code change. The items below change behaviour you might depend on.

1. An unrecognised error has no code ​

resolveErrorDetail() used to return the error's message only for an error nothing recognised, except on CF Workflows, which recorded framework_internal_error with retryable: false. Now every runtime records { message }, with no code and no retryable, when nothing on the error or its cause chain is classified.

Action: code that branched on errorDetail.code === 'framework_internal_error' to mean "unknown" should treat a missing code the same way.

To classify an error yourself without the fallback, use the new classifyKnownError(error). It returns a HelixError for a recognised error and undefined otherwise. classifyError() is unchanged.

2. A wrapper keeps the wrapped error's classification ​

When the classification comes from a cause, the resolved detail keeps the wrapper's message (the failure that surfaced) and takes the wrapped error's code, category and retryable. The wrapped error's message becomes errorDetail.cause.

typescript
const inner = new HelixError({
  message: 'rate limited',
  code: 'provider_rate_limited',
  retryable: true,
});
resolveErrorDetail(new Error('D1 write failed', { cause: inner }));
// { message: 'D1 write failed', cause: 'rate limited',
//   code: 'provider_rate_limited', category: 'provider', retryable: true }

3. Aborts and refusals are typed ​

  • Hard abort: errorDetail is { code: 'framework_cancelled', category: 'framework', retryable: false } on runtime-js, the CF DO and CF Workflows. A soft interrupt still ends interrupted with no error.
  • Workspace refusal: assertRuntimeSupportsWorkspaces() throws a HelixError with code framework_not_supported and retryable: false (Temporal, DBOS and CF Workflows).
  • CF Workflows refuses early. execute(), resume() and retry() reject a workspace-declaring agent before anything is written. Before, the workflow started and returned a late failed result.

Action: match on error.code (or HelixError.isInstance(error)) instead of the message.

4. CF Workflows: transient LLM errors retry ​

A retryable or unclassified LLM error, thrown or returned by the adapter, is rethrown so step.do retries the LLM step under your llmStepRetry (default 3 retries). The retried attempt's partial output is superseded by its attempt id, so a client that re-anchors on data-attempt-superseded (the useHelixChat resync) shows only the retried text.

  • When retries run out, the run fails with the provider's own message and errorDetail. It used to fail with "Agent stopped without completing (no structured output)".
  • A classified non-retryable error (for example provider_auth_error) fails at once.
  • The error stream chunk is written only for a terminal error: a client stops reading at one, so a per-attempt error chunk would hide the retry.

Action: budget for up to llmStepRetry.limit + 1 model calls per step on a transient provider failure. Tune llmStepRetry in WorkflowOptions if that is too many.

5. CF Workflows: handle.result() waits ​

result() no longer gives up after five minutes, and waits while an instance is operator-paused. A failed status read is not cached. Give result() your own deadline if you need one:

typescript
const result = await Promise.race([
  handle.result(),
  new Promise<never>((_, reject) => setTimeout(() => reject(new Error('timed out')), 600_000)),
]);

6. CF Workflows: finishWith runs after sub-agents ​

In a step that calls both a sub-agent and a finishWith tool, the finishWith tool now runs after the sub-agent, as on runtime-js and the CF DO, so it sees the sub-agent's effects.

Deploy note: the finishWith step names changed (…-phase2-finishwith-…). A workflow instance in flight across the deploy that had already run its finishWith may run it again. Drain in-flight instances first if your finishWith tool is not idempotent.

7. CF Workflows: WorkflowInstance has no abort() ​

The WorkflowInstance binding type drops abort(), which real Workflows instances never had. The handles terminate the owner instance. Remove abort from any hand-written binding mocks.

Reference ​

Released under the MIT License.