Skip to content

Typed chat resume and submit failures (DI-39) ​

This guide covers the release in which a chat resume (a handleChatStream request that carries client-tool or approval results) answers a typed failure as a typed HTTP error, as a chat send already does, and agent-server's POST /start answers a typed entry failure the same way. It is a major release of @helix-agents/ai-sdk (37.0.0) and of @helix-agents/agent-server (3.0.0): both change a wire answer.

The send half shipped earlier, in CF-D2 MR 1 (!316): see CF Workflows terminal settle → Over HTTP: a typed error, never started. A chat send whose execute() fails with a HelixError (a completion-support refusal included) already answers { error, code } with core typedErrorResponse's status. This release makes the resume path match it.

What changed ​

handleChatStream: a resume or submit that fails typed ​

Every handleChatStream host is affected: agent-server POST /chat, the Express adapter, createCloudflareChatHandler (the CF Durable Object chat route) and any custom host (a CF Workflows chat host is one).

A chat resume whose …BeforeNow
executor.submitToolResult() rejects with a HelixError (e.g. the CF DO's refusal)200 SSE, data-resume-rejected, no code{ error, code }, core status, no resume
executor.resume() rejects with a HelixError (e.g. runtime-js's refusal)200 SSE, data-resume-rejected, no code{ error, code }, core status
executor.resume() returns a handle with preCommitFailure (CF Workflows)200 stream of the handle{ error, code }, core status
submit or resume rejects with a plain Error200 SSE, data-resume-rejectedunchanged
resume fails after the submit already woke the run (the CF DO's auto-resume race)200 stream of the woken rununchanged

The status is core httpStatusForHelixError's: 400 for a completion-support refusal (framework_completion_strategy_unavailable / framework_completion_schema_unsupported) or validation_error, 409 for a non-retryable state_* error (e.g. a pre-commit state_history_incomplete), 503 for a retryable error, else 500. The body is exactly { error, code } with the message scrubbed, plus the X-Session-Id header, before any stream opens.

A typed error replaces any per-intent rejections. When one intent's submit fails typed, the request ends there with that typed error alone: rejections gathered for other intents (a plain Error submit failure, a stale intent) are not added to it. The request failed as a whole; re-read the snapshot and re-send.

Never 503 once an answer is stored. When this request already stored a submitted result (its submit returned accepted), a retryable error is answered with its non-retryable row (409 / 500): re-sending the request would submit the answer again.

A chat resume refused after its answer was stored (runtime-js). A runtime-js submit never refuses: it stores the answer on the pending call (pendingClientToolCalls[id].submittedResult), then the resume is refused with the 400. The stored answer is kept. The next resume on an adapter that can finish the agent (the model fixed, or its capabilities declared) continues from it. The CF Durable Object refuses the submit itself, so there nothing is written and the call stays pending.

agent-server POST /start and POST /resume ​

POST /start and POST /resume answer an entry that failed typed with core typedErrorResponse, as the CF Durable Object's /start does:

  • AgentServer.startAgent and AgentServer.resumeAgent reject a handle whose entry failed before its run-start commit (preCommitFailure) with helixErrorFromDetail(handle.preCommitFailure) instead of answering { sessionId, streamId, runId } (started). On /resume the old answer made the caller attach to the session's stream, which still held the previous turn. The thrown error keeps the handle's detail: resolveErrorDetail(error) returns preCommitFailure verbatim.
  • A thrown HelixError answers { error, code } with the core status instead of 500{ code: 'INTERNAL_ERROR', errorCode }.
  • A retryable error answers 503 only when a re-sent request can succeed:
    • on /start, only while the failed entry left no session behind. Once the session exists (the entry's run committed, or runtime-js / CF Workflows created a fresh session before a pre-commit failure), a re-sent /start is refused by startAgent's pre-check, so the error gets its non-retryable row (409 for a state_* error, else 500);
    • on /resume, only while the session's current run did not move during the call (the resume committed no new run and its message is not in the log).
  • The session or run read behind that decision answers "committed" when it fails, and is logged (warn) through the server's logger, now a public read-only AgentServer.logger property.
  • A completion-support refusal already answered the typed 400; it is unchanged.

What to check ​

  • A client that read data-resume-rejected for these failures now receives a non-2xx response. A stock useChat / useHelixChat ends the request in status: 'error'; read the code with parseHelixChatError(chat.error). A stock client needs no code change.
  • A custom transport that treated every chat POST as a 200 stream must handle a 4xx / 5xx JSON body { error, code } on a resume, as it already does for a send (409 state_already_running, the typed send failures).
  • A caller of agent-server POST /start or POST /resume that matched 500 { code: 'INTERNAL_ERROR', errorCode } for a typed executor failure now gets the core status and { error, code } (the code in code, not errorCode). A /resume that used to answer 200 for a resume that failed before its run-start commit now answers that typed error. Follow a 503 with a retry; a 409 / 500 means re-read the session first.
  • Persisted state is unchanged by this release.

Released under the MIT License.