Skip to content

Cloudflare store and reader truth ​

This guide covers the release in which the Cloudflare D1 and memory migrations run on one claim-first runner, a Durable Object stream manager reports an unknown stream as absent, and core gains the closed-owner reconciliation planner and executor, which the CF Workflows readers now use. It is a minor release of @helix-agents/core, @helix-agents/store-cloudflare and @helix-agents/memory-cloudflare (0.x packages: minor marks an observable behaviour change) and a major release of @helix-agents/runtime-cloudflare (already ≥ 1.0, so a breaking behaviour change is a major bump).

No schema change, no deploy drain. No migration version was added and no workflow param or instance id changed. Call runMigration() / initialize() at boot as before.

Migrations are safe to run concurrently ​

runMigration() (@helix-agents/store-cloudflare) and CloudflareMemoryStore.initialize() (@helix-agents/memory-cloudflare) now run on core's SQL migration runner. Each migration is one batch whose first statement records its version row.

  • Concurrent boots no longer fail. Two isolates that both see an older version used to both batch the next migration, and the second failed with duplicate column name. Now one claims the version and the other continues from what it recorded. A data statement in a migration never runs twice.
  • Version reads throw on failure. getAgentsMigrationVersion() and getMemoryMigrationVersion() return 0 only when the version table does not exist yet. A failed read (a D1 outage, a permission error) now throws instead of being reported as an unmigrated database. What to check: code that called either function and treated 0 as "run the migrations" now sees the rejection; catch it where a retry is the right answer.
  • The runner's ADD COLUMN / DROP COLUMN guards compare column names case-insensitively, as SQLite does.
  • dropAllTables() now also drops __agents_abandoned_run_starts.
  • The D1 run-start commit pins the current run's lease_until and owner_token together with its status, so a JS run whose lease was renewed between the store's two reads is no longer superseded on the stale value.

A stream that does not exist reads as absent on Cloudflare ​

DurableObjectStreamManager.getStreamInfo(streamId) now returns null when the stream Durable Object reports the stream does not exist, like the memory and Redis managers (new contract case G5.info-null-on-not-found in @helix-agents/core/testing). A stream a run has claimed counts as existing even before its first chunk. The in-process DOStreamManager (Durable Object runtime) still never answers null; it opts out of the case, tracked as RM-71.

What you can observe:

  • A route built on createResumableStreamHandler answers 404 { error: 'Stream not found' } for an unknown stream id. It used to answer a 500 after the handler's reader retries.
  • An admin terminate of a session whose stream was never written no longer writes a run_terminated chunk to it.
  • A passive attach (a second tab, a refresh) during a fresh CF Workflows session's run-start window no longer gets 409 stream_already_attached: the stream exists before the claim, so the attach streams the run.
  • A suspended run whose own K5 landed is finished. committedTerminalStatus and finishCommittedRun now accept any live run (a suspended_* run included), not only a running one. The CF Durable Object wake path and DBOS settleDeadOwnerRun therefore finish such a run (ending or failing its stream) where they used to leave it live or take the fallback status.
  • A failed read or init throws. DurableObjectStreamManager.getStreamInfo and initStream now throw StreamConnectionError when the stream Durable Object answers /info or /init with a non-2xx status (its onRequest catch's 500, or a proxy's 502). They used to return null and false, which read a failed call as "no such stream" and "already existed". null from getStreamInfo now means ONLY "the stream does not exist", and false from initStream means only "it already existed". The error's reason is the DO's error text, else the status; a non-2xx answer whose body is not JSON (plain text, an HTML error page) also becomes a StreamConnectionError, with the status and an excerpt of the body (as is a JSON body without an error field). The methods that throw are initStream, the writer's write, getStreamInfo, endStream, failStream, pauseStream, claimStream, resumeStream, reactivateStream, cleanupToStep and resetStream. createReader and createResumableReader still return null and getAllChunks, getChunksFromStep and getChunksFromSequence still return [] on a non-2xx answer (FU-DO-STREAM-READERS-READ-FAILURE-AS-ABSENT). Core's resolveCommitSequence maps the throw to the typed, retryable state_checkpoint_sequence_unavailable. ResumableStreamHandler.handleRequest does not catch it: it rejects with the StreamConnectionError, and the host framework decides the status (it used to answer a false 404 "Stream not found").

What to check: code that called getStreamInfo() on a DurableObjectStreamManager and assumed a non-null result for an unknown id; handle null as "no such stream". Code that treated a null or false as a soft failure of the call must catch StreamConnectionError instead.

New core APIs ​

All are exported from @helix-agents/core and documented in the core reference.

CF Workflows readers reconcile a closed instance ​

@helix-agents/runtime-cloudflare is a major release: the CF Workflows executor's readers (each handle's result(), getHandle(), canResume() and abort()) and the parent's consult of a child now reconcile a run whose instance closed without its terminal commit through core's closed-owner planner and executor. A reader acts only once it has proven the instance closed; a probe that fails reconciles nothing. See Instances closed outside the product.

What you can observe:

  • A committed suspension is reported as suspended, not failed. A run whose instance was terminated, errored or disappeared after its suspension commit (the session paused, awaiting client calls or children) used to make result() answer { status: 'failed', error: 'Unknown error' } (or the engine's message), and nothing settled the run. result() now answers { status: 'suspended_client_tool' | 'suspended_awaiting_children' | 'suspended_step_partial', runId, suspended }, the run reads suspended_* (written when it still read running), its own active stream is paused, and canResume() is true: submit the results and resume() as usual. A run killed before its suspension commit is settled failed with how the instance closed, as a running run was before. abort() is unchanged: it fails such a run with its own cancellation.
  • The reconnect handle's suspension payload. A handle from getHandle() after the instance closed now reports a suspension from the session: runId is the run's own id (it was the session id; its completed, interrupted and failed results now carry the run's own id too, the session id only for a handle built before the session had any run), suspended.stepId is present whenever the session records one (the workflow's output omitted it for some suspensions), and suspended.toolCallIds lists every pending client call, including ones already answered but not yet resumed. A reconnect handle built before the session had any run reconciles and reports nothing for a later run.
  • A stream left active or paused behind an ended run is ended or failed by the next reader, from the run's committed outcome. It used to stay live until the session's next entry. A host that only serves the chat route still waits for the next send (FU-CFW-CHAT-READER-SETTLE-CLOSED-OWNER).
  • A parent no longer records a resumable child as dead. When a parent's wait for a child's completion event ends with none and the child's instance is closed, the parent reconciles the child as a reader would and takes the path of a delivered event: a child whose suspension committed is a suspended child (the parent cascades suspended_awaiting_children; it used to record ChildRunDiedError), a child whose terminal commit landed answers its committed outcome, and a child with nothing committed is settled failed with how its instance closed and answers that settled failure (errored, AgentError, the persisted errorDetail's message) — the same answer whoever settled it (the consult, a later consult or an executor reader). ChildRunDiedError is gone. A failure write that did not land re-waits (the next consult retries it); a permanent error (no workflow binding to probe the child's instance, a broken invariant, or a corrupt stored row: InvalidStoredStateError, CorruptMessageRowError, UnknownCommitKindError) answers a typed errored (its error.name the error's class) instead of a timeout. A corrupt row used to be read as "nothing known yet", so the parent waited out its whole subAgentEventTimeout. The consult re-reads the child's run once its instance is proven closed, so a child an executor reader settled meanwhile is answered at once. A child whose outcome is durable is answered while its stream Durable Object is down. A companion wait still records a suspended companion as a failure (FU-CFW-COMPANION-WAIT-SUSPENDED-CHILD-AS-FAILED).
  • A handle kept across deleteSession no longer reports a made-up entry failure. deleteSession removes the session's runs, so a closed instance with no run of its own is now read against the session. With the session deleted, result() reports the instance's recorded outcome (a completed output stays completed; an aborted run's cancelled output stays a framework_cancelled failure), and an instance that recorded none (errored, terminated, complete without a result) reports a typed, non-retryable state_session_not_found failure. A session deleted during the entry wait gets the same answer: no preCommitFailure for a recorded non-failure outcome, otherwise the recorded failure (possibly a committed run's, since whether the entry committed can no longer be read) or state_session_not_found; a failed session read there keeps the wait going (cfw.entry.session_read_failed).
  • A stream that cannot be read or written never hides a run's outcome. Readers plan a run whose instance closed without reading its stream status when that read fails (logged cfw.settle.stream_read_failed), and core's reconciliation makes the stream write after the run-level write best-effort (reconcile.stream_unavailable; a later reader repairs the stream): a run whose terminal commit landed completed is reported completed while its stream Durable Object is down (it used to report failed), and the workflow's own error path no longer fails such a run or notifies its parent errored. A transient store error in a reader's settle is logged and the reader answers from what is persisted; a permanent one (a non-retryable HelixError, or a corrupt stored row: InvalidStoredStateError, CorruptMessageRowError, UnknownCommitKindError, also when a D1 D1StateError wraps one) is logged at error level (code: the HelixError code, else the error's class name) and rethrown. A bare D1StateError is a failed D1 call and stays transient.
  • The entry fail-safe retries its abandon. abandonRunStart is made up to 3 times (25 ms × attempt backoff) before the call falls back to terminating the instance, so a transient store error no longer reaches the unfenced terminate. A terminate() that throws because the instance had just closed on its own now counts as stopped: the call re-checks the run's start, returns the handle of a run that started, and settles a stall as usual (ownerStopped: true). It used to report ownerStopped: false and settle nothing. A stalled run whose owner was stopped is settled through the closed-owner reconciliation, so a terminal commit the instance landed before it stopped is finished (the call returns its handle) instead of being overwritten by the stall's failure, and a committed suspension is reported rather than failed. A non-retryable abandon error is not retried.
  • The pre-commit failure write is retried, and the caller keeps the entry's own failure. The instance's error-run-exists and error-pre-commit steps retry a failed read or save (3 retries each, about 47 s at most per step and 94 s for both; the write used to swallow a failed read). When the entry's own steps fail fast this ends inside the 120 s entryOutcomeFailSafeMs; an entry that hangs in its own steps (check-existing, run-start) can outlast it, and the caller then gets the typed abandon / outcome-unknown error. When either step's retries run out, the instance still returns the entry's classified failure as preCommitFailure (an exhausted write is logged at error level; an exhausted run check skips the write and is logged at warn). An instance that closed with no run and no failure output now yields a preCommitFailure derived from the engine status (errored → the engine's message, no code; terminated → framework_cancelled; complete → a message, no code; each says the entry committed no run), and the handle's result() reports the same detail; it used to carry none, so a chat send attached to the session's previous stream.
  • A never-run session's pre-commit failure writes only the session (user ruling R43). A brand-new session whose first entry fails before its run-start commit is recorded failed through core failSessionWithoutRun, under the idle fence, with no stream write (it used to fail or restore the session's stream, which could fail the next turn's live stream) and no hook. Temporal still writes the stream on this path (FU-TEMPORAL-NEVER-RUN-PRE-COMMIT-STREAM-WRITE).
  • terminateRun() resolves once the terminate landed. Its run_terminated stream write is now best-effort and logged (cfw.admin_terminate.stream_end_failed); a failed stream read or write used to reject the call after the instance was already terminated.
  • Typed stream errors at the entry. A failed initStream while a run's stream is made live is now the typed, retryable transport_error (a failed head read is the retryable state_checkpoint_sequence_unavailable; each with the original error as cause); and when a resume or reactivate fails and its re-read fails too, the original error is kept.

What to check: code that treated a CF Workflows result() of failed as final for a session that was waiting on a client tool or a child should handle the suspended_* statuses (an exhaustive switch over AgentResult.status already does). Code that read runId from a reconnect handle's suspended result as the session id should use sessionId. Code that matched ChildRunDiedError in a parent's tool result should match the child's settled failure (an AgentError with the instance-closure message) instead; a child killed after it suspended now makes the parent suspend.

Released under the MIT License.