Cloudflare store and reader truth
This guide covers the release in which the Cloudflare D1 and memory migrations run on one claim-first runner, a Durable Object stream manager reports an unknown stream as absent, and core gains the closed-owner reconciliation planner and executor, which the CF Workflows readers now use. It is a minor release of @helix-agents/core, @helix-agents/store-cloudflare and @helix-agents/memory-cloudflare (0.x packages: minor marks an observable behaviour change) and a major release of @helix-agents/runtime-cloudflare (already ≥ 1.0, so a breaking behaviour change is a major bump).
No schema change, no deploy drain. No migration version was added and no workflow param or instance id changed. Call runMigration() / initialize() at boot as before.
Migrations are safe to run concurrently
runMigration() (@helix-agents/store-cloudflare) and CloudflareMemoryStore.initialize() (@helix-agents/memory-cloudflare) now run on core's SQL migration runner. Each migration is one batch whose first statement records its version row.
- Concurrent boots no longer fail. Two isolates that both see an older version used to both batch the next migration, and the second failed with
duplicate column name. Now one claims the version and the other continues from what it recorded. A data statement in a migration never runs twice. - Version reads throw on failure.
getAgentsMigrationVersion()andgetMemoryMigrationVersion()return0only when the version table does not exist yet. A failed read (a D1 outage, a permission error) now throws instead of being reported as an unmigrated database. What to check: code that called either function and treated0as "run the migrations" now sees the rejection; catch it where a retry is the right answer. - The runner's
ADD COLUMN/DROP COLUMNguards compare column names case-insensitively, as SQLite does. dropAllTables()now also drops__agents_abandoned_run_starts.- The D1 run-start commit pins the current run's
lease_untilandowner_tokentogether with its status, so a JS run whose lease was renewed between the store's two reads is no longer superseded on the stale value.
A stream that does not exist reads as absent on Cloudflare
DurableObjectStreamManager.getStreamInfo(streamId) now returns null when the stream Durable Object reports the stream does not exist, like the memory and Redis managers (new contract case G5.info-null-on-not-found in @helix-agents/core/testing). A stream a run has claimed counts as existing even before its first chunk. The in-process DOStreamManager (Durable Object runtime) still never answers null; it opts out of the case, tracked as RM-71.
What you can observe:
- A route built on
createResumableStreamHandleranswers 404{ error: 'Stream not found' }for an unknown stream id. It used to answer a 500 after the handler's reader retries. - An admin terminate of a session whose stream was never written no longer writes a
run_terminatedchunk to it. - A passive attach (a second tab, a refresh) during a fresh CF Workflows session's run-start window no longer gets 409
stream_already_attached: the stream exists before the claim, so the attach streams the run. - A suspended run whose own K5 landed is finished.
committedTerminalStatusandfinishCommittedRunnow accept any live run (asuspended_*run included), not only arunningone. The CF Durable Object wake path and DBOSsettleDeadOwnerRuntherefore finish such a run (ending or failing its stream) where they used to leave it live or take the fallback status. - A failed read or init throws.
DurableObjectStreamManager.getStreamInfoandinitStreamnow throwStreamConnectionErrorwhen the stream Durable Object answers/infoor/initwith a non-2xx status (itsonRequestcatch's 500, or a proxy's 502). They used to returnnullandfalse, which read a failed call as "no such stream" and "already existed".nullfromgetStreamInfonow means ONLY "the stream does not exist", andfalsefrominitStreammeans only "it already existed". The error'sreasonis the DO'serrortext, else the status; a non-2xx answer whose body is not JSON (plain text, an HTML error page) also becomes aStreamConnectionError, with the status and an excerpt of the body (as is a JSON body without anerrorfield). The methods that throw areinitStream, the writer'swrite,getStreamInfo,endStream,failStream,pauseStream,claimStream,resumeStream,reactivateStream,cleanupToStepandresetStream.createReaderandcreateResumableReaderstill returnnullandgetAllChunks,getChunksFromStepandgetChunksFromSequencestill return[]on a non-2xx answer (FU-DO-STREAM-READERS-READ-FAILURE-AS-ABSENT). Core'sresolveCommitSequencemaps the throw to the typed, retryablestate_checkpoint_sequence_unavailable.ResumableStreamHandler.handleRequestdoes not catch it: it rejects with theStreamConnectionError, and the host framework decides the status (it used to answer a false 404 "Stream not found").
What to check: code that called getStreamInfo() on a DurableObjectStreamManager and assumed a non-null result for an unknown id; handle null as "no such stream". Code that treated a null or false as a soft failure of the call must catch StreamConnectionError instead.
New core APIs
All are exported from @helix-agents/core and documented in the core reference.
runSqlMigrations,readSqlMigrationVersion,splitSqlStatements,classifySqlStatementand their types: the shared claim-first migration runner.planClosedOwnerReconciliation: the pure planner that decides what to do with a run whose owner is provably closed.reconcileClosedOwnerRun: executes a plan, fenced on the run.isSuspensionCommitted,suspendedStatusOf,suspendedResultOf: the one definition of a committed suspension, the suspended status the session describes, and thesuspendedpayload.committedTerminalStatusOf: the pure form ofcommittedTerminalStatus.isSuspendedRunStatus,ClosedOwnerPlanNoneReason,ClosedOwnerNoneReason,ClosedOwnerStreamOutcome: the suspended-status guard (SuspendedRunStatusis now derived intypes/session.ts) and the reconciliation outcome types. EveryClosedOwnerPlancarries itsrunId;reconcileClosedOwnerRunreportsstreamonfinished,suspendedandstream_repaired, and areasononnone.repair_streamalways returnsstream_repaired, so checkstream.isForeignStream: whether a child session writes into its parent's stream.isPermanentStoreError: true for a store-path error a retry cannot fix (a non-retryableHelixError, or a corrupt-row error:InvalidStoredStateError,CorruptMessageRowError,UnknownCommitKindError, matched by name).failSessionWithoutRunandFailSessionWithoutRunInput: records a session that never had a runfailed, under the idle fence, writing no stream and firing no hook (user ruling R43). CF Workflows uses it for a never-run pre-commit failure.
CF Workflows readers reconcile a closed instance
@helix-agents/runtime-cloudflare is a major release: the CF Workflows executor's readers (each handle's result(), getHandle(), canResume() and abort()) and the parent's consult of a child now reconcile a run whose instance closed without its terminal commit through core's closed-owner planner and executor. A reader acts only once it has proven the instance closed; a probe that fails reconciles nothing. See Instances closed outside the product.
What you can observe:
- A committed suspension is reported as suspended, not failed. A run whose instance was terminated, errored or disappeared after its suspension commit (the session
paused, awaiting client calls or children) used to makeresult()answer{ status: 'failed', error: 'Unknown error' }(or the engine's message), and nothing settled the run.result()now answers{ status: 'suspended_client_tool' | 'suspended_awaiting_children' | 'suspended_step_partial', runId, suspended }, the run readssuspended_*(written when it still readrunning), its ownactivestream is paused, andcanResume()is true: submit the results andresume()as usual. A run killed before its suspension commit is settledfailedwith how the instance closed, as arunningrun was before.abort()is unchanged: it fails such a run with its own cancellation. - The reconnect handle's suspension payload. A handle from
getHandle()after the instance closed now reports a suspension from the session:runIdis the run's own id (it was the session id; itscompleted,interruptedandfailedresults now carry the run's own id too, the session id only for a handle built before the session had any run),suspended.stepIdis present whenever the session records one (the workflow's output omitted it for some suspensions), andsuspended.toolCallIdslists every pending client call, including ones already answered but not yet resumed. A reconnect handle built before the session had any run reconciles and reports nothing for a later run. - A stream left
activeorpausedbehind an ended run is ended or failed by the next reader, from the run's committed outcome. It used to stay live until the session's next entry. A host that only serves the chat route still waits for the next send (FU-CFW-CHAT-READER-SETTLE-CLOSED-OWNER). - A parent no longer records a resumable child as dead. When a parent's wait for a child's completion event ends with none and the child's instance is closed, the parent reconciles the child as a reader would and takes the path of a delivered event: a child whose suspension committed is a
suspendedchild (the parent cascadessuspended_awaiting_children; it used to recordChildRunDiedError), a child whose terminal commit landed answers its committed outcome, and a child with nothing committed is settledfailedwith how its instance closed and answers that settled failure (errored,AgentError, the persistederrorDetail's message) — the same answer whoever settled it (the consult, a later consult or an executor reader).ChildRunDiedErroris gone. A failure write that did not land re-waits (the next consult retries it); a permanent error (no workflow binding to probe the child's instance, a broken invariant, or a corrupt stored row:InvalidStoredStateError,CorruptMessageRowError,UnknownCommitKindError) answers a typederrored(itserror.namethe error's class) instead of a timeout. A corrupt row used to be read as "nothing known yet", so the parent waited out its wholesubAgentEventTimeout. The consult re-reads the child's run once its instance is proven closed, so a child an executor reader settled meanwhile is answered at once. A child whose outcome is durable is answered while its stream Durable Object is down. A companion wait still records a suspended companion as a failure (FU-CFW-COMPANION-WAIT-SUSPENDED-CHILD-AS-FAILED). - A handle kept across
deleteSessionno longer reports a made-up entry failure.deleteSessionremoves the session's runs, so a closed instance with no run of its own is now read against the session. With the session deleted,result()reports the instance's recorded outcome (acompletedoutput stayscompleted; an aborted run'scancelledoutput stays aframework_cancelledfailure), and an instance that recorded none (errored, terminated, complete without a result) reports a typed, non-retryablestate_session_not_foundfailure. A session deleted during the entry wait gets the same answer: nopreCommitFailurefor a recorded non-failure outcome, otherwise the recorded failure (possibly a committed run's, since whether the entry committed can no longer be read) orstate_session_not_found; a failed session read there keeps the wait going (cfw.entry.session_read_failed). - A stream that cannot be read or written never hides a run's outcome. Readers plan a run whose instance closed without reading its stream status when that read fails (logged
cfw.settle.stream_read_failed), and core's reconciliation makes the stream write after the run-level write best-effort (reconcile.stream_unavailable; a later reader repairs the stream): a run whose terminal commit landedcompletedis reportedcompletedwhile its stream Durable Object is down (it used to reportfailed), and the workflow's own error path no longer fails such a run or notifies its parenterrored. A transient store error in a reader's settle is logged and the reader answers from what is persisted; a permanent one (a non-retryableHelixError, or a corrupt stored row:InvalidStoredStateError,CorruptMessageRowError,UnknownCommitKindError, also when a D1D1StateErrorwraps one) is logged at error level (code: theHelixErrorcode, else the error's class name) and rethrown. A bareD1StateErroris a failed D1 call and stays transient. - The entry fail-safe retries its abandon.
abandonRunStartis made up to 3 times (25 ms × attempt backoff) before the call falls back to terminating the instance, so a transient store error no longer reaches the unfenced terminate. Aterminate()that throws because the instance had just closed on its own now counts as stopped: the call re-checks the run's start, returns the handle of a run that started, and settles a stall as usual (ownerStopped: true). It used to reportownerStopped: falseand settle nothing. A stalled run whose owner was stopped is settled through the closed-owner reconciliation, so a terminal commit the instance landed before it stopped is finished (the call returns its handle) instead of being overwritten by the stall's failure, and a committed suspension is reported rather than failed. A non-retryable abandon error is not retried. - The pre-commit failure write is retried, and the caller keeps the entry's own failure. The instance's
error-run-existsanderror-pre-commitsteps retry a failed read or save (3 retries each, about 47 s at most per step and 94 s for both; the write used to swallow a failed read). When the entry's own steps fail fast this ends inside the 120 sentryOutcomeFailSafeMs; an entry that hangs in its own steps (check-existing,run-start) can outlast it, and the caller then gets the typed abandon / outcome-unknown error. When either step's retries run out, the instance still returns the entry's classified failure aspreCommitFailure(an exhausted write is logged at error level; an exhausted run check skips the write and is logged at warn). An instance that closed with no run and no failure output now yields apreCommitFailurederived from the engine status (errored→ the engine's message, no code;terminated→framework_cancelled;complete→ a message, no code; each says the entry committed no run), and the handle'sresult()reports the same detail; it used to carry none, so a chat send attached to the session's previous stream. - A never-run session's pre-commit failure writes only the session (user ruling R43). A brand-new session whose first entry fails before its run-start commit is recorded
failedthrough corefailSessionWithoutRun, under the idle fence, with no stream write (it used to fail or restore the session's stream, which could fail the next turn's live stream) and no hook. Temporal still writes the stream on this path (FU-TEMPORAL-NEVER-RUN-PRE-COMMIT-STREAM-WRITE). terminateRun()resolves once the terminate landed. Itsrun_terminatedstream write is now best-effort and logged (cfw.admin_terminate.stream_end_failed); a failed stream read or write used to reject the call after the instance was already terminated.- Typed stream errors at the entry. A failed
initStreamwhile a run's stream is made live is now the typed, retryabletransport_error(a failed head read is the retryablestate_checkpoint_sequence_unavailable; each with the original error ascause); and when a resume or reactivate fails and its re-read fails too, the original error is kept.
What to check: code that treated a CF Workflows result() of failed as final for a session that was waiting on a client tool or a child should handle the suspended_* statuses (an exhaustive switch over AgentResult.status already does). Code that read runId from a reconnect handle's suspended result as the session id should use sessionId. Code that matched ChildRunDiedError in a parent's tool result should match the child's settled failure (an AgentError with the instance-closure message) instead; a child killed after it suspended now makes the parent suspend.