Migrating to Completion Strategies and Preserved Thinking
Overview
Anthropic's Claude Opus 5.5, Sonnet 5.5, Fable 5.1 and Mythos 5.1 reject forced tool use (HTTP 400 tool_choice: type "tool" and "any" are not supported for this model.) and run preserved thinking: a thinking block is bound to the exact system prompt, tool set and earlier messages it was produced with, and a request whose prefix changed is rejected. Before this release, every forced-completion call of an outputSchema agent on those models failed (it forces a named tool, narrows the tools and trims the system prompt), the 400 was swallowed as a "miss", and the run ended framework_forced_completion_exhausted with the provider's error lost. Multi-turn sessions also replayed thinking unfaithfully (blocks merged, empty and redacted blocks dropped).
This release (@helix-agents/core 0.50, @helix-agents/llm-vercel, the JS, Cloudflare DO and Cloudflare Workflows runtimes, the stores, @helix-agents/ai-sdk and @helix-agents/tracing-langfuse) adds:
- Completion strategies. A forced call is enforced either by a named tool choice (
forced_tool, unchanged) or by the provider's native structured output (native_output), chosen per call from capabilities the LLM adapter reports for the model.@helix-agents/llm-vercelships a capability table for Anthropic models and acapabilitiesoverride. See Finishing Agents → Completion strategies. - Preserved thinking. Responses are recorded as ordered
partsand replayed verbatim; on a model whose thinking is bound to history, a changed system prompt or tool set is refused with a typed error before the call. See Preserved thinking.
defineAgent and AgentConfig do not change, and there is no new agent option.
Who is affected:
| You use | Effect |
|---|---|
An outputSchema agent on Opus 5.5, Sonnet 5.5, Fable 5.1 or Mythos 5.1 (JS, CF DO, CF Workflows) | Through @ai-sdk/anthropic (direct, or a LiteLLM proxy behind it): forced completion now works (native_output), nothing to configure for a known model id. Bedrock, Vertex, AI Gateway and OpenRouter ids now resolve to the right profile, but native_output through those providers' own AI SDK packages is unverified (warning, FU-COMPLETION-BEDROCK-VERTEX-NATIVE-OUTPUT in follow-ups) |
An outputSchema agent on a Claude model id the table does not know (a proxy alias such as a LiteLLM sonnet-latest, or a model newer than your package) | Refused at execute() / resume() / retry() with framework_completion_strategy_unavailable (JS, CF DO, CF Workflows). Add a capabilities override |
| Any other model (OpenAI, Google, …, or Claude up to Opus 5 / Sonnet 5) | No change to requests, except that a thinking session with multi-block, empty or redacted reasoning now replays it verbatim (§2); the behaviour changes in §3 still apply |
A custom LLMAdapter | Compiles and behaves as before (legacy); see §5 to opt in |
| Temporal or DBOS | Forced completion on the 5.5 generation still fails until MR 2 / MR 3; see §7 |
1. Configure proxy aliases and new models
What changed
The adapter maps the model to a capability profile (legacy, anthropic-structured, anthropic-5.5). An Anthropic model with no row in the table is unresolved: the framework cannot know whether a forced call would work, so an agent with outputSchema on it is refused before its run starts, typed and with nothing written. An agent without outputSchema is never refused and runs as before, and neither is one with maxCompletionRetries: 0 (it never makes a forced call). On runtime-js and the Cloudflare DO an agent with an llmConfigOverride is not refused at entry either (the override may pick the forced call's model); its forced call fails with the same code if the model that call uses cannot finish. Over HTTP the refusal is a typed 400 { error, code } (code framework_completion_strategy_unavailable or framework_completion_schema_unsupported) from both the Cloudflare DO (/start, /resume, /retry, /submit-tool-result) and @helix-agents/agent-server (/start, /resume, submit-tool-result), never a 500. A refused submit writes no result and leaves its call pending.
A model counts as Anthropic when its id starts with claude- or its provider id contains anthropic. So a LiteLLM alias served through createAnthropic({ baseURL }) (provider anthropic.messages, id sonnet-latest) is an unknown Anthropic model and is refused. Give a custom-named provider a name containing anthropic (createAnthropic({ name: 'anthropic-litellm', baseURL })); under { name: 'litellm' } an alias is not recognised as Anthropic and every forced call on a 5.5-generation backend fails at the provider.
What you need to do
Declare the alias, or a model newer than the table, on the adapter:
import { modelIdOf } from '@helix-agents/core';
import { VercelAIAdapter, resolveAnthropicCapabilities } from '@helix-agents/llm-vercel';
const LITELLM_ALIASES: Record<string, string> = { 'sonnet-latest': 'claude-sonnet-5-5' };
const adapter = new VercelAIAdapter({
capabilities: (config) => {
const alias = modelIdOf(config.model);
const real = alias ? LITELLM_ALIASES[alias] : undefined;
return real ? resolveAnthropicCapabilities(real) : undefined; // undefined → built-in table
},
});See Vercel adapter → Overriding capabilities and the completion strategies example. Install @ai-sdk/anthropic >=3.0.38 (now an optional peer dependency of @helix-agents/llm-vercel) for Anthropic models.
2. If you do nothing
- Models other than the 5.5 generation send the same requests as before, including forced calls (the
anthropic-structuredprofile still uses the named tool choice); only the hidden correction message's wording changed (§3). One exception: on any model with thinking enabled (not only the 5.5 generation), a session whose responses had several reasoning blocks, or empty or redacted ones, now replays each block verbatim and in order instead of merging them into one and dropping the empty and redacted ones, so the bodies of that session's later work requests change (the intended preserved-thinking fix). - 5.5-generation
outputSchemaagents with a known model id served through@ai-sdk/anthropicstart finishing correctly on JS, CF DO and CF Workflows. Through Bedrock, Vertex, AI Gateway or OpenRouter providers the id resolves toanthropic-5.5, but whether native structured output works through those providers is unverified (see the Vercel adapter warning). outputSchemaagents on an unknown Claude id (aliases, unreleased-in-the-table models) start failing at entry withframework_completion_strategy_unavailable; before, they ran and failed (5.5 generation) or worked (older models behind an alias). Add the override.- The behaviour changes below apply.
3. Behaviour changes on every model
Provider errors inside a forced call end the run (JS, CF DO, CF Workflows)
A provider error on a forced call used to be a recoverable miss: another correction message, another forced call, and finally framework_forced_completion_exhausted with the provider's code lost. Now it fails the run with the provider's own errorDetail (for example provider_invalid_request, provider_auth_error). Transient errors are retried by the runtime's own mechanism instead of a forced attempt: the adapter's maxRetries on runtime-js and the CF DO, the step.do retry (llmStepRetry) on CF Workflows. If you assert on framework_forced_completion_exhausted after a provider failure, assert on the provider's code instead. Temporal and DBOS keep the old behaviour (§7).
The hidden correction message is reworded
The persisted hidden [completion] message now reads:
[completion] Provide your final answer now (attempt N) in the required structured format. Do not call any other tool and do not reply with plain prose.
(previously it named the tool to call). It is the same on every runtime and model, because it is persisted before the next call's strategy is known. Recorded-replay LLM fixtures keyed on the request body that cover a forced call need re-recording. Finish semantics do not change: on models that support it, enforcement is still the named tool choice.
Skills and memory fail closed
- A skills provider whose
listSkillsthrows, or whosegetSkillthrows for apreloadSkillsname, used to be logged and skipped (the step ran without the catalog or the body). It now fails the step with the retryableframework_skills_unavailableon every runtime except DBOS. Durable runtimes retry the step (Temporal activity retry, CF Workflowsstep.do); runtime-js and the CF DO fail the run, andretry()re-runs it. DBOS keeps the old behaviour until MR 3: it resolves the catalog in the workflow body, outside any step, so it logs a warn and runs the turn without the catalog and without any preloaded body (one failing preloaded name used to skip only that body) (FU-COMPLETION-DBOS-SKILLS-RECOVERY-STRAND). An unknownpreloadSkillsname (getSkillreturnsnull) is still skipped with a warning.resolveSkillsCatalognow rejects instead of returning''. - On runtime-js and the CF DO, a throwing
memoryManager.buildToolsused to drop the memory tools silently; it now fails the step with a retryableframework_internal_error.
Silently running without part of the prompt would change the prompt prefix, which a 5.5-generation model rejects for the rest of the session, and it hid the failure on every model.
Prompt-prefix check on 5.5-generation models (JS, CF DO, CF Workflows)
Once a 5.5-generation model has produced signed thinking in a session, a call whose rendered system prompt or tool set differs from the one that thinking was produced with (for that model) is refused before the API call with framework_prompt_prefix_changed (non-retryable), naming which changed. The API would reject it with an opaque 400 anyway. Typical causes: a systemPrompt function whose output changes between steps, a deploy that changes prompt text, tool descriptions or schemas, skills or workspace config while sessions are live. Keep them stable for a session's lifetime, put changing context in an appended message, or start a new session. A beforeLLMCall hook that edits earlier messages is unsupported on these models and is not detected. Other models are not checked.
Other observable changes
- New error codes:
framework_completion_strategy_unavailable,framework_completion_schema_unsupported,framework_prompt_prefix_changed,framework_skills_unavailable. See Error handling. native_outputforced calls stream notext_delta(the JSON answer is not shown live; thinking still streams), and the persisted assistant message carries a completion-tool call withmetadata.helixNativeCompletion: true. A native miss (text that does not parse, or amax_tokenscut-off) is persisted withmetadata.hidden: true, soconvertToUIMessagesleaves it out after a reload as well; it still goes to the model on the next call (its signed thinking must replay).forced_completionchunk:attempt_startedgains optionalstrategyandcapabilityProfile. No new enum values, so older readers just ignore them.- Hooks:
beforeLLMCall/afterLLMCallpayloads gain optionalcompletionStrategyandcapabilityProfile;toolChoiceis absent undernative_output. Key onexecutionPhase === 'forced_completion'to detect forced calls. - Langfuse: generation metadata gains
executionPhase,completionStrategyandcapabilityProfile. AssistantMessage.parts(ordered reasoning / text / tool-call parts) is recorded for responses with reasoning, andmetadata.helixPrefix(the prompt fingerprint) on messages from 5.5-generation models.@helix-agents/ai-sdkdoes not forwardhelixPrefixto UI message metadata;stripThinkingalso removesreasoningparts.
4. Removed and changed APIs (@helix-agents/core)
| Before | After |
|---|---|
framework_completion_tool_choice_unsupported error code (documented, never raised) | Removed. A model that rejects forced tool choice uses native_output; one with no strategy fails framework_completion_strategy_unavailable |
isTerminalLLMStepError(result, { forcedCall }) | isTerminalLLMStepError(result): a forced call's error ends the run too |
prepareForcedCompletionCall({ state, effectiveTools }) | prepareForcedCompletionCall({ state, effectiveTools, resolution, modelLabel }); may also fail with framework_completion_strategy_unavailable |
ForcedCompletionDirective = { tools, toolChoice, reasoningMode, executionPhase } | A union discriminated by strategy (forced_tool / native_output); apply it with forcedDirectiveGenerateInput |
classifyForcedCompletionResult made an error step a recoverable miss | It is nonRecoverable with the provider's errorDetail |
ForcedCompletionResultClassification nonRecoverable.code: ErrorCode (required) | code?: ErrorCode: absent for an unclassified provider error, which carries no code anywhere (never an invented one). A failed ForcedCompletionTransition's errorCode (already optional) is then absent too, and FailForcedCompletionInput.code is optional |
resolveSkillsCatalog returned '' when the provider threw | Rejects with framework_skills_unavailable |
PlanStepProcessingOptions.forcedCompletion: { expectedToolName } | Also { expectedToolName, strategy: 'native_output', callIdSalt } |
Added: ModelCapabilities, CapabilityResolution, LEGACY_CAPABILITIES, LEGACY_RESOLUTION, resolveCompletionCapabilities, modelIdOf, modelLabelOf, CompletionStrategy, COMPLETION_STRATEGIES, selectCompletionStrategy, checkCompletionSupport, isCompletionSupportRefusal, resolveForcedCallTransition, forcedDirectiveGenerateInput, assertForcedToolsUnchanged, FORCED_NATIVE_MIN_OUTPUT_TOKENS, nativeCompletionCallId, computePromptFingerprint, checkPrefixStability, stampPromptFingerprint, LLMOutputFormat, AssistantMessagePart (+ schema), COMMON_METADATA_KEYS.PROMPT_PREFIX / NATIVE_COMPLETION, and FaithfulProviderSimulator (@helix-agents/core/testing). See the core reference.
5. Custom adapters and bring-your-own-loop code
- Custom
LLMAdapters need no change: withoutresolveCapabilitiesthey arelegacy(named tool choice), as before. To support a model that rejects forced tool choice, implementresolveCapabilitiesandcheckOutputFormat, honouroutputFormat/minOutputTokens, and return the answer as a plaintextresult; see Custom adapters. Adapters for models with history-bound thinking should capture and replay orderedparts. - Adapter wrappers (an object that forwards
generateStepto a real adapter) must also forwardresolveCapabilitiesandcheckOutputFormat, or every model reads aslegacy. - Bring-your-own loops that run forced completion: resolve capabilities per forced call, call
checkCompletionSupportbefore starting a run, pass the strategy and a replay-stablecallIdSalttoplanStepProcessing, useresolveForcedCallTransition, and runcheckPrefixStability/stampPromptFingerprintaround the LLM call. See Completion strategies (cross-runtime). - Tests:
new MockLLMAdapter(responses, { capabilities, capabilityProfile })simulates a model with those capabilities (rejects forced tool choice, signs and verifies scriptedreasoningblocks).
6. Deploying
- No store migration.
partsis an optional field inMessageSchema(messages are JSON on every store),helixPrefixis message metadata, andForcedCompletionStatedoes not change. - Rolling deploys on preserved-thinking models. A process still on the old version strips
partswhen it reads history and replays the merged legacy form, which an account with preserved-thinking enforcement rejects with a 400 until every process runs the new version. Keep the mixed-version window short for 5.5-generation traffic. - A config-changing deploy (prompt text, tool descriptions or schemas, skills, workspace) now fails live 5.5-generation sessions with
framework_prompt_prefix_changedon their next call instead of an opaque 400. Start new sessions after such a deploy. - Cloudflare Workflows: the completion-support decision is recorded in a
step.do(<agentType>-completion-support), so a replay reuses it and a later change to how anoutputSchemaagent's model resolves (a capability-table change, a removed override) never fails an instance that was admitted. An instance started before this release has no recorded decision: it makes the check fresh on its next replay, so for those keep the resolution stable across the deploy, or drain affected instances first. - Temporal / DBOS: no new activities, steps or workflow branches in this release.
7. Temporal and DBOS (known gaps)
Temporal (MR 2) and DBOS (MR 3) are not migrated yet. On those runtimes:
- every forced call uses the
legacyrequest whatever the adapter reports, so forced completion on Opus 5.5, Sonnet 5.5, Fable 5.1 and Mythos 5.1 still fails (HTTP 400 on each attempt, thenframework_forced_completion_exhausted); - a forced call's provider error stays a recoverable miss (the step result crosses a serialization boundary that loses its classification);
- there is no early refusal and no prompt-prefix check; hook payloads and chunks carry no strategy fields;
- DBOS additionally drops
providerOptions(thinking, effort) andmaxOutputTokensacross its step boundary.
What does apply there: the reworded correction message, skills fail-closed on Temporal (the typed code survives the activity boundary as an ApplicationFailure; DBOS keeps a failing skills provider non-fatal until MR 3), the removed error code, and ordered parts capture and replay (it is adapter-side). Run outputSchema agents on 5.5-generation models on runtime-js or a Cloudflare runtime until MR 2 / MR 3 land. See Finishing Agents → Temporal and DBOS.