Repository navigation
feat(ai, ai-harness): add the harness stack with replay and adapter parity - #1555
Open
AlemTuzlak wants to merge 314 commits into
Open
AlemTuzlak wants to merge 314 commits into
AlemTuzlak wants to merge 314 commits into
Conversation
…nto feat/harness-p10-goal
…pauses or stops A message typed during the last model call of a goal turn runs first, as a steer that came too late, and pauses the goal. The goal's next turn was already queued and still ran once. ctx.session.prompt() now returns the turn, and the goal plugin cancels its queued turn when a user message pauses the goal, on /goal stop, and when a new goal replaces it. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
… clock The timeout tests use a 1 ms budget, so the branch they end in depends on how many milliseconds pass. The coverage number moved between runs. New tests fix the clock and reach each of the five timeout exits every run. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…n state in the snapshot Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…/command/setConfig, and connection state Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…session for any UI Adds @tanstack/store ^0.11.1 and the browser-safe ./view export. Three examples listed @tanstack/store ^0.8.0 without importing it; they move to ^0.11.1 so the workspace keeps one version (sherif). Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…any TUI on a session view The CLI no longer depends on ink or react. An interactive terminal uses line mode unless runCli gets a ui function, which receives a ready session view and resolves when the user quits. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…ove the old event fold In a terminal, line mode opens sign-in links in the browser. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
The coverage gate failed because the new session view code had paths that no test ran. These tests run them. - view-paths.test.ts: a fake SessionViewSource forces each path of createSessionView. It covers failed reads at start, a failed or ended event stream, connection state, failed actions, approvals, on() handlers, the snapshot read coalesce, and dispose. - view-reduce.test.ts: the reducer branches that return the same state, nested and sibling child agents, content-part tool results, run errors, and transcript edge cases. - client-reads.test.ts: the transcript and describe routes answer 400 and 403, and a failed read throws with its route and status. - goal.test.ts: selectGoal gives null for state that is not a goal. - view-fixtures.ts: event and snapshot helpers that the view tests share. No production code changed. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
… screen in the example Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…nto feat/harness-p12-cli
Every direct TanStack Store dependency in the workspace is now on the latest release (store, react-store, and solid-store 0.11.1). Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
… agent
@tanstack/ai:
- The chat middleware context has subagentName and parentSubagentRunId.
chat() takes both as options. The bound ctx.chat sets them.
- A child that calls ctx.chat({ subagents }) gives its own children the
host chatMiddleware and generationMiddleware before their own, also
when the binding has no budget.
- A nested child gets parentSubagentRunId on its run input. Its
SUBAGENT_STARTED does not carry it in the stream of the chat that
started it: attributeChunk adds it one level up, as before.
@tanstack/ai-harness:
- Plugins can contribute agentMiddleware: chat middleware for subagents,
background agents, and their children. Not for the lead turn.
- The lead turn passes run plugins' agentMiddleware and
generationMiddleware to its agents too.
Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Add a "Middleware in every agent" section to the plugins page, with a usage tracker that puts one middleware in middleware and agentMiddleware. Link it from the harness subagents page. The chat subagents page names the new subagentName and parentSubagentRunId context fields. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…harness-p13-agent-middleware
The usage counter now sits in agentMiddleware too, so /usage covers the lead turn, subagents, background agents, and their children. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
The final tool input check kept a missing or literal null argument and checked it against the schema, so a tool with no required fields did not run when the model sent an empty tool_use block. That broke the fix for issue #265, and its E2E test failed. No input is {} again. A null that the schema rejects is checked as {}. Other values, such as 2 or [1,2], must still fit the schema. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
The final tool input check now runs onBeforeToolCall before a client tool is dispatched, so otelMiddleware opened an execute_tool span for a tool the server does not run. The client-tool wait E2E test then saw two root spans. otelMiddleware skips the span for a known tool with no execute. The unit test is back to main's chat and iteration spans. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
The client-tool input error scenarios sent message: 42. The final input check now coerces scalar types, so 42 became "42" and passed, and the client ran the tool. The scenarios now omit the required message, which is invalid with or without coercion. The sandbox persistence spec now expects the runId and source that the chat run records on the saved assistant message. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Replay now gives an unanswered tool call from an earlier turn a "No result provided" result. The provider-free adapter counted every tool message, so after Stop the next user turn answered with text and ended before the race test could press Stop again.
… replay tests The #265 fix runs a tool with no input or a literal null as {}. Five adapter replay-parity tests still expected those inputs to be rejected. They now check that the tool runs with {}, or that the call goes to the client. Other scalars and malformed JSON are still rejected. Mistral ends whitespace-only arguments with a parse error before the input check, so that case still runs 0 times. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…turn summary usage Two new functions had no test, so the Coverage job saw function coverage for ai-compaction drop from 100% to 97.43%: - contentText in conversationSummarizer, for content that is a list of parts. - The addUsage callback of the after-turn check, for a summarizer that reports usage. Each new test fails when its function is broken. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Brings in main at 8ce7ad9. Conflicts and how they were solved: - Anthropic usage: keep the branch's total-input promptTokens and add main's message_start fallback for counts the closing delta leaves out. main's new usage tests now expect the total input. - Anthropic replay: main sends {} for invalid JSON tool arguments. The branch moved that code into appendToolCallBlocks, so the fix goes there. - Anthropic options: thinking and effort are chat({ reasoning }) on this branch, so claude-sonnet-5-5 takes max_tokens without sampling options and no output_config. main's type tests and E2E route follow that. - chat(): main records an afterModel interrupt turn with addTerminalAssistantMessages, which already saves the mid-conversation record. The old helper is gone. - Stream processor: keep both sides' new state. serializeToolInput is no longer used; main's parseToolArguments stays. - Route tree: the branch's tree plus main's seven new routes. main's tree imported the Sonnet 5.5 route from the wrong file. - Lockfile: the branch's lockfile plus main's entry for @tanstack/ai-isolate-e2b. A text merge moved zod for some packages. - Docs and dates: keep the newer content and the newer updatedAt. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
A catalog model whose id is not in the adapter's own table got no
reasoning at all, because each adapter looked up its generated table by
model id. About 500 catalog models lost their thinking this way.
Each first-party text adapter now takes `reasoning?: ModelReasoning` in
its config. When it is set, it wins over the table for the request
fields and for the levels `chat({ reasoning })` takes. `false` sends no
reasoning field. When it is not set, nothing changes.
- Anthropic (anthropicText, createAnthropicChat, anthropicVertexText)
- OpenAI (openaiText, createOpenaiChat) and azureOpenaiText
- Gemini (geminiText, createGeminiChat), AI Studio and Vertex mode
- Bedrock createBedrockConverse
- Mistral (mistralText, createMistralText)
- Cloudflare (cloudflareText, createCloudflareText)
The factories take any model id string. A config with `reasoning` types
`chat({ reasoning })` with every level (ConfigReasoning in @tanstack/ai);
the adapter moves a level the model does not have to the nearest one.
@tanstack/ai-models adds `modelReasoning(record)`, the one rule that
turns a record into the config value.
Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…ogle-vertex modelCost now follows pi's two rules that it did not have: - Tiers. A record's `cost.tiers` holds higher prices above an input size, for example long context. A call uses the tier with the highest `inputTokensAbove` below its input (uncached + cache read + cache write). The tiers come from models.dev `cost.tiers`. A price that models.dev leaves out is 0, as in pi's catalog. - 1-hour cache writes. `TokenCounts.cacheWrite1h` is the part of `cacheWrite` with a 1-hour retention. It costs 2x the input price. Usage gets `promptTokensDetails.cacheWrite1hTokens`. The Anthropic adapter fills it from `cache_creation.ephemeral_1h_input_tokens`, and the Bedrock Converse adapter from the 1-hour `cacheDetails`. The google-vertex catalog listed Claude models with the Gemini wire, which the Gemini adapter cannot call. They are left out, as in pi. The fetch script keeps the tiers. The models.dev snapshot gets the tiers of today's models.dev for the models it already has. Regenerating also picks up 11 gateway models that the snapshot already had. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Brings in main at a32782c: activity messages and the ActivityStore, video bytes through generation persistence, stream backpressure, and the structured output and MCP fixes. Conflicts and how they were solved: - Stream processor: keep main's activity-id guards. The branch already moves the structured-output mark in remapProvisionalMessage, which is main's #1634 fix. The thinking-step lookup skips activity messages. - chat(): keep both new methods; main sets the activities first, then the branch's providerMessages rule. Activity chunks no longer get the assistant call metadata, so stored activity records keep their own. - Persistence: keep the branch's streaming-id tracking and main's activity save after each snapshot. - Sandbox test: keep the branch's MemorySnapshotPersistence type. - @tanstack/ai: add main's fast-json-patch next to the branch's typebox. The lockfile is the branch's plus main's fast-json-patch entries. - Route tree: the branch's tree plus main's two activity routes. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
OpenAI Responses gives each answer item an id and a phase (commentary or final_answer). The adapter kept item ids only for tool calls and reasoning, so a replayed answer lost its phase. The adapter now puts each message item's id and phase on RUN_FINISHED, and chat() keeps them on the assistant message in `metadata.tanstack.responseItems`. A same-model replay sends each text block back as its item, with the id and the phase. Another model, or a message whose blocks do not match its items, gets the plain text. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…ules - Reasoning: a model the adapter does not list takes the `reasoning` config, with the factories that accept it. - Model catalog: build an adapter from a record with modelReasoning, the adapter for each `api`, and the tier and 1-hour cache write prices. - Prompt caching: usage reports `cacheWrite1hTokens`. - OpenAI: answer items keep their id and phase on the next request. Changesets for the reasoning config, the cost rules, and the Vertex Claude records. An E2E wire test checks that a gateway id outside the Anthropic list gets its thinking budget from the config, and none with `reasoning: false`. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
The foreign *_CHUNK spec from main expected only the sender's metadata on a message that ends in RUN_ERROR. This branch also marks that message with `tanstack.stopReason: 'error'`, so a replay leaves the failed batch out. The sender's metadata is still there. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Add examples/harness-cli-batman as a themed copy of the harness CLI. It uses a bat splash, a yellow header, and a separate save folder. Voice input times out after 45 seconds and Esc cancels it.
…kfile The batman example's install re-resolved part of the workspace: zod moved to 4.6.5 for @modelcontextprotocol/server, and @google/genai got a second copy. That broke the ai-mcp type check (two zod copies in the server types) and the ai-vertex factory tests (the mocked SDK was the other copy). The example has the same dependencies as examples/harness-cli, so its lockfile entry now resolves the same way (zod 4.3.6, @tanstack/react-store 0.11.1). The rest of the lockfile is as before the example. A frozen install passes, and ai-mcp test:types, ai-vertex test:lib, and the example's test:types pass. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Brings in main at 32e56bb (Version Packages #1630): the released versions, CHANGELOGs, and the changesets that release used. No conflicts. The lockfile is unchanged and a frozen lockfile check passes. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Flue's shape check found 195 Anthropic Messages models where our
thinking fields differ from pi's. Gateway ids with dots
(anthropic/claude-opus-4.7) got budget thinking, older Claude on
OpenRouter got adaptive thinking, and most other models on the wire got
adaptive thinking because the adapter sent adaptive whenever a model had
no token budget.
Adaptive or budget from the record:
- ModelReasoning.adaptive: true sends adaptive thinking, false sends
budget thinking. Without it, the adapter decides from the id, as before.
- Adaptive thinking without a level map sends pi's default effort.
- modelReasoning(record) sets adaptive from compat.forceAdaptiveThinking
for anthropic-messages records.
- The catalog rule is pi's: Claude 4.6 and later (dash or dot ids) think
adaptively. Fireworks takes pi's list. All 291 shared records now have
pi's flag. Seven Fireworks records take Fireworks' own levels (pi), and
Vercel gemma-4-31b-it reasons, as at Google and OpenRouter.
Mid-conversation effort (pi supportsMidConvoEffort):
- ModelReasoning.midConversationEffort, from the new
compat.supportsMidConvoEffort (5 models, pi's list).
- The adapter sends adaptive thinking with block_binding and a fixed
output_config.effort high, an effort system message before each
earlier answer and one at the end, the two betas, and no temperature.
- Each answer keeps its effort in metadata.tanstack.reasoningEffort.
- The automatic cache marker skips the trailing effort message.
Mid-conversation channels: midConversationChannels takes
{ tools?, systemPrompts? } to turn on one channel only.
Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
- The config reasoning wire route sends four record configs: adaptive (Vercel anthropic/claude-sonnet-4.6), budget (Vercel openai/gpt-5), mid-conversation effort (OpenRouter anthropic/claude-opus-5.5), and reasoning: false. The spec checks the thinking fields, the effort message, and the two betas in the anthropic-beta header. - Docs: adaptive and midConversationEffort in the reasoning config, the thinking shape from modelReasoning(record), a section for a level change during a conversation, and the object form of midConversationChannels. - Changeset for @tanstack/ai, @tanstack/ai-anthropic, @tanstack/ai-models. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
The reasoning drift test (test:maintainer) failed: the catalog now says
Vercel google/gemma-4-31b-it reasons, and the ai-vercel-gateway table
said it does not. models.dev has no reasoning for it at Vercel, but
Vercel's own model list tags it `reasoning`, and Google, OpenRouter, and
pi 0.87.1 list it.
The ai-vercel-gateway target gets the same override, and its table entry
is the one the sync script writes ({ budget: false }). The other drift
that the sync script finds (models added with main) is left for its own
sync.
Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Coverage failed on this PR for ai-isolate-e2b (branches 81.48% vs
82.22%), although the package did not change. The false side of
`recent.includes('heap out of memory')` in `onStderr` was hit only by
the memory test, and only when the host read V8's first GC block as its
own stderr chunk. On a Linux runner the writes usually merge into one
chunk, so that side is missed.
The new test writes ordinary stderr and exits with code 2. It must fail
as E2BExecutionError with the stderr tail, not as MemoryLimitError. It
fails when every stderr chunk sets `outOfMemory`, and it also covers the
stderr tail in the exit message.
The thinking-shape changeset also lists @tanstack/ai-vercel-gateway
(the Gemma 4 31B fix in 6656c1d).
Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
3 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Review the combined harness and replay stack in this PR. The harness keeps one agent conversation open across turns, restarts, and model changes. This combined stack adds the harness, CLI, dashboard, MCP in both directions, durable sessions, host hooks, and compaction. It also keeps saved history usable across adapters and checks final tool input before execution or client dispatch.
🎯 Changes
Compatibility changes
modelOptionswithchat({ reasoning }). The migration guide remains part of this PR.sequential: trueortoolExecution: 'sequential'. Prompt caching is on by default. UsepromptCache: 'none'to disable it.promptTokens. Assistant block order can produce more AG-UI rows. Source metadata separates the requested model from the response model.reason, andwithCompactionaddscompactNext. Durable turns keepLogRecordsCapability.Saved subagent cards use a separate view of the transcript. Ordinary text and ambiguous hosts stay in that view. Both host creators preserve unrelated metadata. They choose an unused ID for missing IDs and explicit duplicate IDs, while keeping a unique explicit ID.
Replay changes provider input. It keeps the saved transcript unchanged. History without source metadata retains its legacy source treatment.
Existing harness stack
This remains the single combined PR for phases P0 to P14. It replaces #1551 and the closed phase PRs. Their descriptions remain the review record.
harnessText.--mcp, and/mcpon--serve.The owner changes also keep media, provider keys, durable tool steps, host hooks, turn leases, skills, prompt caching, routing, and per-turn overrides. Thinking, text, and tool calls keep their block order. Mid-conversation tool and prompt changes preserve the cached start where the model supports them.
MCP in both directions. The client keeps connectors and browser sign-in.
createHarnessMcpServerexposes chat, approvals, agents, and commands to MCP clients. The CLI serves stdio MCP or authenticated HTTP MCP. MCP inputs keep parsed template variables and per-answer MIME types.createMCPClientkeepstoolName, request timeouts, and stricttoolFilterchecks. Missing or repeated filtered names raiseMCPToolFilterError. Tool metadata keeps titles and frozen annotations. Final input validation runs before an accepted call reaches its tool.Compaction.
withCompaction({ countTokens: 'usage' })uses reported usage.compactNext(threadId)requests compaction at the next model call. Durable compaction writes session-log records throughLogRecordsCapability. The host view rebuilds the context from those records.Compaction reports its reason, summary usage, and errors. Summary options keep turn cuts, updates to previous summaries, and bounded tool output. These opt-in features remain part of the owner scope.
Latest owner update preserved
The eight owner commits through
c9f8e5f901f82e5dcba3dbd505d7ee2ab9e110e3stay in the stack. Shared threads run each chat input and command as the principal returned byauthorize.Principal.tenantIdkeeps credential scopes apart. Reads use the user's key before a shared tenant key. Commands save only through their own sender's credentials. Different senders do not join a running turn by default. Client context stays untrusted; server context wins when objects merge.The harness run route now supports
useChat, transcript hydration, interrupted phases, and cursor-based rejoin after reload. Repeated run IDs keep their receipt. A conflicting sender or already-used run record returns a conflict. Invalid resolve input leaves the interrupt pending. A resolve sent as a turn finishes can start immediately after it.Sign-in wait uses
credentials.require(id, { wait: true })inside a chat tool. The stopped sender's sign-in resumes that turn.lifetime: 'turn'names a plugin that lives for one turn;'run'remains an alias.TurnInfocarries message, input context, principal, input ID, and overrides. A plugin may supply each turn's adapter, soHarnessConfig.adapteris optional in that case.harnessTextaccepts explicitinputModalities.The view separates approvals, client tools, and sign-in requests.
ClientToolCall,view.on('clientTool', ...), and its resolve/fail actions answer client tools. The CLI accepts a caller-owned host and principal, keeps that host open, and keeps its bearer authorization gate.fakeTextkeeps parent run identity, distinct tool IDs between adapter instances, and exact optional typing. All eight new owner changesets, shared-thread docs, sandbox-provider docs, and their navigation entries remain.MCP sender isolation. The new per-sender connection cache now encodes tenant and user as a JSON tuple. Names that contain
:cannot reuse another sender's authenticated client or tools. A new real-SDK/HTTP regression covers two names that collided before, distinct scoped credentials, separate tools, and same-sender reuse. This new regression was source-reviewed and was not run locally under the user's push instruction. The owner changeset already includes the@tanstack/ai-mcppatch.Other owner fixes retained. The stack keeps MCP interrupt-kind handling, the Linear OAuth issuer, unique exported tool names, v2 connector credentials, background-agent lease recovery, routed handoff hooks and child resume, interrupted-session recovery, and MCP sessions needed for elicitation. Workspace approval, lazy code-mode tools, skills, Cloudflare example, model catalog, and live plugin commands remain part of the owner scope.
Usage, resumable agents, thread settings, fork, and reset
These five owner features come from the pi-durable comparison. Each one is opt-in or read-only for existing callers.
session.usage()andsnapshot().usagetotal each model call by model and by sender, with the cost that providers report. A durable host keeps oneharness.usagerecord per call. Aharness.usageevent fires on each call.agents.start(agent, input, { resume: true })needs a durable host. The host that takes over runs the agent again from its saved transcript.ctx.step.do(name, fn)keeps side effects from running twice.durability.maxAttemptscaps the runs.defineHarness({ models })andsession.configure()store the model, reasoning, instructions, tools, plugins, and working folder of a thread. A client can change only the fields inexpose.settings, and none by default.host.fork()copies a thread up to a message into a new thread, with its settings and media.session.reset(note?)starts a fresh model context and keeps the full transcript, with a marker.e677b7516stays as it is.Replay and adapter parity
No result providederrors. Keep system changes after tool results.RUN_ERRORwith the provider finish reason. Knowncontent_filterbehavior stays the same.status: 'error', including an empty error string.x-api-key.allowEmptySignaturepermits unsigned same-source thinking on configured gateways. The default remains false.tools: []for tool history without active tools. Wrappers retain that empty list. The no-history request stays unchanged.Final validation. Standard Schema runs first on the exact input. When a safe input schema permits a coercion retry, the authored schema checks the retry result. A successful transform keeps its value. Raw JSON Schema uses full schema checks. Pending or denied calls do not reach tool hooks, execution, or client dispatch.
The only new runtime dependency for this replay work is approved
typeboxat^1.3.27, locked to1.3.34. TypeBox uses its check engine when dynamic evaluation is unavailable. Source review confirms this fallback. Focused controls check validation behavior. They do not prove a deployed Worker.Own-key preservation covers schema copies, nullable widening maps, Standard Schema root copies, Gemini parameter sanitation, and Azure deployment lookup. OpenRouter keeps raw arguments separate from normalized input across all three terminal paths. Scalar, array, null, and missing input keep their contracts.
The host selector still returns a message index or
-1. Presence guards make that return type explicit. This repairs type inference; it does not change caller sentinels or claim a runtimeundefineddefect.Integration. The candidate preserves the earlier 16 owner commits and the eight newer owner commits through
c9f8e5f901f82e5dcba3dbd505d7ee2ab9e110e3, then retains the approved merge with pinned main6d8e6485f92c98a6f2200471e2aff21ab50be013. Main's cancellation and malformed-JSON controls remain. The route registry includes the owner and main routes. Model metadata keeps main's catalog and the owner's modality and reasoning policy.Main merge and data-loss fixes
fdaeabe8fmerges main at7fb4a5f7f. That brings in #1620, live durable streams and one hydrate GET in Strict Mode. These commits fix the data-loss issues and the CI failures that showed up after the merge:581efd816rebasedropped it when the turn ended.af7274d6f4c12b1e89,e677b7516bf329d5e3,809259930nullarguments runs with{}again (#265). The final input check made this fail. Other scalars and malformed JSON are still rejected.809259930updates the five adapter replay tests that expected the rejection.c1e9fdad8execute_toolspan for a client tool. The server does not run it.15868e693ai-compactionfunctions drop from 100% to 97.43%.The test commits
71d5e933c,f02f0ddcd, andf07b75e0akeep the MCP and E2E tests true under the new behavior. Five patch changesets cover the five fixes above. The tools-test adapter now counts only the tool results of the current turn. Replay gives an unanswered call from an earlier turn aNo result providedresult, so the old count ended the run after Stop too early.Adapters from catalog records
A catalog model whose id is not in the adapter's own table got no reasoning field. That was 582 of the 709 catalog reasoning models on first-party wires. Now:
reasoning?: ModelReasoning: Anthropic (incl. Vertex), OpenAI and Azure, Gemini (AI Studio and Vertex), Bedrock Converse, Mistral, and Cloudflare. It wins over the table for the request fields and for the levelschat({ reasoning })takes.falsesends nothing. The factories take any model id.@tanstack/ai-modelsaddsmodelReasoning(record).modelCostuses the models.devcost.tiers(long context) and pricescacheWrite1hat 2 times the input price. Usage addspromptTokensDetails.cacheWrite1hTokens, filled by Anthropic and Bedrock Converse.google-vertexcatalog leaves out the Claude models, which had the Gemini wire.idandphaseinmetadata.tanstack.responseItems, and a same-model replay sends them back.The remaining 66 models with no reasoning field are Bedrock Converse models that are not Claude. pi also sends Converse thinking fields only for Claude.
This batch also merges
maintwo times, at8ce7ad97canda32782c3e. The merges keep the activity messages, video persistence, stream backpressure, and the Anthropicmax_tokensusage, replay, and Sonnet 5.5 fixes frommain. Activity chunks no longer get the assistant call metadata.Anthropic thinking shape from the record
Flue compared the Anthropic thinking fields with pi's for every
anthropic-messagesreasoning record. 195 records were different. Gateway ids with dots (anthropic/claude-opus-4.7) got budget thinking, older Claude on OpenRouter got adaptive thinking, and most other models on the wire got adaptive thinking. Now:ModelReasoning.adaptivepicks the shape.modelReasoning(record)sets it fromcompat.forceAdaptiveThinking, and the catalog flag now equals pi's on all 291 shared records. Adaptive thinking without a level map sends pi's default effort.ModelReasoning.midConversationEffort(from the newcompat.supportsMidConvoEffort, pi's 5 models) puts the level into the messages, so a level change keeps the cached start. Each answer keeps its level inmetadata.tanstack.reasoningEffort.midConversationChannelstakes{ tools?, systemPrompts? }.google/gemma-4-31b-itreasons.@tanstack/ai-vercel-gatewaygets the same Gemma fix (6656c1d71), so the reasoning drift test agrees. Vercel's own model list tags itreasoning.618d71121adds anai-isolate-e2btest for stderr without an out-of-memory report. That branch side was hit only when V8's first GC block came in its own stderr chunk, so Linux CI missed it on this PR.The local copy of the Flue check finds only group E (top-level
efforton Claude 4.6, accepted in the 2026-09-30 report). Before this batch,53fab0187brought theharness-cli-batmanlockfile entry in line withharness-cli, and280263df4andef76ab6b1mergedmain.Docs and changesets
The replay work updates 13 existing pages and
docs/config.json. It keeps the owner's harness, MCP, compaction, migration, and example docs. Examples cover server and client code where required. OpenAI text examples usegpt-5.5and do not use type-assertion casts.The tool-approval page explains same-run legacy support and the trusted context required for cross-run resume. The adapter-switching page explains the separate card view and conservative host preservation. Only those content edits move their
updatedAtdates to October 5. Other content dates and alladdedAtdates stay unchanged.The OpenAI page now describes API-specific
modelOptionsand reasoning support accurately. That factual correction does not change its date.The replay changeset covers 14 runtime packages: three minor releases and eleven patch releases. It includes the
@tanstack/ai-utilspatch for the actual own-key fix. The owner phase, media, reasoning, compaction, MCP, and latest shared-thread/auth/run/CLI/view changesets remain. The owner input-sender changeset also releases the MCP cache-key correction.This batch adds
docs/harness/usage.md,fork-and-reset.md, andthread-settings.md, and a resumable-agents section insubagents.md. It links them from five neighbor pages. Five new changesets cover usage, resumable agents (with a@tanstack/aiminor forctx.step), thread settings and fork, reset, and the recovered-resolve fix.The thinking shape batch updates
docs/adapters/anthropic.md,docs/chat/reasoning.md, anddocs/models/catalog.md, and adds theanthropic-thinking-shapechangeset (@tanstack/ai,@tanstack/ai-anthropic,@tanstack/ai-modelsminor).✅ Checklist
pnpm run test:pr, or these tests do not apply to this pull request.docs/for this change, or this change is not user-facing.pnpm changeset), or this PR does not change a published package.Docs, changesets, and final source review are complete. The full local-test checkbox stays unchecked because the canonical command failed. The user explicitly authorized push despite the outstanding checks.
🚀 Release Impact
Published runtime packages change. Reviewed changesets are present. Branch CI remains pending for the pushed SHA.
Testing
Commands run. The frozen
pnpm@11.9.0install passed after owner/main integration. Focused regressions reached production code and failed before the fixes. The final focused run passed 32 selected cases across eight native commands: 22 JSON cases, five OpenRouter compatibility controls, and five host cases.The first eight-package type check failed only in persistence. Indexed selector reads inferred
number | undefined. The selector repair preserves its existing index-or--1contract. No caller sentinel changed.pnpm exec nx run-many --target=test:types --projects=@tanstack/ai-persistence,@tanstack/ai-openrouter --parallel=1 --outputStyle=stream --verboseThe follow-up cases overlap the earlier 32 selected cases. Their counts are not additive. Focused results do not replace the full repository gate.
pnpm test:prtest:docsandtest:kiiraoutputUsage, settings, fork, reset, and resumable agents. These checks ran on the merged batch, before the rebase onto
e677b7516:@tanstack/ai: build,vitest run,test:types,test:oxlint,test:build@tanstack/ai-harnessvitest run(one worker)build-stubtimeout under load, and that file passed alone (8 of 8).thread-settings.test.tsafter theexpose.settingsgatetsc --noEmitfor the harness exited 0.kiira checkon the changed docs pagesAfter the rebase onto
e677b7516, no local check ran, at the user's request. The E2E specs are written, and branch CI runs them.Main merge and data-loss fixes. These checks ran on
f02f0ddcd, after the rebase onto6f3ea4ad6:test:libforai,ai-harness,ai-compaction,ai-mcptest:typesfor those four packages andtesting/e2etest:oxlintfor those four packagestesting/e2etest:typesand oxfmt onf07b75e0aOn
f07b75e0a, branch CI passed E2E, Bun, and Preview. Test and Coverage failed in five adapter suites that still expected''andnullto be rejected.809259930fixes them. After that fix,test:lib,test:types, andtest:oxlintpassed locally foropenai-base,ai-ollama,ai-gemini,ai-mistral, andai-openrouter(327, 70, 416, 102, and 303 tests).On
809259930, Test passed and Coverage failed:ai-compactionfunctions dropped from 100% to 97.43% (76 of 78).15868e693adds a test for each of the 2 functions. Each test fails when its function is broken.ai-compactionpasses 107 of 107 tests, with types, lint, and format clean.Local E2E did not run, at the user's request. Branch CI runs it.
Adapters from catalog records. These checks ran on
c0d7bfe21:test:libfor the changed packagesai2499,ai-anthropic264,ai-openai434,ai-gemini423,ai-bedrock166,ai-mistral107,ai-cloudflare42,ai-models33,openai-base330. No failures.test:libfor every dependent of@tanstack/aiandopenai-base@google/genaiinstall, and an unbuilt example dependency.test:types,test:oxlintfor the changed packages andtesting/e2etest:docs,kiira checkLocal E2E did not run. The new spec
adapter-config-reasoning-wire.spec.tsruns in branch CI.Anthropic thinking shape. These checks ran on
57ea69fed:vitest runforai,ai-anthropic,ai-modelstsc --noEmitandoxlint --type-awarefor those packagestest:docs,kiira check,sherif,knipforceAdaptiveThinkingequal on 291 of 291 records.supportsMidConvoEfforton the same 5 models.Local E2E did not run.
adapter-config-reasoning-wire.spec.tsnow checks four record configs, and branch CI runs it.On
6656c1d71, branch CI passed Test, E2E (after one rerun ofgeneration-persistence-resume.spec.ts, a reload race frommain), Bun, and Preview. Coverage failed only onai-isolate-e2b. The new e2b test passed 5 of 5 local runs (29 of 29 tests), withtscandoxlintclean.Manual test.
pnpm --filter @tanstack/ai-persistence run test:lib --run --maxWorkers=1 tests/subagent-persistence-cards.test.ts tests/subagent-reconstruct.test.ts --no-file-parallelism.pnpm --filter @tanstack/ai-openrouter run test:lib --run --maxWorkers=1 tests/replay-parity.test.ts --no-file-parallelism.pnpm --filter @tanstack/ai-harness exec vitest run tests/usage.test.ts tests/agent-resume.test.ts tests/thread-settings.test.ts tests/fork.test.ts tests/reset.test.ts. They cover the five new features.pnpm --filter @tanstack/ai-anthropic exec vitest run tests/reasoning-request.test.ts tests/mid-conversation-effort.test.ts tests/mid-conversation-changes.test.ts. They cover the thinking shape, the two-turn effort prefix, and the channel object.How this PR makes testing easy. Public tool-boundary tests cover final validation and approval. Host tests cover ordinary text, ambiguous markers, metadata, and occupied IDs. Adapter suites cover request and stream boundaries. Browser fixtures cover saved harness replay, client phases, streamed errors, and local cancellation. The owner's durable compaction and MCP option fixtures remain registered. The new batch adds
usage.test.ts,agent-resume.test.ts,thread-settings.test.ts,fork.test.ts, andreset.test.tsinai-harness,agent-step.test.tsinai, and the E2E specharness-thread-controls.spec.ts.The full canonical run was manually stopped after known failures. Actual root exit1 and captured-only closure are verified. Failures were five lint findings, a Solid test collection error, and a baseline SBX pre-kill detector failure. Their current failure evidence stays separate from the parity controls. The user explicitly authorizes commit/push despite outstanding local checks. All further local checks and the Solid config repair are cancelled. The two lint repairs were source-reviewed but not rerun. The eight unchanged live Docker/SBX files and all browser E2E checks are left for branch CI. Its results remain pending until the new pushed head exists.
Limits. Provider HTTP mocks prove request and response handling. They do not prove live-provider credentials or service behavior. Browser fixtures define the UI checks for branch CI. No current browser pass is claimed. Source review of TypeBox's no-eval route does not prove a Worker deployment. Local verification uses the installed native Node 24.3 executable, not the repository's Node 24.8 pin. Coverage remains CI-only.
Risk / rollback
A
reasoningconfig wins over the data of an adapter, so a wrong record sends wrong thinking fields. Without the config, nothing changes. Long-context tiers raisemodelCostfor large inputs, as the providers bill them. A record config now also picks adaptive or budget thinking. WithmidConversationEffort, every request thinks adaptively with a fixedoutput_config.effort: 'high'and sends notemperature.Validation can reject input that previously reached a tool. Edited approval input is checked again before dispatch. Missing trusted context rejects cross-run legacy phases and keeps the client phase pending.
Replay changes provider input and preserves the saved transcript. Genuine response identity depends on what the provider reports. Adapters without a generation ID omit it.
Owner limitations remain visible. The stdio MCP loop lacks an automated end-to-end test. The owner report names three compaction limits. This PR now fixes the empty summary and the late host record. Reporting before the durable write is still open.
The new batch is opt-in, with two exceptions. First, a host without a log now saves the transcript before each model call, which is one more write per call. Second,
snapshot()has a newusagefield. A clientconfigureinput is refused unlessexpose.settingslists its fields.Revert the final merge commit to undo the combined change. Opt-in host, gateway, and compaction settings can also be disabled.
Public API change
The owner stack adds the harness, CLI, dashboard, model catalog, and MCP server interfaces described above. Existing adapter factories remain usable. The replay work adds these approved caller surfaces:
MessageSourcehasprovider,api, andmodel. Optional message and saved-run fields carry source, response identity, and failed or aborted status.TextAdapterexposes optional readonlyapiandprovider.RunFinishedEventadds optionalresponseId.StructuredOutputResultadds optionalresponseIdandmodel.AnthropicClientConfigextends the SDK options with optionalapiKey; the SDK suppliesauthToken.AnthropicTextConfigaddsoauth,allowEmptySignature, andprovider. Existing factory and injected-client forms remain.azureOpenaiText,AzureOpenAITextAdapter, andAzureOpenAITextConfig. Configuration coversapiKey,baseURL,resourceName,apiVersion,deploymentName, anddeploymentNameMap.session.usage(),snapshot().usage(SessionUsage,UsageCounts), and theharness.usageevent.AgentStartOptions.resume,ctx.step.do(name, fn)in agent code, and optionalSubagentBinding.stepin@tanstack/ai.defineHarness({ models, expose: { settings } }),session.configure(),session.settings(),client.configure(), and theconfigureinput op.host.fork(harness, { threadId, newThreadId, at?, principal? }),session.reset(note?),client.reset(), and theresetinput op.reasoning?: ModelReasoningon the first-party text adapter configs,ConfigReasoningin@tanstack/ai, and the*ModelIdand*TextAdapterFortypes of each adapter package.modelReasoning(record),ModelReasoning,ModelCostRates.tiers,ModelCostTier, andTokenCounts.cacheWrite1hin@tanstack/ai-models.PromptTokensDetails.cacheWrite1hTokens, andresponseItemsonTanStackMessageMetadataandTanStackRunMetadata.ModelReasoning.adaptiveandModelReasoning.midConversationEffort(in@tanstack/aiand@tanstack/ai-models),ModelCompat.supportsMidConvoEffort, andreasoningEffortonTanStackMessageMetadataandTanStackRunMetadata.midConversationChannelsonAnthropicTextConfigalso takes{ tools?: boolean; systemPrompts?: boolean }.@tanstack/ai/adapter-internalsexportsReplayMessages,ReplayToolIdRule,transformMessagesForReplay,hashToolCallId,sanitizeUnicode, andsanitizeJsonArguments. These support adapter conversion; they add no application callback.Before
After: Azure deployment
Other caller examples remain in the adapter, tool, harness, and migration docs. No migration API or private-key editing procedure is added.
The latest owner P0–P14, shared-thread, MCP, compaction, auth, CLI, and run-route scope stays in this description. Historical owner checks are not presented as tests of this final merged candidate.
Pushed head:
618d71121. It adds the Anthropic thinking shape batch (a0eb57968,57ea69fed), the Gemma drift fix (6656c1d71), and the e2b coverage test (618d71121) on top ofef76ab6b1, with plain pushes, no force. Branch CI passed every check on this head, Coverage included.