Skip to content

feat(ai, ai-harness): add the harness stack with replay and adapter parity - #1555

Open
AlemTuzlak wants to merge 314 commits into
mainfrom
feat/harness-p14-mcp-server
Open

AlemTuzlak wants to merge 314 commits into
mainfrom
feat/harness-p14-mcp-server

Conversation

@AlemTuzlak

@AlemTuzlak AlemTuzlak commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Review the combined harness and replay stack in this PR. The harness keeps one agent conversation open across turns, restarts, and model changes. This combined stack adds the harness, CLI, dashboard, MCP in both directions, durable sessions, host hooks, and compaction. It also keeps saved history usable across adapters and checks final tool input before execution or client dispatch.

🎯 Changes

Compatibility changes

Area Caller impact
Reasoning The owner stack replaces provider reasoning fields in modelOptions with chat({ reasoning }). The migration guide remains part of this PR.
Default execution and caching Server tools run in parallel. Dependent tools need sequential: true or toolExecution: 'sequential'. Prompt caching is on by default. Use promptCache: 'none' to disable it.
Usage and saved history Anthropic, Bedrock, and Claude Code report total input in promptTokens. Assistant block order can produce more AG-UI rows. Source metadata separates the requested model from the response model.
Host and compaction Host joins stop at a refused or cancelled steer. Old fold checkpoints rebuild once. Compaction events add reason, and withCompaction adds compactNext. Durable turns keep LogRecordsCapability.
Model ids The first-party factories take any model id string, so a misspelled or retired id is no longer a type error. It fails at the provider instead.
Tool validation and durable resume Final validation follows middleware. Invalid input can reject a call that previously ran. Older same-run phases work when saved approval and tool bindings match. Cross-run phases need trusted saved approval context. Without it, resume is rejected and the client phase stays pending. Malformed context is rejected in both cases.

Saved subagent cards use a separate view of the transcript. Ordinary text and ambiguous hosts stay in that view. Both host creators preserve unrelated metadata. They choose an unused ID for missing IDs and explicit duplicate IDs, while keeping a unique explicit ID.

Replay changes provider input. It keeps the saved transcript unchanged. History without source metadata retains its legacy source treatment.

Existing harness stack

This remains the single combined PR for phases P0 to P14. It replaces #1551 and the closed phase PRs. Their descriptions remain the review record.

Phase PR Behavior
P0 #1513 Subagents return values and call activities.
P1 #1515 Sessions, typed agents, and plugins.
P2 #1518 Resume crashed turns and serve sessions to clients.
P3 #1519 Commands, settings, auth, and first-party plugins.
P4 #1520 Subagent tree limits and harness children.
P5 #1521 Build artifacts, worker mode, and remote harnessText.
P6 #1522 Dashboard and runnable example.
P7 #1523 MCP connectors with browser sign-in and runtime tools.
P8 #1524 Code mode with pluggable isolates.
P9 #1538 Delegate to coding agents.
P10 #1540 Continue work until a goal judge accepts the result.
P11 #1546 Live session state for any UI.
P12 #1549 CLI with a replaceable UI.
P13 #1550 Middleware and usage for every agent run.
P14 #1554 Harness MCP server, --mcp, and /mcp on --serve.

The owner changes also keep media, provider keys, durable tool steps, host hooks, turn leases, skills, prompt caching, routing, and per-turn overrides. Thinking, text, and tool calls keep their block order. Mid-conversation tool and prompt changes preserve the cached start where the model supports them.

MCP in both directions. The client keeps connectors and browser sign-in. createHarnessMcpServer exposes chat, approvals, agents, and commands to MCP clients. The CLI serves stdio MCP or authenticated HTTP MCP. MCP inputs keep parsed template variables and per-answer MIME types.

createMCPClient keeps toolName, request timeouts, and strict toolFilter checks. Missing or repeated filtered names raise MCPToolFilterError. Tool metadata keeps titles and frozen annotations. Final input validation runs before an accepted call reaches its tool.

Compaction. withCompaction({ countTokens: 'usage' }) uses reported usage. compactNext(threadId) requests compaction at the next model call. Durable compaction writes session-log records through LogRecordsCapability. The host view rebuilds the context from those records.

Compaction reports its reason, summary usage, and errors. Summary options keep turn cuts, updates to previous summaries, and bounded tool output. These opt-in features remain part of the owner scope.

Latest owner update preserved

The eight owner commits through c9f8e5f901f82e5dcba3dbd505d7ee2ab9e110e3 stay in the stack. Shared threads run each chat input and command as the principal returned by authorize. Principal.tenantId keeps credential scopes apart. Reads use the user's key before a shared tenant key. Commands save only through their own sender's credentials. Different senders do not join a running turn by default. Client context stays untrusted; server context wins when objects merge.

The harness run route now supports useChat, transcript hydration, interrupted phases, and cursor-based rejoin after reload. Repeated run IDs keep their receipt. A conflicting sender or already-used run record returns a conflict. Invalid resolve input leaves the interrupt pending. A resolve sent as a turn finishes can start immediately after it.

Sign-in wait uses credentials.require(id, { wait: true }) inside a chat tool. The stopped sender's sign-in resumes that turn. lifetime: 'turn' names a plugin that lives for one turn; 'run' remains an alias. TurnInfo carries message, input context, principal, input ID, and overrides. A plugin may supply each turn's adapter, so HarnessConfig.adapter is optional in that case. harnessText accepts explicit inputModalities.

The view separates approvals, client tools, and sign-in requests. ClientToolCall, view.on('clientTool', ...), and its resolve/fail actions answer client tools. The CLI accepts a caller-owned host and principal, keeps that host open, and keeps its bearer authorization gate. fakeText keeps parent run identity, distinct tool IDs between adapter instances, and exact optional typing. All eight new owner changesets, shared-thread docs, sandbox-provider docs, and their navigation entries remain.

MCP sender isolation. The new per-sender connection cache now encodes tenant and user as a JSON tuple. Names that contain : cannot reuse another sender's authenticated client or tools. A new real-SDK/HTTP regression covers two names that collided before, distinct scoped credentials, separate tools, and same-sender reuse. This new regression was source-reviewed and was not run locally under the user's push instruction. The owner changeset already includes the @tanstack/ai-mcp patch.

Other owner fixes retained. The stack keeps MCP interrupt-kind handling, the Linear OAuth issuer, unique exported tool names, v2 connector credentials, background-agent lease recovery, routed handoff hooks and child resume, interrupted-session recovery, and MCP sessions needed for elicitation. Workspace approval, lazy code-mode tools, skills, Cloudflare example, model catalog, and live plugin commands remain part of the owner scope.

Usage, resumable agents, thread settings, fork, and reset

These five owner features come from the pi-durable comparison. Each one is opt-in or read-only for existing callers.

  • Usage totals. session.usage() and snapshot().usage total each model call by model and by sender, with the cost that providers report. A durable host keeps one harness.usage record per call. A harness.usage event fires on each call.
  • Resumable background agents. agents.start(agent, input, { resume: true }) needs a durable host. The host that takes over runs the agent again from its saved transcript. ctx.step.do(name, fn) keeps side effects from running twice. durability.maxAttempts caps the runs.
  • Thread settings. defineHarness({ models }) and session.configure() store the model, reasoning, instructions, tools, plugins, and working folder of a thread. A client can change only the fields in expose.settings, and none by default.
  • Fork and reset. host.fork() copies a thread up to a message into a new thread, with its settings and media. session.reset(note?) starts a fresh model context and keeps the full transcript, with a marker.
  • Fixes. A resolve that a crash interrupted continues with its routed plan, sender, and context. On a host without a log, the transcript is saved before each model call, so a tool result stays when the next model call fails. The used-up interrupt rule from e677b7516 stays as it is.

Replay and adapter parity

Handoff item Final source behavior
1. Foreign-model history Record the provider, API, and requested model. Foreign replay removes incompatible signatures and redacted thinking. Readable thinking becomes text, and tool IDs stay paired.
2. Replay cleanup Omit failed or aborted assistant batches and their results from provider input. Fill unanswered calls with No result provided errors. Keep system changes after tool results.
3. Claude thinking on Bedrock Keep reasoning text, signatures, redacted bytes, indexes, and block order. Replay signed reasoning before the matching tool call.
4. Unknown finish reasons Emit RUN_ERROR with the provider finish reason. Known content_filter behavior stays the same.
5. Images in tool results Chat Completions sends tool text followed by user images. Mistral keeps images in the tool message. Bedrock uses native tool-result image blocks.
6. Bedrock error status An error property sets status: 'error', including an empty error string.
7. Lone Unicode surrogates Remove lone surrogates from outgoing text and decoded JSON strings. Keep valid pairs, saved raw arguments, signatures, URLs, and bytes.
8. Final argument validation Validate after middleware and before execution or client dispatch. Keep Standard Schema transforms and full JSON Schema checks with coercion. Preserve own JSON keys and ignore inherited lookup entries across all ten runtime paths.
9. Response identity Expose genuine response IDs and response models where the provider supplies them. Keep the requested source model separate. Do not invent a generation ID.
10. Chat Completions content Null or missing text adds no text and still permits tool calls. Object or array text emits the specified error. Keep main's malformed-argument recovery and stream drain.
11. Anthropic Bearer and OAuth Support Bearer credentials, environment precedence, and OAuth identity and betas through the existing SDK. Bearer auth omits x-api-key.
12. Azure OpenAI Add Azure Responses support through the existing OpenAI SDK. Keep endpoint, version, credentials, and deployment-name precedence explicit. Deployment maps use own entries only.
13. Gateway thinking allowEmptySignature permits unsigned same-source thinking on configured gateways. The default remains false.
14. Empty tools with tool history Chat Completions sends tools: [] for tool history without active tools. Wrappers retain that empty list. The no-history request stays unchanged.

Final validation. Standard Schema runs first on the exact input. When a safe input schema permits a coercion retry, the authored schema checks the retry result. A successful transform keeps its value. Raw JSON Schema uses full schema checks. Pending or denied calls do not reach tool hooks, execution, or client dispatch.

The only new runtime dependency for this replay work is approved typebox at ^1.3.27, locked to 1.3.34. TypeBox uses its check engine when dynamic evaluation is unavailable. Source review confirms this fallback. Focused controls check validation behavior. They do not prove a deployed Worker.

Own-key preservation covers schema copies, nullable widening maps, Standard Schema root copies, Gemini parameter sanitation, and Azure deployment lookup. OpenRouter keeps raw arguments separate from normalized input across all three terminal paths. Scalar, array, null, and missing input keep their contracts.

The host selector still returns a message index or -1. Presence guards make that return type explicit. This repairs type inference; it does not change caller sentinels or claim a runtime undefined defect.

Integration. The candidate preserves the earlier 16 owner commits and the eight newer owner commits through c9f8e5f901f82e5dcba3dbd505d7ee2ab9e110e3, then retains the approved merge with pinned main 6d8e6485f92c98a6f2200471e2aff21ab50be013. Main's cancellation and malformed-JSON controls remain. The route registry includes the owner and main routes. Model metadata keeps main's catalog and the owner's modality and reasoning policy.

Main merge and data-loss fixes

fdaeabe8f merges main at 7fb4a5f7f. That brings in #1620, live durable streams and one hydrate GET in Strict Mode. These commits fix the data-loss issues and the CI failures that showed up after the merge:

Commit Fix
581efd816 A host record from the last model call of a turn stays in the log. Before, rebase dropped it when the turn ended.
af7274d6f An empty summary fails the compaction. Before, it replaced the history with nothing. On the overflow path the turn fails with the overflow error.
4c12b1e89, e677b7516 A queued resume turn keeps its parent run. An interrupt is not offered again after its tool ran, including when a later model call fails.
bf329d5e3, 809259930 A tool with no input or null arguments runs with {} again (#265). The final input check made this fail. Other scalars and malformed JSON are still rejected. 809259930 updates the five adapter replay tests that expected the rejection.
c1e9fdad8 The OTel middleware records no execute_tool span for a client tool. The server does not run it.
15868e693 Tests for two compaction functions that had none: part text in the summary prompt, and summary usage after the turn. Coverage saw ai-compaction functions drop from 100% to 97.43%.

The test commits 71d5e933c, f02f0ddcd, and f07b75e0a keep the MCP and E2E tests true under the new behavior. Five patch changesets cover the five fixes above. The tools-test adapter now counts only the tool results of the current turn. Replay gives an unanswered call from an earlier turn a No result provided result, so the old count ended the run after Stop too early.

Adapters from catalog records

A catalog model whose id is not in the adapter's own table got no reasoning field. That was 582 of the 709 catalog reasoning models on first-party wires. Now:

  • Reasoning from the config. Each first-party text adapter takes reasoning?: ModelReasoning: Anthropic (incl. Vertex), OpenAI and Azure, Gemini (AI Studio and Vertex), Bedrock Converse, Mistral, and Cloudflare. It wins over the table for the request fields and for the levels chat({ reasoning }) takes. false sends nothing. The factories take any model id. @tanstack/ai-models adds modelReasoning(record).
  • Prices. modelCost uses the models.dev cost.tiers (long context) and prices cacheWrite1h at 2 times the input price. Usage adds promptTokensDetails.cacheWrite1hTokens, filled by Anthropic and Bedrock Converse.
  • Catalog. The google-vertex catalog leaves out the Claude models, which had the Gemini wire.
  • Responses replay. Each answer item keeps its id and phase in metadata.tanstack.responseItems, and a same-model replay sends them back.

The remaining 66 models with no reasoning field are Bedrock Converse models that are not Claude. pi also sends Converse thinking fields only for Claude.

This batch also merges main two times, at 8ce7ad97c and a32782c3e. The merges keep the activity messages, video persistence, stream backpressure, and the Anthropic max_tokens usage, replay, and Sonnet 5.5 fixes from main. Activity chunks no longer get the assistant call metadata.

Anthropic thinking shape from the record

Flue compared the Anthropic thinking fields with pi's for every anthropic-messages reasoning record. 195 records were different. Gateway ids with dots (anthropic/claude-opus-4.7) got budget thinking, older Claude on OpenRouter got adaptive thinking, and most other models on the wire got adaptive thinking. Now:

  • Adaptive or budget from the record. ModelReasoning.adaptive picks the shape. modelReasoning(record) sets it from compat.forceAdaptiveThinking, and the catalog flag now equals pi's on all 291 shared records. Adaptive thinking without a level map sends pi's default effort.
  • Mid-conversation effort. ModelReasoning.midConversationEffort (from the new compat.supportsMidConvoEffort, pi's 5 models) puts the level into the messages, so a level change keeps the cached start. Each answer keeps its level in metadata.tanstack.reasoningEffort.
  • One channel at a time. midConversationChannels takes { tools?, systemPrompts? }.
  • Catalog data. Seven Fireworks records take Fireworks' own levels (pi), and Vercel google/gemma-4-31b-it reasons. @tanstack/ai-vercel-gateway gets the same Gemma fix (6656c1d71), so the reasoning drift test agrees. Vercel's own model list tags it reasoning.
  • Coverage. 618d71121 adds an ai-isolate-e2b test for stderr without an out-of-memory report. That branch side was hit only when V8's first GC block came in its own stderr chunk, so Linux CI missed it on this PR.

The local copy of the Flue check finds only group E (top-level effort on Claude 4.6, accepted in the 2026-09-30 report). Before this batch, 53fab0187 brought the harness-cli-batman lockfile entry in line with harness-cli, and 280263df4 and ef76ab6b1 merged main.

Docs and changesets

The replay work updates 13 existing pages and docs/config.json. It keeps the owner's harness, MCP, compaction, migration, and example docs. Examples cover server and client code where required. OpenAI text examples use gpt-5.5 and do not use type-assertion casts.

The tool-approval page explains same-run legacy support and the trusted context required for cross-run resume. The adapter-switching page explains the separate card view and conservative host preservation. Only those content edits move their updatedAt dates to October 5. Other content dates and all addedAt dates stay unchanged.

The OpenAI page now describes API-specific modelOptions and reasoning support accurately. That factual correction does not change its date.

The replay changeset covers 14 runtime packages: three minor releases and eleven patch releases. It includes the @tanstack/ai-utils patch for the actual own-key fix. The owner phase, media, reasoning, compaction, MCP, and latest shared-thread/auth/run/CLI/view changesets remain. The owner input-sender changeset also releases the MCP cache-key correction.

This batch adds docs/harness/usage.md, fork-and-reset.md, and thread-settings.md, and a resumable-agents section in subagents.md. It links them from five neighbor pages. Five new changesets cover usage, resumable agents (with a @tanstack/ai minor for ctx.step), thread settings and fork, reset, and the recovered-resolve fix.

The thinking shape batch updates docs/adapters/anthropic.md, docs/chat/reasoning.md, and docs/models/catalog.md, and adds the anthropic-thinking-shape changeset (@tanstack/ai, @tanstack/ai-anthropic, @tanstack/ai-models minor).

✅ Checklist

  • I have followed the steps in the Contributing guide.
  • I have tested code changes locally with pnpm run test:pr, or these tests do not apply to this pull request.
  • I fully understand the code in this pull request, including any code generated with AI assistance.
  • Docs: I updated docs/ for this change, or this change is not user-facing.
  • Changeset: I added a changeset (pnpm changeset), or this PR does not change a published package.

Docs, changesets, and final source review are complete. The full local-test checkbox stays unchecked because the canonical command failed. The user explicitly authorized push despite the outstanding checks.

🚀 Release Impact

  • This change affects published code, and I have generated a changeset.
  • This change is docs/CI/dev-only (no release).

Published runtime packages change. Reviewed changesets are present. Branch CI remains pending for the pushed SHA.

Testing

Commands run. The frozen pnpm@11.9.0 install passed after owner/main integration. Focused regressions reached production code and failed before the fixes. The final focused run passed 32 selected cases across eight native commands: 22 JSON cases, five OpenRouter compatibility controls, and five host cases.

The first eight-package type check failed only in persistence. Indexed selector reads inferred number | undefined. The selector repair preserves its existing index-or--1 contract. No caller sentinel changed.

Current focused evidence Result
Persistence's full host-card and reconstruction files 30 of 30 passed. This includes the positive no-host transfer control.
Strengthened OpenRouter replay controls 12 selected cases passed. Exact serialized nested-null checks cover each required terminal path.
pnpm exec nx run-many --target=test:types --projects=@tanstack/ai-persistence,@tanstack/ai-openrouter --parallel=1 --outputStyle=stream --verbose Native exit 0 after the selector repair. Seven packages passed in the earlier eight-package run.

The follow-up cases overlap the earlier 32 selected cases. Their counts are not additive. Focused results do not replace the full repository gate.

Required final evidence State
Nine existing paths reproduced on clean pinned main, with ordinary-input controls Complete: 25 intended failures and nine ordinary passing controls on pinned main. All protected runtime bytes stayed unchanged. The first Base fixture error is excluded; its corrected run supplies valid proof.
Same-bytes original repro on the integrated candidate Complete: exact5525 bytes passed3/3, native exit0; archived exact hash and feature source removed.
Exact fresh-main SDK/HTTP fixture bytes on the feature candidate Complete: Base2 + Mistral2 + OpenRouter12 passed. Hash-guard removal verified for all three temporary copies.
Full literal pnpm test:pr Known failures; manually stopped with actual root exit1 and captured-only closure verified. Includes live provider tests. No full canonical green or unfinished-target pass is claimed.
Focused local normal checks through existing scripts/Nx Cancelled by explicit user push override. Source-reviewed command packet was not run. Two lint repairs were not retested; Solid config is unchanged.
Fresh final test:docs and test:kiira output Further local checks cancelled by explicit user push override; branch CI pending.
E2E branch CI and the dedicated replay/validation browser sources Local E2E waived at the user's request. Review branch CI for the final pushed head; no current browser pass is claimed.
Final combined owner-c9/MCP regression and full harness checks No further local run at the user's request. Branch CI pending for the final pushed SHA.

Usage, settings, fork, reset, and resumable agents. These checks ran on the merged batch, before the rebase onto e677b7516:

Check Result
@tanstack/ai: build, vitest run, test:types, test:oxlint, test:build 2404 of 2404 tests passed. Every other step exited 0.
@tanstack/ai-harness vitest run (one worker) 830 of 831 passed. The one failure was a build-stub timeout under load, and that file passed alone (8 of 8).
thread-settings.test.ts after the expose.settings gate 15 of 15 passed. tsc --noEmit for the harness exited 0.
kiira check on the changed docs pages No errors.

After the rebase onto e677b7516, no local check ran, at the user's request. The E2E specs are written, and branch CI runs them.

Main merge and data-loss fixes. These checks ran on f02f0ddcd, after the rebase onto 6f3ea4ad6:

Check Result
test:lib for ai, ai-harness, ai-compaction, ai-mcp 2409, 835, 105, and 413 tests passed. No failures.
test:types for those four packages and testing/e2e All exited 0.
test:oxlint for those four packages Exit 0. Only the warnings that were already there.
testing/e2e test:types and oxfmt on f07b75e0a Exit 0.

On f07b75e0a, branch CI passed E2E, Bun, and Preview. Test and Coverage failed in five adapter suites that still expected '' and null to be rejected. 809259930 fixes them. After that fix, test:lib, test:types, and test:oxlint passed locally for openai-base, ai-ollama, ai-gemini, ai-mistral, and ai-openrouter (327, 70, 416, 102, and 303 tests).

On 809259930, Test passed and Coverage failed: ai-compaction functions dropped from 100% to 97.43% (76 of 78). 15868e693 adds a test for each of the 2 functions. Each test fails when its function is broken. ai-compaction passes 107 of 107 tests, with types, lint, and format clean.

Local E2E did not run, at the user's request. Branch CI runs it.

Adapters from catalog records. These checks ran on c0d7bfe21:

Check Result
test:lib for the changed packages ai 2499, ai-anthropic 264, ai-openai 434, ai-gemini 423, ai-bedrock 166, ai-mistral 107, ai-cloudflare 42, ai-models 33, openai-base 330. No failures.
test:lib for every dependent of @tanstack/ai and openai-base Pass, except local environment cases: live Docker sandbox tests, Windows-only Solid and harness build tests, a duplicate @google/genai install, and an unbuilt example dependency.
test:types, test:oxlint for the changed packages and testing/e2e Exit 0.
test:docs, kiira check No broken links. 1806 snippets passed.
Catalog wire check (every reasoning record through the thinking builder of its adapter) 66 of 709 send no field, all Bedrock models that are not Claude. Before: 582.

Local E2E did not run. The new spec adapter-config-reasoning-wire.spec.ts runs in branch CI.

Anthropic thinking shape. These checks ran on 57ea69fed:

Check Result
vitest run for ai, ai-anthropic, ai-models 2499, 273, and 36 tests passed. No failures.
tsc --noEmit and oxlint --type-aware for those packages Exit 0. Only the warnings that were already there.
test:docs, kiira check, sherif, knip No broken links. 1807 snippets passed. No new findings.
Local copy of the Flue shape check (1746 record and level pairs) Only group E is different (24 pairs on 4 records).
Catalog flags against the pi snapshot forceAdaptiveThinking equal on 291 of 291 records. supportsMidConvoEffort on the same 5 models.

Local E2E did not run. adapter-config-reasoning-wire.spec.ts now checks four record configs, and branch CI runs it.

On 6656c1d71, branch CI passed Test, E2E (after one rerun of generation-persistence-resume.spec.ts, a reload race from main), Bun, and Preview. Coverage failed only on ai-isolate-e2b. The new e2b test passed 5 of 5 local runs (29 of 29 tests), with tsc and oxlint clean.

Manual test.

  1. Run pnpm --filter @tanstack/ai-persistence run test:lib --run --maxWorkers=1 tests/subagent-persistence-cards.test.ts tests/subagent-reconstruct.test.ts --no-file-parallelism.
  2. Run pnpm --filter @tanstack/ai-openrouter run test:lib --run --maxWorkers=1 tests/replay-parity.test.ts --no-file-parallelism.
  3. Review the final branch's E2E CI. Local E2E is not run at the user's request.
  4. Read the adapter-switching and tool-approval docs for saved-history and approval compatibility.
  5. Run pnpm --filter @tanstack/ai-harness exec vitest run tests/usage.test.ts tests/agent-resume.test.ts tests/thread-settings.test.ts tests/fork.test.ts tests/reset.test.ts. They cover the five new features.
  6. Run pnpm --filter @tanstack/ai-anthropic exec vitest run tests/reasoning-request.test.ts tests/mid-conversation-effort.test.ts tests/mid-conversation-changes.test.ts. They cover the thinking shape, the two-turn effort prefix, and the channel object.

How this PR makes testing easy. Public tool-boundary tests cover final validation and approval. Host tests cover ordinary text, ambiguous markers, metadata, and occupied IDs. Adapter suites cover request and stream boundaries. Browser fixtures cover saved harness replay, client phases, streamed errors, and local cancellation. The owner's durable compaction and MCP option fixtures remain registered. The new batch adds usage.test.ts, agent-resume.test.ts, thread-settings.test.ts, fork.test.ts, and reset.test.ts in ai-harness, agent-step.test.ts in ai, and the E2E spec harness-thread-controls.spec.ts.

The full canonical run was manually stopped after known failures. Actual root exit1 and captured-only closure are verified. Failures were five lint findings, a Solid test collection error, and a baseline SBX pre-kill detector failure. Their current failure evidence stays separate from the parity controls. The user explicitly authorizes commit/push despite outstanding local checks. All further local checks and the Solid config repair are cancelled. The two lint repairs were source-reviewed but not rerun. The eight unchanged live Docker/SBX files and all browser E2E checks are left for branch CI. Its results remain pending until the new pushed head exists.

Limits. Provider HTTP mocks prove request and response handling. They do not prove live-provider credentials or service behavior. Browser fixtures define the UI checks for branch CI. No current browser pass is claimed. Source review of TypeBox's no-eval route does not prove a Worker deployment. Local verification uses the installed native Node 24.3 executable, not the repository's Node 24.8 pin. Coverage remains CI-only.

Risk / rollback

A reasoning config wins over the data of an adapter, so a wrong record sends wrong thinking fields. Without the config, nothing changes. Long-context tiers raise modelCost for large inputs, as the providers bill them. A record config now also picks adaptive or budget thinking. With midConversationEffort, every request thinks adaptively with a fixed output_config.effort: 'high' and sends no temperature.

Validation can reject input that previously reached a tool. Edited approval input is checked again before dispatch. Missing trusted context rejects cross-run legacy phases and keeps the client phase pending.

Replay changes provider input and preserves the saved transcript. Genuine response identity depends on what the provider reports. Adapters without a generation ID omit it.

Owner limitations remain visible. The stdio MCP loop lacks an automated end-to-end test. The owner report names three compaction limits. This PR now fixes the empty summary and the late host record. Reporting before the durable write is still open.

The new batch is opt-in, with two exceptions. First, a host without a log now saves the transcript before each model call, which is one more write per call. Second, snapshot() has a new usage field. A client configure input is refused unless expose.settings lists its fields.

Revert the final merge commit to undo the combined change. Opt-in host, gateway, and compaction settings can also be disabled.

Public API change

The owner stack adds the harness, CLI, dashboard, model catalog, and MCP server interfaces described above. Existing adapter factories remain usable. The replay work adds these approved caller surfaces:

Surface Added contract
Source and response identity Exported MessageSource has provider, api, and model. Optional message and saved-run fields carry source, response identity, and failed or aborted status. TextAdapter exposes optional readonly api and provider. RunFinishedEvent adds optional responseId. StructuredOutputResult adds optional responseId and model.
Anthropic configuration AnthropicClientConfig extends the SDK options with optional apiKey; the SDK supplies authToken. AnthropicTextConfig adds oauth, allowEmptySignature, and provider. Existing factory and injected-client forms remain.
Azure Responses Root exports add azureOpenaiText, AzureOpenAITextAdapter, and AzureOpenAITextConfig. Configuration covers apiKey, baseURL, resourceName, apiVersion, deploymentName, and deploymentNameMap.
Usage totals session.usage(), snapshot().usage (SessionUsage, UsageCounts), and the harness.usage event.
Resumable agents AgentStartOptions.resume, ctx.step.do(name, fn) in agent code, and optional SubagentBinding.step in @tanstack/ai.
Thread settings defineHarness({ models, expose: { settings } }), session.configure(), session.settings(), client.configure(), and the configure input op.
Fork and reset host.fork(harness, { threadId, newThreadId, at?, principal? }), session.reset(note?), client.reset(), and the reset input op.
Reasoning config reasoning?: ModelReasoning on the first-party text adapter configs, ConfigReasoning in @tanstack/ai, and the *ModelId and *TextAdapterFor types of each adapter package.
Catalog and cost modelReasoning(record), ModelReasoning, ModelCostRates.tiers, ModelCostTier, and TokenCounts.cacheWrite1h in @tanstack/ai-models.
Usage and replay PromptTokensDetails.cacheWrite1hTokens, and responseItems on TanStackMessageMetadata and TanStackRunMetadata.
Thinking shape ModelReasoning.adaptive and ModelReasoning.midConversationEffort (in @tanstack/ai and @tanstack/ai-models), ModelCompat.supportsMidConvoEffort, and reasoningEffort on TanStackMessageMetadata and TanStackRunMetadata.
Anthropic channels midConversationChannels on AnthropicTextConfig also takes { tools?: boolean; systemPrompts?: boolean }.
Adapter utilities @tanstack/ai/adapter-internals exports ReplayMessages, ReplayToolIdRule, transformMessagesForReplay, hashToolCallId, sanitizeUnicode, and sanitizeJsonArguments. These support adapter conversion; they add no application callback.

Before

import { chat } from '@tanstack/ai'
import { openaiText } from '@tanstack/ai-openai'

const stream = chat({
  adapter: openaiText('gpt-5.5'),
  messages: [{ role: 'user', content: 'Hello!' }],
})

for await (const chunk of stream) {
  if (chunk.type === 'TEXT_MESSAGE_CONTENT') console.log(chunk.delta)
}

After: Azure deployment

import { chat } from '@tanstack/ai'
import { azureOpenaiText } from '@tanstack/ai-openai'

const stream = chat({
  adapter: azureOpenaiText('gpt-5.5', {
    resourceName: 'my-resource',
    apiKey: process.env.AZURE_OPENAI_API_KEY,
    apiVersion: 'v1',
    deploymentNameMap: { 'gpt-5.5': 'production-chat' },
  }),
  messages: [{ role: 'user', content: 'Hello!' }],
})

for await (const chunk of stream) {
  if (chunk.type === 'TEXT_MESSAGE_CONTENT') console.log(chunk.delta)
}

Other caller examples remain in the adapter, tool, harness, and migration docs. No migration API or private-key editing procedure is added.

The latest owner P0–P14, shared-thread, MCP, compaction, auth, CLI, and run-route scope stays in this description. Historical owner checks are not presented as tests of this final merged candidate.

Pushed head: 618d71121. It adds the Anthropic thinking shape batch (a0eb57968, 57ea69fed), the Gemma drift fix (6656c1d71), and the e2b coverage test (618d71121) on top of ef76ab6b1, with plain pushes, no force. Branch CI passed every check on this head, Coverage included.

…pauses or stops

A message typed during the last model call of a goal turn runs first, as a
steer that came too late, and pauses the goal. The goal's next turn was already
queued and still ran once. ctx.session.prompt() now returns the turn, and the
goal plugin cancels its queued turn when a user message pauses the goal, on
/goal stop, and when a new goal replaces it.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
… clock

The timeout tests use a 1 ms budget, so the branch they end in depends on how
many milliseconds pass. The coverage number moved between runs. New tests fix
the clock and reach each of the five timeout exits every run.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…session for any UI

Adds @tanstack/store ^0.11.1 and the browser-safe ./view export. Three examples
listed @tanstack/store ^0.8.0 without importing it; they move to ^0.11.1 so the
workspace keeps one version (sherif).

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…any TUI on a session view

The CLI no longer depends on ink or react. An interactive terminal uses line
mode unless runCli gets a ui function, which receives a ready session view and
resolves when the user quits.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…ove the old event fold

In a terminal, line mode opens sign-in links in the browser.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
The coverage gate failed because the new session view code had paths
that no test ran. These tests run them.

- view-paths.test.ts: a fake SessionViewSource forces each path of
  createSessionView. It covers failed reads at start, a failed or ended
  event stream, connection state, failed actions, approvals, on()
  handlers, the snapshot read coalesce, and dispose.
- view-reduce.test.ts: the reducer branches that return the same
  state, nested and sibling child agents, content-part tool results,
  run errors, and transcript edge cases.
- client-reads.test.ts: the transcript and describe routes answer 400
  and 403, and a failed read throws with its route and status.
- goal.test.ts: selectGoal gives null for state that is not a goal.
- view-fixtures.ts: event and snapshot helpers that the view tests share.

No production code changed.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Every direct TanStack Store dependency in the workspace is now on the latest
release (store, react-store, and solid-store 0.11.1).

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
… agent

@tanstack/ai:
- The chat middleware context has subagentName and parentSubagentRunId.
  chat() takes both as options. The bound ctx.chat sets them.
- A child that calls ctx.chat({ subagents }) gives its own children the
  host chatMiddleware and generationMiddleware before their own, also
  when the binding has no budget.
- A nested child gets parentSubagentRunId on its run input. Its
  SUBAGENT_STARTED does not carry it in the stream of the chat that
  started it: attributeChunk adds it one level up, as before.

@tanstack/ai-harness:
- Plugins can contribute agentMiddleware: chat middleware for subagents,
  background agents, and their children. Not for the lead turn.
- The lead turn passes run plugins' agentMiddleware and
  generationMiddleware to its agents too.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Add a "Middleware in every agent" section to the plugins page, with a
usage tracker that puts one middleware in middleware and
agentMiddleware. Link it from the harness subagents page. The chat
subagents page names the new subagentName and parentSubagentRunId
context fields.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
The usage counter now sits in agentMiddleware too, so /usage covers the lead
turn, subagents, background agents, and their children.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
The final tool input check kept a missing or literal null argument and
checked it against the schema, so a tool with no required fields did not
run when the model sent an empty tool_use block. That broke the fix for
issue #265, and its E2E test failed.

No input is {} again. A null that the schema rejects is checked as {}.
Other values, such as 2 or [1,2], must still fit the schema.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
The final tool input check now runs onBeforeToolCall before a client
tool is dispatched, so otelMiddleware opened an execute_tool span for a
tool the server does not run. The client-tool wait E2E test then saw two
root spans.

otelMiddleware skips the span for a known tool with no execute. The
unit test is back to main's chat and iteration spans.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
The client-tool input error scenarios sent message: 42. The final input
check now coerces scalar types, so 42 became "42" and passed, and the
client ran the tool. The scenarios now omit the required message, which
is invalid with or without coercion.

The sandbox persistence spec now expects the runId and source that the
chat run records on the saved assistant message.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Replay now gives an unanswered tool call from an earlier turn a
"No result provided" result. The provider-free adapter counted every
tool message, so after Stop the next user turn answered with text and
ended before the race test could press Stop again.
… replay tests

The #265 fix runs a tool with no input or a literal null as {}. Five
adapter replay-parity tests still expected those inputs to be rejected.
They now check that the tool runs with {}, or that the call goes to the
client. Other scalars and malformed JSON are still rejected.

Mistral ends whitespace-only arguments with a parse error before the
input check, so that case still runs 0 times.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…turn summary usage

Two new functions had no test, so the Coverage job saw function coverage
for ai-compaction drop from 100% to 97.43%:

- contentText in conversationSummarizer, for content that is a list of
  parts.
- The addUsage callback of the after-turn check, for a summarizer that
  reports usage.

Each new test fails when its function is broken.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
@github-actions github-actions Bot added waiting-on: maintainer The ball is in the maintainers’ court merge-conflicts Conflicts with the base branch — needs a rebase waiting-on: author Waiting for the author to respond or update and removed waiting-on: author Waiting for the author to respond or update waiting-on: maintainer The ball is in the maintainers’ court labels Oct 5, 2026
AlemTuzlak and others added 13 commits October 6, 2026 12:05
Brings in main at 8ce7ad9. Conflicts and how they were solved:

- Anthropic usage: keep the branch's total-input promptTokens and add
  main's message_start fallback for counts the closing delta leaves out.
  main's new usage tests now expect the total input.
- Anthropic replay: main sends {} for invalid JSON tool arguments. The
  branch moved that code into appendToolCallBlocks, so the fix goes there.
- Anthropic options: thinking and effort are chat({ reasoning }) on this
  branch, so claude-sonnet-5-5 takes max_tokens without sampling options
  and no output_config. main's type tests and E2E route follow that.
- chat(): main records an afterModel interrupt turn with
  addTerminalAssistantMessages, which already saves the mid-conversation
  record. The old helper is gone.
- Stream processor: keep both sides' new state. serializeToolInput is no
  longer used; main's parseToolArguments stays.
- Route tree: the branch's tree plus main's seven new routes. main's tree
  imported the Sonnet 5.5 route from the wrong file.
- Lockfile: the branch's lockfile plus main's entry for
  @tanstack/ai-isolate-e2b. A text merge moved zod for some packages.
- Docs and dates: keep the newer content and the newer updatedAt.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
A catalog model whose id is not in the adapter's own table got no
reasoning at all, because each adapter looked up its generated table by
model id. About 500 catalog models lost their thinking this way.

Each first-party text adapter now takes `reasoning?: ModelReasoning` in
its config. When it is set, it wins over the table for the request
fields and for the levels `chat({ reasoning })` takes. `false` sends no
reasoning field. When it is not set, nothing changes.

- Anthropic (anthropicText, createAnthropicChat, anthropicVertexText)
- OpenAI (openaiText, createOpenaiChat) and azureOpenaiText
- Gemini (geminiText, createGeminiChat), AI Studio and Vertex mode
- Bedrock createBedrockConverse
- Mistral (mistralText, createMistralText)
- Cloudflare (cloudflareText, createCloudflareText)

The factories take any model id string. A config with `reasoning` types
`chat({ reasoning })` with every level (ConfigReasoning in @tanstack/ai);
the adapter moves a level the model does not have to the nearest one.

@tanstack/ai-models adds `modelReasoning(record)`, the one rule that
turns a record into the config value.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…ogle-vertex

modelCost now follows pi's two rules that it did not have:

- Tiers. A record's `cost.tiers` holds higher prices above an input size,
  for example long context. A call uses the tier with the highest
  `inputTokensAbove` below its input (uncached + cache read + cache write).
  The tiers come from models.dev `cost.tiers`. A price that models.dev
  leaves out is 0, as in pi's catalog.
- 1-hour cache writes. `TokenCounts.cacheWrite1h` is the part of
  `cacheWrite` with a 1-hour retention. It costs 2x the input price.

Usage gets `promptTokensDetails.cacheWrite1hTokens`. The Anthropic adapter
fills it from `cache_creation.ephemeral_1h_input_tokens`, and the Bedrock
Converse adapter from the 1-hour `cacheDetails`.

The google-vertex catalog listed Claude models with the Gemini wire, which
the Gemini adapter cannot call. They are left out, as in pi.

The fetch script keeps the tiers. The models.dev snapshot gets the tiers
of today's models.dev for the models it already has. Regenerating also
picks up 11 gateway models that the snapshot already had.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Brings in main at a32782c: activity messages and the ActivityStore,
video bytes through generation persistence, stream backpressure, and the
structured output and MCP fixes. Conflicts and how they were solved:

- Stream processor: keep main's activity-id guards. The branch already
  moves the structured-output mark in remapProvisionalMessage, which is
  main's #1634 fix. The thinking-step lookup skips activity messages.
- chat(): keep both new methods; main sets the activities first, then
  the branch's providerMessages rule. Activity chunks no longer get the
  assistant call metadata, so stored activity records keep their own.
- Persistence: keep the branch's streaming-id tracking and main's
  activity save after each snapshot.
- Sandbox test: keep the branch's MemorySnapshotPersistence type.
- @tanstack/ai: add main's fast-json-patch next to the branch's typebox.
  The lockfile is the branch's plus main's fast-json-patch entries.
- Route tree: the branch's tree plus main's two activity routes.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
OpenAI Responses gives each answer item an id and a phase (commentary or
final_answer). The adapter kept item ids only for tool calls and
reasoning, so a replayed answer lost its phase.

The adapter now puts each message item's id and phase on RUN_FINISHED,
and chat() keeps them on the assistant message in
`metadata.tanstack.responseItems`. A same-model replay sends each text
block back as its item, with the id and the phase. Another model, or a
message whose blocks do not match its items, gets the plain text.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…ules

- Reasoning: a model the adapter does not list takes the `reasoning`
  config, with the factories that accept it.
- Model catalog: build an adapter from a record with modelReasoning, the
  adapter for each `api`, and the tier and 1-hour cache write prices.
- Prompt caching: usage reports `cacheWrite1hTokens`.
- OpenAI: answer items keep their id and phase on the next request.

Changesets for the reasoning config, the cost rules, and the Vertex
Claude records. An E2E wire test checks that a gateway id outside the
Anthropic list gets its thinking budget from the config, and none with
`reasoning: false`.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
The foreign *_CHUNK spec from main expected only the sender's metadata
on a message that ends in RUN_ERROR. This branch also marks that message
with `tanstack.stopReason: 'error'`, so a replay leaves the failed batch
out. The sender's metadata is still there.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Add examples/harness-cli-batman as a themed copy of the harness CLI.
It uses a bat splash, a yellow header, and a separate save folder.
Voice input times out after 45 seconds and Esc cancels it.
…kfile

The batman example's install re-resolved part of the workspace: zod moved
to 4.6.5 for @modelcontextprotocol/server, and @google/genai got a second
copy. That broke the ai-mcp type check (two zod copies in the server
types) and the ai-vertex factory tests (the mocked SDK was the other
copy).

The example has the same dependencies as examples/harness-cli, so its
lockfile entry now resolves the same way (zod 4.3.6,
@tanstack/react-store 0.11.1). The rest of the lockfile is as before the
example. A frozen install passes, and ai-mcp test:types, ai-vertex
test:lib, and the example's test:types pass.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Brings in main at 32e56bb (Version Packages #1630): the released
versions, CHANGELOGs, and the changesets that release used. No conflicts.
The lockfile is unchanged and a frozen lockfile check passes.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Flue's shape check found 195 Anthropic Messages models where our
thinking fields differ from pi's. Gateway ids with dots
(anthropic/claude-opus-4.7) got budget thinking, older Claude on
OpenRouter got adaptive thinking, and most other models on the wire got
adaptive thinking because the adapter sent adaptive whenever a model had
no token budget.

Adaptive or budget from the record:
- ModelReasoning.adaptive: true sends adaptive thinking, false sends
  budget thinking. Without it, the adapter decides from the id, as before.
- Adaptive thinking without a level map sends pi's default effort.
- modelReasoning(record) sets adaptive from compat.forceAdaptiveThinking
  for anthropic-messages records.
- The catalog rule is pi's: Claude 4.6 and later (dash or dot ids) think
  adaptively. Fireworks takes pi's list. All 291 shared records now have
  pi's flag. Seven Fireworks records take Fireworks' own levels (pi), and
  Vercel gemma-4-31b-it reasons, as at Google and OpenRouter.

Mid-conversation effort (pi supportsMidConvoEffort):
- ModelReasoning.midConversationEffort, from the new
  compat.supportsMidConvoEffort (5 models, pi's list).
- The adapter sends adaptive thinking with block_binding and a fixed
  output_config.effort high, an effort system message before each
  earlier answer and one at the end, the two betas, and no temperature.
- Each answer keeps its effort in metadata.tanstack.reasoningEffort.
- The automatic cache marker skips the trailing effort message.

Mid-conversation channels: midConversationChannels takes
{ tools?, systemPrompts? } to turn on one channel only.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
- The config reasoning wire route sends four record configs: adaptive
  (Vercel anthropic/claude-sonnet-4.6), budget (Vercel openai/gpt-5),
  mid-conversation effort (OpenRouter anthropic/claude-opus-5.5), and
  reasoning: false. The spec checks the thinking fields, the effort
  message, and the two betas in the anthropic-beta header.
- Docs: adaptive and midConversationEffort in the reasoning config, the
  thinking shape from modelReasoning(record), a section for a level
  change during a conversation, and the object form of
  midConversationChannels.
- Changeset for @tanstack/ai, @tanstack/ai-anthropic, @tanstack/ai-models.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
@github-actions github-actions Bot added waiting-on: maintainer The ball is in the maintainers’ court and removed waiting-on: author Waiting for the author to respond or update merge-conflicts Conflicts with the base branch — needs a rebase labels Oct 6, 2026
The reasoning drift test (test:maintainer) failed: the catalog now says
Vercel google/gemma-4-31b-it reasons, and the ai-vercel-gateway table
said it does not. models.dev has no reasoning for it at Vercel, but
Vercel's own model list tags it `reasoning`, and Google, OpenRouter, and
pi 0.87.1 list it.

The ai-vercel-gateway target gets the same override, and its table entry
is the one the sync script writes ({ budget: false }). The other drift
that the sync script finds (models added with main) is left for its own
sync.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Coverage failed on this PR for ai-isolate-e2b (branches 81.48% vs
82.22%), although the package did not change. The false side of
`recent.includes('heap out of memory')` in `onStderr` was hit only by
the memory test, and only when the host read V8's first GC block as its
own stderr chunk. On a Linux runner the writes usually merge into one
chunk, so that side is missed.

The new test writes ordinary stderr and exits with code 2. It must fail
as E2BExecutionError with the stderr tail, not as MemoryLimitError. It
fails when every stderr chunk sets `outOfMemory`, and it also covers the
stderr tail in the exit message.

The thinking-shape changeset also lists @tanstack/ai-vercel-gateway
(the Gemma 4 31B fix in 6656c1d).

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

waiting-on: maintainer The ball is in the maintainers’ court

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants