diff --git a/.changeset/adapter-reasoning-config.md b/.changeset/adapter-reasoning-config.md new file mode 100644 index 0000000000..473256abca --- /dev/null +++ b/.changeset/adapter-reasoning-config.md @@ -0,0 +1,20 @@ +--- +'@tanstack/ai': minor +'@tanstack/ai-anthropic': minor +'@tanstack/ai-openai': minor +'@tanstack/ai-gemini': minor +'@tanstack/ai-bedrock': minor +'@tanstack/ai-mistral': minor +'@tanstack/ai-cloudflare': minor +'@tanstack/ai-models': minor +--- + +Give an adapter the model's reasoning data in its config. A model id that is not in the adapter's own list (a gateway id, a new model, or a catalog id) got no reasoning field before. Now `reasoning?: ModelReasoning` in the config wins over the adapter's table, for the request fields and for the levels that `chat({ reasoning })` takes. `false` sends no reasoning field. Without it, nothing changes. + +The factories take any model id string: `anthropicText`, `createAnthropicChat`, `anthropicVertexText`, `openaiText`, `createOpenaiChat`, `azureOpenaiText`, `geminiText`, `createGeminiChat`, `createBedrockConverse`, `mistralText`, `createMistralText`, `cloudflareText`, and `createCloudflareText`. A misspelled id is no longer a type error. + +`@tanstack/ai-models` adds `modelReasoning(record)`, which turns a catalog record into the config value: + +```ts +createAnthropicChat(record.id, apiKey, { reasoning: modelReasoning(record) }) +``` diff --git a/.changeset/ai-models-catalog.md b/.changeset/ai-models-catalog.md new file mode 100644 index 0000000000..532cfc0717 --- /dev/null +++ b/.changeset/ai-models-catalog.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-models': minor +--- + +New package: a runtime catalog of 1,932 models across 38 providers, generated from models.dev and the OpenRouter and Vercel AI Gateway model lists. `getProviders`, `getModels`, and `getModel` return each model's id, input modalities, context window, output limit, prices, reasoning levels, and request quirks (`compat`). `supportedReasoningLevels`, `clampReasoningLevel`, and `modelCost` work on a record, and `generatedAt` is the time of the data. Each provider also has its own subpath, for example `@tanstack/ai-models/deepseek`. diff --git a/.changeset/ai-models-no-vertex-claude.md b/.changeset/ai-models-no-vertex-claude.md new file mode 100644 index 0000000000..6adc44c716 --- /dev/null +++ b/.changeset/ai-models-no-vertex-claude.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-models': patch +--- + +Leave the Claude models out of the `google-vertex` catalog. They had the Gemini wire (`api: 'google-vertex'`), which the Gemini adapter cannot call. Use `anthropicVertexText` from `@tanstack/ai-anthropic/vertex` for Claude on Vertex. diff --git a/.changeset/ai-null-tool-input.md b/.changeset/ai-null-tool-input.md new file mode 100644 index 0000000000..2445fb5c58 --- /dev/null +++ b/.changeset/ai-null-tool-input.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai': patch +--- + +Run a tool again when the model sends no input or a literal `null` for a tool with no required fields (an empty tool_use block, issue #265). The final tool input check rejected these, so the tool did not run. Now no input is `{}`, and a `null` that the schema rejects is checked as `{}`. Other values must still fit the schema. diff --git a/.changeset/anthropic-thinking-shape.md b/.changeset/anthropic-thinking-shape.md new file mode 100644 index 0000000000..cd0abb67a6 --- /dev/null +++ b/.changeset/anthropic-thinking-shape.md @@ -0,0 +1,16 @@ +--- +'@tanstack/ai': minor +'@tanstack/ai-anthropic': minor +'@tanstack/ai-models': minor +'@tanstack/ai-vercel-gateway': patch +--- + +Send the Anthropic thinking shape that the model record gives. + +- `ModelReasoning.adaptive`: `true` sends adaptive thinking with the effort, `false` sends budget thinking. Without it, the adapter picks from the model id, as before. Adaptive thinking without a level map now sends pi's default effort (`minimal` and `low` give `low`, `medium` gives `medium`, the rest give `high`). +- `ModelReasoning.midConversationEffort`: the level goes into the messages as an effort `system` message, with `block_binding`, a fixed `output_config.effort: 'high'`, and the `mid-conversation-output-config-2026-07-01` and `thinking-binding-controls-2026-08-01` betas. Each answer keeps its level in `metadata.tanstack.reasoningEffort`, so a level change keeps the cached start. On Anthropic's own API, `anthropicText` and `createAnthropicChat` turn it on for `claude-fable-5-1`, `claude-opus-5`, and `claude-opus-5-5` without a `reasoning` config. A custom `baseURL`, `fetch`, client, or `ANTHROPIC_BASE_URL` turns that default off. +- `midConversationChannels` takes `{ tools?: boolean; systemPrompts?: boolean }` to turn on one channel only. + +`modelReasoning(record)` sets `adaptive` from `compat.forceAdaptiveThinking` and `midConversationEffort` from the new `compat.supportsMidConvoEffort` for `anthropic-messages` records. The catalog flags now equal pi's: Claude 4.6 and later think adaptively on every provider (also with dot ids such as `anthropic/claude-opus-4.7`), and the other Vercel AI Gateway models think with a budget. + +Vercel `google/gemma-4-31b-it` reasons, in the catalog and in `@tanstack/ai-vercel-gateway`. Vercel's own model list tags it `reasoning`. Seven Fireworks records take Fireworks' own effort levels. diff --git a/.changeset/block-order.md b/.changeset/block-order.md new file mode 100644 index 0000000000..9d89f0835d --- /dev/null +++ b/.changeset/block-order.md @@ -0,0 +1,12 @@ +--- +'@tanstack/ai': minor +'@tanstack/ai-anthropic': minor +--- + +Keep the order of Claude's thinking, text, and tool calls in one answer. + +- Claude can answer with thinking, a tool call, more thinking, and another tool call. The stream, the stored messages, the client wire, and the next request to Claude keep that order. +- `ModelMessage` has a new optional `blockOrder` field. The library writes it only when the order is not the default (thinking, then text, then tool calls), so a message in the default order does not change. `orderedAssistantBlocks(message)` gives adapter authors the blocks in order. +- Each Claude thinking block ends with its own `REASONING_MESSAGE_END` and `REASONING_END`. +- Text after a second thinking block starts a new text part. +- On the client wire, an answer in another order goes out as ordered rows. A row that exists only for the order has `metadata.tanstack.continues`, and the server joins it back into one message. A `UIMessage` that holds two model calls with a tool result between them reaches the server as assistant, tool, assistant, for every adapter. diff --git a/.changeset/cloudflare-binding-fetch.md b/.changeset/cloudflare-binding-fetch.md new file mode 100644 index 0000000000..940b0b9262 --- /dev/null +++ b/.changeset/cloudflare-binding-fetch.md @@ -0,0 +1,9 @@ +--- +'@tanstack/ai-cloudflare': minor +--- + +Add `cloudflareBindingFetch({ binding, vendor, gateway })`. It sends the requests of the Anthropic and OpenAI text adapters (`createAnthropicChat`, `createOpenaiChat`) through the Workers AI binding (`env.AI.run`) to the AI Gateway `anthropic/...` and `openai/...` models, so a Worker can use Claude and GPT with no provider key. + +- `vendor: 'anthropic'` carries Anthropic Messages requests, and `vendor: 'openai'` carries OpenAI Responses requests. +- The request body goes to `env.AI.run('/', body)`. Headers such as `anthropic-beta` go along as `extraHeaders`, and the SDK's auth and transport headers are dropped. +- A Cloudflare error body is rewrapped so the SDK reports its message. diff --git a/.changeset/compaction-background.md b/.changeset/compaction-background.md new file mode 100644 index 0000000000..666c20292c --- /dev/null +++ b/.changeset/compaction-background.md @@ -0,0 +1,11 @@ +--- +'@tanstack/ai-compaction': minor +--- + +Prepare a compaction summary in the background with `withCompaction({ background: { atTokens } })`. + +- When the count at a model call is over `atTokens` and not over `maxTokens`, the strategy runs on a copy of the messages. The model call does not wait for it. +- The result waits in the metadata store (namespace `@tanstack/ai-compaction:background`), or in memory without one. It applies at the first model call of the next run: as a `tanstack.compaction` record on a durable host, or as the checkpoint. +- A model call over `maxTokens` while the summary runs waits for it and applies it. It compacts inline only when the list is still over `maxTokens`. +- A result whose kept message is gone, because a newer compaction cut further, is dropped. `compaction:ended` and `onCompact` report it with `stale: true`. A failed summary is reported with `error` and applies nothing. +- `CompactionReason` has a new value, `'background'`. Setup throws when `atTokens` is not below `maxTokens`. diff --git a/.changeset/compaction-empty-summary.md b/.changeset/compaction-empty-summary.md new file mode 100644 index 0000000000..43b03fca34 --- /dev/null +++ b/.changeset/compaction-empty-summary.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-compaction': patch +--- + +Fail a compaction when the summarizer returns an empty or whitespace-only summary. Before, `summarizeOldest` replaced the history with an empty summary, so the history was lost. Now the compaction fails like any failed strategy: the history stays, `compaction:ended` (or `onCompact` after the turn) carries the error, and no durable record is written. With `continueOnError`, the run sends the history as it is. diff --git a/.changeset/compaction-usage-and-durable.md b/.changeset/compaction-usage-and-durable.md new file mode 100644 index 0000000000..abc6da631c --- /dev/null +++ b/.changeset/compaction-usage-and-durable.md @@ -0,0 +1,12 @@ +--- +'@tanstack/ai-compaction': minor +--- + +Count with real usage, compact on demand, and keep compaction in a harness log. + +- `countTokens: 'usage'` counts with the usage the provider reported for the last call, plus an estimate of the newer messages. +- `compactNext(threadId)` compacts at the next model call, for example after a context overflow. `auto: false` turns off compaction at `maxTokens` and keeps `compactNext` and the overflow check. +- With `countTokens: 'usage'` and `durable: true`, a harness turn checks once more after its last model call. When the usage passed `maxTokens` or `contextWindow`, compaction runs right away. `contextWindow` also counts with `auto: false`. +- `durable: true` writes a `tanstack.compaction` record into a harness session log. Set `project: { version, record: projectCompaction }` on the host. +- `compaction:ended` and `onCompact` report `reason`, the summary `usage`, and an `error`. `continueOnError` sends the full messages when the strategy fails. +- `summarizeOldest({ cut: 'turn' })` gives the start of a cut turn its own summary. `conversationSummarizer` writes a structured summary, updates the older summary, and takes a `details` hook. diff --git a/.changeset/context-overflow.md b/.changeset/context-overflow.md new file mode 100644 index 0000000000..6558114335 --- /dev/null +++ b/.changeset/context-overflow.md @@ -0,0 +1,8 @@ +--- +'@tanstack/ai': minor +--- + +Add `isContextOverflow({ error, usage, finishReason, contextWindow, provider })`. It tells you that a model call failed, or ended early, because the input did not fit in the model's context window, so you can compact the history and retry. + +- It knows the overflow error messages of about 25 providers (Anthropic, OpenAI, Gemini, Bedrock, Mistral, xAI, Groq, OpenRouter, Ollama, and more), and it ignores rate-limit and throttle errors. +- With `contextWindow`, it also finds a silent overflow (the input is bigger than the window) and a length stop with no output and a full window. diff --git a/.changeset/core-provider-keys.md b/.changeset/core-provider-keys.md new file mode 100644 index 0000000000..8280582361 --- /dev/null +++ b/.changeset/core-provider-keys.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai': minor +--- + +`keyedAdapter(provider, create)` wraps an adapter factory that needs a provider key, and `isKeyedAdapter` tells it apart from a plain adapter. Agents get `ctx.keys` (`get`, `require`, and `adapter`) from the new `SubagentBinding.keys`, or else from each provider's `env` names. diff --git a/.changeset/dashboard-pair-again-and-open-thread.md b/.changeset/dashboard-pair-again-and-open-thread.md new file mode 100644 index 0000000000..ab24599015 --- /dev/null +++ b/.changeset/dashboard-pair-again-and-open-thread.md @@ -0,0 +1,10 @@ +--- +'@tanstack/ai-dashboard': minor +--- + +The dashboard host recovers from a restart, and the page opens threads. + +- **Pair again.** After a dashboard restart, a saved host token gets a `401`. With `onPairingCode`, the host pairs again and calls `onToken` with the new token. Without it, the new `onError` option gets the error. `connection.token` holds the current token. +- **`onError`** also gets an error when the host cannot handle a frame, for example when a thread cannot open. The connection stays. +- **Open a thread.** A host that sets `allowRemoteStart` gets an "Open a thread" form in its row. The host sends the flag in its hello, and for other hosts the open route answers `403` with `remote_start_disabled`. A message sent right after the open waits for the thread. +- **Refusals stay visible.** A refused input shows as "Refused:" with the reason, also in a view that you open later. A refused thread keeps its harness name in the sessions list. diff --git a/.changeset/fake-text-adapter.md b/.changeset/fake-text-adapter.md new file mode 100644 index 0000000000..fdd7fe623f --- /dev/null +++ b/.changeset/fake-text-adapter.md @@ -0,0 +1,9 @@ +--- +'@tanstack/ai': minor +--- + +Add `fakeText()` at the new `@tanstack/ai/testing` subpath: a text adapter that answers from a script, for tests of `chat()`, tools, and middleware with no network and no API key. + +- Queue answers with `setResponses` and `appendResponses`. An answer is `{ text?, thinking?, toolCalls?, finishReason?, error? }`, or a function that gets `{ request, state }`. An empty queue fails with "No more fake responses queued". +- Token usage is estimated as `ceil(characters / 4)`. With `cache: true`, the part of a request that matches the thread's previous request counts as cached tokens. +- Options: `model`, `input` (the adapter's input modalities), `contextWindow`, and `tokensPerSecond` for stream pacing. The stream stops when the request aborts. diff --git a/.changeset/fake-text-fixes.md b/.changeset/fake-text-fixes.md new file mode 100644 index 0000000000..6ae92927c4 --- /dev/null +++ b/.changeset/fake-text-fixes.md @@ -0,0 +1,9 @@ +--- +'@tanstack/ai': patch +--- + +Fix three problems in `fakeText` from `@tanstack/ai/testing`. + +- `fakeText()` type-checks as a text adapter with `exactOptionalPropertyTypes: true`. `inputModalities` is now an optional property. +- `RUN_STARTED` carries `parentRunId` when the call has one, as real adapters do. +- Tool call ids are unique across `fakeText()` instances. The default id is now `fake-call---`. diff --git a/.changeset/harness-agent-messages.md b/.changeset/harness-agent-messages.md new file mode 100644 index 0000000000..e2da19d767 --- /dev/null +++ b/.changeset/harness-agent-messages.md @@ -0,0 +1,11 @@ +--- +'@tanstack/ai-harness': minor +'@tanstack/ai': minor +--- + +- A background agent run takes messages. `start()` on `session.agents.`, `session.agent(name)` and plugin `ctx.agents` returns an `AgentRun` with `send(message, { mode })`. A steer (the default) joins the run's next model call. A follow-up runs the agent again after the run ends, on the run's own transcript, and its receipt names the new run. A steer from another sender runs as a follow-up for that sender. +- `session.agentRuns()` lists the agent runs, and `session.agentRun(id)` finds one. On a durable host they also read the log, so they work after a restart. +- Every background run keeps its messages in its own thread, `::`. +- On a durable host, a run with `resume: true` that a host stop cut gets the messages it had not taken. Without `resume`, its waiting messages settle `aborted`. +- Clients send messages with the `agentMessage` input, `client.sendToAgent()`, or `send()` on an agent item of `createSessionView`. Only runs of agents in `expose.agents` take them. +- `@tanstack/ai`: agent code can start agents with `ctx.agents.start()` when a host binds it (`SubagentBinding.agents`). Without a host, it throws. A harness counts each child against the subagent tree budget, records its parent run, and cancels children before their parent. diff --git a/.changeset/harness-agent-restart.md b/.changeset/harness-agent-restart.md new file mode 100644 index 0000000000..20704cb450 --- /dev/null +++ b/.changeset/harness-agent-restart.md @@ -0,0 +1,7 @@ +--- +'@tanstack/ai-harness': patch +--- + +A background agent run no longer stays `running` after its host stops. The agent run now holds a lease, like a chat turn. When the lease expires, the next host ends the run `failed` and adds a note to the transcript. On a durable host, it also settles the input and, for `start(input, { wake: true })`, starts a new turn. The run does not run again, because it has no checkpoints. + +A background agent that fails now also writes a note, and with `wake: true` it starts a new turn, the same as an agent that finishes. diff --git a/.changeset/harness-auth-wait.md b/.changeset/harness-auth-wait.md new file mode 100644 index 0000000000..6fd2943054 --- /dev/null +++ b/.changeset/harness-auth-wait.md @@ -0,0 +1,11 @@ +--- +'@tanstack/ai': patch +'@tanstack/ai-harness': minor +--- + +A turn can wait for a sign-in and then go on. + +- `ctx.credentials.require(id, { wait: true })`, and `token({ wait: true })` in an `oauthConnector`, stop the running chat turn with an interrupt whose `reason` is `'auth_required'`. Its payload has the connector id. +- When the credential is saved (the `connect:` command, `/connect`, or `ctx.credentials.set`), the same turn goes on and the tool runs again. No new prompt is needed. On a durable host this also works after a restart. +- A resolve with `status: 'cancelled'` gives up: the tool fails as it does without `wait`. +- In `@tanstack/ai`, a tool that throws the MCP input-required shape can set the interrupt `reason`. MCP interrupts still use `'mcp_input'`. diff --git a/.changeset/harness-cli-host-principal.md b/.changeset/harness-cli-host-principal.md new file mode 100644 index 0000000000..847c04b490 --- /dev/null +++ b/.changeset/harness-cli-host-principal.md @@ -0,0 +1,9 @@ +--- +'@tanstack/ai-harness-cli': minor +--- + +`runCli(harness, { host, principal })` can use your own host and user. + +- `host`: the CLI uses this host and does not close it. Without it, the CLI makes its own host and closes it at the end. +- `principal`: every session the CLI opens runs as this principal, and `--serve` gives it to each request with the token. The default is still `{ id: 'cli' }`. +- `host` and `persistence` together fail with "Give runCli a host or persistence, not both." diff --git a/.changeset/harness-connect-iss.md b/.changeset/harness-connect-iss.md new file mode 100644 index 0000000000..b1078a3245 --- /dev/null +++ b/.changeset/harness-connect-iss.md @@ -0,0 +1,8 @@ +--- +'@tanstack/ai-harness': patch +'@tanstack/ai-mcp': patch +--- + +`/connect` now works with MCP servers that send `iss` on the sign-in callback (RFC 9207), such as Linear. `startLoopbackReceiver().waitForCode()` resolves `{ code, iss }`, and `mcpConnector` passes `iss` on to the MCP SDK. A sign-in error names the service, for example `Notion: Sign-in timed out.` + +The session view clears the pending sign-in of a connector when its `connect:` command ends. diff --git a/.changeset/harness-consume-ran-resume.md b/.changeset/harness-consume-ran-resume.md new file mode 100644 index 0000000000..f8a2c0bbdd --- /dev/null +++ b/.changeset/harness-consume-ran-resume.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-harness': patch +--- + +Do not offer an approval or a client-tool interrupt again after its tool ran. When a resumed turn failed, stopped, or hit its time limit after the tool ran, the interrupt came back, and answering it again ran the tool a second time. Now the harness marks it answered in the interrupt store, and a durable host without an interrupt store skips it when it rebuilds from the log. A resume that stops before its tool runs stays open for a retry. diff --git a/.changeset/harness-durable-host-gaps.md b/.changeset/harness-durable-host-gaps.md new file mode 100644 index 0000000000..c41314f0c2 --- /dev/null +++ b/.changeset/harness-durable-host-gaps.md @@ -0,0 +1,18 @@ +--- +'@tanstack/ai-harness': minor +'@tanstack/ai': patch +--- + +Fixes and options for steering and recovery. + +- A steer that joins a turn no longer removes the per-call `providerMessages` change of an earlier middleware. The model call gets that change and the steer message after it. +- Recovery checks the final answer first. A turn whose final answer is in the log settles `completed`, also when an abort was asked, no attempt is left, or its time limit passed. +- New `durability.interruptedToolResult`. Recovery gives this text to a `replay: 'never'` tool call that a crash cut, as the content and the error of its tool message. Without it, the result is the same as before. +- The `durability.recover` hook can return `{ action: 'run', overrides }`. The recovered turn then runs with these `TurnOverrides` (adapter, reasoning, promptCache, tools). The log does not keep them. +- New `session.recover()` and `host.recover(threadId?)` recover the log again, as `open()` does. A turn that another host held at open runs once its lease expires. The work that the session runs or queues stays as it is, and two calls do not run one input twice. +- The log keeps the model retry count of a turn (a `harness.turn.retry` record). An attempt that recovery runs starts from the stored count, so `turn.onModelError` gets the same `retries` as before the crash. A finished tool phase still sets it back to 0. The recover hook gets it as `input.retries`. An older log without the record works as before. +- New `session.continue(options?)`, `client.continue(options?)`, and the `{ op: 'continue' }` input start a turn from the stored transcript, with no new message. They queue, check the `inputId` for a duplicate, and recover after a restart, as a prompt does. When the transcript does not end with a user or a tool message, the input is rejected with `nothing_to_continue`. +- New `TurnAdditions.ephemeral`. `turn.beforeFinish` (and `turn.onJoin`) can return messages that only the next model call gets, after its context. The transcript and the log never keep them. A `beforeFinish` that returns only ephemeral messages also continues the turn, and counts for `maxFinishCycles`. +- New `session.close({ recoverable: true })` and `host.close({ recoverable: true })`. The running turns and agent runs stop as on a crash: no settlement, no aborted run state, and no later log write. Their leases end at once, so the next host that opens the thread runs them again. A plain `close()` and a user cancel still settle `aborted`. +- Recovery never runs a tool call of an answer that stopped at the output limit (finish reason `length`). `chat()` now keeps the finish reason of each assistant message in `metadata.tanstack.finishReason`. After a crash, each call of such an answer gets an error result with the new `durability.truncatedToolResult` text, also for a `replay: 'safe'` tool. Before, a crash could run the calls of a cut answer that kept its tool calls (a provider search, then thinking). +- New `durability.continueCutOff`, off by default. When a crash cuts an answer, recovery adds the answer text from the log as an assistant message, then a user message with a note, in one append. The model then continues the answer. `{ note }` sets the note. An attempt with no answer text runs again as before. `close({ recoverable: true })` now waits until the events from before the close are in the log. diff --git a/.changeset/harness-durable-log.md b/.changeset/harness-durable-log.md new file mode 100644 index 0000000000..1646e8b7ef --- /dev/null +++ b/.changeset/harness-durable-log.md @@ -0,0 +1,23 @@ +--- +'@tanstack/ai-persistence': minor +'@tanstack/ai-harness': minor +--- + +Add a durable session log. A harness session can now keep its events, transcript, inputs, and tool steps in one append-only log, so a turn survives a crash, a retry does not run an input twice, and a tool does not repeat a finished side effect. + +`@tanstack/ai-persistence` adds the `LogStore` contract: `append` writes a batch at a position, all or nothing, and rejects with `LogConflictError` when another writer took that position. `read` and `subscribe` complete it. Put your store in `stores.log`. `defineLogStore` types an adapter, `memoryLogStore()` is the reference store, and `runPersistenceConformance` runs the log cases when `stores.log` is present. + +`@tanstack/ai-harness` turns on durable mode when the host gets `stores.log` and `stores.runs`: + +- `prompt`, `steer`, `followUp`, and `resolve` take an optional `inputId`, and so do `prompt`, `steer`, and `followUp` of `createHarnessClient`. The same id with the same payload does not run again. `Operation.receipt` resolves when the input is stored, and `session.settled(inputId)` gives how an input ended, also after a restart. +- `defineHarness({ durability: { maxAttempts, timeoutMs } })` limits each input. The default is 10 attempts and no timeout. +- `durableTool(definition, execute)` gives a tool `step.do(name, fn)`. After a crash, a finished step returns its stored value and does not run again. The tool's `append(records)` adds host records to the same append as the tool batch. +- `session.append(records)` adds host records to the log. The host option `project: { record, version }` folds host records into the model context. `logMessageStore` reads the transcript of a log outside a session. +- Host options `coalesceMs` (default 100 ms in durable mode) and `lease: { ttlMs, renewMs }`. Clients get a `harness.input.settled` event when an input ends. +- A `busy: 'steer'` prompt joins the running turn in the order it arrived, and it settles with that turn. + +Breaking changes in the harness: + +- `HarnessPersistence` is now a union of the current stores and the log stores. +- After a crash, the note for a tool that may have run now sets `error`, so the model sees a tool error. +- A steer that a turn did not reach is now answered in the same turn, not in a new turn. diff --git a/.changeset/harness-host-generic.md b/.changeset/harness-host-generic.md new file mode 100644 index 0000000000..6b303e426c --- /dev/null +++ b/.changeset/harness-host-generic.md @@ -0,0 +1,13 @@ +--- +'@tanstack/ai-harness': minor +'@tanstack/ai-persistence': minor +--- + +Harness host features for frameworks that build on the harness: + +- `host.open(harness, { threadId, logId })`: sessions that share one log. A record can name another session with `thread`, and `createHarnessHost({ reduce })` folds every host record of a log (`host.logState(logId)`). +- `durableTool(definition, execute, { replay: 'never' })`. +- `defineHarness({ turn: { onModelError, beforeFinish, maxFinishCycles, canJoin, onJoin } })`, `durability.recover`, `retryTransientErrors()`, and `isTransientModelError()`. +- `stores.leases` (`LeaseStore`): turn leases from your own system, in place of `stores.runs`. +- Durable tools that a middleware returns from `onConfig` get `step` and `append`. +- A waiting steer stops the join when it has an abort request, and `cancel` on it settles it `aborted`. diff --git a/.changeset/harness-input-sender-context.md b/.changeset/harness-input-sender-context.md new file mode 100644 index 0000000000..779bd73393 --- /dev/null +++ b/.changeset/harness-input-sender-context.md @@ -0,0 +1,20 @@ +--- +'@tanstack/ai-harness': minor +'@tanstack/ai-persistence': patch +'@tanstack/ai-mcp': patch +--- + +Each harness input carries its sender and its context. + +- `prompt`, `steer`, `followUp` and `resolve` take `principal`. Without it, the input belongs to the principal that opened the session. `createHarnessHandler` and `handleHarnessSocket` run each input, and each command, as the principal that `authorize` returned for its request. +- The log keeps the sender of each input. The turn's tools, `ctx.session.principal`, `ctx.credentials` and `ctx.keys` use the turn's sender. A turn that waits for a sign-in continues only when its sender signs in. +- `turn.canJoin` gets the `principal` of each waiting input and the `turnPrincipal` of the running turn, so it can keep one person's message out of another person's turn. The `routing.router` context gets `principal` and `context`. +- Without `turn.canJoin`, a waiting message joins the running turn only when the same person sent both. A message from another person waits and runs as its own turn, as its sender. +- A command acts for the person who runs it: its `session.principal` is that person, and `session.prompt` from a command runs as that person. The wake turn after a background agent runs as the person who started the agent. +- `Principal.tenantId` goes into the credential scope, so two organizations' credentials for one user stay apart. +- A credential read finds the user's own credential first, then the tenant's (saved without a `userId`), so one key can serve a whole organization. Saves and deletes change the user's own credential only. +- `prompt`, `steer` and `followUp` take `context`: JSON data from the client, stored with the input and kept after a restart. `POST run` fills it from AG-UI `forwardedProps`. Tools get it merged with `HarnessConfig.context`, and the harness value wins for a key in both. +- A plugin `adapter(turn)` gets `TurnInfo`: the turn's `message`, `context`, `principal`, `inputId` and `overrides`. A plugin with `lifetime: 'turn'` gets the same as `ctx.turn` in `setup`. +- `CommandContext` has `principal` and `credentials` of the principal that ran the command. +- In `@tanstack/ai-persistence`, a stored input's `principal` can carry `tenantId`. +- `mcpConnector` keeps its sign-in state, client and tools per sender, so one person's MCP sign-in is not used for another person's turn. diff --git a/.changeset/harness-interrupt-restart.md b/.changeset/harness-interrupt-restart.md new file mode 100644 index 0000000000..f936d958de --- /dev/null +++ b/.changeset/harness-interrupt-restart.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-harness': patch +--- + +A resolve after a restart now continues the turn that stopped for an interrupt. Before, the waiting interrupts lived only in memory, so the next host refused the resolve with `no_pending_interrupts`. The session now keeps them in `stores.metadata`, with the agent cards of a routed turn, and the next host that opens the thread reads them back. A resolve removes the stored copy. Without a metadata store, the interrupts stay in memory, as before. diff --git a/.changeset/harness-keep-host-records.md b/.changeset/harness-keep-host-records.md new file mode 100644 index 0000000000..52313b3d58 --- /dev/null +++ b/.changeset/harness-keep-host-records.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-harness': patch +--- + +Keep a host record that lands during the last model call of a turn. At the end of a turn, `withPersistence` tags older messages with the run, so the engine save rewrites them, and the log rebase then kept only the engine list. A host message that a record had folded in after the engine's last save was lost. Now the engine list still wins for the older messages, and the host messages that came after stay after it. diff --git a/.changeset/harness-media.md b/.changeset/harness-media.md new file mode 100644 index 0000000000..5151e0647d --- /dev/null +++ b/.changeset/harness-media.md @@ -0,0 +1,14 @@ +--- +'@tanstack/ai-harness': minor +'@tanstack/ai-harness-cli': minor +'@tanstack/ai-mcp': minor +'@tanstack/ai-acp': minor +--- + +`@tanstack/ai-harness`: a harness turn can take images, audio, video, and documents, and the harness keeps the media that agents make. Sessions have `putMedia`, `getMedia`, `loadMedia`, and `mediaUrl`, and `mediaPart(record)` sends a stored file with a prompt. The transcript keeps a small `harness-media:` URL, and only the model call gets the bytes. Every `ctx.generateImage`, `ctx.generateSpeech`, `ctx.generateAudio`, and `ctx.generateVideo({ stream: true })` result is saved and published as a `harness.media` event, and the records are saved on the message for history. `defineHarness({ media })` sets `maxBytes`, `kinds`, `accepts`, and `transcribe`, and a file the model cannot read stops the turn with a clear error. `createHarnessHandler` serves uploads, signed media URLs (`mediaSecret`), and file bytes with Range support. The browser client has `upload`, `mediaUrl`, and `loadMedia`, and the session view has media parts and `view.send(text, attachments)`. + +`@tanstack/ai-harness-cli`: `@path` in a message attaches a file. Generated media is saved to `./-media` (`--media-dir` to change it), and `--mcp` lets path attachments read the working folder. + +`@tanstack/ai-mcp`: the harness MCP server's `chat` tool takes `path`, `url`, and `data` attachments (`filePaths` sets the folders a path may read). Results carry the media of the turn: small images and audio inline, other files as `harness-media://` links that `resources/read` serves. + +`@tanstack/ai-acp`: the ACP agent takes image, audio, and resource blocks, and advertises its prompt capabilities. diff --git a/.changeset/harness-p0-agent-results.md b/.changeset/harness-p0-agent-results.md new file mode 100644 index 0000000000..af26a929b6 --- /dev/null +++ b/.changeset/harness-p0-agent-results.md @@ -0,0 +1,6 @@ +--- +'@tanstack/ai': minor +'@tanstack/ai-persistence': patch +--- + +A subagent can now return a plain value and call any activity. `defineAgent`'s `run` can return a promise of any value, such as an image result. The value arrives on `SUBAGENT_FINISHED.result`, and the parent model gets it as the tool result. A very long string in that result (for example a base64 image) reaches the parent model as a short note. `run` also gets the activities on `ctx` (`ctx.chat`, `ctx.generateImage`, `ctx.generateVideo`, and the rest), which fill in the thread id, a run id, and the abort signal, plus `ctx.forward` for a plain `chat()` call. The optional `produces` field says what an agent makes. `RunRecord` gets optional fields for harness sessions (`kind`, `activity`, `agent`, `result`, `artifacts`, `principal`, `leaseOwner`, `leaseExpiresAt`, `checkpoint`), and the memory run stores keep them. diff --git a/.changeset/harness-p1-core.md b/.changeset/harness-p1-core.md new file mode 100644 index 0000000000..28c75291c4 --- /dev/null +++ b/.changeset/harness-p1-core.md @@ -0,0 +1,9 @@ +--- +'@tanstack/ai-harness': minor +'@tanstack/ai': minor +'@tanstack/ai-persistence': minor +--- + +New package `@tanstack/ai-harness`. `defineHarness` takes the same options as `chat()`, plus typed `agents`, `plugins`, and a `busy` policy. `createHarnessHost().open(harness, { threadId })` opens a long-lived session. A session runs chat turns with `prompt`, takes messages during a turn with `steer` and `followUp`, answers approvals with `resolve`, and cancels work with `cancel`. `session.agents..run(input)` runs a typed agent from code, and `start(input, { wake: true })` runs it in the background. Every operation streams AG-UI events with cursors. `definePlugin` adds tools, prompts, chat middleware, generation middleware, and agents, with capabilities that plugins provide to each other and resources that close with the session or the turn. + +`@tanstack/ai` exports `CapabilityRegistry`, `runAgentStream`, `createSubagentId`, `compactForModel`, and the `SubagentsBag` type for harness hosts. `@tanstack/ai-persistence` adds the optional `inbox` store (`InboxStore`, `defineInboxStore`) that keeps accepted session inputs across restarts. The memory backend and the conformance testkit cover it. diff --git a/.changeset/harness-p10-goal.md b/.changeset/harness-p10-goal.md new file mode 100644 index 0000000000..bf27c56dca --- /dev/null +++ b/.changeset/harness-p10-goal.md @@ -0,0 +1,7 @@ +--- +'@tanstack/ai-harness': minor +--- + +Add the `goal({ judge })` plugin to `@tanstack/ai-harness/plugins`. `/goal ` sets a goal and starts a turn. After each turn, the `judge` model reads the goal and the end of the transcript and decides if the goal is met. If it is not met, the plugin starts the next turn, up to `maxRounds` turns (20 by default). The loop also stops when a turn waits for approval or fails, and when the user sends a message. `/goal` shows the status, `/goal stop` ends the goal, and `/goal resume` continues it. The goal is plugin state, so it survives a restart. The `GoalMet` event tells other plugins and clients that the goal is met. + +A plugin's `ctx.session.prompt(text)` now returns the turn, so the plugin can cancel it before it starts. The turn always waits in the queue, also on a harness with `busy: 'reject'`. diff --git a/.changeset/harness-p11-session-view.md b/.changeset/harness-p11-session-view.md new file mode 100644 index 0000000000..845ea5b33b --- /dev/null +++ b/.changeset/harness-p11-session-view.md @@ -0,0 +1,7 @@ +--- +'@tanstack/ai-harness': minor +--- + +Add `createSessionView(session | client)` at `@tanstack/ai-harness/view`. It keeps a live TanStack Store of everything a UI shows: messages with streaming text, tool calls, and child agents, plus approvals, questions, sign-ins, status, background agents, commands, settings, tools, and plugin state. It has actions (`send`, `command`, `setConfig`, `cancel`, `approve`, `reject`) and typed events (`view.on('approval', ...)`, `view.on(GoalMet, ...)`). It works with any UI library through the TanStack Store adapters or `store.subscribe`, and it works in the browser with `createHarnessClient`. + +New reads for it: `session.transcript()`, `session.describe()`, `snapshot().plugins`, and on the client `transcript()`, `describe()`, `answer()`, `command()`, `setConfig()`, and `events({ onConnection })`. `createHarnessHandler` serves `GET .../transcript` and `GET .../describe`. `selectGoal(state)` reads the goal from a view. diff --git a/.changeset/harness-p12-cli.md b/.changeset/harness-p12-cli.md new file mode 100644 index 0000000000..148499ae0c --- /dev/null +++ b/.changeset/harness-p12-cli.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-harness-cli': minor +--- + +The CLI has no UI library now: it does not depend on `ink` or `react`. An interactive terminal uses line mode, which now prints from a session view and opens sign-in links in the browser. `runCli(harness, { ui })` runs your own screen with any TUI library: `ui` gets a ready `createSessionView` view and resolves when the user quits. An Ink screen is in `examples/harness-cli/src/tui.tsx`. diff --git a/.changeset/harness-p13-agent-middleware.md b/.changeset/harness-p13-agent-middleware.md new file mode 100644 index 0000000000..ea64d5500d --- /dev/null +++ b/.changeset/harness-p13-agent-middleware.md @@ -0,0 +1,8 @@ +--- +'@tanstack/ai': minor +'@tanstack/ai-harness': minor +--- + +`@tanstack/ai`: the chat middleware context has `subagentName`, the agent name of a child that runs through `ctx.chat`, and `parentSubagentRunId`, the id of the child that started a nested child. `chat()` takes both as options. A child that calls `ctx.chat({ subagents })` now gives its own children the host middleware (`subagents.binding.chatMiddleware` and `generationMiddleware`) before their own, also when the binding has no budget. + +`@tanstack/ai-harness`: plugins can contribute `agentMiddleware`, chat middleware for every agent run: subagents the lead model calls (coding agents too), background agents, and their children. It does not run in the lead turn. Put the same middleware in `middleware` and `agentMiddleware` to see every model call. Run plugins' `generationMiddleware` now also reaches the agents of their turn. The first-party `usage()` plugin now counts the model calls of every agent, not only the lead turn. diff --git a/.changeset/harness-p14-mcp-server.md b/.changeset/harness-p14-mcp-server.md new file mode 100644 index 0000000000..91a5dceab1 --- /dev/null +++ b/.changeset/harness-p14-mcp-server.md @@ -0,0 +1,8 @@ +--- +'@tanstack/ai-mcp': minor +'@tanstack/ai-harness-cli': minor +--- + +`@tanstack/ai-mcp`: add `createHarnessMcpServer` at `@tanstack/ai-mcp/harness`. It serves a harness as an MCP server, so Claude Code, Claude Desktop, Cursor, or another agent can use it. The tools are `chat`, `steer`, `cancel`, `approve`, `reject`, `resolve`, `answer`, and `status`, plus `agent_` for each exposed agent and `command_` for each plugin command. Every tool takes an optional `threadId`. With `approvals: 'ask'` (the default), the client asks the user about each approval when it supports elicitation. Otherwise the approvals come back in the result. `approvals: 'auto'` approves every tool call. Each interrupt in a result has a `kind`: `approval`, `client-tool`, or `generic`. `resolve` answers every kind in one call: `approved` for an approval, and `payload` for the others. + +`@tanstack/ai-harness-cli`: `--mcp` serves the harness as an MCP server over stdio, and `--yes` approves every tool call in MCP mode. `--serve` also serves MCP at `/mcp`, behind the same bearer token. Both need `@tanstack/ai-mcp`, which is an optional peer dependency. diff --git a/.changeset/harness-p2-protocols.md b/.changeset/harness-p2-protocols.md new file mode 100644 index 0000000000..cbdc0bbe0b --- /dev/null +++ b/.changeset/harness-p2-protocols.md @@ -0,0 +1,14 @@ +--- +'@tanstack/ai-harness': minor +'@tanstack/ai-harness-cli': minor +'@tanstack/ai-acp': minor +'@tanstack/ai': minor +--- + +Harness sessions now survive a crash and talk to clients. + +- **Resume.** A session holds a lease on each running turn, saves the transcript around each tool phase, and records tool calls that have no result yet. When a host opens a session whose turn lost its lease, it continues the turn in a new run and emits `harness.operation.resumed`. `toolDefinition({ replay: 'safe' | 'never' })` in `@tanstack/ai` decides whether an unfinished tool runs again. +- **Protocol.** `createHarnessHandler` serves capabilities, standard AG-UI runs, the session event stream with cursors, control inputs with receipts, and snapshots. `authorize` is required. `handleHarnessSocket` serves the same session tier over a WebSocket. `createHarnessClient` in `@tanstack/ai-harness/client` is a typed, reconnecting client. +- **`harnessText`** runs a harness as the text adapter of another `chat()` call. +- **ACP v2.** `@tanstack/ai-acp/agent` adds `createAcpAgent` and `serveAcp`, which serve a harness as an ACP v2 agent (experimental, like the draft protocol). `@tanstack/ai-acp` now uses `@agentclientprotocol/sdk` 1.5. The existing ACP client adapters keep working. +- **New package `@tanstack/ai-harness-cli`.** `runCli(harness)` gives a harness an Ink terminal UI, a print mode with exit codes, NDJSON output, `--acp`, and `--serve`. diff --git a/.changeset/harness-p3-plugins.md b/.changeset/harness-p3-plugins.md new file mode 100644 index 0000000000..7f03d928e3 --- /dev/null +++ b/.changeset/harness-p3-plugins.md @@ -0,0 +1,15 @@ +--- +'@tanstack/ai-harness': minor +'@tanstack/ai-harness-cli': minor +'@tanstack/ai-persistence': minor +'@tanstack/ai': minor +--- + +Harness plugins can now do everything a Claude Code style agent needs. + +- **Plugin API.** `setup` can return `commands` (`defineCommand`, with typed input), `config` (`configOption.select`, `.boolean`, `.text`, `.number`), `contribute` for extension points (`createExtensionPoint`), and `adapter` to pick the main model. The setup context adds `collect`, typed events (`createPluginEvent`, `emit`, `on`), `state` (saved in the metadata store, published as `STATE_SNAPSHOT`), `config`, `credentials`, and `session` (`ask`, `prompt`, `transcript`, `setConfig`). Prompts can be functions called for each turn. +- **Session.** `session.command`, `session.commands`, `session.setConfig`, `session.config`, `session.answer`, and `session.inspect`. The protocol accepts `command`, `answer`, and `config` inputs. +- **Auth.** `oauthConnector` adds `connect:` and `disconnect:` commands and gives tools a fresh token. The OAuth runner (`loopbackLogin`, `deviceLogin`, `createPkce`, `exchangeCode`, `refreshCredential`) uses PKCE S256 and a single-use loopback on `127.0.0.1`. Tools that need a missing credential emit `harness.auth_required`. +- **First-party plugins** in `@tanstack/ai-harness/plugins`: `permissions` (with `default`, `plan`, `acceptEdits`, and `bypass` modes and the `PermissionRules` extension point), `todos`, `modelPicker`, `projectInstructions`, `fileCommands`, `compact`, and `usage`. `workspaceTools` is in the Node entry `@tanstack/ai-harness/plugins/coding`. +- **`@tanstack/ai-persistence`** adds the optional `credentials` store (`CredentialStore`, `defineCredentialStore`) and compare-and-set on the memory metadata store. **`@tanstack/ai`** adds optional `getVersioned` and `setIf` to `MetadataStore`. +- **CLI.** Plugin commands, `/config`, `/connect`, and `/disconnect`. Questions from plugins are answered on the next line. Sign-in links open in the browser. diff --git a/.changeset/harness-p4-subagents.md b/.changeset/harness-p4-subagents.md new file mode 100644 index 0000000000..3f64256a15 --- /dev/null +++ b/.changeset/harness-p4-subagents.md @@ -0,0 +1,11 @@ +--- +'@tanstack/ai': minor +'@tanstack/ai-harness': minor +--- + +Subagent trees now have limits, and a harness can run agents from code and call other harnesses. + +- **`subagents.limits`** in `@tanstack/ai`: `maxDepth`, `maxConcurrent`, `maxCalls`, and `timeoutMs` for the whole tree of children the model starts through tools. The tree shares one budget (`SubagentBudget`), which a child's `ctx.chat({ subagents })` passes on, so a child cannot reset it. A refused start reaches the model as a tool error. +- **`ctx.agents`** in harness plugins: `run`, `start` (with `wake`), and `group` with `onFailure: 'cancel-siblings' | 'collect'`. Every child of a group settles before the group returns. +- **`harnessAgent(harness)`** turns a harness into an agent for `subagents.agents`. `defineHarness` takes an optional `description`. +- A harness applies default limits (`DEFAULT_SUBAGENT_LIMITS`: depth 2, 3 at once, 12 per tree) when `subagents.limits` is not set, and agents started from code count against them. diff --git a/.changeset/harness-p5-build.md b/.changeset/harness-p5-build.md new file mode 100644 index 0000000000..0a6e2a896d --- /dev/null +++ b/.changeset/harness-p5-build.md @@ -0,0 +1,9 @@ +--- +'@tanstack/ai-harness': minor +--- + +Run a harness outside your server. + +- **`@tanstack/ai-harness/build`**: `buildHarness` bundles a harness with Bun into a worker artifact (`harness.js`) and writes `harness.manifest.json` (name, agents, plugins, requirements, sha256 digest). With `compile`, it also builds a single executable. `artifactText(dir)` runs the artifact as worker processes and uses it as a text adapter, after it checks the digest. `readManifest` reads and checks a manifest. +- **`@tanstack/ai-harness/worker`**: `runHarnessWorker` serves a harness as session-tier frames over NDJSON on stdin and stdout. +- **`harnessText({ url, token })`** uses a harness served by `createHarnessHandler` (or `runCli --serve`) on another machine as a text adapter. diff --git a/.changeset/harness-p6-dashboard.md b/.changeset/harness-p6-dashboard.md new file mode 100644 index 0000000000..dd71ffb871 --- /dev/null +++ b/.changeset/harness-p6-dashboard.md @@ -0,0 +1,8 @@ +--- +'@tanstack/ai-dashboard': minor +'@tanstack/ai-harness-cli': minor +--- + +New package `@tanstack/ai-dashboard`: a self-hosted dashboard for harness sessions. `npx @tanstack/ai-dashboard` starts it (`startDashboard` in code). Agents dial out with `connectDashboard` from `@tanstack/ai-dashboard/connect`, pair with a one-time code, and get a revocable host token. The web app lists hosts and sessions, streams messages and tool calls, shows approval cards and plugin questions, and sends prompts, steers, and stops. It installs as a web app on a phone. Inputs for an offline host wait at the dashboard and reach it on reconnect. + +`@tanstack/ai-harness-cli` adds `--dashboard `, which pairs on first use (or reads `HARNESS_DASHBOARD_TOKEN`) and keeps the session connected. diff --git a/.changeset/harness-p7-connectors.md b/.changeset/harness-p7-connectors.md new file mode 100644 index 0000000000..17ea459419 --- /dev/null +++ b/.changeset/harness-p7-connectors.md @@ -0,0 +1,11 @@ +--- +'@tanstack/ai-mcp': minor +'@tanstack/ai-harness': minor +'@tanstack/ai-persistence': minor +--- + +`@tanstack/ai-mcp/connector` adds `mcpConnector`: a harness plugin that signs the user in to a remote MCP server (Notion, Linear, and others) with OAuth, then gives the model the server's tools. `/connect ` registers a client, signs in with PKCE on a `127.0.0.1` loopback, and keeps the token in the credential store. Tools the server does not mark read-only ask for approval. + +`@tanstack/ai-harness` adds `discoverTools` to plugins, for tools found at run time, and `startLoopbackReceiver` for OAuth redirects. A harness turn now runs up to 50 model calls by default (`agentLoopStrategy` still overrides it). + +`@tanstack/ai-persistence`: an `oauth` credential can keep the OAuth `client` it was issued to. diff --git a/.changeset/harness-p8-code-mode.md b/.changeset/harness-p8-code-mode.md new file mode 100644 index 0000000000..7c5d58f318 --- /dev/null +++ b/.changeset/harness-p8-code-mode.md @@ -0,0 +1,8 @@ +--- +'@tanstack/ai-code-mode': minor +'@tanstack/ai-harness': minor +--- + +`@tanstack/ai-code-mode/harness` adds `codeMode({ driver })`: a harness plugin that gives the model one `execute_typescript` tool. The model writes a program that calls several tools, and the program runs in the isolate of any `@tanstack/ai-isolate-*` driver. Read-only tools move into code mode, including MCP tools found after sign-in. Tools that need approval, edits, and commands stay normal tool calls. Tool names that are not JavaScript identifiers (for example `notion-search`) get a safe name inside the program. + +`@tanstack/ai-harness` adds `prepareTools` to plugins, to change the tool list of each turn. `PermissionRules` and `decidePermission` are now also exported from the package root. diff --git a/.changeset/harness-p9-coding-agents.md b/.changeset/harness-p9-coding-agents.md new file mode 100644 index 0000000000..62a57e4b40 --- /dev/null +++ b/.changeset/harness-p9-coding-agents.md @@ -0,0 +1,13 @@ +--- +'@tanstack/ai-sandbox': minor +'@tanstack/ai-harness': minor +'@tanstack/ai-harness-cli': minor +'@tanstack/ai-acp': minor +'@tanstack/ai-dashboard': minor +--- + +`@tanstack/ai-sandbox/harness` adds `codingAgents({ sandbox, agents, workspace? })`: a harness plugin that gives the lead model one tool per coding agent (Claude Code, Codex, Grok Build, or any ACP agent). Each agent runs in the sandbox, keeps its own session per thread (also after a restart), and starts read-only in the harness `plan` mode. `workspace: 'shared'` (default) runs one agent at a time in one sandbox per thread. `'per-agent'` gives each agent its own sandbox. `/fresh [agent]` starts new sessions. + +`@tanstack/ai-harness` plugins can contribute `subagents`: agents the model can call as tools. + +Child agent work is now visible: the CLI shows each child's tool calls and a finish line with the start of its answer, ACP editors get the child's tool calls, and the dashboard shows a block per child. diff --git a/.changeset/harness-provider-keys.md b/.changeset/harness-provider-keys.md new file mode 100644 index 0000000000..f95bdc073c --- /dev/null +++ b/.changeset/harness-provider-keys.md @@ -0,0 +1,12 @@ +--- +'@tanstack/ai-harness': minor +'@tanstack/ai-harness-cli': patch +--- + +Users of a shipped harness can connect their own model provider in the app, with no `.env` file. The new `providerKeys({ providers })` plugin in `@tanstack/ai-harness/plugins` adds `/connect `, `/disconnect `, and `/keys`. `/connect` runs the provider's `signIn` when it has one, or else asks the user to paste the key. A provider can set `keyUrl`, the page where the user makes a key: `/connect` opens it in the browser before it asks. The key is saved in the credential store of the user. `/keys` and the plugin state show where each key comes from: a saved key (the last 4 characters only), the env var, or no key. + +`defineHarness({ adapter })`, the `modelPicker` choices, `compact({ adapter })`, and `goal({ judge })` now accept a `keyedAdapter(...)`. The session builds it with the user's key just before each call. The saved key comes first, then the provider's env var. When there is no key, the turn stops with `harness.auth_required`, so the user sees `/connect `. Agents get the same keys as `ctx.keys`, and plugins get them as `ctx.keys` on their setup context. + +The `usage()` plugin state has `contextTokens`: the prompt tokens of the last lead model call, so a UI can show how full the context is. + +A `Question` can set `secret: true`. The view gives it as `ViewQuestion.secret`, and the session does not keep a secret answer in the inbox. In a terminal, the CLI line mode hides what the user types for a secret question, and it never prints the answer. diff --git a/.changeset/harness-recovered-resolve.md b/.changeset/harness-recovered-resolve.md new file mode 100644 index 0000000000..9731f9ea0b --- /dev/null +++ b/.changeset/harness-recovered-resolve.md @@ -0,0 +1,6 @@ +--- +'@tanstack/ai-harness': patch +--- + +- A resolve that a crash interrupted now continues the interrupted turn after the restart. Before, it ran without the id of the interrupted run, so `chat()` refused it, and it lost the routed plan, the sender, and the context of that turn. +- On a host without a log, the transcript is saved before each model call. So when the model call after an approved tool fails, the result of that tool stays in the transcript. diff --git a/.changeset/harness-reset.md b/.changeset/harness-reset.md new file mode 100644 index 0000000000..2344a2f14d --- /dev/null +++ b/.changeset/harness-reset.md @@ -0,0 +1,9 @@ +--- +'@tanstack/ai-harness': minor +--- + +`session.reset(note?)` starts a fresh model context and keeps the transcript. + +- From the next turn, the model sees only the note (when given) and what comes after the reset. `transcript()` keeps every message, plus a marker message with `metadata.harness.reset`. The session view shows the marker as a notice. +- A reset is an input: it runs once per `inputId`, and it waits for a running turn to end. While the thread waits for interrupts, it is refused with `pending_interrupts`. +- Clients send `{ op: 'reset', note }` or call `client.reset(note)`. A `harness.reset` event fires when it applies. diff --git a/.changeset/harness-resolve-fixes.md b/.changeset/harness-resolve-fixes.md new file mode 100644 index 0000000000..c5f1196b53 --- /dev/null +++ b/.changeset/harness-resolve-fixes.md @@ -0,0 +1,9 @@ +--- +'@tanstack/ai-harness': patch +--- + +Fix three bugs in how a harness session answers interrupts and steers. + +- A resolve whose payload fails validation (for example a client tool result that does not match its `outputSchema`) no longer drops the pause. The resolve still settles `failed`, but the interrupt stays pending, so a correct resolve continues the turn. +- A resolve sent as soon as a client reads `RUN_FINISHED` with interrupts is now accepted. Before, it was rejected with `no_pending_interrupts` or `busy` while the turn was still ending. The resolve now runs right after the interrupted turn ends. +- `prompt(text, { busy: 'steer' })` that joins the running turn now returns the running turn's id in its receipt, as `steer()` does. diff --git a/.changeset/harness-resumable-agents.md b/.changeset/harness-resumable-agents.md new file mode 100644 index 0000000000..7c77ee3018 --- /dev/null +++ b/.changeset/harness-resumable-agents.md @@ -0,0 +1,12 @@ +--- +'@tanstack/ai': minor +'@tanstack/ai-harness': minor +--- + +A background agent can continue on the host that takes over after a crash. + +- `agents.start(agent, input, { resume: true })` needs a durable host (`stores.log`). Without one, it throws at start. Without `resume`, an agent that a stopped host left running still fails, as before. +- The next host runs the agent again with the same input. Its model calls continue from the saved transcript, and its finished tool calls do not run again. +- `ctx.step.do(name, fn)` in agent code keeps the result of `fn` in the log, so a side effect does not run again when the agent runs again. +- `durability.maxAttempts` caps the runs. After the cap, the run settles `failed` with the code `attempts_exhausted`. +- In `@tanstack/ai`, the agent run context has `step`, and `SubagentBinding` takes an optional `step`. Without a binding, `step.do` runs `fn` each time. diff --git a/.changeset/harness-routing.md b/.changeset/harness-routing.md new file mode 100644 index 0000000000..768942ad57 --- /dev/null +++ b/.changeset/harness-routing.md @@ -0,0 +1,8 @@ +--- +'@tanstack/ai': minor +'@tanstack/ai-harness': minor +--- + +`@tanstack/ai`: a router pick can give an agent its input. Anywhere a pick has a name, it can have `{ name, input }`, so `subagents.router` can start an agent with `inputSchema`. `chat()` checks the input against the schema, and `run` reads it as `ctx.input`. A resume keeps each input of the plan. `subagentRoute` gets `needsInput(result)`, which lists the picked agents with an `inputSchema`, and `pick(result, { inputs })`, which puts each input on its name. `inputs` is typed from each schema. + +`@tanstack/ai-harness`: `defineHarness` takes `routing: { router, order, strategy, limits, sandbox }`. The router runs at the start of each new turn. It sends the turn to the main model, or to root agents: the harness `agents` and the `agents` of plugins. Its context has the fields of the `subagents.router` context, plus `session`, `input`, `operationId`, `inputId`, and `adapter`. With `strategy: 'handoff'`, the main model answers after the agents, in the same turn, and the `turn` hooks run for it. A resolve continues the routed plan without a new router call. A routed turn holds a run lease, and its text is the text of the agents. A `subagents.router` turn gets the same lease and text, and a child that stops for an interrupt resumes on `resolve`. `HarnessRouting` and `HarnessRouterContext` are exported. diff --git a/.changeset/harness-run-route.md b/.changeset/harness-run-route.md new file mode 100644 index 0000000000..35a2e954a9 --- /dev/null +++ b/.changeset/harness-run-route.md @@ -0,0 +1,11 @@ +--- +'@tanstack/ai-harness': minor +--- + +`ChatClient` and `useChat` work with the harness handler's `POST run` route. + +- A turn started by `POST run` runs as the request's `runId`, so a `ChatClient` can answer its approvals and client tools. A retry with the same `runId` runs once. The same `runId` with another message gets 409. +- `GET run?threadId=` answers with the transcript, the running turn, and the waiting interrupts, as the hydrate data that a `ChatClient` with `persistence: true` reads. +- `GET run?runId=` streams the turn with that run id from its first event, with each event's cursor as its SSE id. A reloaded `useChat` with `persistence: true` joins a running turn this way, and `Last-Event-ID` resumes the stream. `canAccess` decides who can join. +- A question from `ctx.session.ask` during a turn arrives on that turn's stream. +- Each `snapshot().activeOperations` entry has `startedCursor`, so `events({ from: startedCursor })` reads a running operation from its first event. diff --git a/.changeset/harness-sessions-and-coding-tools.md b/.changeset/harness-sessions-and-coding-tools.md new file mode 100644 index 0000000000..949de55aa3 --- /dev/null +++ b/.changeset/harness-sessions-and-coding-tools.md @@ -0,0 +1,20 @@ +--- +'@tanstack/ai-harness': minor +'@tanstack/ai-persistence': minor +'@tanstack/ai-sandbox': minor +'@tanstack/ai-mcp': minor +'@tanstack/ai-code-mode': minor +--- + +Session lists, coding tools, and safer clients for the harness. + +- **Session index.** `host.sessions` lists, renames, deletes, and forks sessions (`fork(harness, threadId, { before | through })`). A fork of a thread that another harness runs throws, and the HTTP route answers `409`. `host.fork` takes `at: null` to copy no messages. `host.events()` is a status feed (`running`, `waiting`, `idle`). Over HTTP: `GET sessions`, `POST sessions`, and `GET host-events`. Subagent runs show up as child sessions of their thread. `@tanstack/ai-persistence` adds the `SessionIndexStore` contract, a memory store, and conformance tests. +- **Waiting inputs.** `session.inputs()` lists them. The `cancelInput` and `setDelivery` inputs cancel one or move it between steer and queue. +- **Coding tools** in the new Node entry `@tanstack/ai-harness/plugins/coding`: `workspaceTools` (moved here) gives `read_file` (text pages, images, PDFs), `write_file`, `edit_file` or `patch` (`editStyle`), `list_files`, `grep`, `bash` (with `background`), `webfetch`, and `websearch`. A `WorkspaceBackend` runs them on the host or in a sandbox (`sandboxWorkspaceBackend` from `@tanstack/ai-sandbox/harness`). Also new here: `snapshots()` with `/undo` and `/redo` (they restore only the files of the turn, and need git 2.26 or newer), and `formatter()`. On Windows, commands do not run a program from the working folder by its name. `@vscode/ripgrep` and `turndown` are optional peers. +- **More plugins** in `@tanstack/ai-harness/plugins`: `agents()` (profiles, Markdown agents, `/agent`, and a step limit that also refuses tool calls at the limit), `question()`, `title()`, and `boundToolOutput`. `projectInstructions()` reads up to the repo root and adds an environment block. `usage()` shows `session.usage()` and prices only the calls with no provider cost. +- **Permissions.** Rules can match the paths and commands that a call touches (`PermissionResources`). Questions take once, always, or reject with a message. The rules apply in subagent runs too, and to the real path of a link. `grep` skips the files that the rules protect. `webfetch` asks by default. `PermissionDecisionCapability` lets other plugins ask what `permissions()` decides, and code mode uses it. +- **Clients set only what you expose.** `defineHarness({ expose: { config, commands } })`. A client config input or command outside these lists gets `not_exposed`, so a client cannot switch the permission mode to `bypass`. Nothing is exposed by default, and server code is not affected. `GET describe` lists only the exposed commands and config keys, and a session view on a client shows a notice for a refused one. The harness MCP server also makes tools only for exposed commands. +- **`mcp()` harness plugin** in `@tanstack/ai-mcp/harness`: MCP servers by name, tools named `_`, `/mcp` status, and `codeMode: true`. `timeoutMs` limits the connect step too. +- **Code mode** moves a tool marked `metadata.codeMode` only when `permissions()` allows it in plan mode, and never moves a tool that declares permission resources. +- **`defineHarness({ toolExecution })`** passes `'parallel'` or `'sequential'` to the `chat()` call of every turn. +- **Plugin API.** `ctx.session.note(text, { wake })`, `ctx.session.entry()` and `updateEntry(patch)`, `prepareTools({ tools, model })`, and the `WorkspaceHooks` extension point. `ctx.collect()` now works with every array method. diff --git a/.changeset/harness-skills.md b/.changeset/harness-skills.md new file mode 100644 index 0000000000..b0c9ebd23c --- /dev/null +++ b/.changeset/harness-skills.md @@ -0,0 +1,11 @@ +--- +'@tanstack/ai-harness': minor +'@tanstack/ai-skills': minor +'@tanstack/ai-code-mode': minor +--- + +Skills in a harness, live commands, and lazy code-mode tools. + +- `@tanstack/ai-skills/harness`: the new `skills({ dirs })` plugin gives the harness model the skills of a list of folders, and makes each skill a `/` command that starts a turn with that skill. A skill with the name of another command is `/skill:`, and `/skills` lists them. The plugin watches the folders, so a skill that you add or remove while the session runs changes the commands at once. +- `@tanstack/ai-harness`: `ctx.commands` (`has`, `set`, `delete`, `ready`) lets a plugin add and remove its commands while the session runs. Each change sends a `harness.commands.changed` event, and session views read the command list again. +- `@tanstack/ai-code-mode`: `codeMode({ lazy: true })` makes each tool that moves into code mode lazy. The prompt lists only the tool names, and the model asks `discover_tools` for the signatures before it calls them. diff --git a/.changeset/harness-thread-settings-fork.md b/.changeset/harness-thread-settings-fork.md new file mode 100644 index 0000000000..97bdcbdc5f --- /dev/null +++ b/.changeset/harness-thread-settings-fork.md @@ -0,0 +1,12 @@ +--- +'@tanstack/ai-harness': minor +--- + +Stored settings per thread, and a fork of a thread. + +- `defineHarness({ models: { fast, strong } })` names the models a thread can pick. +- `session.configure({ model, reasoning, instructions, tools, plugins, cwd })` stores settings for the thread, from the next turn. `null` clears a field. An unknown model, plugin or bad value gets a rejected receipt. Turn `overrides` still win over the settings. +- A client can change only the fields in `defineHarness({ expose: { settings } })`. An input with another field is refused with `not_exposed`, and with no list a client can change no field. Server code that calls `session.configure()` can change every field. +- `session.settings()` reads them, `describe()` lists them with the model names, and a `harness.settings.changed` event fires on each change. Clients send `{ op: 'configure', settings }` or call `client.configure(settings)`. +- `workspaceTools` work in the thread's `cwd`, and refuse a folder outside their root. +- `host.fork(harness, { threadId, newThreadId, at?, principal? })` copies the transcript up to the message `at` into a new thread, with its stored settings, plugin config and media. It does not copy plugin state, pending interrupts, queued inputs or usage totals. diff --git a/.changeset/harness-turn-lifetime.md b/.changeset/harness-turn-lifetime.md new file mode 100644 index 0000000000..653d0988f5 --- /dev/null +++ b/.changeset/harness-turn-lifetime.md @@ -0,0 +1,9 @@ +--- +'@tanstack/ai-harness': minor +--- + +Plugins and models per turn. + +- A plugin with `lifetime: 'turn'` is set up for each chat turn and disposed when the turn ends. `'run'` still works the same way, as an old name for `'turn'`. +- `adapter` in `defineHarness` is optional. A plugin `adapter()` or the turn's `overrides.adapter` can give the model. A turn with no model fails with a clear error. +- `harnessText(harness, { inputModalities })` sets the input kinds the harness takes, for a harness without an adapter. diff --git a/.changeset/harness-usage-totals.md b/.changeset/harness-usage-totals.md new file mode 100644 index 0000000000..ddae9eb3bd --- /dev/null +++ b/.changeset/harness-usage-totals.md @@ -0,0 +1,10 @@ +--- +'@tanstack/ai-harness': minor +--- + +A harness session keeps usage totals for its thread. + +- `session.usage()` and `snapshot().usage` return `{ total, byModel, bySender }`. Each entry has `calls`, the token counts, and `cost` when a provider reported one. +- Every model call counts: the turn's own calls, its subagents, and background agents, for the sender of the turn or the person who started the agent. +- A durable host keeps a `harness.usage` record per call in the log. Other hosts keep the totals in `stores.metadata`, or in memory without one. +- A `harness.usage` event fires on each call, so a UI can show a live total. diff --git a/.changeset/harness-view-client-tools.md b/.changeset/harness-view-client-tools.md new file mode 100644 index 0000000000..488d2b1f57 --- /dev/null +++ b/.changeset/harness-view-client-tools.md @@ -0,0 +1,9 @@ +--- +'@tanstack/ai-harness': minor +--- + +`createSessionView` from `@tanstack/ai-harness/view` now handles client tools and sign-ins that a turn waits for. + +- A client tool (a harness tool with no server implementation) is in the new `state.clientTools` list, not in `approvals`. Each item has `id`, `toolCallId`, `tool`, `args`, `resolve(output)` and `fail(message)`. A `clientTool` event fires for each new item. +- The view sends one resume when every open approval and client tool has an answer. +- A turn that waits for a sign-in (`credentials.require(id, { wait: true })`) shows in `state.signIns`, not in `approvals`, also after a reload. diff --git a/.changeset/harness-work-claims.md b/.changeset/harness-work-claims.md new file mode 100644 index 0000000000..80c9e4ceea --- /dev/null +++ b/.changeset/harness-work-claims.md @@ -0,0 +1,10 @@ +--- +'@tanstack/ai-persistence': minor +'@tanstack/ai-harness': minor +--- + +A host can continue the work of hosts that stopped, without anyone opening the thread. + +- `@tanstack/ai-persistence` has a new optional store, `WorkClaimStore` (`claim`, `release`, `listExpired`), with `defineWorkClaimStore`, conformance cases, and a memory version in `memoryPersistence()`. +- With `stores.workClaims`, a harness session claims its thread while the thread has work, renews the claim, and releases it when the thread is idle. A thread that waits for an approval or a sign-in is idle. A failed claim write is a `harness.plugin.warning` event, and the work goes on. +- `host.resumePending({ harnesses, close?, limit? })` opens each thread whose claim expired, and the session recovers its pending work. Two hosts that sweep at once never take the same thread. A swept session closes when its work ends (`close: 'whenIdle'`, the default, unless the app opened the thread too) or stays open (`'never'`). Call it at boot, and from a cron job or a Durable Object alarm. diff --git a/.changeset/log-records-capability.md b/.changeset/log-records-capability.md new file mode 100644 index 0000000000..b23d5fc935 --- /dev/null +++ b/.changeset/log-records-capability.md @@ -0,0 +1,6 @@ +--- +'@tanstack/ai': minor +'@tanstack/ai-harness': minor +--- + +Add `LogRecordsCapability`. A chat middleware can append records to the durable session log of its run. A durable harness host provides it to the chat run of each session turn. Agent runs do not get it. diff --git a/.changeset/mcp-connect-timeout.md b/.changeset/mcp-connect-timeout.md new file mode 100644 index 0000000000..2aeb5c2988 --- /dev/null +++ b/.changeset/mcp-connect-timeout.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-mcp': minor +--- + +`createMCPClient({ requestOptions })` now also applies to the connect handshake. A server that does not answer while it connects fails after `timeout`, not after the SDK default of 60 seconds. If you set a short `timeout` for tool calls, a slow server can now fail at connect. On a `stdio` server, connect can take up to two times `timeout`. diff --git a/.changeset/mcp-resource-template-args.md b/.changeset/mcp-resource-template-args.md new file mode 100644 index 0000000000..016c2abdde --- /dev/null +++ b/.changeset/mcp-resource-template-args.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-mcp': minor +--- + +A resource template can parse its variables. With `resourceDefinition({ uriTemplate, argsSchema })`, `argsSchema.parse` runs on the template variables before `read(uri, variables, ctx)` gets them. A read result `{ text | blob, mimeType }` sets the MIME type of that answer. diff --git a/.changeset/mcp-tool-options.md b/.changeset/mcp-tool-options.md new file mode 100644 index 0000000000..d81e484c3d --- /dev/null +++ b/.changeset/mcp-tool-options.md @@ -0,0 +1,16 @@ +--- +'@tanstack/ai-mcp': minor +--- + +New `createMCPClient` options: + +- `toolName: (tool) => string` sets the name the model sees for each tool. It wins over `prefix`. `metadata.mcp.serverToolName` keeps the server's own name, so `callTool()` and MCP Apps widget calls still reach the tool. Two tools with the same final name throw `DuplicateToolNameError`. +- `requestOptions: { timeout, resetTimeoutOnProgress }` is sent with tool lists, tool calls, resource requests, and prompt requests. The connect handshake, `subscriptions/listen`, and task status polls keep the SDK defaults. A spec 2026 `tools/call` uses `timeout` and gets no progress notifications. +- `toolFilter` also takes a list of server tool names. The client keeps exactly those tools, in list order. A missing or repeated name throws the new `MCPToolFilterError`, which names them and the available tools. + +`toolName` and `requestOptions` also work per server in `createMCPClients`, in `mcpConnector`, and in MCP Apps widget calls. + +Two default changes: + +- A server input schema without `properties` now gets `properties: {}`, and a schema without `type` gets `type: 'object'`. A full schema does not change. +- `metadata.mcp.annotations` is now a frozen copy of the server's annotations, not the server's own object. diff --git a/.changeset/mid-conversation-changes.md b/.changeset/mid-conversation-changes.md new file mode 100644 index 0000000000..b942a2e2c7 --- /dev/null +++ b/.changeset/mid-conversation-changes.md @@ -0,0 +1,16 @@ +--- +'@tanstack/ai': minor +'@tanstack/openai-base': minor +'@tanstack/ai-openai': minor +'@tanstack/ai-anthropic': minor +--- + +Keep the prompt cache when tools or system prompts are added during a conversation. + +- Before each model call, `chat()` compares the tools and the system prompts with the earlier calls. On a model with a mid-conversation channel, an added tool, or a system prompt added at the end, goes out in the conversation. The start of the request stays the same. +- `openaiText` on `gpt-5.4-mini`, `gpt-5.5`, `gpt-5.6-luna`, `gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-6-astra`, `gpt-6-luna`, and `gpt-6-sol`: an added tool is an `additional_tools` input item, and an added prompt is a `developer` message. The Responses adapter of `@tanstack/openai-base` sends these when a subclass sets `midConversationChannels`. +- `anthropicText` on `claude-opus-4-8`, `claude-opus-5`, `claude-opus-5-5`, `claude-fable-5`, and `claude-fable-5-1`: every request with tools has the `mid-conversation-tool-changes-2026-07-01` beta and one placeholder tool, unless a provider tool such as `webSearchTool()` takes part. An added tool has `defer_loading: true` and a `tool_addition` block in a `system` message. +- Every other model and adapter sends the same request as before. +- The channels are on by default only with the provider's own API. A custom `baseURL`, a custom `fetch`, an injected client, or the `OPENAI_BASE_URL` / `ANTHROPIC_BASE_URL` environment variable turns the default off. Set `midConversationChannels: true` to turn the channels on anyway. Set `false` to turn them off anywhere. +- With prompt caching on (the default), the automatic Claude tool cache marker goes on the last start tool, not on the placeholder or an added tool. +- New optional fields: `ModelMessage.midConversationChange`, `TextOptions.midConversationChanges`, and `TextAdapter.midConversationChannels`, with the types `MidConversationChange`, `MidConversationChanges`, and `MidConversationChannels`. `splitMidConversationChanges` helps adapter authors read the changes. diff --git a/.changeset/model-cost-tiers.md b/.changeset/model-cost-tiers.md new file mode 100644 index 0000000000..98d11b8396 --- /dev/null +++ b/.changeset/model-cost-tiers.md @@ -0,0 +1,10 @@ +--- +'@tanstack/ai-models': minor +'@tanstack/ai-event-client': minor +'@tanstack/ai-anthropic': patch +'@tanstack/ai-bedrock': patch +--- + +Price long context and 1-hour cache writes in `modelCost`. A record's `cost.tiers` (from models.dev) holds higher prices above an input size, and a call uses the tier with the highest `inputTokensAbove` below its input. `TokenCounts.cacheWrite1h`, the part of `cacheWrite` with a 1-hour retention, costs 2 times the input price. + +Usage reports that part as `promptTokensDetails.cacheWrite1hTokens`. The Anthropic adapter reads it from `cache_creation.ephemeral_1h_input_tokens`, and the Bedrock Converse adapter from the 1-hour `cacheDetails`. diff --git a/.changeset/openrouter-sign-in.md b/.changeset/openrouter-sign-in.md new file mode 100644 index 0000000000..1be7b413c2 --- /dev/null +++ b/.changeset/openrouter-sign-in.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-openrouter': minor +--- + +Add `openrouterSignIn()` at `@tanstack/ai-openrouter/pkce`: a local app, such as a CLI, signs the user in to OpenRouter in the browser and gets an API key, with no key to copy. diff --git a/.changeset/otel-client-tool-span.md b/.changeset/otel-client-tool-span.md new file mode 100644 index 0000000000..1492e6ebb8 --- /dev/null +++ b/.changeset/otel-client-tool-span.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai': patch +--- + +Do not record an `execute_tool` span for a client tool in `otelMiddleware`. Middleware now sees a client tool's input before dispatch, so the middleware opened a span for a tool that the server does not run. A client tool runs in the client, so the server records only the chat and iteration spans, as before. diff --git a/.changeset/persistence-keep-change-record.md b/.changeset/persistence-keep-change-record.md new file mode 100644 index 0000000000..099c275f24 --- /dev/null +++ b/.changeset/persistence-keep-change-record.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-persistence': patch +--- + +A message store now keeps the mid-conversation change record of an assistant message when a client sends the same message again without it. With `useChat`, the prompt cache then holds across turns on models that receive tool and prompt changes as changes. diff --git a/.changeset/prompt-caching-default.md b/.changeset/prompt-caching-default.md new file mode 100644 index 0000000000..0567be4e64 --- /dev/null +++ b/.changeset/prompt-caching-default.md @@ -0,0 +1,26 @@ +--- +'@tanstack/ai': minor +'@tanstack/ai-harness': minor +'@tanstack/ai-anthropic': minor +'@tanstack/ai-bedrock': minor +'@tanstack/ai-openai': minor +'@tanstack/ai-models': minor +'@tanstack/ai-openrouter': minor +'@tanstack/ai-mistral': minor +'@tanstack/ai-claude-code': minor +--- + +Prompt caching is now on by default. Each `chat()` call asks the provider to cache the stable start of the request (system prompt, tools, and earlier messages), so repeated requests cost less and answer sooner. + +- New `chat()` option `promptCache`: `'none' | 'short' | 'long'`, or `{ retention, key }`. The default is `'short'`. The key is `promptCache.key`, else the `threadId` or `conversationId` that you give. +- To turn it off, pass `promptCache: 'none'`. Do this for a large prompt that you send to Claude one time only, because Claude bills the cache write. +- Each adapter sends its own fields: + - Claude (Anthropic, Bedrock, OpenRouter) gets cache markers. + - OpenAI gets `prompt_cache_key`, and `prompt_cache_retention` or `prompt_cache_options` on `'long'`. + - Mistral gets `prompt_cache_key`. + - OpenRouter gets `sessionId`. +- A manual `cache_control`, `cachePoint`, `prompt_cache_key`, or `sessionId` wins over the automatic value. +- Harness sessions cache by default too. `defineHarness({ promptCache })` sets the default, and `host.open(harness, { threadId, promptCache })` overrides it for one session. +- `usage.promptTokens` is now the total input on Anthropic, Bedrock, and Claude Code: the uncached tokens plus the cache reads and writes. Before, it was the uncached tokens only. The cache parts stay in `promptTokensDetails`. +- New usage counts: cache reads on Mistral and on the OpenRouter Responses adapter. The harness `usage()` plugin also counts cache reads and writes. +- New compat field `supportsExplicitPromptCacheMode` on `openaiCompatible` and on `ModelCompat`. diff --git a/.changeset/provider-input-modalities-a.md b/.changeset/provider-input-modalities-a.md new file mode 100644 index 0000000000..660f89d774 --- /dev/null +++ b/.changeset/provider-input-modalities-a.md @@ -0,0 +1,8 @@ +--- +'@tanstack/ai-openai': patch +'@tanstack/ai-anthropic': patch +'@tanstack/ai-gemini': patch +'@tanstack/ai-mistral': patch +--- + +The text adapters set `inputModalities` from their model metadata, so a known model gives the input kinds it reads and an unknown model gives `undefined`. diff --git a/.changeset/provider-input-modalities-b.md b/.changeset/provider-input-modalities-b.md new file mode 100644 index 0000000000..a0efb244a9 --- /dev/null +++ b/.changeset/provider-input-modalities-b.md @@ -0,0 +1,9 @@ +--- +'@tanstack/ai-groq': patch +'@tanstack/ai-byteplus': patch +'@tanstack/ai-grok': patch +'@tanstack/ai-openrouter': patch +'@tanstack/ai-llmgateway': patch +--- + +The text adapters set `inputModalities` from their model metadata, so a known model gives the input kinds it reads and an unknown model gives `undefined`. diff --git a/.changeset/reasoning-option-anthropic.md b/.changeset/reasoning-option-anthropic.md new file mode 100644 index 0000000000..b1f737c01c --- /dev/null +++ b/.changeset/reasoning-option-anthropic.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-anthropic': minor +--- + +Breaking: set reasoning with `chat({ reasoning })`. `thinking`, `effort`, and `output_config.effort` leave `modelOptions`, with their option types. Claude 4.7 and later get adaptive thinking with `output_config.effort`, Claude 4.6 gets adaptive thinking with `effort`, older models get a thinking token budget, and `off` disables thinking. diff --git a/.changeset/reasoning-option-bedrock.md b/.changeset/reasoning-option-bedrock.md new file mode 100644 index 0000000000..a4b601e04a --- /dev/null +++ b/.changeset/reasoning-option-bedrock.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-bedrock': minor +--- + +Breaking: set reasoning with `chat({ reasoning })`. `reasoning_effort` (Chat Completions) and `reasoning` (Responses) leave `modelOptions`. On Converse, Claude gets budget thinking in `additionalModelRequestFields`. diff --git a/.changeset/reasoning-option-byteplus.md b/.changeset/reasoning-option-byteplus.md new file mode 100644 index 0000000000..2dee6d902e --- /dev/null +++ b/.changeset/reasoning-option-byteplus.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-byteplus': minor +--- + +Breaking: set reasoning with `chat({ reasoning })`. `thinking`, `reasoning_effort`, `BytePlusThinkingOption`, and `BytePlusReasoningEffort` are removed. The adapter sends `thinking.type`, plus `reasoning_effort` when the level has an effort. diff --git a/.changeset/reasoning-option-claude-code.md b/.changeset/reasoning-option-claude-code.md new file mode 100644 index 0000000000..834e9e3ac5 --- /dev/null +++ b/.changeset/reasoning-option-claude-code.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-claude-code': minor +--- + +Add `chat({ reasoning })`: effort models get `--effort`, and budget models get `MAX_THINKING_TOKENS`. diff --git a/.changeset/reasoning-option-cloudflare.md b/.changeset/reasoning-option-cloudflare.md new file mode 100644 index 0000000000..2af64f90a5 --- /dev/null +++ b/.changeset/reasoning-option-cloudflare.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-cloudflare': minor +--- + +Breaking: set reasoning with `chat({ reasoning })`. `modelOptions.reasoning_effort` is removed; `off` sends `null`. `chat_template_kwargs` stays in `modelOptions`. diff --git a/.changeset/reasoning-option-codex.md b/.changeset/reasoning-option-codex.md new file mode 100644 index 0000000000..fab04d0f9b --- /dev/null +++ b/.changeset/reasoning-option-codex.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-codex': minor +--- + +Breaking: set reasoning with `chat({ reasoning })`. `modelOptions.modelReasoningEffort` is removed. The adapter config's `modelReasoningEffort` stays as the default. diff --git a/.changeset/reasoning-option-core.md b/.changeset/reasoning-option-core.md new file mode 100644 index 0000000000..4352d9e063 --- /dev/null +++ b/.changeset/reasoning-option-core.md @@ -0,0 +1,7 @@ +--- +'@tanstack/ai': minor +--- + +Add `reasoning` to `chat()`: a level (`off`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max`) or `{ level, summary, budgetTokens }`. An adapter declares each model's levels, so the option is typed per model, and `budgetTokens` only goes to models that think with a token budget. Leaving it out keeps the provider default. Middleware can change it in `onConfig`. + +New exports: `REASONING_LEVELS`, `supportedReasoningLevels`, `clampReasoningLevel`, and the types `ReasoningLevel`, `ReasoningMap`, `ModelReasoning`, `ReasoningCapability`, `ReasoningOption`, `ReasoningOptionFor`, `ReasoningRequest`, and `AdapterReasoning`. Adapter authors get `resolveReasoning`, `reasoningValue`, `reasoningBudget`, and `DEFAULT_REASONING_BUDGETS` from `@tanstack/ai/adapter-internals`. `BaseTextAdapter` and `TextAdapter` take a new last type parameter for the reasoning levels. diff --git a/.changeset/reasoning-option-gemini.md b/.changeset/reasoning-option-gemini.md new file mode 100644 index 0000000000..2810d9e6a5 --- /dev/null +++ b/.changeset/reasoning-option-gemini.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-gemini': minor +--- + +Breaking: set reasoning with `chat({ reasoning })`. `thinkingConfig` leaves the chat `modelOptions`, and `GeminiThinkingOptions` is removed. Gemini 3 gets `thinkingLevel`, Gemini 2.5 gets `thinkingBudget`, and the Interactions adapter gets `thinking_level`. Image generation keeps its own `thinkingConfig`. diff --git a/.changeset/reasoning-option-grok.md b/.changeset/reasoning-option-grok.md new file mode 100644 index 0000000000..0c53ee0b3b --- /dev/null +++ b/.changeset/reasoning-option-grok.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-grok': minor +--- + +Breaking: set reasoning with `chat({ reasoning })`. `modelOptions.reasoning`, `GrokReasoning`, and `GrokReasoningEffort` are removed. `grok-build-0.1` takes no reasoning, because the xAI API refuses it. diff --git a/.changeset/reasoning-option-groq.md b/.changeset/reasoning-option-groq.md new file mode 100644 index 0000000000..19c24e7bfe --- /dev/null +++ b/.changeset/reasoning-option-groq.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-groq': minor +--- + +Breaking: set reasoning with `chat({ reasoning })`. `modelOptions.reasoning_effort` is removed. Qwen 3 gets `default` or `none`. `reasoning_format` and `include_reasoning` stay in `modelOptions`. diff --git a/.changeset/reasoning-option-llmgateway.md b/.changeset/reasoning-option-llmgateway.md new file mode 100644 index 0000000000..e596e09258 --- /dev/null +++ b/.changeset/reasoning-option-llmgateway.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-llmgateway': minor +--- + +Breaking: set reasoning with `chat({ reasoning })`. `modelOptions.reasoning_effort` is removed; the adapter sends the level as `reasoning_effort`. diff --git a/.changeset/reasoning-option-lovable.md b/.changeset/reasoning-option-lovable.md new file mode 100644 index 0000000000..a527404b68 --- /dev/null +++ b/.changeset/reasoning-option-lovable.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-lovable': minor +--- + +Breaking: set reasoning with `chat({ reasoning })`. `reasoning` and `include_reasoning` leave `modelOptions`. The adapters send `reasoning_effort` or `reasoning.effort`. diff --git a/.changeset/reasoning-option-mistral.md b/.changeset/reasoning-option-mistral.md new file mode 100644 index 0000000000..a4fc441def --- /dev/null +++ b/.changeset/reasoning-option-mistral.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-mistral': minor +--- + +Add `chat({ reasoning })`: Mistral Small and Medium get `reasoning_effort`, and Magistral gets `prompt_mode: "reasoning"`. diff --git a/.changeset/reasoning-option-ollama.md b/.changeset/reasoning-option-ollama.md new file mode 100644 index 0000000000..c90a3eae0d --- /dev/null +++ b/.changeset/reasoning-option-ollama.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-ollama': minor +--- + +Breaking: set reasoning with `chat({ reasoning })`. `modelOptions.think` is removed. A model name that the package does not list gets the on/off toggle. diff --git a/.changeset/reasoning-option-openai-base.md b/.changeset/reasoning-option-openai-base.md new file mode 100644 index 0000000000..72f7b7e926 --- /dev/null +++ b/.changeset/reasoning-option-openai-base.md @@ -0,0 +1,5 @@ +--- +'@tanstack/openai-base': minor +--- + +The Chat Completions and Responses base adapters take a reasoning type parameter and a protected `modelReasoning(model)` hook. When a subclass returns the model's reasoning data, the base sends `chat({ reasoning })` as `reasoning_effort` (Chat Completions) or `reasoning: { effort, summary }` (Responses). The Chat Completions base also gets `includeUsageInStream` and `requestHeaders` hooks. diff --git a/.changeset/reasoning-option-openai.md b/.changeset/reasoning-option-openai.md new file mode 100644 index 0000000000..1173446b2e --- /dev/null +++ b/.changeset/reasoning-option-openai.md @@ -0,0 +1,7 @@ +--- +'@tanstack/ai-openai': minor +--- + +Breaking: set reasoning with `chat({ reasoning })`. `modelOptions.reasoning` and the `OpenAIReasoningOptions` types are removed; the adapter sends the level as `reasoning.effort` with a summary. + +`openaiCompatible` takes `compat` for a provider's request quirks (thinking format, developer role, `maxTokensField`, the DeepSeek `reasoning_content` replay, `store`, strict tools, session headers, and Anthropic cache markers), and each model entry takes `reasoning` (its level map) and its own `compat`. diff --git a/.changeset/reasoning-option-opencode.md b/.changeset/reasoning-option-opencode.md new file mode 100644 index 0000000000..4ff94d3f45 --- /dev/null +++ b/.changeset/reasoning-option-opencode.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-opencode': minor +--- + +Add `chat({ reasoning })`: the adapter writes the model's reasoning options into the OpenCode server config. diff --git a/.changeset/reasoning-option-openrouter.md b/.changeset/reasoning-option-openrouter.md new file mode 100644 index 0000000000..20b6dfefcd --- /dev/null +++ b/.changeset/reasoning-option-openrouter.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-openrouter': minor +--- + +Breaking: set reasoning with `chat({ reasoning })`. `modelOptions.reasoning` and the `ReasoningOptions` type are removed. `off` goes out as `effort: "none"`. The Responses adapter also takes `budgetTokens`. diff --git a/.changeset/reasoning-option-vercel-gateway.md b/.changeset/reasoning-option-vercel-gateway.md new file mode 100644 index 0000000000..2cf6c1381f --- /dev/null +++ b/.changeset/reasoning-option-vercel-gateway.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-vercel-gateway': minor +--- + +Breaking: set reasoning with `chat({ reasoning })`. `reasoning` and `include_reasoning` leave `modelOptions`. Chat Completions gets AI Gateway's `reasoning` object, and Responses gets `reasoning.effort`. diff --git a/.changeset/replay-adapter-parity.md b/.changeset/replay-adapter-parity.md new file mode 100644 index 0000000000..e5bb8d2e72 --- /dev/null +++ b/.changeset/replay-adapter-parity.md @@ -0,0 +1,28 @@ +--- +'@tanstack/ai': minor +'@tanstack/ai-utils': patch +'@tanstack/ai-openai': minor +'@tanstack/ai-anthropic': minor +'@tanstack/openai-base': patch +'@tanstack/ai-bedrock': patch +'@tanstack/ai-byteplus': patch +'@tanstack/ai-gemini': patch +'@tanstack/ai-grok': patch +'@tanstack/ai-mistral': patch +'@tanstack/ai-ollama': patch +'@tanstack/ai-openrouter': patch +'@tanstack/ai-persistence': patch +'@tanstack/ai-harness': patch +--- + +Record assistant and run source metadata, provider generation IDs, and reported response models. Preserve metadata through saved history and prepare valid requests when a conversation switches providers or APIs. Clean failed turns and unanswered tool calls without changing the saved transcript. + +Validate final tool input after middleware. Coerce raw JSON Schema arguments and return tool errors for invalid input. Preserve authored Standard Schema transforms and raw provider arguments. Check approval and client-output resume bindings, and preserve JSON-compatible approved edits across resume. + +Keep pending harness interrupts available after a rejected resume response, so a valid response can retry. Preserve stored approval bindings across client-output phases. + +Add Azure OpenAI Responses support and Anthropic Bearer/OAuth authentication. Add an option to replay unsigned thinking through Anthropic-protocol gateways. + +Preserve ordered thinking replay and tool-result images where providers support them. Send Bedrock tool error status, reject invalid Chat Completions content and unknown finish reasons, remove lone Unicode surrogates from outgoing text, and include an empty tools list when tool history requires it. + +Update BytePlus's adapter comments to describe source-aware signature replay through the shared adapter. diff --git a/.changeset/responses-answer-items.md b/.changeset/responses-answer-items.md new file mode 100644 index 0000000000..0f6f8bc51a --- /dev/null +++ b/.changeset/responses-answer-items.md @@ -0,0 +1,6 @@ +--- +'@tanstack/ai': minor +'@tanstack/openai-base': patch +--- + +Keep each OpenAI Responses answer item's `id` and `phase` (`commentary` or `final_answer`) on the assistant message, in `metadata.tanstack.responseItems`. A same-model replay sends each item back with its `id` and `phase`. Another model gets the plain text, as before. diff --git a/.changeset/responses-function-call-namespace.md b/.changeset/responses-function-call-namespace.md new file mode 100644 index 0000000000..3603655621 --- /dev/null +++ b/.changeset/responses-function-call-namespace.md @@ -0,0 +1,5 @@ +--- +'@tanstack/openai-base': patch +--- + +Send the `namespace` of a function call back to the Responses API. A tool that came through `additional_tools` is called in a namespace, and without it the next request failed with `400 Missing namespace for function_call`. The adapter now keeps the namespace in the tool call metadata and sends it with the replayed `function_call` item. diff --git a/.changeset/text-adapter-input-modalities.md b/.changeset/text-adapter-input-modalities.md new file mode 100644 index 0000000000..f614ac62b1 --- /dev/null +++ b/.changeset/text-adapter-input-modalities.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai': minor +--- + +Text adapters can declare `inputModalities`, the input kinds the model reads, at runtime. A `BaseTextAdapter` subclass sets it from its model metadata, and `undefined` means not known. diff --git a/.changeset/tool-choice-adapters.md b/.changeset/tool-choice-adapters.md new file mode 100644 index 0000000000..5cc812a196 --- /dev/null +++ b/.changeset/tool-choice-adapters.md @@ -0,0 +1,17 @@ +--- +'@tanstack/ai-anthropic': patch +'@tanstack/ai-bedrock': patch +'@tanstack/ai-gemini': patch +'@tanstack/ai-grok': patch +'@tanstack/ai-mistral': patch +'@tanstack/ai-openrouter': patch +--- + +Send `chat({ toolChoice })` to each provider. + +- **Anthropic and Bedrock Converse.** Claude Fable 5.1, Mythos 5.1, Opus 5.5, and Sonnet 5.5 reject a forced tool, and every Claude model rejects one while thinking is on. On those, `'required'` and a named tool fall back to `auto`. +- **Bedrock Converse** has no `none`. `'none'` sends no tool config, except after tool calls in the history: then it sends the tools with `auto`, because Bedrock needs a tool config there. +- **Bedrock Converse, tool history with no tools.** When the request has no tools, the tool calls and results in the history go to the model as text. Before, Bedrock answered with a 400. +- **Bedrock Converse, structured output.** On the Claude models that cannot force a tool, structured output uses the native JSON schema output (`outputConfig.textFormat`) instead of a forced tool. +- **Gemini** maps the choice to `functionCallingConfig` (`AUTO`, `NONE`, `ANY`, and `allowedFunctionNames`). With only provider tools, such as Google Search, it sends no tool config. +- **Grok, Mistral, and OpenRouter** (Chat Completions and Responses) send their own `tool_choice` shape. diff --git a/.changeset/tool-choice-and-subagent-tool.md b/.changeset/tool-choice-and-subagent-tool.md new file mode 100644 index 0000000000..9033ce617a --- /dev/null +++ b/.changeset/tool-choice-and-subagent-tool.md @@ -0,0 +1,11 @@ +--- +'@tanstack/ai': minor +'@tanstack/openai-base': minor +--- + +New `chat()` options for coding agents. + +- **`toolChoice`.** `chat({ toolChoice })` takes `'auto'`, `'none'`, `'required'`, or `{ type: 'tool', name }`. A middleware can change it for one model call. A tool choice in `modelOptions` wins. No tool choice goes out when the request has no tools. +- **`replaceResult`.** `onAfterToolCall` can return `{ type: 'replaceResult', result }` to change the result that the client and the next model call see. Several middlewares chain. +- **One subagent tool.** With `subagents: { tool: 'single' }`, the model picks an agent by name in one `subagent` tool. An agent with an `inputSchema` takes `input`, and an agent without one takes an optional `prompt`. A `sessionId` continues an earlier child of the same agent (with `withPersistence`), and `background` runs the child in the background. `LoadChild` can return the name of that agent. +- **`@tanstack/openai-base`** sends `tool_choice` from Chat Completions and Responses, and exports `toChatCompletionsToolChoice` and `toResponsesToolChoice`. diff --git a/.changeset/tool-choice-cli-adapters.md b/.changeset/tool-choice-cli-adapters.md new file mode 100644 index 0000000000..1956f46c76 --- /dev/null +++ b/.changeset/tool-choice-cli-adapters.md @@ -0,0 +1,12 @@ +--- +'@tanstack/ai-acp': patch +'@tanstack/ai-claude-code': patch +'@tanstack/ai-codex': patch +'@tanstack/ai-grok-build': patch +'@tanstack/ai-opencode': patch +--- + +The CLI-style adapters follow `chat({ toolChoice })` where they can. + +- **Claude Code.** `'none'` and a named tool turn off the built-in tools. The adapter then bridges none of your tools, or only the named one. `'required'` logs a warning. +- **Codex, OpenCode, Grok Build, and ACP-compatible adapters.** `'none'` and a named tool limit only your bridged tools. The built-in tools of the agent stay on, so these values and `'required'` log a warning. diff --git a/.changeset/turn-overrides.md b/.changeset/turn-overrides.md new file mode 100644 index 0000000000..ad608c1ff5 --- /dev/null +++ b/.changeset/turn-overrides.md @@ -0,0 +1,13 @@ +--- +'@tanstack/ai': minor +'@tanstack/ai-harness': minor +--- + +One prompt can now run with its own settings. + +- `session.prompt(message, { overrides })` and `session.followUp(message, { overrides })` take `TurnOverrides`: an `adapter`, `reasoning`, a `promptCache`, and extra `tools` for that one turn. They apply to every model call of the turn: the tool loop, `onModelError` retries, `beforeFinish` cycles, and joins. The next turn uses the defaults again. +- The overrides are kept in memory only. A queued turn keeps them, a steer that joins a turn uses that turn's overrides, and a turn that recovery runs again after a restart uses the defaults. +- A middleware that rebuilds the tools in `onConfig` keeps the override tools. A durable tool among them gets `step` and `append`. +- `HarnessConfig.reasoning` sets the default reasoning of every turn. +- The plugin `adapter` picker now gets the turn: `{ operationId, inputId, overrides }`. A picker with no parameter still works. +- `ChatMiddlewareConfig` has `promptCache`, next to `reasoning`. `onConfig` can change the prompt cache of the next model call. diff --git a/.changeset/workspace-outside.md b/.changeset/workspace-outside.md new file mode 100644 index 0000000000..1a41e367b1 --- /dev/null +++ b/.changeset/workspace-outside.md @@ -0,0 +1,5 @@ +--- +'@tanstack/ai-harness': minor +--- + +`workspaceTools({ root, outside: 'ask' })`: a path outside the workspace asks the user first, and a yes allows that folder for the rest of the session. The `bypass` mode allows it without a question. `list_files` and `grep` take an optional `path`, so the agent can search an allowed folder. Without the option, a path outside the workspace is refused, as before. diff --git a/CLAUDE.md b/CLAUDE.md index bbacdbd9e9..90d1e0d3aa 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -359,7 +359,7 @@ OPENAI_API_KEY=sk-... pnpm --filter @tanstack/ai-e2e record | New provider adapter | Add provider to `feature-support.ts` + `test-matrix.ts`. Existing feature tests auto-run. | | New feature (e.g., new generation type) | Add feature to types, feature config, support matrix. Create fixture + spec file. | | Bug fix in chat/streaming | Add a test case to `chat.spec.ts` or `tools-test/` that reproduces the bug. | -| Tool system change | Add scenario to `tools-test-scenarios.ts` + test in `tools-test/` specs. | +| Tool system change | Add scenario to `src/lib/tools-test-tools.ts` + test in `tools-test/` specs. | | Middleware change | Add test to `middleware.spec.ts` with appropriate scenario. | | Client-side change (useChat, etc.) | Add test covering the observable behavior change. | diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index b1c1c52d49..77847554c0 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -185,7 +185,7 @@ Tests are included in typecheck. `vite.config.ts` / `vitest.config.ts` are not | New provider adapter | Add provider to `feature-support.ts` + `test-matrix.ts`. Tests auto-run. | | New feature (e.g. new generation type) | Add to types, feature config, support matrix, fixture, spec file. | | Chat / streaming bug fix | Test case in `chat.spec.ts` or `tools-test/`. | -| Tool system change | Scenario in `tools-test-scenarios.ts` + spec. | +| Tool system change | Scenario in `src/lib/tools-test-tools.ts` + spec. | | Middleware change | Test in `middleware.spec.ts`. | | Client-side change (useChat etc.) | Test covering the observable behavior change. | diff --git a/docs/adapters/acp-compatible.md b/docs/adapters/acp-compatible.md index b5b18dc4b9..535fa7a389 100644 --- a/docs/adapters/acp-compatible.md +++ b/docs/adapters/acp-compatible.md @@ -312,6 +312,19 @@ const pi = acpCompatible({ `chat()`-provided tools bridged into the agent are always auto-approved, regardless of mode. +## Tool choice + +`chat({ toolChoice })` limits only the tools that the adapter bridges into the agent. ACP has no field to turn off the built-in tools of the agent or to force a tool call. The agent decides when it calls a tool. For the values, see [Choose when the model calls a tool](../tools/tools#choose-when-the-model-calls-a-tool). + +| Value | What the adapter does | +| --- | --- | +| `'auto'` | Bridges all of your tools. | +| `'none'` | Bridges none of your tools. | +| `{ type: 'tool', name }` | Bridges only the named tool. The agent does not have to call it. | +| `'required'` | Bridges all of your tools. The agent does not have to call one. | + +Each value except `'auto'` logs a warning the first time an adapter instance gets it. The warning starts with the harness `name`, for example `pi:`. + ## Session Resume On every run the adapter emits the harness session id as a CUSTOM event named `.session-id` (e.g. `pi.session-id`). Thread that id back through `modelOptions.sessionId` on the next call and the harness resumes the session — only the trailing user message is sent, since the agent already holds the prior context: diff --git a/docs/adapters/anthropic.md b/docs/adapters/anthropic.md index 4a36dbc589..9a6b635517 100644 --- a/docs/adapters/anthropic.md +++ b/docs/adapters/anthropic.md @@ -71,6 +71,59 @@ const config: Omit = { const adapter = createAnthropicChat("claude-sonnet-4-6", process.env.ANTHROPIC_API_KEY!, config); ``` +## Bearer and OAuth tokens + +Use `authToken` for a Bearer token. The adapter sends `Authorization: Bearer` and omits `x-api-key`: + +```typescript +import { chat } from '@tanstack/ai' +import { anthropicText } from '@tanstack/ai-anthropic' + +const stream = chat({ + adapter: anthropicText('claude-sonnet-5-5', { + authToken: process.env.ANTHROPIC_AUTH_TOKEN, + }), + messages: [{ role: 'user', content: 'Hello!' }], +}) + +for await (const chunk of stream) { + if (chunk.type === 'TEXT_MESSAGE_CONTENT') console.log(chunk.delta) +} +``` + +Without explicit credentials, the adapter reads the environment in this order: + +1. `ANTHROPIC_AUTH_TOKEN`. +2. `ANTHROPIC_OAUTH_TOKEN`. +3. `ANTHROPIC_API_KEY`. + +Explicit `authToken` or `apiKey` takes precedence over environment credentials. When both explicit values exist, `authToken` takes precedence. + +OAuth tokens containing `sk-ant-oat` are detected automatically. An environment `ANTHROPIC_OAUTH_TOKEN` also selects OAuth when `ANTHROPIC_AUTH_TOKEN` is absent. Set `oauth: true` to select OAuth explicitly. + +OAuth requests include the Claude Code identity system block, CLI identity headers, and the `claude-code-20250219` and `oauth-2025-04-20` betas. A Bearer token alone does not select OAuth. An injected SDK client owns its credentials. Adapter OAuth options still control the request identity. + +## Replay unsigned gateway thinking + +Some Anthropic-protocol gateways return readable thinking without a signature. Enable replay for those replies with `allowEmptySignature`: + +```typescript +import { anthropicText } from '@tanstack/ai-anthropic' + +const gateway = anthropicText('claude-sonnet-5-5', { + baseURL: 'https://gateway.example.com', + apiKey: process.env.GATEWAY_API_KEY, + provider: 'my-anthropic-gateway', + allowEmptySignature: true, +}) + +console.log(gateway.provider) +``` + +`allowEmptySignature` defaults to `false`. It permits ordinary unsigned thinking for matching-source history. Redacted thinking still requires its provider data. Foreign history uses the normal replay rules. + +Set `provider` to identify the gateway separately from direct Anthropic. Source matching compares the provider, API, and requested model. See [Keep saved history when you switch](../advanced/runtime-adapter-switching#keep-saved-history-when-you-switch). + ## Claude on Vertex Use `@tanstack/ai-anthropic/vertex` when Claude must run on Vertex AI. That @@ -233,81 +286,88 @@ A streamed response that stops at `max_tokens` ends in a `RUN_ERROR` with `code: One exception: structured output (`chat({ outputSchema })`) on models that use the non-streaming finalization path clamps this default to ~21K tokens. The Anthropic SDK rejects a non-streaming request whose `max_tokens` could exceed its 10-minute timeout, so the full ceiling can't be used there. Streaming chat is unaffected. To raise the structured-output ceiling toward a model's true max, stream the response. -### Thinking (Extended Thinking) +### Thinking -Enable extended thinking with a token budget. This allows Claude to show its reasoning process, which is streamed as `thinking` chunks: +Set how hard Claude thinks with `reasoning` on `chat()`. The adapter turns the level into the right thinking fields for the model: -```typescript ignore -modelOptions: { - thinking: { - type: "enabled", - budget_tokens: 2048, // Maximum tokens for thinking - }, -} +```typescript +import { chat } from "@tanstack/ai"; +import { anthropicText } from "@tanstack/ai-anthropic"; + +const stream = chat({ + adapter: anthropicText("claude-sonnet-5"), + messages: [{ role: "user", content: "Plan a database migration." }], + reasoning: "xhigh", +}); ``` -**Note:** `budget_tokens` must be less than `modelOptions.max_tokens` — set `max_tokens` high enough to leave room for the visible response alongside the thinking budget, or the request is rejected. +What the adapter sends for each kind of model: + +- **Claude 4.7 and later, Sonnet 5, Fable 5**: adaptive thinking, with the level as `output_config.effort`. +- **Claude Opus 4.6 and Sonnet 4.6**: adaptive thinking, with the level as `effort`. +- **Haiku 4.5, Sonnet 4.5, Opus 4.5, Opus 4.1**: thinking with a token budget. Set it with `reasoning: { level: "high", budgetTokens: 8000 }`. The adapter raises `max_tokens` when it is below the budget. +- **`off`**: thinking disabled. `claude-fable-5` and `claude-sonnet-5-5` always think, so their types do not take `off`. + +The thinking text streams back as thinking parts. Pass `summary: false` to keep it hidden: `reasoning: { level: "high", summary: false }`. See [Reasoning](../chat/reasoning) for the levels and how a level the model does not have moves to the nearest one. + +#### Change the level during a conversation -### Adaptive Thinking (Claude 4.6+, Sonnet 5, Fable 5) +A new level changes the start of the request, so Claude reads nothing from the cache. On `claude-fable-5-1`, `claude-opus-5`, and `claude-opus-5-5`, the level goes into the messages instead. Then a new level keeps the cached start. -Newer Claude models use adaptive thinking — the model decides when and how -much to think, and depth is tuned with `output_config.effort` instead of a -token budget: +On Anthropic's own API, the adapter does this for these models by default: ```typescript import { chat } from "@tanstack/ai"; import { anthropicText } from "@tanstack/ai-anthropic"; const stream = chat({ - adapter: anthropicText("claude-sonnet-5"), + adapter: anthropicText("claude-opus-5-5"), messages: [{ role: "user", content: "Plan a database migration." }], - modelOptions: { - thinking: { type: "adaptive", display: "summarized" }, - output_config: { effort: "xhigh" }, - max_tokens: 64_000, - }, + reasoning: "low", }); ``` -Per-model rules (enforced by the adapter's types): - -- **`claude-sonnet-5`, `claude-opus-4-8`, `claude-opus-4-7`** — adaptive - thinking with an explicit `{ type: "disabled" }` opt-out. The manual - `{ type: "enabled", budget_tokens }` shape is rejected with a 400, and - the sampling parameters (`temperature`, `top_p`, `top_k`) are not - accepted (on Sonnet 5 the API rejects non-default values; on Opus - 4.7/4.8 the parameters are removed entirely). -- **`claude-fable-5`** — thinking is always on. The only accepted explicit - config is `{ type: "adaptive" }` (both `disabled` and `budget_tokens` - return a 400), and sampling parameters are rejected. -- **`claude-sonnet-5-5`** — the types accept only `{ type: "adaptive" }`. - Both `disabled` and `budget_tokens` return a 400, and so do non-default - sampling values. To turn off up-front thinking, the API takes - `{ type: "between_tools" }`, which the adapter does not type yet. -- **`claude-opus-4-6` / `claude-sonnet-4-6`** — accept - `{ type: "adaptive" }` alongside the deprecated - `{ type: "enabled", budget_tokens }` shape, and still accept sampling - parameters. -- **`display`** defaults to `"omitted"` on Opus 4.7+ and the 5-generation - models — set `"summarized"` to stream the reasoning text. -- **`effort`** accepts `"low" | "medium" | "high" | "xhigh" | "max"`; - `"xhigh"` is available on Claude Opus 4.7+, Claude Sonnet 5, Claude - Sonnet 5.5, and Claude Fable 5. -- **`output_config`** is accepted on Claude Opus 4.7, Opus 4.8, Sonnet 5, - Fable 5, Opus 5, Fable 5.1, Opus 5.5, and Sonnet 5.5. When you also pass - an `outputSchema`, the adapter adds `output_config.format` and keeps the - `effort` you set. +With a custom `baseURL`, a custom `fetch`, or a proxy in `ANTHROPIC_BASE_URL`, it is off, because the endpoint must pass the betas. To turn it on there, set `midConversationEffort: true` in the `reasoning` config. `modelReasoning(record)` sets it for these models in the [model catalog](../models/catalog), also for their OpenRouter ids: + +```typescript +import { chat } from "@tanstack/ai"; +import { createAnthropicChat } from "@tanstack/ai-anthropic"; +import { getModel, modelReasoning } from "@tanstack/ai-models"; + +const record = getModel("openrouter", "anthropic/claude-opus-5.5"); +if (record) { + const stream = chat({ + adapter: createAnthropicChat( + record.id, + process.env.OPENROUTER_API_KEY ?? "", + { baseURL: record.baseUrl, reasoning: modelReasoning(record) }, + ), + messages: [{ role: "user", content: "Plan a database migration." }], + reasoning: "low", + }); +} +``` + +What the adapter sends with `midConversationEffort`: + +- `thinking` with `type: "adaptive"` and `block_binding`, and `output_config.effort: "high"`, on every request. +- At the end of the messages, a `system` message with no content and `output_config.effort` set to the level of this call. +- The same message before each earlier answer of this adapter, with the level of that answer. The answer keeps it in `metadata.tanstack.reasoningEffort`. +- The `anthropic-beta` header with `mid-conversation-output-config-2026-07-01` and `thinking-binding-controls-2026-08-01`. +- No `temperature`. ### Prompt Caching -Cache prompts for better performance and reduced costs: +`chat()` caches Claude prompts by default. It adds `cache_control` markers to the system prompt, the last tool, and the last user message. To send no markers, pass `promptCache: 'none'`. For the retention, the cache key, and the cost, see [Prompt Caching](../advanced/prompt-caching). + +To place a marker yourself, set `cache_control` in the `metadata` of a message part, a system prompt, or a tool. A marker of your own turns the automatic markers off for that request: ```typescript import { chat } from "@tanstack/ai"; import { anthropicText } from "@tanstack/ai-anthropic"; const stream = chat({ - adapter: anthropicText("claude-sonnet-4-6"), + adapter: anthropicText("claude-sonnet-5-5"), messages: [ { role: "user", @@ -327,6 +387,50 @@ const stream = chat({ }); ``` +`modelOptions.cache_control` asks Anthropic to place one marker for the whole request. It also turns the automatic markers off. + +#### Tools and prompts added during a conversation + +On some Claude models, a tool or a system prompt that you add between model calls goes into the conversation, not into `tools` or `system`. The marked start of the request stays the same, so Claude reads it from the cache. + +The models: `claude-opus-4-8`, `claude-opus-5`, `claude-opus-5-5`, `claude-fable-5`, and `claude-fable-5-1`. + +What the adapter sends on these models: + +- `system` keeps the system prompts of the first call, with their `cache_control`. +- In a request with tools, the `anthropic-beta` header has `mid-conversation-tool-changes-2026-07-01`, and `tools` has one placeholder tool, `__tanstack_deferred_placeholder__`. The model must never call it. It keeps Anthropic's hidden setup for added tools inside the cached start. +- An added tool goes to the end of `tools` with `defer_loading: true`. A `system` message lists it in a `tool_addition` block. +- A system prompt that you add goes into a `system` message as a text block. The message comes directly before the next assistant message, or at the end of the messages. +- With a provider tool such as `webSearchTool()` in the first call or in a change, `tools` is the full list and has no placeholder. + +The automatic tool marker goes on the last tool of the first call, not on the placeholder or an added tool. So the marked start does not move when a tool is added. A `cache_control` of your own still wins. + +The channels are on by default with Anthropic's own API. With a custom `baseURL`, a custom `fetch`, or a proxy in the `ANTHROPIC_BASE_URL` environment variable, they are off, and every request is the same as on a model outside the list. An adapter on your own client (`createAnthropicChatWithClient`, `anthropicVertexText`) has no channels. Set `midConversationChannels` to choose: + +- `false`: send the full lists on every call. +- `true`: use the channels with a custom `baseURL`, `fetch`, or `ANTHROPIC_BASE_URL`. Set it only when that endpoint sends the request and the `anthropic-beta` header to Anthropic as they are. +- `{ systemPrompts: true }` or `{ tools: true }`: use only that channel, for an endpoint that passes only one. With `{ systemPrompts: true }`, an added prompt goes into a `system` message, and an added tool goes out in the full `tools` list. + +```typescript +import { anthropicText } from "@tanstack/ai-anthropic"; + +export const fullLists = anthropicText("claude-opus-5-5", { + midConversationChannels: false, +}); + +export const throughProxy = anthropicText("claude-opus-5-5", { + baseURL: "https://llm-proxy.example.com", + midConversationChannels: true, +}); + +export const promptsOnly = anthropicText("claude-opus-5-5", { + baseURL: "https://llm-gateway.example.com", + midConversationChannels: { systemPrompts: true }, +}); +``` + +See [Mid-Conversation Changes](../advanced/mid-conversation-changes) for how the library finds the changes. + ## Summarization Anthropic supports text summarization: diff --git a/docs/adapters/bedrock.md b/docs/adapters/bedrock.md index df409a2961..ab9e4cd532 100644 --- a/docs/adapters/bedrock.md +++ b/docs/adapters/bedrock.md @@ -179,7 +179,9 @@ const adapter = createBedrockText( ### Prompt caching -Add a `cachePoint` to make a prompt prefix eligible for caching. Later requests can read matching tokens at the reduced cache rate. Bedrock bills cache misses at the standard input rate. +`chat()` adds cache points for Claude models by default: one after the system prompt and one at the end of the last user message. Other models get no automatic cache point. To turn this off, pass `promptCache: 'none'`. For the retention and the cost, see [Prompt Caching](../advanced/prompt-caching). + +To choose the places yourself, set `metadata.cachePoint`. A cache point of your own turns the automatic cache points off for that request. Later requests can read matching tokens at the reduced cache rate. Bedrock bills cache misses at the standard input rate. Explicit prompt caching is model-dependent. Use `cachePoint` only with a model that AWS lists as supporting it. The minimum checkpoint size and the TTL options also vary by model. @@ -229,7 +231,18 @@ Tools take the same metadata. Pass `metadata: { cachePoint: { type: 'default' } ### Token usage -`onUsage` and `RUN_FINISHED.usage` report Bedrock's counts as `promptTokens`, `completionTokens`, and `totalTokens`. When a request hits or writes a prompt cache, the cache counts arrive on `promptTokensDetails.cachedTokens` and `promptTokensDetails.cacheWriteTokens`. Bedrock counts only the uncached part of the input in `promptTokens`, so add the two cache counts to it to get the full input size. +`onUsage` and `RUN_FINISHED.usage` report Bedrock's counts as `promptTokens`, `completionTokens`, and `totalTokens`. `promptTokens` is the full input size, cached tokens included. When a request hits or writes a prompt cache, the cache counts arrive on `promptTokensDetails.cachedTokens` and `promptTokensDetails.cacheWriteTokens`. + +### Structured output + +Pass `outputSchema` to `chat()` to get an object that matches your schema. For the full steps, see [One-Shot Extraction](../structured-outputs/one-shot). + +Converse gets the object in one of two ways: + +- **Forced tool**: most models. The adapter forces a `structured_output` tool that takes your schema. +- **Native JSON schema output** (`outputConfig.textFormat`): the Claude models that do not take a forced tool. These are `claude-fable-5-1`, `claude-mythos-5-1`, `claude-opus-5-5`, `claude-sonnet-5-5`, and all Claude models with thinking on. + +If the native answer is not valid JSON, the call fails with an error that names the model. ## Chat Completions API (`api: 'chat'`) @@ -356,6 +369,20 @@ The adapter ships with a hand-seeded snapshot catalog (`src/model-catalog.genera For the full list of models and which API endpoints they support, see the [AWS API compatibility matrix](https://docs.aws.amazon.com/bedrock/latest/userguide/models-api-compatibility.html). +## Replay thinking and tool results + +A Claude conversation with thinking and tools needs its signed thinking on the next request. Converse preserves the reasoning text, signature, and order before the matching `toolUse` blocks. + +Readable `reasoningText.signature` replay applies to Claude models. Same-source redacted encrypted reasoning can replay for other Converse models that support it. Foreign history drops redacted thinking and signatures, and converts readable thinking to text. + +Keep assistant metadata and reasoning parts when you save history. Converse records source provider `amazon-bedrock` and API `bedrock-converse-stream`. The Chat Completions and Responses adapters use `openai-completions` and `openai-responses` API identities. + +Converse tool results preserve images as image blocks and send `status: 'error'` for tool errors. With a text-only model, image results use a text placeholder. The Chat Completions path places supported tool-result images in a following user message. + +Converse accepts tool calls and tool results in the history only when the request has tools. If a request has no tools, the adapter sends them as text, for example `[Tool call lookup_weather({"location":"Paris"})]`. Images in a tool result stay images. Your saved messages do not change. + +An AWS HTTP request ID is a transport ID. It does not become `responseId`. See [Read the provider response identity](../chat/stream-events#read-the-provider-response-identity). + ## Supported Capabilities - Streaming chat completions diff --git a/docs/adapters/byteplus.md b/docs/adapters/byteplus.md index cea9c1d760..2c814c7787 100644 --- a/docs/adapters/byteplus.md +++ b/docs/adapters/byteplus.md @@ -142,7 +142,7 @@ export function Chat() { ### Model options -Ark's chat endpoint is OpenAI-compatible, so sampling parameters keep their OpenAI snake_case names and live in `modelOptions`. `thinking`, `reasoning_effort`, `repetition_penalty` and `service_tier` are the Ark-only additions: +Ark's chat endpoint is OpenAI-compatible, so sampling parameters keep their OpenAI snake_case names and live in `modelOptions`. `repetition_penalty` and `service_tier` are the Ark-only additions. The reasoning level goes in `reasoning`: ```typescript import { chat, toServerSentEventsResponse } from '@tanstack/ai' @@ -158,25 +158,21 @@ export async function POST(request: Request) { temperature: 0.7, top_p: 0.9, max_tokens: 2048, - thinking: { type: 'enabled' }, - reasoning_effort: 'medium', }, + reasoning: 'medium', }) return toServerSentEventsResponse(stream) } ``` -Two constraints the type system can't express, both live-verified as `400`s: - -- `max_tokens` and `max_completion_tokens` are mutually exclusive. -- `reasoning_effort` cannot be combined with `thinking: { type: 'disabled' }`. +`max_tokens` and `max_completion_tokens` are mutually exclusive. The type system can't express this, and Ark returns a `400` when you send both. `service_tier: 'flex'` routes the request to the cheaper offline batch queue with no latency guarantee. ## Reasoning and `encrypted_content` -Seed models reason by default. Reasoning arrives as its own stream of `reasoning_content` deltas and is surfaced as reasoning content rather than answer text, so `useChat` renders it separately from the reply. Turn it off per request: +Seed models reason by default. Reasoning arrives as its own stream of `reasoning_content` deltas and is surfaced as reasoning content rather than answer text, so `useChat` renders it separately from the reply. The adapter sends `reasoning` as Ark's `thinking.type`, plus `reasoning_effort` when the level has an effort. Turn reasoning off per request on a model that can stop thinking: ```typescript import { chat, toServerSentEventsResponse } from '@tanstack/ai' @@ -186,16 +182,16 @@ export async function POST(request: Request) { const { messages } = await request.json() const stream = chat({ - adapter: byteplusText('dola-seed-2-1-turbo-260628'), + adapter: byteplusText('glm-5-2-260617'), messages, - modelOptions: { thinking: { type: 'disabled' } }, + reasoning: 'off', }) return toServerSentEventsResponse(stream) } ``` -`disabled` works everywhere; `auto` is accepted only by `gpt-oss-120b-250805`. `deepseek-v3-2-251201` is the one model that defaults to reasoning *off*. +`off` sends `thinking: { type: 'disabled' }` and no effort, because Ark rejects the pair. `deepseek-v3-2-251201` is the one model that defaults to reasoning *off*. The four "thinking summary" models — `dola-seed-2-1-turbo-260628`, `seed-2-0-lite-260428`, `seed-2-0-mini-260428` and `seed-2-0-pro-260328` — also emit an opaque `encrypted_content` blob alongside the reasoning trace. It is a signature over that trace, and BytePlus's docs ask for it back verbatim on the assistant message in the next turn. diff --git a/docs/adapters/claude-code.md b/docs/adapters/claude-code.md index 0673bb562a..81c9bb976c 100644 --- a/docs/adapters/claude-code.md +++ b/docs/adapters/claude-code.md @@ -189,6 +189,19 @@ const stream = chat({ **Client-side and approval-gated tools are not supported.** The harness executes tools inside a live subprocess, which cannot pause across HTTP requests to wait for a browser round-trip or a human approval. Passing a tool without a server `execute()` implementation — or one marked `needsApproval` — fails fast with a descriptive error. Run those tools outside the harness with a regular provider adapter. +## Tool choice + +`chat({ toolChoice })` sets which tools Claude Code can call. For the values, see [Choose when the model calls a tool](../tools/tools#choose-when-the-model-calls-a-tool). + +| Value | What the adapter does | +| --- | --- | +| `'auto'` | No change. Claude Code can call every tool. | +| `'none'` | Turns off the built-in tools and bridges none of your tools. Claude Code answers in text. | +| `{ type: 'tool', name }` | Turns off the built-in tools and bridges only the named tool. Claude Code can call only this tool, but it does not have to. | +| `'required'` | No change. Claude Code cannot force a tool call, so the adapter logs a warning the first time. | + +For `'none'` and a named tool, the adapter passes `--tools ""` and `--strict-mcp-config`. Then other MCP servers, for example from the workspace or a settings file, do not load. + ## Structured Output Pass `outputSchema` on `chat()`. Claude Code runs one harness turn, uses its native tools, and returns a typed object. The schema JSON is passed to `--json-schema` as inline JSON (the CLI rejects a file path). Tool activity and prose stream as usual. The object arrives as `structured-output.complete`, including when Claude delivers it through its built-in `StructuredOutput` tool. diff --git a/docs/adapters/cloudflare.md b/docs/adapters/cloudflare.md index f06a80fbce..80ea199d46 100644 --- a/docs/adapters/cloudflare.md +++ b/docs/adapters/cloudflare.md @@ -151,6 +151,63 @@ The first argument is the provider slug from your gateway dashboard (`openai`, ` The `ts-react-chat` example does this. Set `CLOUDFLARE_ACCOUNT_ID`, `CLOUDFLARE_API_TOKEN`, and `CLOUDFLARE_AI_GATEWAY_ID` in its `.env`, and the Cloudflare, OpenAI, Anthropic, and Groq models in its picker all go through your gateway. +### Claude and GPT from a Worker + +Inside a Worker, the AI Gateway `anthropic/...` and `openai/...` models use their provider's own API, not Chat Completions. Keep the provider adapter, and give it `cloudflareBindingFetch` as its `fetch`. The binding signs the requests, so the Worker needs no provider key: + +```typescript +import { chat, toServerSentEventsResponse } from "@tanstack/ai"; +import { createAnthropicChat } from "@tanstack/ai-anthropic"; +import { cloudflareBindingFetch } from "@tanstack/ai-cloudflare"; +import type { Ai } from "@cloudflare/workers-types"; + +interface Env { + AI: Ai; +} + +export default { + async fetch(request: Request, env: Env) { + const { messages } = await request.json(); + + // The SDK needs a key value. The binding does not use it. + const adapter = createAnthropicChat("claude-opus-5-5", "cloudflare-binding", { + fetch: cloudflareBindingFetch({ + binding: env.AI, + vendor: "anthropic", + gateway: { id: "default" }, + }), + }); + + const stream = chat({ adapter, messages, reasoning: "high" }); + return toServerSentEventsResponse(stream); + }, +}; +``` + +- `vendor: "anthropic"` sends Anthropic Messages requests to `anthropic/`. Use it with `createAnthropicChat`. +- `vendor: "openai"` sends OpenAI Responses requests to `openai/`. Use it with `createOpenaiChat`. +- The adapter keeps all of its options: `chat({ reasoning })`, tools, `cache_control` for prompt caching, and Anthropic betas, which go out as the `anthropic-beta` header. +- For `@cf/...` models and other gateway vendors, keep `createCloudflareText` with `binding: env.AI`. + +An adapter with a gateway `baseURL` (`cloudflareGateway()`) or with `cloudflareBindingFetch` sends the full tools and system prompts on every call, also on a model with a [mid-conversation channel](../advanced/mid-conversation-changes). AI Gateway then gets the same request as on any other model. To send tools and prompts added during a conversation through the channel, set `midConversationChannels: true`: + +```typescript +import { createAnthropicChat } from "@tanstack/ai-anthropic"; +import { cloudflareBindingFetch } from "@tanstack/ai-cloudflare"; +import type { Ai } from "@cloudflare/workers-types"; + +export function gatewayClaude(env: { AI: Ai }) { + return createAnthropicChat("claude-opus-5-5", "cloudflare-binding", { + fetch: cloudflareBindingFetch({ + binding: env.AI, + vendor: "anthropic", + gateway: { id: "default" }, + }), + midConversationChannels: true, + }); +} +``` + ## Bring your own key Two different things go by this name. Both work. @@ -250,7 +307,7 @@ export default { ## Model options -Sampling and reasoning controls go in `modelOptions`. Reasoning models stream their thinking as `reasoning_content`, which shows up as `REASONING_*` events. +Sampling controls go in `modelOptions`, and the reasoning level goes in `reasoning`. Reasoning models stream their thinking as `reasoning_content`, which shows up as `REASONING_*` events. ```typescript import { chat } from "@tanstack/ai"; @@ -262,12 +319,13 @@ const stream = chat({ modelOptions: { temperature: 0.3, max_tokens: 512, - reasoning_effort: "low", - chat_template_kwargs: { enable_thinking: false }, }, + reasoning: "low", }); ``` +`reasoning` goes out as `reasoning_effort`. `off` sends `null`, which turns reasoning off. Model template settings, such as `chat_template_kwargs`, stay in `modelOptions`. + ## Evaluate Use `cloudflareDecider('typesafe/jev')` with `decide()`. diff --git a/docs/adapters/codex.md b/docs/adapters/codex.md index f838c3b9c6..f9ec5cec66 100644 --- a/docs/adapters/codex.md +++ b/docs/adapters/codex.md @@ -73,7 +73,7 @@ const stream = chat({ | `cwd` | Working directory for the harness session. Defaults to `process.cwd()`. | | `sandboxMode` | Codex sandbox: `'read-only'`, `'workspace-write'`, or `'danger-full-access'`. Default is `'workspace-write'` on local-process and Docker. Default is `'danger-full-access'` on Daytona and Cloudflare, because those providers cannot create a nested bubblewrap namespace. Isolation is then the outer VM plus `defineSandboxPolicy`. | | `approvalPolicy` | Codex approval policy. Defaults to `'never'` — headless runs have no approval UI, so anything else can stall a turn. | -| `modelReasoningEffort` | `'minimal'` \| `'low'` \| `'medium'` \| `'high'` \| `'xhigh'`. | +| `modelReasoningEffort` | The default effort when a call sets no `reasoning`: `'minimal'` \| `'low'` \| `'medium'` \| `'high'`. | | `skipGitRepoCheck` | Skip the harness's git-repo safety check. Defaults to `true` (server adapters routinely point at scratch directories). | | `networkAccessEnabled` | Allow network access inside the `workspace-write` sandbox. | | `webSearchMode` | `'disabled'` \| `'cached'` \| `'live'`. | @@ -86,8 +86,9 @@ const stream = chat({ | `config` | Extra `--config key=value` overrides passed to the Codex CLI (e.g. additional `mcp_servers` entries). | Per-call overrides go through `modelOptions`: `sessionId`, `sandboxMode`, -`approvalPolicy`, `modelReasoningEffort`, `workingDirectory`, -`skipGitRepoCheck`, and `authMode`. +`approvalPolicy`, `workingDirectory`, `skipGitRepoCheck`, and `authMode`. +Set the effort per call with `reasoning` on `chat()`; the adapter sends it as +`model_reasoning_effort`. ## Stateful Sessions @@ -189,6 +190,19 @@ const stream = chat({ **Client-side and approval-gated tools are not supported.** The harness executes tools inside a live subprocess, which cannot pause across HTTP requests to wait for a browser round-trip or a human approval. Passing a tool without a server `execute()` implementation — or one marked `needsApproval` — fails fast with a descriptive error. Run those tools outside the harness with a regular provider adapter. +## Tool choice + +`chat({ toolChoice })` limits only the tools that the adapter bridges into Codex. The built-in Codex tools always stay on, and Codex decides when it calls a tool. For the values, see [Choose when the model calls a tool](../tools/tools#choose-when-the-model-calls-a-tool). + +| Value | What the adapter does | +| --- | --- | +| `'auto'` | Bridges all of your tools. | +| `'none'` | Bridges none of your tools. | +| `{ type: 'tool', name }` | Bridges only the named tool. Codex does not have to call it. | +| `'required'` | Bridges all of your tools. Codex does not have to call one. | + +Each value except `'auto'` logs a warning the first time an adapter instance gets it. + ## Structured Output Pass `outputSchema` on `chat()`. Codex runs one harness turn and constrains the last message with `--output-schema`. Tool activity and assistant text stream as Codex writes them. The last message is also parsed as the schema object and arrives as `structured-output.complete`. diff --git a/docs/adapters/gemini.md b/docs/adapters/gemini.md index 557918ef47..0614cb0661 100644 --- a/docs/adapters/gemini.md +++ b/docs/adapters/gemini.md @@ -39,8 +39,6 @@ Need Gemini on Vertex AI (regional endpoints and Google Cloud credentials)? Use Use `gemini-3.8-flash` for chat with multimodal input, thinking, and built-in tools. It also supports structured output and caching. -For Gemini 3.8 Flash, set `modelOptions.thinkingConfig.thinkingLevel` to `LOW`, `MEDIUM`, or `HIGH`. The Interactions adapter uses `modelOptions.generation_config.thinking_level` with `low`, `medium`, or `high`. Gemini 3.8 Flash does not accept the `minimal` thinking level. - ```typescript import { chat } from "@tanstack/ai"; import { geminiText } from "@tanstack/ai-gemini"; @@ -359,13 +357,13 @@ const stream = chat({ // snake_case generation config distinct from geminiText's camelCase one. generation_config: { - thinking_level: "low", - thinking_summaries: "auto", stop_sequences: [""], }, response_modalities: ["text"], }, + // Sent as generation_config.thinking_level and thinking_summaries. + reasoning: "low", }); ``` @@ -424,16 +422,25 @@ const stream = chat({ ### Thinking -Enable thinking for models that support it: +Set how hard Gemini thinks with `reasoning` on `chat()`: -```typescript ignore -modelOptions: { - thinking: { - includeThoughts: true, - }, -} +```typescript +import { chat } from "@tanstack/ai"; +import { geminiText } from "@tanstack/ai-gemini"; + +const stream = chat({ + adapter: geminiText("gemini-3.8-flash"), + messages: [{ role: "user", content: "Plan a trip to Kyoto." }], + reasoning: "high", +}); ``` +- **Gemini 3 models** take thinking levels. The adapter sends `thinkingConfig.thinkingLevel`. Gemini 3.8 Flash has `low`, `medium`, and `high`. +- **Gemini 2.5 models** take a token budget. The adapter sends `thinkingConfig.thinkingBudget`, from `budgetTokens` or a default for the level. `off` sends a budget of `0`. +- **The Interactions adapter** sends `generation_config.thinking_level`. + +The thinking text streams back as thinking parts. `summary: false` sets `includeThoughts: false`. See [Reasoning](../chat/reasoning). + ### Structured Output Configure structured output format: diff --git a/docs/adapters/grok-build.md b/docs/adapters/grok-build.md index 5b9ecf00f2..402ff079be 100644 --- a/docs/adapters/grok-build.md +++ b/docs/adapters/grok-build.md @@ -190,6 +190,24 @@ tools inside a live process and can't pause across an HTTP round-trip. A tool without a server `execute()` (or marked `needsApproval`) fails fast; run those with a regular provider adapter. +## Tool choice + +`chat({ toolChoice })` limits only the tools that the adapter bridges into +Grok Build. The built-in Grok Build tools always stay on, and Grok Build +decides when it calls a tool. The rules are the same on both protocols. For +the values, see +[Choose when the model calls a tool](../tools/tools#choose-when-the-model-calls-a-tool). + +| Value | What the adapter does | +| --- | --- | +| `'auto'` | Bridges all of your tools. | +| `'none'` | Bridges none of your tools. | +| `{ type: 'tool', name }` | Bridges only the named tool. Grok Build does not have to call it. | +| `'required'` | Bridges all of your tools. Grok Build does not have to call one. | + +Each value except `'auto'` logs a warning the first time an adapter instance +gets it. + ## Durable runs A durable sandbox run (you pass `runs` and `durability` to `withSandbox`) diff --git a/docs/adapters/groq.md b/docs/adapters/groq.md index 0f65d10b83..006d59325f 100644 --- a/docs/adapters/groq.md +++ b/docs/adapters/groq.md @@ -170,14 +170,21 @@ const stream = chat({ ### Reasoning -Enable reasoning for models that support it (e.g., `openai/gpt-oss-120b`, `qwen/qwen3-32b`). This allows the model to show its reasoning process, which is streamed as `thinking` chunks: +Set how hard a reasoning model (`openai/gpt-oss-120b`, `qwen/qwen3-32b`) thinks with `reasoning` on `chat()`: -```typescript ignore -modelOptions: { - reasoning_effort: "medium", // "none" | "default" | "low" | "medium" | "high" -} +```typescript +import { chat } from "@tanstack/ai"; +import { groqText } from "@tanstack/ai-groq"; + +const stream = chat({ + adapter: groqText("openai/gpt-oss-120b"), + messages: [{ role: "user", content: "Hello!" }], + reasoning: "medium", +}); ``` +The adapter sends `reasoning_effort`. Qwen 3 only turns thinking on or off: `high` sends `default`, and `off` sends `none`. To pick how the reasoning text comes back, set `reasoning_format` in `modelOptions`. See [Reasoning](../chat/reasoning). + ## Summarization Summarize long text content: diff --git a/docs/adapters/llmgateway.md b/docs/adapters/llmgateway.md index 017569959b..30297512b7 100644 --- a/docs/adapters/llmgateway.md +++ b/docs/adapters/llmgateway.md @@ -137,12 +137,12 @@ const stream = chat({ modelOptions: { temperature: 0.7, max_completion_tokens: 4096, - reasoning_effort: "high", }, + reasoning: "high", }); ``` -`reasoning_effort` accepts the extended scale `none` / `minimal` / `low` / `medium` / `high` / `xhigh` / `max` in addition to OpenAI's standard tiers — which tiers a model honors depends on the model and provider it is routed to (see the model's page on [llmgateway.io/models](https://llmgateway.io/models)). +`reasoning` goes out as `reasoning_effort`. The types list the levels each model has, from the model catalog. See [Reasoning](../chat/reasoning). Reasoning models stream their thinking as `reasoning_content` deltas, which the adapter surfaces as AG-UI `REASONING_*` events. diff --git a/docs/adapters/mistral.md b/docs/adapters/mistral.md index fde1c10ffe..5fe17128c3 100644 --- a/docs/adapters/mistral.md +++ b/docs/adapters/mistral.md @@ -434,6 +434,14 @@ Creates a Mistral text adapter with an explicit API key. **Returns:** A Mistral text adapter instance. +## Images in tool results + +A tool can return an image for a model that accepts image input. Mistral keeps the text and image URL blocks together in the tool message's content. + +Image-only results use `(see attached image)` as the tool text. Empty results use `(no tool output)`. A text-only model receives an image-omission placeholder without the image blocks. + +Keep tool-result content as content parts when you save history. JSON text that contains base64 image data is still text. See [Tool Definition](../tools/tools#tool-definition). + ## Limitations - **Embeddings**: Use the [Mistral SDK](https://github.com/mistralai/client-ts) directly for `mistral-embed`. diff --git a/docs/adapters/ollama.md b/docs/adapters/ollama.md index 785a088588..a27b048e6a 100644 --- a/docs/adapters/ollama.md +++ b/docs/adapters/ollama.md @@ -159,9 +159,13 @@ export async function POST(request: Request) { **Note:** Tool support varies by model. Models like `llama3`, `mistral`, and `qwen2` generally have good tool calling support. +## Tool choice + +Ollama has no tool choice. The adapter ignores `chat({ toolChoice })`, so the model decides to call a tool or not, also with `'none'` or `'required'`. To stop tool calls for a request, do not pass tools. For the other providers, see [Choose when the model calls a tool](../tools/tools#choose-when-the-model-calls-a-tool). + ## Model Options -Ollama supports various provider-specific options. Unlike the other providers, Ollama nests its sampling and runner parameters inside an `options` object **within** `modelOptions` — `temperature`, `top_p`, and `num_predict` (the token-limit key) all live under `modelOptions.options`: +Ollama supports various provider-specific options. Unlike the other providers, Ollama nests its sampling and runner parameters inside an `options` object **within** `modelOptions`. `temperature`, `top_p`, and `num_predict` (the token-limit key) all live under `modelOptions.options`: ```typescript import { chat } from "@tanstack/ai"; @@ -258,7 +262,7 @@ const result = await embed({ console.log(result.embeddings[0]?.vector); ``` -Known models (`nomic-embed-text`, `mxbai-embed-large`, `all-minilm`, `snowflake-arctic-embed`, `bge-m3`, `embeddinggemma`) get autocomplete, and any other model name is accepted. Pull the model first with `ollama pull nomic-embed-text`. Output dimensions are fixed per model — the top-level `dimensions` option is not supported. +Known models (`nomic-embed-text`, `mxbai-embed-large`, `all-minilm`, `snowflake-arctic-embed`, `bge-m3`, `embeddinggemma`) get autocomplete, and any other model name is accepted. Pull the model first with `ollama pull nomic-embed-text`. Output dimensions are fixed per model, so the top-level `dimensions` option is not supported. See the [Embeddings guide](../embeddings.md) for the full API. @@ -334,7 +338,7 @@ Creates an Ollama text/chat adapter with an explicit host or client config. ### `ollamaSummarize(model)` / `createOllamaSummarize(model, hostOrConfig?)` -Creates an Ollama summarization adapter — same signature shape as the chat adapter. +Creates an Ollama summarization adapter, with the same signature shape as the chat adapter. ## Benefits of Ollama diff --git a/docs/adapters/openai-compatible.md b/docs/adapters/openai-compatible.md index fd027f5ef0..9e341e1f34 100644 --- a/docs/adapters/openai-compatible.md +++ b/docs/adapters/openai-compatible.md @@ -63,7 +63,24 @@ const stream = chat({ }); ``` -`deepseek("deepseek-reasoner")` is valid; `deepseek("gpt-4o")` is a type error — only declared models are accepted. +`deepseek("deepseek-reasoner")` is valid. `deepseek("gpt-5.5")` is a type error. Only declared models are accepted. + +## Replay and stream errors + +Your provider can receive history from a different model or API. The configured `name` identifies the source provider. Chat Completions uses the `openai-completions` API identity. A compatible Responses adapter uses `openai-responses`. + +Keep the assistant's `metadata.tanstack.source` when you save history. The target adapter removes foreign signatures and remaps tool IDs together with their results. See [Keep saved history when you switch](../advanced/runtime-adapter-switching#keep-saved-history-when-you-switch). + +The Chat Completions adapter handles tool history and output as follows: + +- With tool history and no active tools, it sends `tools: []`. +- Tool-result text stays in the tool message. Supported images follow in a user message. +- Image-only results use `(see attached image)`. An empty result uses `(no tool output)`. +- A text-only model receives an image-omission placeholder without the images. + +`delta.content` can be a string, `null`, or absent. An object or array produces `RUN_ERROR`. Unknown finish reasons also produce `RUN_ERROR` with `Provider finish_reason: `. + +Outgoing text removes lone UTF-16 surrogates. These are broken halves of a Unicode character. Valid pairs, such as emoji, stay intact. ## One-Shot Usage @@ -108,6 +125,46 @@ const provider = openaiCompatible({ > Capabilities are enforced at the type level. If a provider rejects a feature at runtime (e.g. tools on a model that doesn't support them), declare that model with `createModel` and omit the unsupported feature so the types stop you from calling it. +## Reasoning Models + +Reasoning models on OpenAI-compatible endpoints do not agree on the wire format. DeepSeek wants `thinking: { type }`, Qwen wants `enable_thinking`, and some servers reject the `developer` role. Tell the adapter what the provider expects with `compat`, and give each reasoning model its levels: + +```typescript +import { chat } from "@tanstack/ai"; +import { openaiCompatible } from "@tanstack/ai-openai/compatible"; + +const deepseek = openaiCompatible({ + baseURL: "https://api.deepseek.com", + apiKey: process.env.DEEPSEEK_API_KEY!, + compat: { + thinkingFormat: "deepseek", + maxTokensField: "max_tokens", + requiresReasoningContentOnAssistantMessages: true, + }, + models: [ + { + name: "deepseek-v4-flash", + // Each level and the value the provider takes. `null`: no such level. + reasoning: { off: "none", minimal: null, medium: null, high: "high", max: "max" }, + }, + { name: "deepseek-chat", reasoning: false }, + ], +}); + +const stream = chat({ + adapter: deepseek("deepseek-v4-flash"), + messages: [{ role: "user", content: "Prove that the square root of 2 is irrational." }], + reasoning: "max", +}); +``` + +- `reasoning` on a model takes a level map, `true` for every level up to `high`, or `false` for a model that does not reason. The `reasoning` option on `chat()` is then typed to those levels. +- `thinkingFormat` picks the request shape: `openai`, `deepseek`, `zai`, `qwen`, `qwen-chat-template`, `chat-template`, `baseten`, `openrouter`, `together`, `string-thinking`, or `ant-ling`. +- `compat` on a model entry overrides the provider's `compat` for that model. +- `requiresReasoningContentOnAssistantMessages` sends the earlier thinking back on each assistant turn, which DeepSeek needs. + +The [model catalog](../models/catalog) has the levels and `compat` for many providers, from the same data these fields use. + ## Configuration `openaiCompatible` accepts every OpenAI SDK `ClientOptions` field besides `apiKey`/`baseURL` (which are required and promoted to the top level). The most useful are `defaultHeaders` and `defaultQuery`, for providers that need extra auth or routing parameters: @@ -139,7 +196,7 @@ import { openaiCompatible } from "@tanstack/ai-openai/compatible"; const provider = openaiCompatible({ baseURL: "https://my-resource.openai.azure.com/openai/v1", apiKey: process.env.AZURE_OPENAI_API_KEY!, - models: ["gpt-4o"], + models: ["gpt-5.5"], api: "responses", // default is "chat-completions" }); ``` @@ -229,22 +286,19 @@ const litellm = openaiCompatible({ ## Azure OpenAI -Azure uses a resource-scoped URL and a separate API-version. Use the `/openai/v1` endpoint with `defaultQuery` for the version and `defaultHeaders` for the `api-key` header: +Use `azureOpenaiText` for Azure's Responses API. It configures the `api-key` header, endpoint, API version, and deployment mapping: ```typescript -import { openaiCompatible } from "@tanstack/ai-openai/compatible"; +import { azureOpenaiText } from '@tanstack/ai-openai' -const azure = openaiCompatible({ - name: "azure", - baseURL: "https://YOUR_RESOURCE.openai.azure.com/openai/v1", - apiKey: process.env.AZURE_OPENAI_API_KEY!, // also sent as Bearer; Azure accepts the api-key header below - models: ["gpt-4o"], // your Azure deployment name - defaultQuery: { "api-version": "2026-01-01-preview" }, - defaultHeaders: { "api-key": process.env.AZURE_OPENAI_API_KEY! }, -}); +const azure = azureOpenaiText('gpt-5.5', { + resourceName: 'my-resource', + apiKey: process.env.AZURE_OPENAI_API_KEY, + deploymentName: 'production-chat', +}) ``` -> Confirm the current `api-version` and endpoint shape in Azure's documentation — Azure's API surface evolves independently of OpenAI's. +See [Azure OpenAI](./openai#azure-openai) for environment variables and configuration precedence. ## Example: With Tools diff --git a/docs/adapters/openai.md b/docs/adapters/openai.md index d816924139..e6614bd4e4 100644 --- a/docs/adapters/openai.md +++ b/docs/adapters/openai.md @@ -2,12 +2,11 @@ title: OpenAI id: openai-adapter order: 1 -description: "Use OpenAI models with TanStack AI — GPT-4o, GPT-5, DALL-E image generation, TTS, and Whisper transcription via @tanstack/ai-openai." +description: "Use OpenAI models with TanStack AI: GPT-5.5 chat, image generation, speech, and transcription through @tanstack/ai-openai." keywords: - tanstack ai - openai - - gpt-4o - - gpt-5 + - gpt-5.5 - dall-e - whisper - openai tts @@ -15,7 +14,7 @@ keywords: - chatgpt --- -The OpenAI adapter provides access to OpenAI's models, including GPT-4o, GPT-5, image generation (DALL-E), text-to-speech (TTS), and audio transcription (Whisper). +Use the OpenAI adapter for GPT-5.5 chat, image generation, speech, and audio transcription. > Using a third-party provider that speaks the OpenAI API (DeepSeek, Moonshot/Kimi, Together, Fireworks, a local LM Studio/vLLM server, …)? See the [OpenAI-Compatible Adapter](./openai-compatible) for a generic `openaiCompatible({ baseURL, apiKey, models })` factory. @@ -41,7 +40,7 @@ import { chat } from "@tanstack/ai"; import { openaiText } from "@tanstack/ai-openai"; const stream = chat({ - adapter: openaiText("gpt-5.2"), + adapter: openaiText("gpt-5.5"), messages: [{ role: "user", content: "Hello!" }], }); ``` @@ -66,7 +65,7 @@ import { chat } from "@tanstack/ai"; import { openaiChatCompletions } from "@tanstack/ai-openai"; const stream = chat({ - adapter: openaiChatCompletions("gpt-5.2"), + adapter: openaiChatCompletions("gpt-5.5"), messages: [{ role: "user", content: "Hello!" }], }); ``` @@ -77,7 +76,7 @@ With an explicit API key: import { chat } from "@tanstack/ai"; import { createOpenaiChatCompletions } from "@tanstack/ai-openai"; -const adapter = createOpenaiChatCompletions("gpt-5.2", process.env.OPENAI_API_KEY!, { +const adapter = createOpenaiChatCompletions("gpt-5.5", process.env.OPENAI_API_KEY!, { // organization, baseURL, headers — all optional }); @@ -87,7 +86,61 @@ const stream = chat({ }); ``` -Both adapters work identically with [Structured Outputs](../structured-outputs/overview) — including `stream: true` — and accept the same `modelOptions` (temperature, top_p, max_tokens, stop, …). The reasoning section below applies to `openaiText`; `openaiChatCompletions` accepts `modelOptions.reasoning.effort` but cannot stream summary text. +Both adapters support [Structured Outputs](../structured-outputs/overview), including `stream: true`. Their `modelOptions` follow the selected API. The reasoning section below applies to `openaiText`. `openaiChatCompletions` cannot stream reasoning summary text. + +## Azure OpenAI + +Use `azureOpenaiText` when your OpenAI model runs on Azure. The adapter uses Azure's Responses API and `api-key` authentication. + +```typescript +import { chat } from '@tanstack/ai' +import { azureOpenaiText } from '@tanstack/ai-openai' + +const stream = chat({ + adapter: azureOpenaiText('gpt-5.5', { + resourceName: 'my-resource', + apiKey: process.env.AZURE_OPENAI_API_KEY, + apiVersion: 'v1', + deploymentNameMap: { 'gpt-5.5': 'production-chat' }, + }), + messages: [{ role: 'user', content: 'Hello!' }], +}) + +for await (const chunk of stream) { + if (chunk.type === 'TEXT_MESSAGE_CONTENT') console.log(chunk.delta) +} +``` + +The requested model stays `gpt-5.5` in `metadata.tanstack.source.model`. Azure receives `production-chat` as the deployment. A provider-reported response model is stored separately in `metadata.tanstack.model`. + +You can configure Azure through the environment: + +```sh +AZURE_OPENAI_API_KEY=your-key +AZURE_OPENAI_RESOURCE_NAME=my-resource +AZURE_OPENAI_API_VERSION=v1 +AZURE_OPENAI_DEPLOYMENT_NAME_MAP=gpt-5.5=production-chat +``` + +Endpoint selection uses this order: + +1. Explicit `baseURL`. +2. Explicit `resourceName`. +3. `AZURE_OPENAI_BASE_URL`. +4. `AZURE_OPENAI_RESOURCE_NAME`. + +Azure resource URLs use `/openai/v1`. A custom proxy URL keeps its configured path. Provide an endpoint or resource name. The adapter cannot infer one. + +An explicit `apiKey` overrides `AZURE_OPENAI_API_KEY`. `apiVersion` overrides `AZURE_OPENAI_API_VERSION`. The default is `v1`. + +Deployment selection uses this order: + +1. `deploymentName`. +2. The model's entry in `deploymentNameMap`, when that map is supplied. +3. The environment map, when no explicit map is supplied. +4. The requested model. + +An explicit map does not merge with the environment map. Azure's source provider and API are both `azure-openai-responses`. ## Basic Usage - Custom API Key @@ -95,7 +148,7 @@ Both adapters work identically with [Structured Outputs](../structured-outputs/o import { chat } from "@tanstack/ai"; import { createOpenaiChat } from "@tanstack/ai-openai"; -const adapter = createOpenaiChat("gpt-5.2", process.env.OPENAI_API_KEY!, { +const adapter = createOpenaiChat("gpt-5.5", process.env.OPENAI_API_KEY!, { // ... your config options }); @@ -226,7 +279,7 @@ export async function POST(request: Request) { if (!apiKey) return byokMissing(openaiByok); const stream = chat({ - adapter: createOpenaiChat("gpt-6-astra", apiKey), + adapter: createOpenaiChat("gpt-5.5", apiKey), messages: params.messages, threadId: params.threadId, runId: params.runId, @@ -259,7 +312,7 @@ const config: Omit = { baseURL: "https://api.openai.com/v1", // Optional, for custom endpoints }; -const adapter = createOpenaiChat("gpt-5.2", process.env.OPENAI_API_KEY!, config); +const adapter = createOpenaiChat("gpt-5.5", process.env.OPENAI_API_KEY!, config); ``` ### Tools that cannot use strict mode @@ -277,7 +330,7 @@ To get strict mode back, change the schema so that the reason goes away. If you ```typescript import { createOpenaiChat } from "@tanstack/ai-openai"; -const adapter = createOpenaiChat("gpt-6-astra", process.env.OPENAI_API_KEY!, { +const adapter = createOpenaiChat("gpt-5.5", process.env.OPENAI_API_KEY!, { strictFallbackWarning: false, }); ``` @@ -294,7 +347,7 @@ export async function POST(request: Request) { const { messages } = await request.json(); const stream = chat({ - adapter: openaiText("gpt-5.2"), + adapter: openaiText("gpt-5.5"), messages, }); @@ -326,7 +379,7 @@ export async function POST(request: Request) { const { messages } = await request.json(); const stream = chat({ - adapter: openaiText("gpt-5.2"), + adapter: openaiText("gpt-5.5"), messages, tools: [getWeather], }); @@ -344,7 +397,7 @@ import { chat } from "@tanstack/ai"; import { openaiText } from "@tanstack/ai-openai"; const stream = chat({ - adapter: openaiText("gpt-5.2"), + adapter: openaiText("gpt-5.5"), messages: [{ role: "user", content: "Hello!" }], modelOptions: { temperature: 0.7, @@ -358,18 +411,77 @@ const stream = chat({ ### Reasoning -Enable reasoning for models that support it (e.g., GPT-5, O3). This allows the model to show its reasoning process, which is streamed as `thinking` chunks: +Set how hard a reasoning model (GPT-5 and later, the o-series) thinks with `reasoning` on `chat()`: + +```typescript +import { chat } from "@tanstack/ai"; +import { openaiText } from "@tanstack/ai-openai"; + +const stream = chat({ + adapter: openaiText("gpt-5.5"), + messages: [{ role: "user", content: "Plan a database migration." }], + reasoning: "high", +}); +``` + +The adapter sends the level as `reasoning.effort`, with `summary: "auto"` so the reasoning summary streams back as thinking parts. Pass `reasoning: { level: "high", summary: false }` to skip the summary. The types list only the levels the model has. See [Reasoning](../chat/reasoning). + +For a model that this package does not list, pass the model's reasoning data as `reasoning` in the config. See [A model the adapter does not list](../chat/reasoning#a-model-the-adapter-does-not-list). + +### Answers on the next request + +OpenAI gives each answer item an `id` and a `phase`: `commentary` for text before a tool call, and `final_answer` for the answer. The adapter keeps both in the assistant message, in `metadata.tanstack.responseItems`. When the same model gets that message again, for example on the next turn, the adapter sends each item back with its `id` and `phase`. Another model gets the plain text. -```typescript ignore -modelOptions: { - reasoning: { - effort: "medium", // "none" | "minimal" | "low" | "medium" | "high" - summary: "detailed", // "auto" | "detailed" (optional) +### Prompt caching + +`chat()` sends `prompt_cache_key` by default, set to the `threadId` that you pass. OpenAI uses the key to send requests with the same start to the same cache. With `promptCache: 'long'`, `chat()` also asks for the long retention. See [Prompt Caching](../advanced/prompt-caching). + +To use your own key, set it in `modelOptions`. Your value wins over the automatic one: + +```typescript +import { chat } from "@tanstack/ai"; +import { openaiText } from "@tanstack/ai-openai"; + +const stream = chat({ + adapter: openaiText("gpt-5.5"), + messages: [{ role: "user", content: "Hello!" }], + modelOptions: { + prompt_cache_key: "acme-support", }, -} +}); +``` + +`modelOptions.prompt_cache_retention` also wins over the automatic `prompt_cache_retention`. + +#### Tools and prompts added during a conversation + +On some models, a tool or a system prompt that you add between model calls goes into the conversation, not into `tools` or `instructions`. The start of the request stays the same, so OpenAI can read it from its prompt cache. The `prompt_cache_key` above does not change. + +`gpt-5.5` supports these channels. For other models, see `OPENAI_MODEL_MID_CONVERSATION_CHANNELS` in the [local capability catalog](https://github.com/TanStack/ai/blob/main/packages/ai-openai/src/model-meta.ts). Models outside that catalog send the full lists. + +What the adapter sends on these models: + +- `tools` and `instructions` keep the tools and the system prompts of the first call. +- An added tool goes into `input` as `{ type: "additional_tools", role: "developer", tools: [...] }`. +- A system prompt that you add at the end of the list goes into `input` as a `developer` message, at its place in the conversation. +- With a provider tool such as `webSearchTool()` in the first call or in a change, `tools` is the full list for that request. + +The channels are on by default with OpenAI's own API. With a custom `baseURL` or `fetch`, or a proxy in the `OPENAI_BASE_URL` environment variable, they are off, and every call sends the full lists. Set `midConversationChannels` to choose: + +- `false`: send the full lists on every call. +- `true`: use the channels with a custom `baseURL`, `fetch`, or `OPENAI_BASE_URL`. Set it only when that endpoint sends the request to OpenAI as it is. + +```typescript +import { openaiText } from "@tanstack/ai-openai"; + +const fullLists = openaiText("gpt-5.5", { midConversationChannels: false }); +const throughProxy = openaiText("gpt-5.5", { + baseURL: "https://llm-proxy.example.com/v1", + midConversationChannels: true, +}); ``` -When reasoning is enabled, the model's reasoning process is streamed separately from the response text and appears as a collapsible thinking section in the UI. +See [Mid-Conversation Changes](../advanced/mid-conversation-changes) for how the library finds the changes. ## Summarization @@ -380,7 +492,7 @@ import { summarize } from "@tanstack/ai"; import { openaiSummarize } from "@tanstack/ai-openai"; const result = await summarize({ - adapter: openaiSummarize("gpt-5-mini"), + adapter: openaiSummarize("gpt-5.5"), text: "Your long text to summarize...", maxLength: 100, style: "concise", // "concise" | "bullet-points" | "paragraph" @@ -578,7 +690,7 @@ Creates an OpenAI text adapter against the Responses API (`/v1/responses`) using **Parameters:** -- `model` - OpenAI chat model id (e.g. `"gpt-5.2"`, `"gpt-4o-mini"`) +- `model` - OpenAI chat model ID, such as `"gpt-5.5"`. - `config?.organization` - Organization ID (optional) - `config?.baseURL` - Custom base URL (optional) @@ -645,7 +757,7 @@ import { openaiText } from "@tanstack/ai-openai"; import { webSearchTool } from "@tanstack/ai-openai/tools"; const stream = chat({ - adapter: openaiText("gpt-5.2"), + adapter: openaiText("gpt-5.5"), messages: [{ role: "user", content: "What's new in AI this week?" }], tools: [webSearchTool({ type: "web_search" })], }); @@ -665,7 +777,7 @@ import { openaiText } from "@tanstack/ai-openai"; import { webSearchPreviewTool } from "@tanstack/ai-openai/tools"; const stream = chat({ - adapter: openaiText("gpt-5.2"), + adapter: openaiText("gpt-5.5"), messages: [{ role: "user", content: "Latest news about TypeScript" }], tools: [ webSearchPreviewTool({ @@ -696,7 +808,7 @@ import { openaiText } from "@tanstack/ai-openai"; import { fileSearchTool } from "@tanstack/ai-openai/tools"; const stream = chat({ - adapter: openaiText("gpt-5.2"), + adapter: openaiText("gpt-5.5"), messages: [{ role: "user", content: "What does the handbook say about PTO?" }], tools: [ fileSearchTool({ @@ -721,7 +833,7 @@ import { openaiText } from "@tanstack/ai-openai"; import { imageGenerationTool } from "@tanstack/ai-openai/tools"; const stream = chat({ - adapter: openaiText("gpt-5.2"), + adapter: openaiText("gpt-5.5"), messages: [{ role: "user", content: "Draw a logo for my app" }], tools: [ imageGenerationTool({ @@ -746,7 +858,7 @@ import { openaiText } from "@tanstack/ai-openai"; import { codeInterpreterTool } from "@tanstack/ai-openai/tools"; const stream = chat({ - adapter: openaiText("gpt-5.2"), + adapter: openaiText("gpt-5.5"), messages: [{ role: "user", content: "Analyse this CSV and plot a chart" }], tools: [ codeInterpreterTool({ type: "code_interpreter", container: { type: "auto" } }), @@ -768,7 +880,7 @@ import { openaiText } from "@tanstack/ai-openai"; import { mcpTool } from "@tanstack/ai-openai/tools"; const stream = chat({ - adapter: openaiText("gpt-5.2"), + adapter: openaiText("gpt-5.5"), messages: [{ role: "user", content: "List my GitHub issues" }], tools: [ mcpTool({ @@ -820,7 +932,7 @@ import { openaiText } from "@tanstack/ai-openai"; import { localShellTool } from "@tanstack/ai-openai/tools"; const stream = chat({ - adapter: openaiText("gpt-5.6"), + adapter: openaiText("gpt-5.5"), messages: [{ role: "user", content: "Run the test suite and summarise failures" }], tools: [localShellTool()], }); @@ -852,7 +964,7 @@ import { openaiText } from "@tanstack/ai-openai"; import { shellTool } from "@tanstack/ai-openai/tools"; const stream = chat({ - adapter: openaiText("gpt-5.6"), + adapter: openaiText("gpt-5.5"), messages: [{ role: "user", content: "Count lines in all JS files" }], tools: [shellTool({ environment: { type: "local" } })], }); @@ -903,7 +1015,7 @@ export async function POST(request: Request) { const { messages } = await request.json(); const stream = chat({ - adapter: openaiText("gpt-5.2"), + adapter: openaiText("gpt-5.5"), messages, tools: [ shellTool({ @@ -936,7 +1048,7 @@ import { openaiText } from "@tanstack/ai-openai"; import { applyPatchTool } from "@tanstack/ai-openai/tools"; const stream = chat({ - adapter: openaiText("gpt-5.6"), + adapter: openaiText("gpt-5.5"), messages: [{ role: "user", content: "Fix the import paths in src/index.ts" }], tools: [applyPatchTool()], }); @@ -978,7 +1090,7 @@ import { openaiText } from "@tanstack/ai-openai"; import { customTool } from "@tanstack/ai-openai/tools"; const stream = chat({ - adapter: openaiText("gpt-5.2"), + adapter: openaiText("gpt-5.5"), messages: [{ role: "user", content: "Look up order #1234" }], tools: [ customTool({ diff --git a/docs/adapters/opencode.md b/docs/adapters/opencode.md index c0b75d1900..fc7fc55b4d 100644 --- a/docs/adapters/opencode.md +++ b/docs/adapters/opencode.md @@ -196,6 +196,19 @@ const stream = chat({ **Client-side and approval-gated tools are not supported.** The harness executes tools inside a live process, which cannot pause across HTTP requests to wait for a browser round-trip or a human approval. Passing a tool without a server `execute()` implementation — or one marked `needsApproval` — fails fast with a descriptive error. Run those tools outside the harness with a regular provider adapter. +## Tool choice + +`chat({ toolChoice })` limits only the tools that the adapter bridges into OpenCode. The built-in OpenCode tools (`bash`, `edit`, and the others) always stay on, and OpenCode decides when it calls a tool. For the values, see [Choose when the model calls a tool](../tools/tools#choose-when-the-model-calls-a-tool). + +| Value | What the adapter does | +| --- | --- | +| `'auto'` | Bridges all of your tools. | +| `'none'` | Bridges none of your tools. | +| `{ type: 'tool', name }` | Bridges only the named tool. OpenCode does not have to call it. | +| `'required'` | Bridges all of your tools. OpenCode does not have to call one. | + +Each value except `'auto'` logs a warning the first time an adapter instance gets it. + ## Structured Output Pass `outputSchema` on `chat()`. OpenCode has no native schema flag. The adapter adds the JSON Schema to the prompt and parses the last assistant text (markdown fences are stripped). Tool activity still streams. The object arrives as `structured-output.complete`. diff --git a/docs/adapters/openrouter.md b/docs/adapters/openrouter.md index 3fadcb646c..72cc72216e 100644 --- a/docs/adapters/openrouter.md +++ b/docs/adapters/openrouter.md @@ -387,7 +387,7 @@ Plugin ids include `web`, `file-parser`, `response-healing`, `moderation`, and ` ### Reasoning -`reasoning` is OpenRouter's unified reasoning configuration: +Set how hard the routed model thinks with `reasoning` on `chat()`: ```typescript import { chat } from "@tanstack/ai"; @@ -396,13 +396,11 @@ import { openRouterText } from "@tanstack/ai-openrouter"; const stream = chat({ adapter: openRouterText("anthropic/claude-sonnet-5"), messages: [{ role: "user", content: "Hello!" }], - modelOptions: { - reasoning: { effort: "high" }, - }, + reasoning: "high", }); ``` -`effort` accepts `"none"`, `"minimal"`, `"low"`, `"medium"`, `"high"`, `"xhigh"`, and `"max"`. To switch reasoning off for a request, pass `reasoning: { enabled: false }`; the adapter sends it as `effort: "none"` because the SDK's request schema drops `enabled`. +The adapter sends the level as OpenRouter's `reasoning.effort`. `off` sends `effort: "none"`. The Responses adapter (`openRouterResponsesText`) also takes `budgetTokens` and sends it as `reasoning.max_tokens`. See [Reasoning](../chat/reasoning). ### Session and metadata diff --git a/docs/advanced/byok.md b/docs/advanced/byok.md index 69a036a6c7..c8e5a73118 100644 --- a/docs/advanced/byok.md +++ b/docs/advanced/byok.md @@ -182,6 +182,35 @@ byok.setServerCoverage(true); Then a send with no pasted key still POSTs. The relay uses the env key. If that is also empty, the relay returns `byokMissing` (401). The client sets `snapshot.prompt`. +## Keys inside agents + +An agent that calls a model needs a key too. Wrap the adapter in `keyedAdapter`. Before the call, the agent makes it with `ctx.keys`: + +```typescript +import { defineAgent, keyedAdapter } from "@tanstack/ai"; +import { createOpenaiChat } from "@tanstack/ai-openai"; +import { openaiByok } from "@tanstack/ai-openai/byok"; + +const gpt = keyedAdapter(openaiByok, (key) => + createOpenaiChat("gpt-6-astra", key), +); + +export const writer = defineAgent({ + name: "writer", + description: "Writes a short draft", + run: async (ctx) => + ctx.chat({ + adapter: await ctx.keys.adapter(gpt), + messages: [{ role: "user", content: "Write a short draft." }], + stream: false, + }), +}); +``` + +- Without a host, `ctx.keys` reads the env var of the provider, here `OPENAI_API_KEY`. If it is empty, the call throws `Missing OpenAI API key. Set OPENAI_API_KEY.` +- A host can give its own keys in `subagents.binding.keys`. A child agent gives them to its own children. +- In a harness, `ctx.keys` reads the key that each user saved with `/connect`. See [Connect model providers](../harness/provider-keys). + ## Image, audio, and other providers For other cases: diff --git a/docs/advanced/compaction.md b/docs/advanced/compaction.md index 5b9ab3593b..1c7faea6db 100644 --- a/docs/advanced/compaction.md +++ b/docs/advanced/compaction.md @@ -51,7 +51,7 @@ Pass `strategy` to change how the history shrinks. Three are built in. | Strategy | What it does | Cost | |----------|--------------|------| | `evictOldest` (default) | Drop the oldest messages, leave a marker | No extra model call | -| `summarizeOldest` | Replace the oldest messages with an LLM summary | One summarize call | +| `summarizeOldest` | Replace the oldest messages with an LLM summary | One summarize call (two with `cut: 'turn'`) | | `clearToolResults` | Stub the content of old tool results, keep the messages | No extra model call | ### evictOldest @@ -77,23 +77,23 @@ import { openaiText, openaiSummarize } from "@tanstack/ai-openai"; import { withCompaction, summarizeOldest } from "@tanstack/ai-compaction"; import type { ModelMessage } from "@tanstack/ai"; -async function summarizeHistory(messages: Array): Promise { +async function summarizeHistory(messages: Array) { const text = messages .map((m) => `${m.role}: ${typeof m.content === "string" ? m.content : ""}`) .join("\n"); - const { summary } = await summarize({ - adapter: openaiSummarize("gpt-5.5"), + // Return the whole result. Compaction reports its `usage`. + return summarize({ + adapter: openaiSummarize("gpt-6.1-sol"), text, }); - return summary; } export async function POST(request: Request) { const { messages } = await request.json(); const stream = chat({ - adapter: openaiText("gpt-5.5"), + adapter: openaiText("gpt-6.1-sol"), messages, middleware: [ withCompaction({ @@ -107,6 +107,8 @@ export async function POST(request: Request) { } ``` +`summarize` can return the summary text, or an object with `summary` and `usage`. Return the `summarize()` result as it is, so compaction can report what the summary cost. + ### clearToolResults Best for agent loops. Tool output (file reads, command output) is usually most of the tokens. This strategy replaces the content of old tool results with a stub and keeps every message and its tool-call pairing in place. The conversation shape does not change. @@ -123,7 +125,7 @@ withCompaction({ ### Write your own -A strategy is a function. It gets the messages and the budget, and returns the rewritten messages, or `null` to change nothing. It runs only when the estimate is over `maxTokens`. +A strategy is a function. It gets the messages and the budget, and returns the rewritten messages, or `null` to change nothing. It runs when the count is over `maxTokens`, and when `compactNext` forces it. Then the count can be under `maxTokens`. ```typescript import { withCompaction } from "@tanstack/ai-compaction"; @@ -170,22 +172,146 @@ withCompaction({ | Option | Type | Default | Description | |--------|------|---------|-------------| -| `maxTokens` | `number` | - | **Required.** Compact when the estimated tokens across `messages` pass this. | +| `maxTokens` | `number` | - | **Required.** Compact when the token count of `messages` passes this. See `countTokens`. | | `strategy` | `CompactionStrategy` | `evictOldest()` | How to shrink the messages. | | `estimateTokens` | `(message: ModelMessage) => number` | characters / 4 | Per-message token estimate. Pass a real tokenizer if you need exact counts. | | `strategyKey` | `string` | built-in strategy identity | Stable checkpoint identity. Set it for custom strategies, custom estimators, or a custom eviction marker. Change it when your `summarize` function can change. | -| `onCompact` | `(info: CompactionInfo) => void` | - | Runs after each compaction. `info` is `{ before, after, messagesBefore, messagesAfter }` (token and message counts). | +| `onCompact` | `(info: CompactionInfo) => void` | - | Runs after each compaction. `info` is `{ before, after, messagesBefore, messagesAfter, reason, usage, error, stale }`. `error` is set when the check after a harness turn fails or a background summary fails. `stale` is set when a background summary no longer fits. | +| `countTokens` | `'estimate' \| 'usage'` | `'estimate'` | `'usage'` counts with the usage the provider reported for the last call. See [Count with real usage](#count-with-real-usage). | +| `auto` | `boolean` | `true` | `false` turns off compaction at `maxTokens`. `compactNext` and the overflow check still run. | +| `contextWindow` | `number` | - | The model's context window. With `countTokens: 'usage'` and `durable: true` in a harness with a log, compaction also runs after a turn whose usage passed it, even with `auto: false`. | +| `durable` | `boolean` | `false` | In a harness with a log, write the result into the session log. See [Compact a harness session](../harness/compaction). | +| `continueOnError` | `boolean` | `false` | When the strategy fails, send the full messages and go on. By default the run fails. | +| `background` | `{ atTokens: number }` | - | Prepare the summary when the count passes `atTokens`, and apply it at the next run. `atTokens` must be below `maxTokens`. See [Prepare the summary in the background](#prepare-the-summary-in-the-background). | ### Strategy options | Strategy | Options | |----------|---------| | `evictOldest` | `keepRecentTokens` (default `maxTokens / 2`), `marker` | -| `summarizeOldest` | `summarize` (**required**), `keepRecentTokens`, `summaryRole` (default `assistant`) | +| `summarizeOldest` | `summarize` (**required**), `keepRecentTokens`, `summaryRole` (default `assistant`), `cut` (default `'message'`) | | `clearToolResults` | `keepRecentToolResults` (default `3`), `stub` | The token count is a rough `characters / 4` estimate. It is good enough to trigger on, not exact. Pass `estimateTokens` for provider-accurate counts. +## Count with real usage + +The `characters / 4` estimate drifts on long agent runs. Code, JSON, and other languages do not split into tokens like English text. Then compaction runs too late and the call fails, or it runs too early and drops context you still need. + +Set `countTokens: 'usage'`. After each model call, compaction saves the usage that the provider reported. The next check starts from that number and estimates only the messages that came after it. + +```ts group=compaction-usage +import { chat, toServerSentEventsResponse } from '@tanstack/ai' +import { openaiText } from '@tanstack/ai-openai' +import { withCompaction } from '@tanstack/ai-compaction' + +export async function POST(request: Request) { + const { messages } = await request.json() + + const stream = chat({ + adapter: openaiText('gpt-6.1-sol'), + messages, + middleware: [withCompaction({ maxTokens: 100_000, countTokens: 'usage' })], + }) + + return toServerSentEventsResponse(stream) +} +``` + +- Before the first model call there is no usage yet, so compaction uses the estimate. +- With a [`metadata` store](#compaction-and-persistence), the usage is saved. The next request and a restarted server start from the real count. +- The strategies still use the estimate to choose which messages to keep. + +## Set the limit from the context window + +Most models publish a context window. Keep room for the answer: set `maxTokens` to the window minus a reserve. + +```ts group=compaction-usage +import { summarizeOldest } from '@tanstack/ai-compaction' +import { summarize } from '@tanstack/ai' +import { openaiSummarize } from '@tanstack/ai-openai' + +const contextWindow = 200_000 +const reserveTokens = 20_000 + +const windowed = withCompaction({ + maxTokens: contextWindow - reserveTokens, + contextWindow, + countTokens: 'usage', + strategy: summarizeOldest({ + keepRecentTokens: 8_000, + summarize: (messages) => + summarize({ + adapter: openaiSummarize('gpt-6.1-sol'), + text: messages + .map((m) => `${m.role}: ${typeof m.content === 'string' ? m.content : ''}`) + .join('\n'), + }), + }), +}) +``` + +| You want | Set | +|---|---| +| Compact before the context passes the window minus a reserve | `maxTokens: contextWindow - reserveTokens` | +| Keep the newest messages word for word | `keepRecentTokens` on the strategy | +| In a durable harness: no automatic compaction, but recovery from a silent overflow | `auto: false`, `contextWindow`, `countTokens: 'usage'`, `durable: true` | + +Some providers accept a request that is too large. They cut the input and answer without an error. In a [harness with a durable log](../harness/compaction), compaction with `countTokens: 'usage'` and `durable: true` checks once more after the last model call of a turn. If the usage passed `maxTokens` or `contextWindow`, it compacts right away, so the next turn starts small. With `auto: false`, only `contextWindow` counts. The answer is not sent again. The only extra model calls are the summary calls. + +Outside a durable harness, there is no check after a run. The next request compacts before its first model call. + +## Prepare the summary in the background + +A `summarizeOldest` call takes time. When compaction runs at `maxTokens`, the user waits for that call before the model answers. Set `background` to prepare the summary earlier, while the chat goes on: + +```ts group=compaction-background +import { chat, summarize, toServerSentEventsResponse } from '@tanstack/ai' +import { openaiSummarize, openaiText } from '@tanstack/ai-openai' +import { summarizeOldest, withCompaction } from '@tanstack/ai-compaction' + +// Create it once, so a summary from one request reaches the next. +const compaction = withCompaction({ + maxTokens: 100_000, + background: { atTokens: 80_000 }, + strategy: summarizeOldest({ + summarize: (messages) => + summarize({ + adapter: openaiSummarize('gpt-6.1-sol'), + text: messages + .map((m) => `${m.role}: ${typeof m.content === 'string' ? m.content : ''}`) + .join('\n'), + }), + }), +}) + +export async function POST(request: Request) { + const { messages, threadId } = await request.json() + + const stream = chat({ + adapter: openaiText('gpt-6.1-sol'), + messages, + threadId, + middleware: [compaction], + }) + + return toServerSentEventsResponse(stream) +} +``` + +- `atTokens` must be below `maxTokens`. If it is not, `withCompaction` throws. +- When the count at a model call passes `atTokens`, the summary starts on a copy of the messages. The model call does not wait for it. +- The ready summary waits in the [`metadata` store](#compaction-and-persistence), under the namespace `@tanstack/ai-compaction:background`. Without a store, it waits in memory. +- The summary applies at the first model call of the next run on the same `threadId`, never in the middle of a run. +- A model call over `maxTokens` while the summary runs waits for it and applies it. It compacts inline only when the messages are still over `maxTokens`. A cancel of the run stops the wait, not the summary. + +Sometimes a summary does not apply: + +- A newer compaction cut past the message that the summary keeps. The summary is dropped and reported with `stale: true`. +- The summary call failed. Nothing applies, and the failure is reported with `error`. The next model call past `atTokens` starts a new summary. + +`onCompact` always gets these reports. A `compaction:ended` event for a failed summary reaches the stream only while the run that started the summary still runs. + ## What it keeps safe - **The system prompt is never dropped.** `chat()` keeps it separate from `messages`, so compaction only touches the conversation. @@ -193,6 +319,112 @@ The token count is a rough `characters / 4` estimate. It is good enough to trigg - **It runs before every model call.** Compaction skips `init`. It runs on `beforeModel` and `structuredOutput`. Each later call can compact again. - **The canonical transcript stays complete.** Compaction writes provider-only context. Persistence and other middleware still read `ctx.messages`. +## Compact now after an overflow + +A call can still pass the model's context limit and fail. `isContextOverflow` from `@tanstack/ai` tells you that a call failed for this reason. Call `compactNext(threadId)`, then send the request again. The next model call on that thread compacts first, even when the count is under `maxTokens`. + +```ts group=compaction-overflow +import { chat, isContextOverflow } from '@tanstack/ai' +import type { ModelMessage } from '@tanstack/ai' +import { openaiText } from '@tanstack/ai-openai' +import { withCompaction } from '@tanstack/ai-compaction' + +// Create it once, so compactNext reaches the same instance. +const compaction = withCompaction({ maxTokens: 100_000, countTokens: 'usage' }) + +async function answer(threadId: string, messages: Array) { + for (let attempt = 0; attempt < 2; attempt++) { + let text = '' + let overflow = false + for await (const chunk of chat({ + adapter: openaiText('gpt-6.1-sol'), + messages, + threadId, + middleware: [compaction], + })) { + if (chunk.type === 'TEXT_MESSAGE_CONTENT') text += chunk.delta + if (chunk.type === 'RUN_ERROR') { + if (!isContextOverflow({ error: chunk })) throw new Error(chunk.message) + overflow = true + } + } + if (!overflow) return text + compaction.compactNext(threadId) + } + throw new Error('The conversation is too long, even after compaction.') +} +``` + +In a harness, call `compactNext` from `turn.onModelError`. See [Compact a harness session](../harness/compaction#retry-after-a-context-overflow). + +`compactNext` runs the strategy even under `maxTokens`, but a strategy can still find nothing to cut. `composeStrategies` stops when the estimate is under `maxTokens`, and `evictOldest` keeps `maxTokens / 2` by default. For an overflow retry, set a small `keepRecentTokens` on the strategy. + +`isContextOverflow` knows the overflow errors of Anthropic, OpenAI, Gemini, Bedrock, Mistral, xAI, Groq, OpenRouter, Ollama, and more. It ignores rate-limit errors. It takes one object, and every field is optional: + +- `error`: a `RUN_ERROR` event, an `Error`, or a message string. +- `usage`, `finishReason`, and `contextWindow`: some providers accept an overflow and cut the input without an error. With the model's `contextWindow`, a call counts as an overflow when its `usage.promptTokens` is more than the window, or when it stopped with `'length'`, wrote nothing, and its input fills the window. +- `provider`: set `'cerebras'` for Cerebras, which answers an overflow with a bare `400` or `413`. + +## Better summaries + +The basic `summarizeOldest` setup works for a chat. A coding agent needs more. It must keep the goal, the decisions, and the files it touched across many compactions. Three opt-in parts help: + +1. `cut: 'turn'` on `summarizeOldest`: a better summary when the cut falls inside a turn. The older turns get one summary. The start of the current turn gets its own short summary, under `## Turn context`. +2. `conversationSummarizer`: a ready summarizer with a structured prompt (goal, constraints, progress, key decisions, next steps, critical context). When an older summary exists, it updates that summary. It reports the usage of each summary call. +3. `details` on `conversationSummarizer`: your own facts after the summary, for example the files the agent read. The next compaction gets them back as `previousDetails`. + +```ts group=compaction-better +import { openaiText } from '@tanstack/ai-openai' +import { + conversationSummarizer, + summarizeOldest, + withCompaction, +} from '@tanstack/ai-compaction' + +function pathOf(args: string) { + const input: unknown = JSON.parse(args) + return typeof input === 'object' && + input !== null && + 'path' in input && + typeof input.path === 'string' + ? input.path + : undefined +} + +const agentCompaction = withCompaction({ + maxTokens: 180_000, + contextWindow: 200_000, + countTokens: 'usage', + strategy: summarizeOldest({ + cut: 'turn', + keepRecentTokens: 8_000, + summarize: conversationSummarizer({ + adapter: openaiText('gpt-6.1-sol'), + details: ({ messages, previousDetails }) => { + const files = new Set(previousDetails?.split('\n').filter(Boolean)) + for (const message of messages) { + for (const call of message.toolCalls ?? []) { + const path = + call.function.name === 'read_file' + ? pathOf(call.function.arguments) + : undefined + if (path) files.add(path) + } + } + return files.size > 0 ? [...files].join('\n') : undefined + }, + }), + }), +}) +``` + +- In the text sent to the summarizer, each tool result is cut to 2,000 characters. Set `maxToolResultChars` to change it. +- A `summarize` function of your own gets a second argument, `summarize(messages, input)`: + - `previousSummary` and `previousDetails`: the older summary and its details. + - `turnPrefix`: `true` when the call summarizes the start of the current turn. + - `turnPrefixMessages`: the start of the current turn. Only the call that writes the details gets it, so the details cover each dropped message. + - `signal`: aborts when the run aborts. + ## DevTools After a compaction, the chat stream includes three CUSTOM events in order: @@ -201,6 +433,16 @@ After a compaction, the chat stream includes three CUSTOM events in order: event as soon as it is emitted, so a slow `summarizeOldest` call still shows as started on the client while it runs. The state and ended events follow when the strategy returns. + +`compaction:ended` also says why compaction ran and what it cost: + +- `reason`: `'threshold'` (the count passed `maxTokens`), `'forced'` (`compactNext`), or `'background'` (a [background summary](#prepare-the-summary-in-the-background)). +- `usage`: the token usage of the summary calls. +- `error`: the message when the strategy failed. By default the run then fails. Set `continueOnError: true` to send the full messages instead. +- `stale`: `true` when a ready background summary no longer fit the messages and was dropped. + +A compaction after the last model call of a harness turn has `reason: 'after-turn'`. The stream is closed by then, so only `onCompact` reports it. A failed strategy or `onCompact` there does not fail the turn, because the answer is complete. Only a failed write to the log fails it. + TanStack AI DevTools has a Compaction tab on the hook. Each compact shows: - started, state, and ended rows @@ -242,7 +484,10 @@ Set `strategyKey` for custom strategies, custom estimators, or custom marker functions. Change `strategyKey` when your `summarize` function can change. Without a metadata store or safe key, compaction stays stateless. +With `durable: true` in a harness with a durable log, compaction writes its result into the session log instead. A restarted session then sees the same context. See [Compact a harness session](../harness/compaction). + ## Next steps +- [Compact a harness session](../harness/compaction): keep the compaction in the session log, and retry after an overflow - [Middleware](./middleware): the full hook reference and how middleware composes - [Built-in Middleware](./built-in-middleware): ready-made middleware that ships in `@tanstack/ai` diff --git a/docs/advanced/extend-adapter.md b/docs/advanced/extend-adapter.md index a2cf47db76..2dbb6ba9da 100644 --- a/docs/advanced/extend-adapter.md +++ b/docs/advanced/extend-adapter.md @@ -209,3 +209,155 @@ chat({ messages: [{ role: 'user', content: 'Analyze this...' }] }) ``` + +## Add reasoning to your adapter + +`chat({ reasoning })` reaches an adapter as `options.reasoning`: a level, a `summary` flag, and an optional `budgetTokens`. To support it: + +1. Give each model its reasoning data as a `ModelReasoning`: a map from each level to the value your provider takes, and whether it takes a token budget. +2. Declare the levels on the adapter type, so `chat()` checks them. The last type parameter of `BaseTextAdapter` and the `openai-base` adapters is `{ levels; budget }`. +3. Turn the request into your provider's field. `resolveReasoning` clamps the level to the model and looks up the value. + +An adapter built on `@tanstack/openai-base` only needs `modelReasoning`. The base then sends `reasoning_effort` on Chat Completions, or `reasoning.effort` on Responses: + +```typescript +import OpenAI from "openai"; +import { OpenAIBaseChatCompletionsTextAdapter } from "@tanstack/openai-base"; +import type { + DefaultMessageMetadataByModality, + Modality, + ModelReasoning, +} from "@tanstack/ai"; + +const MY_MODEL_REASONING: Record = { + "my-model": { + map: { off: "none", minimal: null, low: "low", medium: "medium", high: "high" }, + budget: false, + }, +}; + +class MyTextAdapter extends OpenAIBaseChatCompletionsTextAdapter< + "my-model", + Record, + ReadonlyArray, + DefaultMessageMetadataByModality, + ReadonlyArray, + { levels: "off" | "low" | "medium" | "high"; budget: false } +> { + protected override modelReasoning(model: string) { + return MY_MODEL_REASONING[model]; + } +} + +export const myText = () => + new MyTextAdapter( + "my-model", + "my-provider", + new OpenAI({ apiKey: process.env.MY_API_KEY, baseURL: "https://api.example.com/v1" }), + ); +``` + +An adapter with another wire format calls `resolveReasoning` from `@tanstack/ai/adapter-internals` in its request code and sends the result its own way. + +Models added with `extendAdapter` take no `reasoning` option, because the adapter has no reasoning data for them. + +## Keep block order and mid-conversation changes + +Two optional fields reach your adapter. If your adapter ignores them, it sends the same request as before. + +### Send assistant blocks in their order + +A model can answer with thinking, a tool call, more thinking, then text. A `ModelMessage` keeps `thinking`, `content`, and `toolCalls` in separate fields. The optional `blockOrder` map says how they mix. `orderedAssistantBlocks` gives you the blocks in that order: + +```typescript +import { orderedAssistantBlocks } from "@tanstack/ai"; +import type { ModelMessage } from "@tanstack/ai"; + +type WireBlock = + | { type: "reasoning"; text: string; signature?: string } + | { type: "text"; text: string } + | { type: "call"; id: string; name: string; args: string }; + +export function orderedWireBlocks(message: ModelMessage): Array | undefined { + const blocks = orderedAssistantBlocks(message); + if (!blocks) return undefined; + return blocks.map((block): WireBlock => { + if (block.type === "thinking") { + return { type: "reasoning", text: block.thinking.content, signature: block.thinking.signature }; + } + if (block.type === "text") return { type: "text", text: block.text }; + return { + type: "call", + id: block.toolCall.id, + name: block.toolCall.function.name, + args: block.toolCall.function.arguments, + }; + }); +} +``` + +- The function returns `undefined` when the message has no map, or when the map does not match the message. Then send your default order. +- The library writes `blockOrder` only when the order is not the default. The default order is all thinking, then the text, then the tool calls. + +### Send mid-conversation changes + +Tools or system prompts can grow between model calls. If your provider has a mid-conversation channel, it can take the change inside the conversation. Then the cached start of the request stays the same. To support it: + +1. Set `midConversationChannels` on the adapter to the channels that your provider has: `{ tools, systemPrompts }`. Before each call, the `chat()` engine compares the lists. It passes the result as `options.midConversationChanges`. +2. Resolve the change with `splitMidConversationChanges`. It gives the start lists and the changes by message index. It returns `undefined` when a name or a count does not match the current lists. +3. Send the start lists at the top of the request. Send each change directly before the message at its index. A change at `options.messages.length` goes at the end. + +An adapter that extends the Responses adapter of `@tanstack/openai-base` only sets the field. The base then sends `additional_tools` and `developer` messages: + +```typescript +import { OpenAIBaseResponsesTextAdapter } from "@tanstack/openai-base"; + +export class MyResponsesAdapter extends OpenAIBaseResponsesTextAdapter<"my-model"> { + override readonly midConversationChannels = { tools: true, systemPrompts: true }; +} +``` + +An adapter with another wire format does the steps itself: + +```typescript +import { splitMidConversationChanges } from "@tanstack/ai"; +import type { ModelMessage, TextOptions } from "@tanstack/ai"; + +type WireTool = { name: string; description: string }; +type WireItem = + | { kind: "message"; message: ModelMessage } + | { kind: "change"; tools: Array; prompts: Array }; + +export function buildRequest( + options: TextOptions, + tools: Array, + prompts: Array, +) { + const changes = options.midConversationChanges; + const split = changes + ? splitMidConversationChanges({ changes, tools, systemPrompts: prompts }) + : undefined; + if (!split) { + // No changes, or names that do not match: send the full lists. + const items = options.messages.map((message): WireItem => ({ kind: "message", message })); + return { tools, system: prompts, items }; + } + const items: Array = []; + const pushChange = (index: number) => { + const change = split.at.get(index); + if (change) items.push({ kind: "change", tools: change.tools, prompts: change.systemPrompts }); + }; + options.messages.forEach((message, index) => { + pushChange(index); + items.push({ kind: "message", message }); + }); + pushChange(options.messages.length); + return { tools: split.startTools, system: split.startSystemPrompts, items }; +} +``` + +- `split.addedTools` lists every added tool, in change order. Use it if your provider also wants the added tools in its tool list. Claude takes them there with `defer_loading`. +- A tool can keep its name and get a new definition. Take every definition from `options.tools`. +- If your provider cannot add some kind of tool later (for example a hosted search tool), send the full tool list for that request. Keep the prompt changes in the conversation. + +See [Mid-Conversation Changes](./mid-conversation-changes) for how the library finds the changes. diff --git a/docs/advanced/mid-conversation-changes.md b/docs/advanced/mid-conversation-changes.md new file mode 100644 index 0000000000..9457996402 --- /dev/null +++ b/docs/advanced/mid-conversation-changes.md @@ -0,0 +1,153 @@ +--- +title: Mid-Conversation Changes +id: mid-conversation-changes +order: 17 +description: "Add tools or system prompts during a conversation and keep the provider's prompt cache. TanStack AI sends the change through the model's own channel on GPT and Claude models that have one." +keywords: + - tanstack ai + - prompt caching + - mid-conversation changes + - dynamic tools + - additional_tools + - tool_addition +--- + +Your agent gets a new tool in the middle of a conversation. A middleware adds it after the first tool call, a [lazy tool](../tools/lazy-tool-discovery) is found, or the user connects an MCP server. `chat()` asks the provider to cache the start of each request ([Prompt Caching](./prompt-caching)), and the tool list is at the start. A new tool changes the start, so the provider cannot use its cache. The next call costs more and starts slower. + +On models with a mid-conversation channel, TanStack AI keeps the start of the request the same and sends the change later in the conversation. You change no code. + +## Add a tool during a run + +This route starts with one tool. A middleware adds a second tool after the first tool call: + +```ts group=mid-conversation-changes +import { chat, toServerSentEventsResponse, toolDefinition } from "@tanstack/ai"; +import type { ChatMiddleware } from "@tanstack/ai"; +import { openaiText } from "@tanstack/ai-openai"; +import { z } from "zod"; + +const getWeather = toolDefinition({ + name: "get_weather", + description: "Get the current weather for a city", + inputSchema: z.object({ city: z.string() }), +}).server(async ({ city }) => ({ city, temperature: 21 })); + +const getForecast = toolDefinition({ + name: "get_forecast", + description: "Get the forecast for a city", + inputSchema: z.object({ city: z.string() }), +}).server(async ({ city }) => ({ city, tomorrow: "sunny" })); + +// Adds get_forecast before the second model call. +const addForecast: ChatMiddleware = { + name: "add-forecast", + onConfig: (ctx, config) => { + if (ctx.phase !== "beforeModel" || ctx.iteration === 0) return; + if (config.tools.some((tool) => tool.name === "get_forecast")) return; + return { tools: [...config.tools, getForecast] }; + }, +}; + +export async function POST(request: Request) { + const { messages } = await request.json(); + const stream = chat({ + adapter: openaiText("gpt-6-astra"), + messages, + tools: [getWeather], + middleware: [addForecast], + }); + return toServerSentEventsResponse(stream); +} +``` + +What each model call sends: + +| Model | First call | Second call | +|---|---|---| +| `gpt-6-astra` | `tools`: `get_weather` | The same `tools`. An `additional_tools` item with `get_forecast` comes after the tool result. | +| `claude-opus-5-5` | `tools`: `get_weather` and a placeholder tool | The same `tools`, then `get_forecast` with `defer_loading: true`. A `system` message with a `tool_addition` block comes after the tool result. | +| A model with no channel, for example `gpt-6.1-sol` | `tools`: `get_weather` | `tools`: `get_weather` and `get_forecast` | + +## What happens on its own + +Before each model call, the library compares the tools and the system prompts with the earlier calls: + +1. The first call is a start point. Its tools and system prompts stay at the start of every later request. +2. A tool that you add is a change. A system prompt that you add at the end of the list is a change too. A change goes out at its place in the conversation, in the model's own format. +3. Any other difference makes a new start point. That request sends the full lists. + +The library saves a small record on the first assistant message of a call that makes a start point or a change, in the `midConversationChange` field. The record holds tool names and short prompt hashes, not tool definitions. Keep the field when you store messages yourself. + +## The models + +| Provider | Models | Channels | +|---|---|---| +| OpenAI (`openaiText`) | `gpt-5.4-mini`, `gpt-5.5`, `gpt-5.6-luna`, `gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-6-astra`, `gpt-6-luna`, `gpt-6-sol` | Tools and system prompts | +| Anthropic (`anthropicText`) | `claude-opus-4-8`, `claude-opus-5`, `claude-opus-5-5`, `claude-fable-5`, `claude-fable-5-1` | Tools and system prompts | + +Every other model and adapter sends the full lists on each call. The channels are on by default only when the adapter talks to the provider's own API. See [Turn it on or off](#turn-it-on-or-off). + +To check an adapter in code, read `midConversationChannels` on it: + +```ts group=mid-conversation-changes +const adapter = openaiText("gpt-6-astra"); +console.log(adapter.midConversationChannels); // { tools: true, systemPrompts: true } +``` + +It is `undefined` on a model with no channel, and on an adapter with the channels off. + +The provider pages show the request shapes: [OpenAI](../adapters/openai#tools-and-prompts-added-during-a-conversation) and [Anthropic](../adapters/anthropic#tools-and-prompts-added-during-a-conversation). + +## See it in the usage numbers + +Log `usage.promptTokensDetails.cachedTokens` for each model call, as in [Track cached tokens and cost](./prompt-caching#track-cached-tokens-and-cost). After a tool is added, the next call still reads the start of the request from the cache. On a model with no channel, the cached tokens drop at that call. + +On Claude, the automatic tool cache marker stays on the last tool of the start point. The placeholder and the added tools get no marker, so the marked start does not move. + +## Turn it on or off + +The channels are on by default when the adapter talks to the provider's own API. A gateway or a proxy can drop or change the extra request fields, so the channels are off by default when the adapter talks to another endpoint: + +- a custom `baseURL` or a custom `fetch` +- a proxy in the `OPENAI_BASE_URL` or `ANTHROPIC_BASE_URL` environment variable + +An Anthropic adapter on your own client (`createAnthropicChatWithClient`, `anthropicVertexText`) has no channels. + +Set `midConversationChannels` in the adapter config to choose: + +- `false`: every call sends the full lists. +- `true`: use the channels with a custom `baseURL` or `fetch`, or with a base URL from the environment. Set it only when that endpoint sends the request to the provider as it is. + +```ts group=mid-conversation-changes +import { anthropicText } from "@tanstack/ai-anthropic"; + +const fullGpt = openaiText("gpt-6-astra", { midConversationChannels: false }); +const proxiedClaude = anthropicText("claude-opus-5-5", { + baseURL: "https://llm-proxy.example.com", + midConversationChannels: true, +}); +``` + +On a model with no channel, `true` changes nothing. + +## Limits + +- These break the cache once: a removed tool, an edited system prompt, a new order of the system prompts, and a tool that keeps its name but gets a new definition. The lazy discovery tool leaves the list when every lazy tool is found, so that call breaks it once too. +- A provider tool, for example `webSearchTool()`, in the start point or in a change: that request sends the full `tools`. System prompt changes still use the channel. +- With a harness or a message store ([`withPersistence`](../persistence/chat-persistence)), the records last across turns and restarts. With plain `useChat` and no message store, they last for one `chat()` call, so each turn starts a new start point. +- The structured-output call at the end of a run with `outputSchema` sends the full request. +- Claude on Vertex or Bedrock, the Chat Completions adapters, and the other adapters on the OpenAI API have no channels. + +## In a harness + +A harness session keeps its messages in the session log. So the records hold across turns, a restart, and a `/model` switch. You set nothing. + +- Tools that a plugin finds before a turn ([Tools that appear later](../harness/mcp#tools-that-appear-later)) go out as a change on a model with a channel. +- After a `/model` switch to a model with no channel, each call sends the full lists. A switch back continues from the stored records. +- A `project` function or a compaction that removes the assistant message with the start point makes a new start point. The cache breaks once. + +## What you have now + +- Tools and system prompts that you add during a conversation keep the prompt cache on GPT and Claude models with a channel. +- The full request on every other model. +- One option, `midConversationChannels`: `false` sends the full request on every call, and `true` uses the channels behind a gateway or a proxy. diff --git a/docs/advanced/middleware.md b/docs/advanced/middleware.md index 8c0ecb2647..9398d68272 100644 --- a/docs/advanced/middleware.md +++ b/docs/advanced/middleware.md @@ -2,7 +2,7 @@ title: Middleware id: middleware order: 1 -description: "Hook into every stage of TanStack AI's chat() lifecycle with middleware — logging, analytics, stream transforms, tool interception, and side effects." +description: "Hook into every stage of TanStack AI's chat() lifecycle with middleware: logging, analytics, stream transforms, tool interception, and side effects." keywords: - tanstack ai - middleware @@ -14,15 +14,15 @@ keywords: - stream transform --- -Middleware lets you hook into every stage of the `chat()` lifecycle — from configuration to streaming, tool execution, usage tracking, and completion. You can observe, transform, or short-circuit behavior at each stage without modifying your adapter or tool implementations. +Middleware lets you hook into every stage of the `chat()` lifecycle, from configuration to streaming, tool execution, usage tracking, and completion. You can observe, transform, or short-circuit behavior at each stage without modifying your adapter or tool implementations. Common use cases include: -- **Logging and observability** — track token usage, tool execution timing, errors -- **Configuration transforms** — inject system prompts, adjust temperature per iteration, filter tools -- **Stream processing** — redact sensitive content, transform chunks, drop unwanted events -- **Tool call interception** — validate arguments, cache results, abort on dangerous calls -- **Side effects** — send analytics, update databases, trigger notifications +- **Logging and observability**: track token usage, tool execution timing, errors +- **Configuration transforms**: inject system prompts, adjust temperature per iteration, filter tools +- **Stream processing**: redact sensitive content, transform chunks, drop unwanted events +- **Tool call interception**: validate arguments, cache or replace results, abort on dangerous calls +- **Side effects**: send analytics, update databases, trigger notifications ## Quick Start @@ -50,7 +50,7 @@ const stream = chat({ ``` > **Just want to see chunks flowing through your middleware during development?** -> Use `debug: { middleware: true }` on your `chat()` call — no custom middleware required. See [Debug Logging](./debug-logging). +> Use `debug: { middleware: true }` on your `chat()` call. No custom middleware is required. See [Debug Logging](./debug-logging). ## Lifecycle Overview @@ -113,7 +113,7 @@ The context's `phase` field tracks where you are in the lifecycle: Called once during `init` (startup) and once per iteration during `beforeModel` (before each model call). On the separate-finalization path, `onConfig` additionally re-fires at the structured-output boundary with `ctx.phase === 'structuredOutput'`, receiving the post-`onStructuredOutputConfig` view of the config. A single-iteration separate-finalization run therefore fires `onConfig` three times (`init` + `beforeModel` + `structuredOutput`). Native-combined output does not add this third call. Use `onConfig` to transform the configuration that the model receives. -Return a **partial** config object with only the fields you want to change — they are shallow-merged with the current config automatically. No need to spread the existing config. +Return a **partial** config object with only the fields you want to change. They are shallow-merged with the current config automatically. No need to spread the existing config. ```typescript import { type ChatMiddleware } from "@tanstack/ai"; @@ -122,7 +122,7 @@ const dynamicTemperature: ChatMiddleware = { name: "dynamic-temperature", onConfig: (ctx, config) => { if (ctx.phase === "init") { - // Add a system prompt at startup — only systemPrompts is overwritten + // Add a system prompt at startup. Only systemPrompts is overwritten return { systemPrompts: [ ...config.systemPrompts, @@ -133,7 +133,7 @@ const dynamicTemperature: ChatMiddleware = { if (ctx.phase === "beforeModel" && ctx.iteration > 0) { // Increase temperature on retries. Sampling params live in the - // provider-native modelOptions object — `temperature` is universal, + // provider-native modelOptions object. `temperature` is universal, // so it's the same key across providers. Spread the existing // modelOptions so other model options stay unchanged. const current = @@ -151,7 +151,7 @@ const dynamicTemperature: ChatMiddleware = { }; ``` -> Sampling parameters (`temperature`, `top_p` / `topP`, the various `max*Tokens` keys) live inside `modelOptions` under each provider's native name — they are no longer root config fields. `temperature` happens to be spelled the same across every provider, so the example above is provider-agnostic; if you mutate a token limit instead, use the provider-native key (e.g. `max_output_tokens` for OpenAI, `num_predict` nested under `modelOptions.options` for Ollama). See [Moving Sampling Options into modelOptions](../migration/sampling-options-to-model-options). +> Sampling parameters (`temperature`, `top_p` / `topP`, the various `max*Tokens` keys) live inside `modelOptions` under each provider's native name. They are not root config fields. `temperature` happens to be spelled the same across every provider, so the example above is provider-agnostic; if you mutate a token limit instead, use the provider-native key (e.g. `max_output_tokens` for OpenAI, `num_predict` nested under `modelOptions.options` for Ollama). See [Moving Sampling Options into modelOptions](../migration/sampling-options-to-model-options). **Config fields you can transform:** @@ -162,21 +162,85 @@ const dynamicTemperature: ChatMiddleware = { | `systemPrompts` | `string[]` | System prompts | | `tools` | `Tool[]` | Available tools | | `metadata` | `Record` | Request metadata | -| `modelOptions` | `Record` | Provider-native options — this is where sampling params (`temperature`, `top_p` / `topP`, the provider's `max*Tokens` key) now live, alongside every other model-specific knob. See [Moving Sampling Options into modelOptions](../migration/sampling-options-to-model-options). | +| `modelOptions` | `Record` | Provider-native options. Sampling params (`temperature`, `top_p` / `topP`, the provider's `max*Tokens` key) live here, next to every other model-specific option. See [Moving Sampling Options into modelOptions](../migration/sampling-options-to-model-options). | +| `reasoning` | `ReasoningRequest \| undefined` | How hard the model thinks at this call. See [Reasoning](../chat/reasoning#change-the-level-in-middleware). | +| `promptCache` | `ResolvedPromptCache \| undefined` | The prompt cache of this call: its `retention` and `key`. See [Change the prompt cache of a call](#change-the-prompt-cache-of-a-call). | +| `toolChoice` | `ToolChoice \| undefined` | How the model uses the tools at the next call. See [Change the tool choice of a call](#change-the-tool-choice-of-a-call). | -When multiple middleware define `onConfig`, the config is **piped** through them in order — each receives the merged config from the previous middleware. +When multiple middleware define `onConfig`, the config is **piped** through them in order. Each receives the merged config from the previous middleware. Return `providerMessages` when a transform must affect only the model call. For compatibility, returning `messages` also updates provider input unless the same result sets `providerMessages` explicitly. +#### Change the prompt cache of a call + +A slow tool, such as a build or a test suite, can pause an agent run for more than 5 minutes. The short Claude cache ends 5 minutes after its last read. Return `promptCache` from `onConfig` to change the cache of the next model call, the same as `reasoning`: + +```typescript +import { type ChatMiddleware } from "@tanstack/ai"; + +const longCacheAfterTools: ChatMiddleware = { + name: "long-cache-after-tools", + onConfig: (ctx, config) => { + if (ctx.phase !== "beforeModel" || ctx.iteration === 0) return; + return { promptCache: { ...config.promptCache, retention: "long" } }; + }, +}; +``` + +- `config.promptCache` is the current value, with its `retention` and its `key`. +- The returned value goes to the adapter on the next model call. It stays for the later calls until a middleware returns another value. +- For the retention values and their cost, see [Prompt Caching](./prompt-caching#pick-the-retention). + +#### Change the tool choice of a call + +An agent run can reach its last model call while the model still calls tools. The run then ends after a tool result, with no answer for the user. Return `toolChoice` from `onConfig` to change how the model uses the tools at the next call: + +```typescript +import { chat, maxIterations, toolDefinition, type ChatMiddleware } from "@tanstack/ai"; +import { openaiText } from "@tanstack/ai-openai"; +import { z } from "zod"; + +const MAX_CALLS = 10; + +const search = toolDefinition({ + name: "search", + description: "Search the docs", + inputSchema: z.object({ query: z.string() }), +}).server(async ({ query }) => ({ hits: [`A page about ${query}`] })); + +// The last call gets no tool calls, so the model answers in text. +const answerOnLastCall: ChatMiddleware = { + name: "answer-on-last-call", + onConfig: (ctx) => { + if (ctx.phase === "beforeModel" && ctx.iteration === MAX_CALLS - 1) { + return { toolChoice: "none" }; + } + }, +}; + +const stream = chat({ + adapter: openaiText("gpt-6.1-sol"), + messages: [{ role: "user", content: "How do I add a tool?" }], + tools: [search], + agentLoopStrategy: maxIterations(MAX_CALLS), + middleware: [answerOnLastCall], +}); +``` + +- `config.toolChoice` is the `chat()` option. +- The returned value applies to the next model call only. The call after it starts again from the `chat()` option. +- Return it in the `beforeModel` phase. A value from the `init` phase does not reach a model call. +- For the values, and for what each provider does with them, see [Choose when the model calls a tool](../tools/tools#choose-when-the-model-calls-a-tool). + ### onStructuredOutputConfig -Called once at the start of the final structured-output adapter call — only when `chat()` was invoked with `outputSchema` **and** `supportsCombinedToolsAndSchema()` does not return `true` for the current model/options. Pipes through middleware in order, like `onConfig`, but with access to the **JSON Schema** being sent to the provider. Use this hook when you need to transform the schema (e.g., inject `$defs`, strip vendor-incompatible keywords) or apply structured-output-specific behavior (e.g., suppress system prompts on the final call). +Called once at the start of the final structured-output adapter call, only when `chat()` was invoked with `outputSchema` **and** `supportsCombinedToolsAndSchema()` does not return `true` for the current model/options. Pipes through middleware in order, like `onConfig`, but with access to the **JSON Schema** being sent to the provider. Use this hook when you need to transform the schema (e.g., inject `$defs`, strip vendor-incompatible keywords) or apply structured-output-specific behavior (e.g., suppress system prompts on the final call). -> Native-combined adapters (modern OpenAI, Claude 4.5+, Gemini 3.x, Grok 4.x — see issue #605) skip the separate finalization call and never invoke this hook. The engine passes the converted schema directly to `chatStream` after `onConfig` runs, so middleware cannot transform the native-combined schema. +> Native-combined adapters (modern OpenAI, Claude 4.5+, Gemini 3.x, and Grok 4.x, per issue #605) skip the separate finalization call and never invoke this hook. The engine passes the converted schema directly to `chatStream` after `onConfig` runs, so middleware cannot transform the native-combined schema. -Return a **partial** `StructuredOutputMiddlewareConfig` with only the fields you want to change — they are shallow-merged with the current config. Return `void` to pass through. +Return a **partial** `StructuredOutputMiddlewareConfig` with only the fields you want to change. They are shallow-merged with the current config. Return `void` to pass through. ```typescript import { type ChatMiddleware } from "@tanstack/ai"; @@ -204,7 +268,7 @@ const injectDefs: ChatMiddleware = { | `providerMessages` | `ModelMessage[]` | Temporary context sent to the final call | | `systemPrompts` | `SystemPrompt[]` | System prompts on the final call | | `metadata` | `Record` | Request metadata | -| `modelOptions` | `Record` | Provider-native options — this is where sampling params (`temperature`, `top_p` / `topP`, the provider's `max*Tokens` key) now live, alongside every other model-specific knob. See [Moving Sampling Options into modelOptions](../migration/sampling-options-to-model-options). | +| `modelOptions` | `Record` | Provider-native options. Sampling params (`temperature`, `top_p` / `topP`, the provider's `max*Tokens` key) live here, next to every other model-specific option. See [Moving Sampling Options into modelOptions](../migration/sampling-options-to-model-options). | | `outputSchema` | `JSONSchema` | JSON Schema being sent to the provider for structured output | **Ordering at the structured-output boundary:** @@ -212,7 +276,7 @@ const injectDefs: ChatMiddleware = { 1. `onStructuredOutputConfig` fires first, piping through every middleware in array order. 2. `onConfig` then re-fires at the same boundary with `ctx.phase === 'structuredOutput'`, receiving the post-`onStructuredOutputConfig` view of the config (minus `outputSchema`). Use `onConfig` for general-purpose transforms that apply to every adapter call; use `onStructuredOutputConfig` when you need access to the schema. -When multiple middleware define `onStructuredOutputConfig`, the config is **piped** through them in order — each receives the merged config from the previous middleware. +When multiple middleware define `onStructuredOutputConfig`, the config is **piped** through them in order. Each receives the merged config from the previous middleware. ### onStart @@ -264,7 +328,7 @@ When multiple middleware define `onChunk`, chunks flow through them in order. If #### Chunk types you'll see -`onChunk` receives every [AG-UI event](https://docs.ag-ui.com/introduction) the run produces — not just text. Narrow on `chunk.type` (a discriminated union) before reading type-specific fields. The common ones: +`onChunk` receives every [AG-UI event](https://docs.ag-ui.com/introduction) the run produces, not only text. Narrow on `chunk.type` (a discriminated union) before reading type-specific fields. The common ones: | `chunk.type` | Meaning | Key fields | |--------------|---------|-----------| @@ -273,18 +337,18 @@ When multiple middleware define `onChunk`, chunks flow through them in order. If | `TOOL_CALL_START` / `TOOL_CALL_ARGS` / `TOOL_CALL_END` | Tool invocation streaming | `toolCallId`, `toolCallName`, `delta` (args), result on end | | `STEP_STARTED` / `STEP_FINISHED` | Thinking / reasoning steps | `delta`, `signature` | | `STATE_SNAPSHOT` / `STATE_DELTA` | Agent state sync | `snapshot`, `delta` | -| `CUSTOM` | Extensibility events (incl. structured-output — see below) | `name`, `value` | +| `CUSTOM` | Extensibility events, including structured output (see below) | `name`, `value` | See the [AG-UI protocol docs](https://docs.ag-ui.com/introduction) for the full event catalogue and exact field shapes. #### Transforming structured-output chunks -There is **no separate `onStructuredOutputChunk` hook** — and you don't need one. When `chat()` is invoked with `outputSchema`, the structured-output chunks (the JSON `TEXT_MESSAGE_CONTENT` deltas, plus the `structured-output.start` / `structured-output.complete` CUSTOM events and any finalization `RUN_ERROR`) flow through the **same `onChunk` hook** as everything else. You transform, expand, or drop them exactly like any other chunk. +There is **no separate `onStructuredOutputChunk` hook**, and you do not need one. When `chat()` is invoked with `outputSchema`, the structured-output chunks (the JSON `TEXT_MESSAGE_CONTENT` deltas, plus the `structured-output.start` / `structured-output.complete` CUSTOM events and any finalization `RUN_ERROR`) flow through the **same `onChunk` hook** as everything else. You transform, expand, or drop them exactly like any other chunk. How you distinguish them depends on which finalization path the adapter takes: - **Separate-finalization adapters** (`supportsCombinedToolsAndSchema()` does not return `true` for the current model/options): `ctx.phase === 'structuredOutput'` during the finalization call. Discriminate on the phase. -- **Native-combined adapters** (modern OpenAI Chat Completions / Responses, Claude 4.5+, Gemini 3.x, Grok 4.x — see issue #605): the schema-constrained JSON is produced on the model's natural final turn, so **`ctx.phase` stays `'modelStream'`** — the `'structuredOutput'` phase never fires. Discriminate on the CUSTOM event name (`structured-output.start` / `structured-output.complete`) instead. +- **Native-combined adapters** (modern OpenAI Chat Completions / Responses, Claude 4.5+, Gemini 3.x, and Grok 4.x, per issue #605): the schema-constrained JSON is produced on the model's natural final turn, so **`ctx.phase` stays `'modelStream'`**. The `'structuredOutput'` phase never fires. Discriminate on the CUSTOM event name (`structured-output.start` / `structured-output.complete`) instead. ```typescript ignore import { type ChatMiddleware } from "@tanstack/ai"; @@ -294,7 +358,7 @@ const redactStructuredOutput: ChatMiddleware = { onChunk: (ctx, chunk) => { // Separate-finalization path: the JSON streams as TEXT_MESSAGE_CONTENT // during the 'structuredOutput' phase. Transform the delta like any - // other text chunk — here, redact anything that looks like an SSN before + // other text chunk. Here, redact anything that looks like an SSN before // it reaches the client. if ( ctx.phase === "structuredOutput" && @@ -319,11 +383,11 @@ const redactStructuredOutput: ChatMiddleware = { }; ``` -> Why is there `onStructuredOutputConfig` but no `onStructuredOutputChunk`? Because the **config** shape genuinely differs at the structured-output boundary — it carries an `outputSchema` field that plain `ChatMiddlewareConfig` doesn't (see [onStructuredOutputConfig](#onstructuredoutputconfig)). **Chunks** are all just `StreamChunk` regardless of phase, so one `onChunk` plus `ctx.phase` (or the CUSTOM event name) covers every case — a parallel chunk hook would be redundant. +> Why is there `onStructuredOutputConfig` but no `onStructuredOutputChunk`? Because the **config** shape differs at the structured-output boundary: it carries an `outputSchema` field that plain `ChatMiddlewareConfig` doesn't (see [onStructuredOutputConfig](#onstructuredoutputconfig)). **Chunks** are all just `StreamChunk` regardless of phase, so one `onChunk` plus `ctx.phase` (or the CUSTOM event name) covers every case. A parallel chunk hook is not necessary. ### onShouldContinue -Called when the engine is deciding whether to start another agent-loop iteration (after a tool phase or between model turns). Combined with AND semantics across middleware **and** with `agentLoopStrategy` — any explicit `false` stops the loop. Return `true`, `void`, or `undefined` to allow continuation. +Called when the engine is deciding whether to start another agent-loop iteration (after a tool phase or between model turns). Combined with AND semantics across middleware **and** with `agentLoopStrategy`: any explicit `false` stops the loop. Return `true`, `void`, or `undefined` to allow continuation. Does **not** abort the run: the stream finishes normally with the current messages. Use `ctx.abort()` only for a hard abort. @@ -510,14 +574,25 @@ capability, then return those fields from `onConfig` when | --- | --- | | `onInterruptBoundary` | Nothing. It can only pause. | | `onInterruptResolution` | Pending-tool policy (`toolResume`) | -| `onConfig` | `messages`, `systemPrompts`, `tools`, `modelOptions`, `metadata` | +| `onConfig` | `messages`, `systemPrompts`, `tools`, `modelOptions`, `metadata`, `reasoning`, `promptCache`, `toolChoice` | The full resume order, plus an example that writes a user note into the system prompt, is in [Apply Answers](../interrupts/apply-answers). ### onBeforeToolCall -Called before each tool executes. The first middleware that returns a non-void decision short-circuits — remaining middleware are skipped for that tool call. +Called before final input validation and tool execution. The first middleware that returns a non-void decision skips the remaining middleware for that call. + +The normal input path is: + +1. Parse the provider's raw JSON arguments. +2. Run `onBeforeToolCall` with the parsed input. +3. Validate the final input against the tool's schema. +4. Execute the server tool, or emit the validated client execution descriptor. + +`transformArgs` replaces the input for the final check. A validation error becomes a tool error. The tool does not execute, and no client execution descriptor is emitted. + +An approval request can show a checked preview before this path. That preview does not replace the raw input for the final check. Outstanding or denied approvals do not run `onBeforeToolCall` or dispatch the tool. Approved `editedArgs` remain raw until middleware and final validation run. See [Tool Approval](../interrupts/tool-approval). ```typescript import { type ChatMiddleware } from "@tanstack/ai"; @@ -560,13 +635,13 @@ The `hookCtx` provides: |-------|------|-------------| | `toolCall` | `ToolCall` | Raw tool call object | | `tool` | `Tool \| undefined` | Resolved tool definition | -| `args` | `unknown` | Parsed arguments | +| `args` | `unknown` | Parsed raw input, before final schema validation | | `toolName` | `string` | Tool name | | `toolCallId` | `string` | Tool call ID | ### onAfterToolCall -Called after each tool execution (or skip). All middleware run — there is no short-circuiting. +Called after each tool execution (or skip). Every middleware runs, in array order. ```typescript import { type ChatMiddleware } from "@tanstack/ai"; @@ -593,9 +668,46 @@ The `info` object provides: | `toolCallId` | `string` | Tool call ID | | `ok` | `boolean` | Whether execution succeeded | | `duration` | `number` | Execution time in milliseconds | -| `result` | `unknown` | Result (when `ok` is true) | +| `result` | `unknown` | Result (when `ok` is true). After a `replaceResult` from an earlier middleware, the new result. | | `error` | `unknown` | Error (when `ok` is false) | +#### Replace a tool result + +A tool can return more than the model needs, for example 500 rows or a long log. The model reads every token of it. Return `{ type: 'replaceResult', result }` to give the model and the stream a different result: + +```typescript +import { type ChatMiddleware } from "@tanstack/ai"; + +const firstTwentyRows: ChatMiddleware = { + name: "first-twenty-rows", + onAfterToolCall: (ctx, info) => { + if (!info.ok || !Array.isArray(info.result)) return; + if (info.result.length <= 20) return; + return { + type: "replaceResult", + result: { + rows: info.result.slice(0, 20), + omitted: info.result.length - 20, + }, + }; + }, +}; +``` + +**Return values:** + +| Return | Effect | +|--------|--------| +| `void` / `undefined` | The result stays the same | +| `{ type: 'replaceResult', result }` | The model gets `result`. The stream sends it in `TOOL_CALL_RESULT`, so the client shows it too. | + +Several middleware can replace the same result: + +- Each middleware gets the result of the middleware before it as `info.result`. The last replacement wins. +- For example, with `middleware: [firstTwentyRows, summarize]`, `summarize` gets the 20 rows. +- A failed call stays failed. The replacement becomes the error that the model reads. +- A `skip` result from `onBeforeToolCall` also goes through this hook. `info.result` is the parsed skip result. + ### Tool hook order in one turn When the model calls several server tools in one turn, the tools run at the same time: @@ -650,10 +762,10 @@ Exactly **one** terminal hook fires per `chat()` invocation. They are mutually e > - `onStructuredOutputConfig` fires before the separate provider call, and `ctx.phase` is `'structuredOutput'` for its chunks. > - `onIteration` does **not** fire for finalization; it only fires for agent-loop iterations. > - `onFinish` fires after finalization completes. Its `info` object reflects the **agent loop's** terminal state. -> - `info.content` — the agent loop's accumulated text. Separate-finalization JSON deltas are **not** included. Middleware can observe the completed result through the `structured-output.complete` CUSTOM event in `onChunk`. -> - `info.usage` — the agent loop's last `RUN_FINISHED.usage`. For a tools-less structured-output run (no agent-loop iteration produces `RUN_FINISHED`), this is `undefined`. To capture finalization tokens, use `onUsage` — that hook fires for **every** `RUN_FINISHED` carrying usage, including the finalization call. -> - `info.finishReason` — the agent loop's last `finishReason`. `null` when no agent-loop iteration produced `RUN_FINISHED` (e.g. a tools-less structured-output run). -> - `info.duration` — wall-clock duration of the entire `chat()` invocation, including finalization. +> - `info.content`: the agent loop's accumulated text. Separate-finalization JSON deltas are **not** included. Middleware can observe the completed result through the `structured-output.complete` CUSTOM event in `onChunk`. +> - `info.usage`: the agent loop's last `RUN_FINISHED.usage`. For a tools-less structured-output run (no agent-loop iteration produces `RUN_FINISHED`), this is `undefined`. To capture finalization tokens, use `onUsage`. That hook fires for **every** `RUN_FINISHED` carrying usage, including the finalization call. +> - `info.finishReason`: the agent loop's last `finishReason`. `null` when no agent-loop iteration produced `RUN_FINISHED` (e.g. a tools-less structured-output run). +> - `info.duration`: wall-clock duration of the entire `chat()` invocation, including finalization. > > **Native-combined output:** Adapters with native-combined support produce the schema-constrained JSON in the regular agent-loop stream. `onStructuredOutputConfig` does not fire, `ctx.phase` remains `'modelStream'`, and `onIteration` fires for the iteration that produces the JSON. The JSON is agent-loop text, so `info.content` includes it. Middleware observes the `structured-output.complete` event in `onChunk` during the same phase. > @@ -689,7 +801,7 @@ The `info` object for `onFinish` (`FinishInfo`): | `finishReason` | `string \| null` | The agent loop's last `finishReason`. `null` when no agent-loop iteration produced `RUN_FINISHED` (e.g. a tools-less `chat({ outputSchema })` run). | | `duration` | `number` | Total run duration in milliseconds, including any structured-output finalization. | | `content` | `string` | The agent loop's accumulated text content. Includes native-combined structured JSON; excludes separate-finalization JSON. Observe the completed result through the `structured-output.complete` CUSTOM event via `onChunk`. | -| `usage` | `{ promptTokens; completionTokens; totalTokens } \| undefined` | **Optional.** The agent loop's last `RUN_FINISHED.usage`. **Does not include finalization tokens** — use `onUsage` to observe those. Always guard with `if (info.usage)` or `info.usage?.`. | +| `usage` | `{ promptTokens; completionTokens; totalTokens } \| undefined` | **Optional.** The agent loop's last `RUN_FINISHED.usage`. **Does not include finalization tokens.** Use `onUsage` to observe those. Always guard with `if (info.usage)` or `info.usage?.`. | ## Context Object @@ -814,13 +926,13 @@ const stream = chat({ | Hook | Composition | Effect of Order | |------|------------|----------------| -| `onConfig` | **Piped** — each receives previous output | Earlier middleware transforms first | -| `onStructuredOutputConfig` | **Piped** — each receives previous output | Earlier middleware transforms first | +| `onConfig` | **Piped**: each receives previous output | Earlier middleware transforms first | +| `onStructuredOutputConfig` | **Piped**: each receives previous output | Earlier middleware transforms first | | `onStart` | Sequential | All run in order | -| `onChunk` | **Piped** — chunks flow through each middleware | If first drops a chunk, later middleware never see it | -| `onBeforeToolCall` | **First-win** — first non-void decision wins | Earlier middleware has priority | -| `onShouldContinue` | **AND** — any explicit `false` stops the loop | Order only affects which middleware runs first when short-circuiting | -| `onAfterToolCall` | Sequential | All run in order | +| `onChunk` | **Piped**: chunks flow through each middleware | If first drops a chunk, later middleware never see it | +| `onBeforeToolCall` | **First-win**: the first non-void decision wins | Earlier middleware has priority | +| `onShouldContinue` | **AND**: any explicit `false` stops the loop | Order only affects which middleware runs first when short-circuiting | +| `onAfterToolCall` | **Piped**: each receives the result of the previous `replaceResult` | The last replacement wins | | `onUsage` | Sequential | All run in order | | `onFinish/onAbort/onError` | Sequential | All run in order | @@ -830,7 +942,7 @@ Middleware often need to **share state**. A provider middleware sets something u ### Creating a capability -A capability is created with `createCapability()('name')` — a **curried** call: +A capability is created with `createCapability()('name')`, a **curried** call: ```typescript import { createCapability } from "@tanstack/ai"; @@ -839,11 +951,11 @@ const counterCapability = createCapability<{ value: number }>()("counter"); const [getCounter, provideCounter] = counterCapability; ``` -The currying is deliberate: you supply the **value type** explicitly (`<{ value: number }>`) while the **name literal** is inferred from the argument (`"counter"`). A single `createCapability('name')` call can't do both — supplying `T` explicitly stops TypeScript inferring the name, collapsing it to `string` and defeating the compile-time coverage check that keys on the literal name. +The currying is deliberate: you supply the **value type** explicitly (`<{ value: number }>`) while the **name literal** is inferred from the argument (`"counter"`). A single `createCapability('name')` call cannot do both. Supplying `T` explicitly stops TypeScript inferring the name, collapsing it to `string` and defeating the compile-time coverage check that keys on the literal name. The returned `counterCapability` is a hybrid value: -- It **destructures to `[get, provide]`** — the two accessors you use inside hooks. +- It **destructures to `[get, provide]`**: the two accessors you use inside hooks. - It **is itself the identity** you list in `requires` / `provides`. There is no separate token to import. The accessors: @@ -851,30 +963,30 @@ The accessors: | Accessor | Behavior | |----------|----------| | `getCounter(ctx)` | Returns the value. **Throws** if the capability was never provided. | -| `getCounter(ctx, { optional: true })` | Returns `TValue \| undefined` — no throw when absent. | +| `getCounter(ctx, { optional: true })` | Returns `TValue \| undefined`. No throw when absent. | | `provideCounter(ctx, value)` | Sets the value for this run. Call it from `setup`. | -Equivalently, the context exposes `ctx.get(capability)`, `ctx.getOptional(capability)`, and `ctx.provide(capability, value)` — pass the capability handle directly. These are typed by the handle you pass (`ctx.get(counterCapability)` returns the value type), so `getCounter(ctx)` and `ctx.get(counterCapability)` are interchangeable — use whichever reads better in your hook. +Equivalently, the context exposes `ctx.get(capability)`, `ctx.getOptional(capability)`, and `ctx.provide(capability, value)`. Pass the capability handle directly. These are typed by the handle you pass (`ctx.get(counterCapability)` returns the value type), so `getCounter(ctx)` and `ctx.get(counterCapability)` are interchangeable. Use whichever reads better in your hook. > **Capability names must be unique across your app.** The compile-time coverage check keys on the name literal (runtime keys on the handle reference), so two capabilities sharing a name will conflate in the type-level check. ### The `setup` hook -Provisioning happens in a dedicated `setup(ctx)` hook. It **runs first** — before any `onConfig` (init), across all middleware in array order — so that by the time the rest of the lifecycle begins, every capability is in place. `setup` receives the stable `ChatMiddlewareContext` (not the mutable config), and may be async. +Provisioning happens in a dedicated `setup(ctx)` hook. It **runs first**, before any `onConfig` (init), across all middleware in array order. So by the time the rest of the lifecycle begins, every capability is in place. `setup` receives the stable `ChatMiddlewareContext` (not the mutable config), and may be async. ### `requires` / `provides` / `optionalRequires` -Three array fields on a middleware declare its capability contract. Each is a `ReadonlyArray` — you list the capability handles themselves: +Three array fields on a middleware declare its capability contract. Each is a `ReadonlyArray`. You list the capability handles themselves: | Field | Meaning | |-------|---------| | `provides` | Capabilities this middleware sets up. Each one **must** be `provide`d inside `setup`, or `chat()` throws after the setup phase. | | `requires` | Capabilities this middleware reads. `chat()` validates (compile time + runtime) that some earlier middleware provides each one. | -| `optionalRequires` | Capabilities used **if present** but not required. Non-gating — never causes a validation error. Read with `getX(ctx, { optional: true })`. | +| `optionalRequires` | Capabilities used **if present** but not required. Non-gating: never causes a validation error. Read with `getX(ctx, { optional: true })`. | ### Array example -Author middleware with `defineChatMiddleware` — it sharpens the `requires` / `provides` tuple types so the coverage check and builder can read them precisely. Here a **provider** sets up a counter in `setup`, and a **consumer** reads it in a hook: +Author middleware with `defineChatMiddleware`. It sharpens the `requires` / `provides` tuple types so the coverage check and builder can read them precisely. Here a **provider** sets up a counter in `setup`, and a **consumer** reads it in a hook: ```typescript import { @@ -917,7 +1029,7 @@ const stream = chat({ }); ``` -If you drop `withCounter` from the array, `chat()` reports a compile-time error at the `middleware` option naming the missing `"counter"` capability — and throws at runtime before the adapter is ever called. +If you drop `withCounter` from the array, `chat()` reports a compile-time error at the `middleware` option naming the missing `"counter"` capability, and throws at runtime before the adapter is ever called. ### Builder example @@ -953,7 +1065,7 @@ const countsChunks = defineChatMiddleware({ const middleware = createChatMiddleware() .use(withCounter) // provides "counter" - .use(countsChunks) // requires "counter" — OK, already provided above + .use(countsChunks) // requires "counter": OK, already provided above .build(); const stream = chat({ @@ -963,7 +1075,7 @@ const stream = chat({ }); ``` -Swap the two `.use()` calls (`.use(countsChunks).use(withCounter)`) and the builder rejects it at the `.use(countsChunks)` line — the consumer is ordered before its provider, so `"counter"` isn't in the provided set yet. +Swap the two `.use()` calls (`.use(countsChunks).use(withCounter)`) and the builder rejects it at the `.use(countsChunks)` line. The consumer is ordered before its provider, so `"counter"` isn't in the provided set yet. ### Validation guarantees @@ -971,13 +1083,13 @@ The capability system fails loudly and early: - **Compile-time coverage.** A required capability that nothing provides surfaces as a type error at the `middleware` option. This is enforced two ways: an **array coverage check** on `middleware: [...]`, and the order-aware **`createChatMiddleware()` builder** (which additionally enforces ordering). - **Runtime coverage.** Even if types are bypassed, `chat()` validates coverage and **throws before the adapter runs** if a required capability is missing. -- **Post-`setup` assertion.** If a middleware declares a capability in `provides` but never calls its `provide` accessor during `setup`, `chat()` throws after the setup phase — you can't silently forget to provision. +- **Post-`setup` assertion.** If a middleware declares a capability in `provides` but never calls its `provide` accessor during `setup`, `chat()` throws after the setup phase. You cannot forget to provision without an error. - **Duplicate provide → last-wins + warning.** If two middleware provide the same capability, the last write wins and a development warning is emitted. - **Unique names.** Capability `name`s must be unique across your app; the compile-time coverage check keys on the name literal (runtime keys on the handle reference). ## Built-in Middleware -TanStack AI ships ready-made middleware for common cases — caching tool results, redacting streamed text, and OpenTelemetry tracing: +TanStack AI ships ready-made middleware for common cases: caching tool results, redacting streamed text, and OpenTelemetry tracing: | Middleware | Import | What it does | |------------|--------|--------------| @@ -1086,27 +1198,33 @@ const auditTrail: ChatMiddleware = { ### Per-Iteration Tool Swapping -Expose different tools at different stages of the agent loop: +Expose different tools at different stages of the agent loop. A `tools` list that `onConfig` returns stays for the later model calls. So keep the full list from the start of the run, and give it back after the first call: ```typescript -import { type ChatMiddleware } from "@tanstack/ai"; - -const toolSwapper: ChatMiddleware = { - name: "tool-swapper", - onConfig: (ctx, config) => { - if (ctx.phase !== "beforeModel") return; +import { type ChatMiddleware, type ChatMiddlewareConfig } from "@tanstack/ai"; - if (ctx.iteration === 0) { - // First iteration: only allow search - return { - tools: config.tools.filter((t) => t.name === "search"), - }; - } - // Later iterations: allow all tools - }, -}; +// A function, so that each chat() call keeps its own full list. +function toolSwapper(): ChatMiddleware { + let allTools: ChatMiddlewareConfig["tools"] = []; + return { + name: "tool-swapper", + onConfig: (ctx, config) => { + if (ctx.phase === "init") { + allTools = config.tools; + return; + } + if (ctx.phase !== "beforeModel") return; + // First model call: only search. Later calls: every tool again. + return ctx.iteration === 0 + ? { tools: allTools.filter((t) => t.name === "search") } + : { tools: allTools }; + }, + }; +} ``` +Pass `middleware: [toolSwapper()]` to `chat()`. + ### Content Filtering Drop or transform chunks before they reach the consumer: @@ -1162,6 +1280,7 @@ import type { ToolCallHookContext, BeforeToolCallDecision, AfterToolCallInfo, + AfterToolCallDecision, IterationInfo, ToolPhaseCompleteInfo, UsageInfo, @@ -1186,8 +1305,9 @@ import type { ## Next Steps -- [Built-in Middleware](./built-in-middleware) — `toolCacheMiddleware`, `contentGuardMiddleware`, `otelMiddleware` +- [Built-in Middleware](./built-in-middleware): `toolCacheMiddleware`, `contentGuardMiddleware`, `otelMiddleware` - [Compaction](./compaction): keep long chats under the context limit with `withCompaction` -- [OpenTelemetry](./otel) — emit traces and metrics via `otelMiddleware`- [Tools](../tools/tools) — Learn about the isomorphic tool system -- [Agentic Cycle](../chat/agentic-cycle) — Understand the multi-step agent loop -- [Streaming](../chat/streaming) — How streaming works in TanStack AI +- [OpenTelemetry](./otel): emit traces and metrics via `otelMiddleware` +- [Tools](../tools/tools): learn about the isomorphic tool system +- [Agentic Cycle](../chat/agentic-cycle): understand the multi-step agent loop +- [Streaming](../chat/streaming): how streaming works in TanStack AI diff --git a/docs/advanced/prompt-caching.md b/docs/advanced/prompt-caching.md new file mode 100644 index 0000000000..12d8c825be --- /dev/null +++ b/docs/advanced/prompt-caching.md @@ -0,0 +1,237 @@ +--- +title: Prompt Caching +id: prompt-caching +order: 4 +description: "chat() caches the stable start of each request by default, so a repeated system prompt, tool list, and history cost less and answer sooner. Pick the retention, set the cache key, or turn it off." +keywords: + - tanstack ai + - prompt caching + - promptCache + - cache_control + - prompt_cache_key + - cached tokens + - token cost +--- + +An agent sends the same start with every request: the system prompt, the tools, and the earlier messages. You pay for those input tokens on every call, and the model reads them again before it answers. + +Prompt caching lets the provider keep that start. The next request with the same start reads it from the cache, so it costs less and the answer starts sooner. `chat()` turns it on by default. + +## Cache a conversation + +Pass the same `threadId` with every request of a conversation. `chat()` uses it as the cache key: + +```typescript +import { + chat, + chatParamsFromRequest, + toServerSentEventsResponse, +} from "@tanstack/ai"; +import { openaiText } from "@tanstack/ai-openai"; + +export async function POST(request: Request) { + const { messages, threadId, runId } = await chatParamsFromRequest(request); + + const stream = chat({ + adapter: openaiText("gpt-6.1-sol"), + systemPrompts: ["You are a support agent for Acme."], + messages, + // The cache key. useChat sends one threadId for each conversation. + threadId, + runId, + }); + + return toServerSentEventsResponse(stream); +} +``` + +From the second request on, the system prompt and the earlier turns come from the cache. + +The cache key is the first of these that you set: + +1. `promptCache.key`. +2. The `threadId` or `conversationId` of the call. +3. Nothing. Claude does not need a key. OpenAI routes requests by key, so a stable `threadId` gives a better hit rate. + +A request reads from the cache only when it starts the same way as an earlier request. Text that changes on every call, such as the current time in the system prompt, ends the match at that point. + +`summarize()` and direct adapter calls do not add cache fields. + +## Pick the retention + +Set `promptCache` on `chat()`: + +| Value | What it does | +|---|---| +| `'short'` | The default. The provider keeps the cache for a short time, 5 minutes on Claude. | +| `'long'` | The provider keeps the cache longer, 1 hour on Claude. For other providers, see [What each provider gets](#what-each-provider-gets). | +| `'none'` | No caching. | + +To set your own key too, pass an object: + +```typescript +import { chat } from "@tanstack/ai"; +import { openaiText } from "@tanstack/ai-openai"; + +const stream = chat({ + adapter: openaiText("gpt-6.1-sol"), + messages: [{ role: "user", content: "Where is my order?" }], + promptCache: { retention: "long", key: "acme-support" }, +}); +``` + +Both fields are optional. Without `retention`, the retention is `'short'`. The types are exported from `@tanstack/ai`: `PromptCacheRetention`, `PromptCacheOptions`, and `ResolvedPromptCache`. + +In a harness, you set `promptCache` one time for every session. See [Build your first harness](../harness/overview#prompt-caching). To change it for one prompt of a session, see [Give one prompt its own settings](../harness/turn-control#give-one-prompt-its-own-settings). + +A middleware can change the retention or the key of one model call. See [Change the prompt cache of a call](./middleware#change-the-prompt-cache-of-a-call). + +## When to turn it off or keep it longer + +On OpenAI, Gemini, and Mistral, a cache write costs nothing extra and a cache read costs less. Keep the default there. + +Claude bills the cache write. This applies to Claude on Anthropic, Bedrock, Vertex, and OpenRouter: + +| Price, compared to normal input | `'short'` (5 minutes) | `'long'` (1 hour) | +|---|---|---| +| Cache write | about 1.25 times | about 2 times | +| Cache read | about 0.1 times | about 0.1 times | + +The newest Claude models read from the cache for less. With `'short'`, the same start sent 2 times in 5 minutes already costs less than no cache. With `'long'`, it costs less from the third request. + +The usage of a call gives the 1-hour part of the cache write as `promptTokensDetails.cacheWrite1hTokens`, on Anthropic and Bedrock. To price it, see [Work out what a call cost](../models/catalog#work-out-what-a-call-cost). + +Two cases change the choice: + +- **A large prompt that you send one time.** For example, one long document per call that is never sent again. Nothing reads the cache, so the write is only extra cost. Use `'none'`. +- **Long pauses.** The short Claude cache ends 5 minutes after its last read. If people often reply after more than 5 minutes, use `'long'`. + +```typescript +import { chat } from "@tanstack/ai"; +import type { ModelMessage } from "@tanstack/ai"; +import { anthropicText } from "@tanstack/ai-anthropic"; + +// One long document per call, never sent again: skip the cache write. +export function reviewDocument(document: string) { + return chat({ + adapter: anthropicText("claude-sonnet-5-5"), + messages: [{ role: "user", content: `Review this contract:\n\n${document}` }], + promptCache: "none", + }); +} + +// People often reply after more than 5 minutes: keep the cache for 1 hour. +export function reply(messages: Array, threadId: string) { + return chat({ + adapter: anthropicText("claude-sonnet-5-5"), + messages, + threadId, + promptCache: "long", + }); +} +``` + +Claude caches a prompt start only when it is longer than a model minimum, from 512 to 4096 tokens. A shorter start is not cached and costs nothing extra. + +## What each provider gets + +Automatic caching adds these fields to the request. + +### Claude models + +| Adapter | `'short'` | `'long'` | +|---|---|---| +| `anthropicText`, Claude on Vertex | `cache_control: { type: 'ephemeral' }` on the system blocks, the last tool, and the last block of the last user message. At most 4 markers. With [mid-conversation tool changes](./mid-conversation-changes) on, the tool marker goes on the last start tool instead, and a mid-conversation `system` message at the end takes the message marker. | The same markers with `ttl: '1h'` | +| `bedrockText` (Converse), Claude models only | A `cachePoint` after the system prompt and at the end of the last user message | Both with `ttl: '1h'` | +| `openRouterText` with an `anthropic/*` model | `sessionId`, plus markers on the system prompt, the last tool, and the last message | The markers with `ttl: '1h'` | +| `openaiCompatible` with `compat.cacheControlFormat: 'anthropic'` (Claude through a gateway) | Markers on the system prompt, the last tool, and the last message | The markers with `ttl: '1h'` | + +### Other providers + +| Adapter | `'short'` | `'long'` | +|---|---|---| +| `openaiText` (Responses) | `prompt_cache_key`, cut to 64 characters | Also `prompt_cache_retention: '24h'`. On gpt-5.6 and later, `prompt_cache_options: { ttl: '30m' }` in its place. | +| `openaiChatCompletions`, `openaiCompatible` | `prompt_cache_key`, only to `api.openai.com` | The key and `prompt_cache_retention: '24h'`, when the provider supports it (`compat.supportsLongCacheRetention`) | +| `openRouterText`, every model | `sessionId`, set to the cache key | The same | +| `mistralText` | `prompt_cache_key` | The same | +| Gemini, Vertex Gemini | Nothing. The provider caches by itself. | The same | + +On gpt-5.6 and later, `'none'` sends `prompt_cache_options: { mode: 'explicit' }`, so OpenAI caches nothing. + +The Workers AI binding (`cloudflareBindingFetch`) passes the cache fields through. To turn caching off there, pass `promptCache: 'none'`. + +## Set the cache yourself + +Your own cache settings win over the automatic ones: + +- **Anthropic**: `metadata.cache_control` on a system prompt, a message part, or a tool, or `modelOptions.cache_control`. Then `chat()` adds no markers to that request. See [Anthropic](../adapters/anthropic#prompt-caching). +- **Bedrock**: `metadata.cachePoint` on a system prompt, a message part, or a tool. Then `chat()` adds no cache points to that request. See [Amazon Bedrock](../adapters/bedrock#prompt-caching). +- **OpenAI**: `modelOptions.prompt_cache_key` and `modelOptions.prompt_cache_retention` replace the automatic values. See [OpenAI](../adapters/openai#prompt-caching). +- **OpenRouter**: `modelOptions.sessionId` replaces the automatic `sessionId`. + +## When the tools change + +The tools are near the start of each request. A tool that a middleware adds during a run changes that start, so the next call cannot read the cache from there on. On GPT and Claude models with a mid-conversation channel, `chat()` keeps the start the same and sends the new tool later in the conversation. A system prompt that you add at the end of the list goes the same way. See [Mid-Conversation Changes](./mid-conversation-changes). + +## Track cached tokens and cost + +`usage.promptTokens` is the full input on every adapter, cached tokens included. Two fields split it: + +- `usage.promptTokensDetails.cachedTokens`: the tokens read from the cache. +- `usage.promptTokensDetails.cacheWriteTokens`: the tokens written to the cache. Claude reports this field. + +Read them in the [`onUsage`](./middleware#onusage) middleware hook. It runs one time for each model call: + +```typescript +import { chat } from "@tanstack/ai"; +import type { ChatMiddleware } from "@tanstack/ai"; +import { anthropicText } from "@tanstack/ai-anthropic"; + +const cacheLog: ChatMiddleware = { + name: "cache-log", + onUsage: (ctx, usage) => { + const read = usage.promptTokensDetails?.cachedTokens ?? 0; + const written = usage.promptTokensDetails?.cacheWriteTokens ?? 0; + console.log( + `Call ${ctx.iteration}: ${usage.promptTokens} input, ${read} read from cache, ${written} written`, + ); + }, +}; + +const stream = chat({ + adapter: anthropicText("claude-sonnet-5-5"), + messages: [{ role: "user", content: "Hello!" }], + threadId: "thread-1", + middleware: [cacheLog], +}); +``` + +The `RUN_FINISHED` event has the same `usage` object. See [Stream Events](../chat/stream-events). + +To get the input cost, price each part with its own rate: + +```typescript +import type { TokenUsage } from "@tanstack/ai"; + +// Prices per token, from the price list of your provider. +interface InputPrices { + input: number; + cacheRead: number; + cacheWrite: number; +} + +export function inputCost(usage: TokenUsage, prices: InputPrices) { + const cached = usage.promptTokensDetails?.cachedTokens ?? 0; + const written = usage.promptTokensDetails?.cacheWriteTokens ?? 0; + const uncached = usage.promptTokens - cached - written; + return ( + uncached * prices.input + + cached * prices.cacheRead + + written * prices.cacheWrite + ); +} +``` + +To see cache reads in a test with no provider, use the `cache` option of the fake model. See [Test with a Fake Model](./testing). + +You now see how much of each request comes from the cache, and what its input costs. diff --git a/docs/advanced/runtime-adapter-switching.md b/docs/advanced/runtime-adapter-switching.md index bf68cf93d1..fc401bcb82 100644 --- a/docs/advanced/runtime-adapter-switching.md +++ b/docs/advanced/runtime-adapter-switching.md @@ -19,31 +19,107 @@ Learn how to build interfaces where users can switch between LLM providers at ru With TanStack AI, the model is passed directly to the adapter factory function. This gives you full type safety and autocomplete at the point of definition: ```typescript -import { chat, toServerSentEventsResponse } from '@tanstack/ai' +import { chat, chatParamsFromRequest, toServerSentEventsResponse } from '@tanstack/ai' import { anthropicText } from '@tanstack/ai-anthropic' import { openaiText } from '@tanstack/ai-openai' -type Provider = 'openai' | 'anthropic' - -// Define adapters with their models - autocomplete works here! +// Each factory retains its model's option types. const adapters = { - anthropic: () => anthropicText('claude-sonnet-4-6'), // ✅ Autocomplete! - openai: () => openaiText('gpt-5.5'), // ✅ Autocomplete! + anthropic: () => anthropicText('claude-sonnet-5-5'), + openai: () => openaiText('gpt-5.5'), } -async function handleRequest(request: Request) { - // In your request handler: - const body = await request.json() - const provider: Provider = body.forwardedProps?.provider || 'openai' +export async function POST(request: Request) { + const params = await chatParamsFromRequest(request) + const provider = params.forwardedProps.provider === 'anthropic' ? 'anthropic' : 'openai' const stream = chat({ adapter: adapters[provider](), - messages: body.messages, + messages: params.messages, + }) + return toServerSentEventsResponse(stream) +} +``` + +## Keep saved history when you switch + +A conversation can contain replies from several providers. Keep each message's metadata when you save or send the history. `chat()` records the source under `metadata.tanstack.source`: + +```json +{ + "provider": "openai", + "api": "openai-responses", + "model": "gpt-5.5" +} +``` + +The provider, wire API, and requested model must all match for same-source replay. `metadata.tanstack.model` holds the provider-reported response model when available. It can differ from the requested model. + +For foreign history, the target adapter converts readable thinking to text and drops foreign signatures and redacted thinking. It also remaps tool call IDs and their results together. Messages without source metadata receive the same-source treatment. + +Core prepares a request copy of the history: + +- Failed or aborted assistant messages and their associated tool results are omitted. +- Unanswered tool calls receive an error result: `No result provided`. +- System messages between a tool call and its results move after the results. + +These changes leave your saved transcript intact. A transport disconnect alone does not mark a provider reply as failed. + +Subagent cards use a separate view of the saved transcript. A host message holds the routed subagents' answers. The card view removes only a host that it can identify. Ordinary assistant text and ambiguous messages stay. A new host uses an unused ID, so existing replies stay. + +Gemini replays saved thinking and ordered tool history through its supported wire fields. This also applies to same-source and source-free history with thinking. Ordinary history without these parts retains its request shape. + +Use the existing request parser on the server to retain message metadata: + +```typescript +import { chat, chatParamsFromRequest, toServerSentEventsResponse } from '@tanstack/ai' +import { anthropicText } from '@tanstack/ai-anthropic' +import { openaiText } from '@tanstack/ai-openai' + +export async function POST(request: Request) { + const params = await chatParamsFromRequest(request) + const adapter = params.forwardedProps.provider === 'anthropic' + ? anthropicText('claude-sonnet-5-5') + : openaiText('gpt-5.5') + + return toServerSentEventsResponse(chat({ + adapter, + messages: params.messages, + threadId: params.threadId, + runId: params.runId, + })) +} +``` + +Send the selected provider from the client: + +```tsx +import { useState } from 'react' +import { fetchServerSentEvents, useChat } from '@tanstack/ai-react' + +export function ProviderChat() { + const [provider, setProvider] = useState('openai') + const { messages, sendMessage } = useChat({ + connection: fetchServerSentEvents('/api/chat'), + forwardedProps: { provider }, }) + + return ( + <> + + +

{messages.length} messages in this conversation

+ + ) } ``` -## Why This Works +The next request uses the selected provider with the same conversation history. + +## Adapter model types Each adapter factory function accepts a model name as its first argument and returns a fully typed adapter: @@ -70,7 +146,7 @@ Here's a complete example showing a multi-provider chat API: ```typescript ignore import { createFileRoute } from '@tanstack/react-router' -import { chat, maxIterations, toServerSentEventsResponse } from '@tanstack/ai' +import { chat, chatParamsFromRequest, toServerSentEventsResponse } from '@tanstack/ai' import { openaiText } from '@tanstack/ai-openai' import { anthropicText } from '@tanstack/ai-anthropic' import { geminiText } from '@tanstack/ai-gemini' @@ -80,8 +156,8 @@ type Provider = 'openai' | 'anthropic' | 'gemini' | 'ollama' // Define adapters with their models const adapters = { - anthropic: () => anthropicText('claude-sonnet-4-6'), - gemini: () => geminiText('gemini-3-flash-preview'), + anthropic: () => anthropicText('claude-sonnet-5-5'), + gemini: () => geminiText('gemini-3.8-flash'), ollama: () => ollamaText('mistral:7b'), openai: () => openaiText('gpt-5.5'), } @@ -91,17 +167,14 @@ export const Route = createFileRoute('/api/chat')({ handlers: { POST: async ({ request }) => { const abortController = new AbortController() - const body = await request.json() - // `forwardedProps` is the AG-UI field set by `useChat({ forwardedProps })`. - // The legacy `body.data.provider` access still works (mirrored on the - // wire for backward compatibility) but `forwardedProps` is preferred. - const provider: Provider = body.forwardedProps?.provider || 'openai' + const params = await chatParamsFromRequest(request) + const selected = params.forwardedProps.provider + const provider: Provider = selected === 'anthropic' || selected === 'gemini' || selected === 'ollama' + ? selected : 'openai' const stream = chat({ adapter: adapters[provider](), - tools: [...], - systemPrompts: [...], - messages: body.messages, + messages: params.messages, abortController, }) @@ -157,8 +230,8 @@ import { anthropicSummarize } from '@tanstack/ai-anthropic' type SummarizeProvider = 'openai' | 'anthropic' const summarizeAdapters: Record ReturnType> = { - openai: () => openaiSummarize('gpt-5.4-mini'), - anthropic: () => anthropicSummarize('claude-sonnet-4-6'), + openai: () => openaiSummarize('gpt-5.5'), + anthropic: () => anthropicSummarize('claude-sonnet-5-5'), } export async function POST(request: Request) { @@ -179,62 +252,51 @@ export async function POST(request: Request) { ## Migration from Switch Statements -If you have existing code using switch statements, here's how to migrate: +You can replace a provider switch with a map of adapter factories. Both forms keep the model on its adapter. ### Before -```typescript ignore -let adapter -let model - -switch (provider) { - case 'anthropic': - adapter = anthropicText() - model = 'claude-sonnet-4-6' - break - case 'openai': - default: - adapter = openaiText() - model = 'gpt-5.5' - break -} +```typescript +import { chat, chatParamsFromRequest, toServerSentEventsResponse } from '@tanstack/ai' +import { anthropicText } from '@tanstack/ai-anthropic' +import { openaiText } from '@tanstack/ai-openai' -const stream = chat({ - adapter: adapter as any, - model: model as any, - messages, -}) +export async function POST(request: Request) { + const params = await chatParamsFromRequest(request) + let adapter + switch (params.forwardedProps.provider) { + case 'anthropic': + adapter = anthropicText('claude-sonnet-5-5') + break + default: + adapter = openaiText('gpt-5.5') + break + } + + return toServerSentEventsResponse(chat({ adapter, messages: params.messages })) +} ``` ### After ```typescript -import { chat, toServerSentEventsResponse } from '@tanstack/ai' +import { chat, chatParamsFromRequest, toServerSentEventsResponse } from '@tanstack/ai' import { anthropicText } from '@tanstack/ai-anthropic' import { openaiText } from '@tanstack/ai-openai' -type AfterProvider = 'openai' | 'anthropic' - const adapters = { - anthropic: () => anthropicText('claude-sonnet-4-6'), + anthropic: () => anthropicText('claude-sonnet-5-5'), openai: () => openaiText('gpt-5.5'), } export async function POST(request: Request) { - const body = await request.json() - const provider: AfterProvider = body.forwardedProps?.provider ?? 'openai' - - const stream = chat({ + const params = await chatParamsFromRequest(request) + const provider = params.forwardedProps.provider === 'anthropic' ? 'anthropic' : 'openai' + return toServerSentEventsResponse(chat({ adapter: adapters[provider](), - messages: body.messages, - }) - - return toServerSentEventsResponse(stream) + messages: params.messages, + })) } ``` -The key changes: - -1. Replace the switch statement with an object of factory functions -2. Each factory function creates an adapter with the model included -3. No more `as any` casts - full type safety! +Each factory retains the selected model's types. The map lets you add a provider in one place. diff --git a/docs/advanced/testing.md b/docs/advanced/testing.md new file mode 100644 index 0000000000..7c26c1bd1c --- /dev/null +++ b/docs/advanced/testing.md @@ -0,0 +1,127 @@ +--- +title: Test with a Fake Model +id: testing +order: 16 +description: "Test chat(), tools, and middleware with no network and no API key. fakeText() from @tanstack/ai/testing answers from a script and estimates token usage." +keywords: + - tanstack ai + - testing + - fake adapter + - mock model + - unit tests + - fakeText +--- + +A test that calls a real model is slow, costs money, needs an API key, and can get a different answer each time. `fakeText()` from `@tanstack/ai/testing` is a text adapter that answers from a script you write. Pass it to `chat()` like any other adapter. + +## Answer a prompt + +Queue the answers, then run `chat()`: + +```ts group=testing +import { chat } from '@tanstack/ai' +import { fakeText } from '@tanstack/ai/testing' + +const fake = fakeText() +fake.setResponses([{ text: 'Hello!' }]) + +let answer = '' +for await (const chunk of chat({ + adapter: fake, + messages: [{ role: 'user', content: 'Hi' }], +})) { + if (chunk.type === 'TEXT_MESSAGE_CONTENT') answer += chunk.delta +} +console.log(answer) // 'Hello!' +``` + +- Each model call takes the next answer from the queue. +- An empty queue ends the call with a `RUN_ERROR`: "No more fake responses queued". +- `appendResponses` adds answers to the queue, and `pendingResponses()` counts what is left. + +## Test a tool + +An answer with `toolCalls` makes `chat()` run your tools. Queue the answer that comes after the tool results too: + +```ts group=testing +import { toolDefinition } from '@tanstack/ai' +import { z } from 'zod' + +const getWeather = toolDefinition({ + name: 'get_weather', + description: 'Get the weather in a city', + inputSchema: z.object({ city: z.string() }), +}).server(async ({ city }) => ({ city, sky: 'sunny' })) + +const weatherFake = fakeText() +weatherFake.setResponses([ + { toolCalls: [{ name: 'get_weather', input: { city: 'Oslo' } }] }, + { text: 'It is sunny in Oslo.' }, +]) + +for await (const chunk of chat({ + adapter: weatherFake, + messages: [{ role: 'user', content: 'Weather in Oslo?' }], + tools: [getWeather], +})) { + if (chunk.type === 'TOOL_CALL_RESULT') console.log(chunk.content) +} +console.log(weatherFake.state.callCount) // 2 +``` + +## Answer from the request + +A function in the queue builds its answer when the call arrives. It gets one object with the `request` and the `state`: + +```ts group=testing +const echo = fakeText() +echo.setResponses([ + ({ request, state }) => ({ + text: `Call ${state.callCount} saw ${request.messages.length} messages.`, + }), +]) +``` + +## Test errors and overflow + +An answer with `error` fails the call. With [`isContextOverflow`](./compaction#compact-now-after-an-overflow), test the code that compacts and retries: + +```ts group=testing +import { isContextOverflow } from '@tanstack/ai' + +const small = fakeText({ contextWindow: 1_000 }) +small.setResponses([{ error: 'prompt is too long: 2000 tokens > 1000 maximum' }]) + +for await (const chunk of chat({ + adapter: small, + messages: [{ role: 'user', content: 'A very long question' }], +})) { + if (chunk.type === 'RUN_ERROR') { + console.log(isContextOverflow({ error: chunk })) // true + } +} +``` + +The fake also reports token usage on each call: `ceil(characters / 4)` over the request and the answer. So a long message is many tokens, like with a real model. + +## Options + +| Option | What it does | +|---|---| +| `model` | The model id. Default `'fake-model'`. | +| `input` | The input kinds the model reads, for example `['text', 'image']`. | +| `contextWindow` | The context window in tokens. Read it back as `fake.contextWindow`. | +| `tokensPerSecond` | Stream the text at this pace, 4 characters per token. | +| `cache` | With a `threadId`, count the part of the request that matches the thread's previous request as cached tokens. | + +## Answer fields + +| Field | What it does | +|---|---| +| `text` | The visible answer. | +| `thinking` | Thinking text, streamed before the answer. | +| `toolCalls` | `{ name, input?, id? }` calls for `chat()` to run. | +| `finishReason` | `'stop'`, `'length'`, `'content_filter'`, or `'tool_calls'`. Default `'tool_calls'` with tool calls, else `'stop'`. | +| `error` | Fail the call with a `RUN_ERROR` that has this message. | + +Your tests now run with no network and no key, and they get the same answers every time. diff --git a/docs/advanced/typed-options.md b/docs/advanced/typed-options.md index f75efe5e22..b1a7f5f89c 100644 --- a/docs/advanced/typed-options.md +++ b/docs/advanced/typed-options.md @@ -34,8 +34,8 @@ const chatOptions = createChatOptions({ // native key. modelOptions: { temperature: 0.3, - reasoning: { effort: 'medium' }, }, + reasoning: 'medium', }) // Later, anywhere in your codebase: @@ -96,9 +96,7 @@ export const supportChatOptions = createChatOptions({ adapter: openaiText('gpt-5.5'), systemPrompts: ['You are a customer-support assistant for Acme Corp.'], tools: [lookupOrder], - modelOptions: { - reasoning: { effort: 'medium' }, - }, + reasoning: 'medium', }) ``` diff --git a/docs/chat/agentic-cycle.md b/docs/chat/agentic-cycle.md index 14e52b84ff..2fd6945d02 100644 --- a/docs/chat/agentic-cycle.md +++ b/docs/chat/agentic-cycle.md @@ -264,3 +264,56 @@ export async function POST(request: Request) { ``` Place this **before** `toolCacheMiddleware` so over-budget skips win over cache hits. See [`onShouldContinue`](../advanced/middleware#onshouldcontinue) for the hook contract. + +### Stop the loop from a tool (middleware recipe) + +Some tools end the work, for example a `final_answer` tool. After the model calls it, the model must not get another turn. Stop the loop with middleware: + +- **`onToolPhaseComplete`** sees every result of the turn. +- **`onShouldContinue`** returns `false` to stop after the tool results. + +```typescript +import { + chat, + toolDefinition, + toServerSentEventsResponse, + type ChatMiddleware, +} from "@tanstack/ai"; +import { openaiText } from "@tanstack/ai-openai"; +import { z } from "zod"; + +const finalAnswer = toolDefinition({ + name: "final_answer", + description: "Give the final answer to the user", + inputSchema: z.object({ answer: z.string() }), +}).server(async ({ answer }) => ({ answer })); + +/** App-owned policy: stop when every result of a turn comes from a stop tool. */ +function stopAfter(toolNames: Array): ChatMiddleware { + let stop = false; + return { + name: "stop-after-tool", + onToolPhaseComplete(_ctx, info) { + stop = + info.results.length > 0 && + info.results.every((result) => toolNames.includes(result.toolName)); + }, + onShouldContinue() { + return stop ? false : undefined; + }, + }; +} + +export async function POST(request: Request) { + const { messages } = await request.json(); + const stream = chat({ + adapter: openaiText("gpt-6.1-sol"), + messages, + tools: [finalAnswer], + middleware: [stopAfter(["final_answer"])], + }); + return toServerSentEventsResponse(stream); +} +``` + +A turn that also calls other tools goes on, so the model can still use their results. diff --git a/docs/chat/reasoning.md b/docs/chat/reasoning.md new file mode 100644 index 0000000000..e6fafd1dd3 --- /dev/null +++ b/docs/chat/reasoning.md @@ -0,0 +1,171 @@ +--- +title: Reasoning +id: reasoning +order: 6 +description: "Set how hard a reasoning model thinks with one reasoning option on chat(), the same on every provider, with per-model type checks." +keywords: + - tanstack ai + - reasoning + - thinking + - reasoning effort + - extended thinking + - thinking budget +--- + +Every provider has its own switch for reasoning: `reasoning.effort` on OpenAI, `thinking` on Anthropic, `thinkingConfig` on Gemini, `think` on Ollama. The `reasoning` option on `chat()` is one switch for all of them. You pick a level, and the adapter sends what the model needs. + +## Set a level + +```typescript +import { chat } from "@tanstack/ai"; +import { openaiText } from "@tanstack/ai-openai"; + +const stream = chat({ + adapter: openaiText("gpt-5.5"), + messages: [{ role: "user", content: "Plan a database migration." }], + reasoning: "high", +}); +``` + +The levels, from least to most thinking, are `off`, `minimal`, `low`, `medium`, `high`, `xhigh`, and `max`. Leave `reasoning` out to keep the provider's default. + +## The types show the model's levels + +Each model has only some of the levels. The types take only the levels of the model you picked: + +```typescript +import { chat } from "@tanstack/ai"; +import { anthropicText } from "@tanstack/ai-anthropic"; + +chat({ + adapter: anthropicText("claude-opus-4-8"), + messages: [{ role: "user", content: "Hello" }], + reasoning: "max", +}); + +chat({ + adapter: anthropicText("claude-haiku-4-5"), + messages: [{ role: "user", content: "Hello" }], + // @ts-expect-error claude-haiku-4-5 has no max level + reasoning: "max", +}); +``` + +A model that does not reason takes no `reasoning` option at all. + +## Hide the thinking text or set a budget + +Pass an object for more control: + +- `level`: the level. +- `summary`: `false` keeps the thinking text out of the stream. The default is `true`. +- `budgetTokens`: a thinking token budget. Only models that think with a budget take it, such as Claude Haiku 4.5 and Gemini 2.5. + +```typescript +import { chat } from "@tanstack/ai"; +import { anthropicText } from "@tanstack/ai-anthropic"; + +const stream = chat({ + adapter: anthropicText("claude-haiku-4-5"), + messages: [{ role: "user", content: "Check this proof." }], + reasoning: { level: "high", budgetTokens: 8000 }, +}); +``` + +Without `budgetTokens`, a budget model gets a budget for the level: 1024 tokens for `minimal`, 2048 for `low`, 8192 for `medium`, and 16384 for `high`. + +## A level from the user + +A level from a settings menu may not be one the model has. The adapter moves it to the nearest level the model has: up first, then down. For example, `minimal` on `claude-opus-4-8` becomes `low`. + +To show only the levels a model has, read them from the [model catalog](../models/catalog): + +```typescript +import { getModel, supportedReasoningLevels } from "@tanstack/ai-models"; + +const model = getModel("anthropic", "claude-opus-4-8"); +const levels = model ? supportedReasoningLevels(model) : []; +// ["low", "medium", "high", "xhigh", "max"] +``` + +## A model the adapter does not list + +An adapter knows the levels of the models in its own list. A gateway id, a new model, or an id from a catalog is not in that list. For such a model, the adapter sends no reasoning field, and the types take no `reasoning` option. + +Give the adapter the model's reasoning data in its config. The factory takes any model id: + +```typescript +import { chat } from "@tanstack/ai"; +import { createAnthropicChat } from "@tanstack/ai-anthropic"; + +const adapter = createAnthropicChat( + "anthropic/claude-sonnet-4.6", + process.env.AI_GATEWAY_API_KEY ?? "", + { + baseURL: "https://ai-gateway.vercel.sh", + reasoning: { + map: { minimal: null, xhigh: null, max: "max" }, + budget: true, + adaptive: true, + }, + }, +); + +const stream = chat({ + adapter, + messages: [{ role: "user", content: "Plan a database migration." }], + reasoning: "high", +}); +``` + +The `reasoning` config has the same shape on every adapter: + +- `false`: the model does not reason. No reasoning field goes out, and the types take no level. +- `map`: the provider value for each level. `null` means that the model does not have the level. A level with no entry passes as its own name, except `xhigh` and `max`. +- `budget`: `true` when the model thinks with a token budget. +- `adaptive` (Anthropic): `true` sends adaptive thinking with the effort, and `false` sends thinking with a token budget. Without it, the adapter picks from the model id. +- `midConversationEffort` (Anthropic): the level goes into the messages, so a new level keeps the cached start. See [Change the level during a conversation](../adapters/anthropic#change-the-level-during-a-conversation). + +With `reasoning` in the config, the types take every level. The adapter moves a level that the model does not have to the nearest one, as in [A level from the user](#a-level-from-the-user). The config wins over the adapter's own data, also for a model in its list. + +These factories take the `reasoning` config: + +| Package | Factories | +| --- | --- | +| `@tanstack/ai-anthropic` | `anthropicText`, `createAnthropicChat`, `anthropicVertexText` | +| `@tanstack/ai-openai` | `openaiText`, `createOpenaiChat`, `azureOpenaiText` | +| `@tanstack/ai-gemini` | `geminiText`, `createGeminiChat`, also with `vertexai: true` | +| `@tanstack/ai-bedrock` | `createBedrockConverse` | +| `@tanstack/ai-mistral` | `mistralText`, `createMistralText` | +| `@tanstack/ai-cloudflare` | `cloudflareText`, `createCloudflareText` | + +To build the config from a catalog record, see [Build an adapter from a record](../models/catalog#build-an-adapter-from-a-record). + +## Change the level in middleware + +Middleware can set the level for a call, for example to lower it on a retry: + +```typescript +import type { ChatMiddleware } from "@tanstack/ai"; + +const lowEffort: ChatMiddleware = { + name: "low-effort", + onConfig: () => ({ reasoning: { level: "low", summary: true } }), +}; +``` + +## What each provider receives + +| Provider | Sent as | +| --- | --- | +| OpenAI, Grok, Bedrock Responses | `reasoning.effort`, with a summary | +| Anthropic | adaptive thinking with the effort, or a thinking budget | +| Gemini | `thinkingConfig.thinkingLevel` or `thinkingConfig.thinkingBudget` | +| OpenRouter, Vercel AI Gateway | the gateway's `reasoning` object | +| Groq, LLM Gateway, Lovable, Cloudflare, Bedrock chat | `reasoning_effort` | +| Mistral | `reasoning_effort` or `prompt_mode` | +| Ollama | `think` | +| Claude Code, Codex | the CLI effort setting | +| `openaiCompatible` | the model's `thinkingFormat`, see [OpenAI-Compatible](../adapters/openai-compatible#reasoning-models) | + +The thinking text streams back as thinking parts. See [Thinking & Reasoning](./thinking-content) to show it in your UI. diff --git a/docs/chat/stream-events.md b/docs/chat/stream-events.md index f34d42c5ce..de4bf9b72f 100644 --- a/docs/chat/stream-events.md +++ b/docs/chat/stream-events.md @@ -39,12 +39,14 @@ Later: On `RUN_FINISHED`, in-process `chat()` still uses TanStack `TokenUsage` (`promptTokens`). The SSE and HTTP wires use the spec `usage` array (`inputTokens`). Read `finishReason` from `metadata.tanstack.finishReason`. Custom servers: see [Event metadata](../protocol/metadata). +`usage.promptTokens` is the full input, cached tokens included. Cache reads are on `usage.promptTokensDetails.cachedTokens`, and cache writes are on `usage.promptTokensDetails.cacheWriteTokens`. On the wire, they are `cachedInputTokens` and `cacheWriteInputTokens`. See [Prompt Caching](../advanced/prompt-caching#track-cached-tokens-and-cost). + ```typescript import { chat } from "@tanstack/ai"; import { openaiText } from "@tanstack/ai-openai"; const stream = chat({ - adapter: openaiText("gpt-5.6"), + adapter: openaiText("gpt-5.5"), messages: [{ role: "user", content: "Hello!" }], }); @@ -59,6 +61,34 @@ for await (const chunk of stream) { } ``` +## Read the provider response identity + +You can correlate a finished reply with the provider's generation. Read its identity from `metadata.tanstack` on the `chat()` stream: + +```typescript +import { chat } from '@tanstack/ai' +import { openaiText } from '@tanstack/ai-openai' + +for await (const chunk of chat({ + adapter: openaiText('gpt-5.5'), + messages: [{ role: 'user', content: 'Hello!' }], +})) { + if (chunk.type !== 'RUN_FINISHED') continue + const metadata = chunk.metadata?.tanstack + console.log(metadata?.responseId, metadata?.model, metadata?.source) +} +``` + +- `responseId`: the provider's generation ID, when supplied. +- `model`: the provider-reported response model, when supplied. +- `source`: the provider, wire API, and requested logical model for that call. + +`source` also identifies provider error events. Assistant messages retain these fields through client conversion and persistence. An Azure deployment can report a different model from `source.model`. + +Direct adapter iterators expose terminal `responseId` and `model` fields at the top level. `chat()` and the wire use `metadata.tanstack`. When the provider supplies no generation ID, the adapter omits `responseId`. An HTTP request ID does not fill that field. + +A direct adapter's `structuredOutput()` result can also supply `responseId` and `model`. The structured-output fallback forwards those fields to its terminal event. The parsed object returned by `chat()` with `outputSchema` stays the schema's data. + ## Threads and runs Two ids frame every stream. They come from the AG-UI protocol, not from a storage layer. @@ -118,7 +148,7 @@ const weatherTool = toolDefinition({ }); const stream = chat({ - adapter: openaiText("gpt-5.6"), + adapter: openaiText("gpt-5.5"), messages: [{ role: "user", content: "What is the weather in Paris?" }], tools: [weatherTool], }); diff --git a/docs/chat/subagents.md b/docs/chat/subagents.md index de2f4698a3..1f44fc3f9b 100644 --- a/docs/chat/subagents.md +++ b/docs/chat/subagents.md @@ -90,6 +90,8 @@ The router can return: - `{ names, order }` - `{ steps }` +Anywhere a name goes, `{ name, input }` can go too, for example `{ name: 'pricer', input: { vendor: 'Acme' } }`. An agent with `inputSchema` needs it. + `subagents.order` is the default for an array. `parallel` starts the names together. `sequence` runs them one after another, and each later child reads the earlier child text. Omit `order` to get `parallel`. `subagentRoute` asks Jev for the order. Jev picks `parallel` when no agent must read text from another agent. The same topic is not a reason to wait. Jev picks `sequence` only when a later agent must read the earlier text, such as research notes and then a draft. @@ -154,7 +156,70 @@ const stream = chat({ }) ``` -**Without a router.** The library adds one synthetic server tool per agent. The main model calls that tool. The public stream still emits `SUBAGENT_STARTED` / `SUBAGENT_FINISHED` (or `SUBAGENT_ERROR`) and nested parts. The UI does not treat spawn as a normal tool card. The child's events stream while the tool runs, and the child's text becomes the tool result. The child reads the conversation as it is at that tool call. +**Give a routed agent its input.** An agent with `inputSchema` gets its input from the router. `route.needsInput(result)` lists the picked agents that need one, each with its schema. Make each input. Then pass them to `route.pick(result, { inputs })`. Here the model writes each input, with the schema as `outputSchema`: + +```ts +import { chat, decide, defineAgent, subagentRoute } from '@tanstack/ai' +import { openaiText } from '@tanstack/ai-openai' +import { typesafeDecider } from '@tanstack/ai-typesafe' +import { z } from 'zod' + +const writer = defineAgent({ + name: 'writer', + description: 'Writes the post', + run: async function* () {}, +}) +const pricer = defineAgent({ + name: 'pricer', + description: 'Compares the plans and prices of one vendor', + inputSchema: z.object({ vendor: z.string() }), + run: (ctx) => + ctx.chat({ + adapter: openaiText('gpt-5.6'), + messages: [{ role: 'user', content: `Compare the plans of ${ctx.input.vendor}` }], + stream: false, + }), +}) +const messages = [ + { role: 'user' as const, content: 'What does Acme cost? Write a post about it.' }, +] + +const stream = chat({ + adapter: openaiText('gpt-5.6'), + messages, + subagents: { + agents: [pricer, writer], + router: async ({ messages: turnMessages, agents }) => { + const route = subagentRoute(agents, { then: ['writer'] }) + const result = await decide({ + adapter: typesafeDecider('jev-latest'), + state: turnMessages, + questions: route.questions, + }) + const inputs = Object.fromEntries( + await Promise.all( + route.needsInput(result).map(async ({ name, inputSchema }) => [ + name, + await chat({ + adapter: openaiText('gpt-5.6'), + messages: turnMessages, + outputSchema: inputSchema, + }), + ]), + ), + ) + return route.pick(result, { inputs }) + }, + }, +}) +``` + +- `inputs` is typed. Each key is the name of an agent with `inputSchema`, and the value has the type of that schema. +- If a picked agent with a schema has no input, `pick` throws `Agent "pricer" needs input. Pass it in inputs.` +- `chat()` checks each input against the schema before the child starts. If an input does not match, `chat()` throws, and the child does not start. +- `run` reads the input as `ctx.input`. A resumed child gets the same input. + +**Without a router.** The library adds one synthetic server tool per agent. The main model calls that tool. The public stream still emits `SUBAGENT_STARTED` / `SUBAGENT_FINISHED` (or `SUBAGENT_ERROR`) and nested parts. The UI does not treat spawn as a normal tool card. The child's events stream while the tool runs, and the child's text becomes the tool result. The child reads the conversation as it is at that tool call. For one tool that can start every agent, see [Give the model one subagent tool](#give-the-model-one-subagent-tool). ## Let the model write the brief @@ -202,10 +267,168 @@ const stream = chat({ - If the input does not match the schema, the model gets a tool error and can call the tool again. The child does not start. - `ctx.messages` still holds the parent conversation. Add parts of it when the child needs more context. - A resumed child gets the same `ctx.input` as the first run. -- `inputSchema` needs tool mode. If `subagents.router` is set and an agent has `inputSchema`, `chat()` throws. +- With `subagents.router`, the router gives the input. See [Route, or let the model pick](#route-or-let-the-model-pick). To show the brief in the UI, see [Show the brief on a card](#show-the-brief-on-a-card). +## Give the model one subagent tool + +Without a router, the model gets one tool per agent, and each call starts a new child. The model cannot send a follow-up to a child that already answered. Set `tool: 'single'`. The model then gets one tool, `subagent`, and names the agent in each call. A call can also continue a child from an earlier turn. + +```ts +import { + chat, + chatParamsFromRequest, + defineAgent, + toServerSentEventsResponse, +} from '@tanstack/ai' +import { openaiText } from '@tanstack/ai-openai' +import { memoryPersistence, withPersistence } from '@tanstack/ai-persistence' +import { z } from 'zod' + +const researcher = defineAgent({ + name: 'researcher', + description: 'Researches one question and returns sourced findings', + inputSchema: z.object({ question: z.string() }), + run: (ctx) => + ctx.chat({ + adapter: openaiText('gpt-6.1-sol'), + messages: [{ role: 'user', content: ctx.input.question }], + }), +}) + +const writer = defineAgent({ + name: 'writer', + description: 'Writes and edits drafts', + run: (ctx) => ctx.chat({ adapter: openaiText('gpt-6.1-sol') }), +}) + +// A continued child loads from this store. Use a durable backend in production. +const persistence = memoryPersistence() + +export async function POST(request: Request) { + const params = await chatParamsFromRequest(request) + const stream = chat({ + adapter: openaiText('gpt-6.1-sol'), + messages: params.messages, + threadId: params.threadId, + runId: params.runId, + middleware: [withPersistence(persistence)], + subagents: { agents: [researcher, writer], tool: 'single' }, + }) + return toServerSentEventsResponse(stream) +} +``` + +The model fills these fields of the tool input: + +- `agent`: the agent name. Each call needs it. +- `input`: the input of an agent with `inputSchema`. That agent needs it. +- `prompt`: the task as text, for an agent without `inputSchema`. The child gets it as the last user message. +- `sessionId`: the `subagentRunId` of an earlier result. The call continues that child. +- `background`: `true` starts the child and returns at once. Only a harness can do this. See [Run agents from a harness](../harness/subagents#give-the-model-one-subagent-tool). + +On the client, each child is a `type: 'subagent'` part, the same as with one tool per agent. `part.subagent.name` is the agent name. + +### Continue a child + +Each result gives the model the `subagentRunId` of the child. The model uses it in a later turn: + +1. The model calls `{ "agent": "writer", "prompt": "Draft a post about tides" }`. +2. It gets `{ "subagentRunId": "subagent-1767225600000-x7k2m9q", "result": "Tides rise twice a day..." }`. +3. In the next turn, it calls `{ "agent": "writer", "sessionId": "subagent-1767225600000-x7k2m9q", "prompt": "Make it shorter" }`. + +The writer gets its stored transcript, then `Make it shorter` as the last user message. + +- `sessionId` needs `withPersistence` on the parent `chat()`. The stored child comes from that store. +- A thread can continue only its own children. An id from another thread gives the tool error `Unknown sessionId "..."`. +- The call must name the agent that the child ran under. Another agent gives the tool error `Session "..." belongs to agent "writer", not "researcher".` +- The stream uses the same `subagentRunId`, so the client shows the new work on the card with that id. + +### Wrong calls + +A wrong call returns to the model as a tool error, and the run continues. For example: + +- `Agent "researcher" takes input, not prompt.` +- `Agent "researcher" needs input.` +- `sessionId needs a persistence store.` +- `background needs a harness host.` + +If an agent or a tool already has the name `subagent`, `chat()` throws before the run starts. + +## Return a value, or call any activity + +A child does not have to be a chat. It can make an image, or return a plain value that the parent model reads as the tool result. Call the activity on `ctx` and return what it gives you: + +```ts group=subagent-values +import { chat, defineAgent } from '@tanstack/ai' +import { openaiImage, openaiText } from '@tanstack/ai-openai' +import { z } from 'zod' + +const heroImage = defineAgent({ + name: 'heroImage', + description: 'Makes a hero image for a post', + produces: 'image', + inputSchema: z.object({ prompt: z.string() }), + run: (ctx) => + ctx.generateImage({ + adapter: openaiImage('gpt-image-2'), + prompt: `${ctx.input.prompt}. Brand colors: blue and white.`, + size: '1536x1024', + }), +}) + +const stream = chat({ + adapter: openaiText('gpt-5.6'), + messages: [{ role: 'user', content: 'Make a hero image about squids' }], + subagents: { agents: [heroImage] }, +}) +``` + +- `ctx.chat`, `ctx.generateImage`, `ctx.generateVideo`, `ctx.generateSpeech`, and the other activities on `ctx` take the same options as the plain functions. They fill in the thread id, a run id, and the abort signal. +- `ctx.chat` also uses the parent conversation (`ctx.messages`) when you do not pass `messages`. +- `run` can return a promise of any value. The value arrives on `SUBAGENT_FINISHED.result`, and the parent model gets it as the tool result. +- A string result also streams as the child's text, so the card shows it. +- A very long string in the result (for example a base64 image) reaches the parent model as a short note, `[omitted 5000 characters]`. The full value stays on `SUBAGENT_FINISHED.result`. +- `produces` says what the agent makes (`'image'`, `'text'`, and so on). It does not change how the agent runs. + +When you call the plain `chat()` instead, spread `ctx.forward` to pass the child's ids and abort controller in one line: + +```ts group=subagent-values +const researcher = defineAgent({ + name: 'researcher', + description: 'Looks up facts', + run: (ctx) => + chat({ + adapter: openaiText('gpt-5.6'), + messages: ctx.messages, + ...ctx.forward, + }), +}) +``` + +## Limit the tree + +A child can have children of its own, and a model can call a child many times. Set `limits` so one request cannot start an unbounded tree: + +```ts group=subagent-values +const limited = chat({ + adapter: openaiText('gpt-5.6'), + messages: [{ role: 'user', content: 'Research squids in depth' }], + subagents: { + agents: [researcher], + limits: { maxDepth: 2, maxConcurrent: 3, maxCalls: 12, timeoutMs: 120_000 }, + }, +}) +``` + +- `maxDepth`: how deep the tree may grow. The first chat is depth 0. +- `maxCalls`: how many children the whole tree may start. +- `maxConcurrent`: how many children one run may have running at once. +- `timeoutMs`: how long one child may run. A child never gets more time than its parent has left. + +The whole tree shares one budget. A child that calls `ctx.chat({ subagents })` passes it on, so a child cannot reset it. A refused start reaches the model as a tool error, for example `subagent limit reached (maxCalls 12)`. + ## Strategy - `exclusive` (default): the chosen child owns the turn. Main does not answer after it. @@ -314,6 +537,8 @@ A child is its own `chat()` call. Put the child's middleware in that call. The p Inside a child, the middleware context has `subagentRunId`. Use it to tell a child run from a top-level run, and to link a child trace to its card. +When the child calls `ctx.chat`, the context also has `subagentName`, the name of the agent. For a nested child, `parentSubagentRunId` is the id of the child that started it. + On a routed turn where a child runs and main does not, the parent's middleware does not run. Only `withPersistence` records that turn. A turn where main runs, including a handoff, runs the parent's middleware as usual. ## Persistence @@ -580,7 +805,9 @@ function Briefs({ message }: { message: UIMessage }) { } ``` -A child that a router started has no `parentToolCallId`, so it has no brief. +A child that a router started has no `parentToolCallId`, so `briefOf` finds no brief for it. + +With `tool: 'single'`, that tool call is a call of the `subagent` tool. The brief is in the `input` field or the `prompt` field of its input. ### Style one child's parts diff --git a/docs/chat/thinking-content.md b/docs/chat/thinking-content.md index b5af832a01..4f07c021da 100644 --- a/docs/chat/thinking-content.md +++ b/docs/chat/thinking-content.md @@ -16,7 +16,7 @@ keywords: Some models expose their internal reasoning as "thinking" content -- Claude with extended thinking, OpenAI o-series models with reasoning, and others. TanStack AI captures this as `ThinkingPart` in messages, streamed to your UI in real-time alongside text and tool calls. -Unsigned thinking stays in the UI. Signed thinking is a `ThinkingPart` with a `signature`. Anthropic extended thinking uses this. The next request sends signed thinking back in the same order as the original response, including around provider-executed tools. The next-turn body puts that signature on spec `encryptedValue` on the `role: "reasoning"` fan-out. Stream events use `REASONING_ENCRYPTED_VALUE`. +Unsigned thinking stays in the UI. Signed thinking is a `ThinkingPart` with a `signature`. Anthropic extended thinking uses this. The next request sends signed thinking back in the same order as the original response, including around tool calls and text. The next-turn body puts that signature on spec `encryptedValue` on the `role: "reasoning"` fan-out. Stream events use `REASONING_ENCRYPTED_VALUE`. ## How It Works @@ -38,11 +38,7 @@ Claude can also send a redacted thinking block. It is encrypted, so it has no te ## Enabling Thinking -How you enable thinking depends on the provider. - -### Anthropic (Extended Thinking) - -Pass the `thinking` option in `modelOptions` with `type: "enabled"` and a `budget_tokens` (minimum 1024). Keep `budget_tokens` below `modelOptions.max_tokens` so there is room for the visible response in addition to the thinking budget: +Turn thinking on with `reasoning` on `chat()`. It works the same on every provider: ```typescript import { chat, toServerSentEventsResponse } from "@tanstack/ai"; @@ -53,60 +49,13 @@ export async function POST(request: Request) { const stream = chat({ adapter: anthropicText("claude-sonnet-4-6"), messages, - modelOptions: { - max_tokens: 32000, - // budget_tokens must be at least 1024 and below max_tokens - thinking: { type: "enabled", budget_tokens: 10000 }, - }, - }); - return toServerSentEventsResponse(stream); -} -``` - -### OpenAI (Reasoning Models) - -OpenAI o-series models (o1, o3, o3-mini, o3-pro) perform reasoning automatically. You can control the depth with the `reasoning` option: - -```typescript -import { chat, toServerSentEventsResponse } from "@tanstack/ai"; -import { openaiText } from "@tanstack/ai-openai"; - -export async function POST(request: Request) { - const { messages } = await request.json(); - const stream = chat({ - adapter: openaiText("o3-mini"), - messages, - modelOptions: { - reasoning: { - effort: "medium", // 'none' | 'minimal' | 'low' | 'medium' | 'high' - summary: "auto", // 'auto' | 'detailed' - }, - }, + reasoning: "high", }); return toServerSentEventsResponse(stream); } ``` -When `reasoning.summary` is set, the adapter streams reasoning summary text as thinking content. Without it, reasoning tokens are still used internally but may not be surfaced depending on the model. - -GPT-5 and later models also support reasoning. Their `reasoning.effort` accepts `"none" | "minimal" | "low" | "medium" | "high"`, and reasoning activates on any non-`none` value: - -```typescript -import { chat, toServerSentEventsResponse } from "@tanstack/ai"; -import { openaiText } from "@tanstack/ai-openai"; - -export async function POST(request: Request) { - const { messages } = await request.json(); - const stream = chat({ - adapter: openaiText("gpt-5.5"), - messages, - modelOptions: { - reasoning: { effort: "high" }, - }, - }); - return toServerSentEventsResponse(stream); -} -``` +The thinking text streams back as `ThinkingPart`s. Some models reason without returning any text; with those, the part stays empty. See [Reasoning](./reasoning) for the levels, token budgets, and hiding the thinking text. ## Rendering in React @@ -150,6 +99,12 @@ The typical streaming order is: `STEP_STARTED` and `STEP_FINISHED` only carry `stepName`. They do not carry thinking text. +Claude can think more than once in one answer: thinking, a tool call, more thinking, then text. Each thinking block streams as its own reasoning message, with its own `REASONING_MESSAGE_END` and `REASONING_END`. Then: + +- `message.parts` has the parts in the order Claude sent them. +- The stored messages keep that order, and the next request sends the blocks back to Claude in it. +- A message that you build yourself goes back in the default order: thinking, then text, then tool calls. + If you use `useChat` from `@tanstack/ai-react` (or the Solid/Vue/Svelte equivalents), your `messages` array updates with both thinking and text parts as they arrive. ## Next Steps diff --git a/docs/community-adapters/guide.md b/docs/community-adapters/guide.md index e70d6f9d48..1c355075bb 100644 --- a/docs/community-adapters/guide.md +++ b/docs/community-adapters/guide.md @@ -92,13 +92,11 @@ Example: ```typescript ignore export type OpenAIChatModelProviderOptionsByName = { [GPT5_2.name]: OpenAIBaseOptions & - OpenAIReasoningOptions & OpenAIStructuredOutputOptions & OpenAIToolsOptions & OpenAIStreamingOptions & OpenAIMetadataOptions [GPT5_2_CHAT.name]: OpenAIBaseOptions & - OpenAIReasoningOptions & OpenAIStructuredOutputOptions & OpenAIToolsOptions & OpenAIStreamingOptions & @@ -109,6 +107,8 @@ export type OpenAIChatModelProviderOptionsByName = { ``` This ensures strict type safety and feature correctness at compile time. +Reasoning is not a provider option: `chat({ reasoning })` owns it. Declare each model's reasoning levels on the adapter instead. See [Add reasoning to your adapter](../advanced/extend-adapter#add-reasoning-to-your-adapter). + ### 5. Define supported input modalities Models typically support different input modalities (e.g. text, images, audio). These must be defined per model to prevent invalid usage. @@ -140,9 +140,9 @@ export interface OpenAIBaseOptions { // Feature fragments that can be stitched per-model /** - * Reasoning options for models + * Tool options for models */ -export interface OpenAIReasoningOptions { +export interface OpenAIToolsOptions { //... } @@ -160,7 +160,6 @@ Models can then opt into only the features they support: ```typescript ignore export type OpenAIChatModelProviderOptionsByName = { [GPT5_2.name]: OpenAIBaseOptions & - OpenAIReasoningOptions & OpenAIStructuredOutputOptions & OpenAIToolsOptions & OpenAIStreamingOptions & @@ -190,6 +189,44 @@ Adapters are implemented per capability, so only implement what your provider su Refer to the [OpenAI adapter](https://github.com/TanStack/ai/blob/main/packages/ai-openai/src/adapters/text.ts) for a complete, end-to-end implementation example. +### Preserve history and response identity + +A user can switch to your adapter with history from another provider. Declare the text adapter's identity so replay can detect that history: + +- `name`: the adapter name. +- `provider`: an optional provider identity. It defaults to `name`. +- `api`: an optional wire API identity. It defaults to the adapter kind. + +Use the actual API selected for each call. Adapters with several APIs must apply replay after that selection. The requested logical model belongs in `source.model`. A mapped deployment does not change it. + +Core cleans failed and orphan history before the adapter call. This also removes results that belong to dropped failed call batches. Your adapter applies source-sensitive replay and the target API's tool-ID rules. + +The internal adapter helper is available through the existing adapter subpath: + +```typescript +import { hashToolCallId, transformMessagesForReplay } from '@tanstack/ai/adapter-internals' +import type { ModelMessage } from '@tanstack/ai' + +export function prepareMessages(messages: ReadonlyArray, model: string) { + return transformMessagesForReplay(messages, { + provider: 'example-provider', + api: 'example-messages', + model, + }, (id, { attempt }) => { + const normalized = id.replace(/[^a-zA-Z0-9_-]/g, '_').slice(0, 64) + return attempt === 0 + ? normalized + : normalized.slice(0, 40) + '_' + hashToolCallId(id) + '_' + attempt + }).messages +} +``` + +That example API permits letters, digits, underscores, and hyphens in IDs. Its maximum length is 64. The `attempt` value handles collisions. Use your API's rules. + +Replay operates on a request copy. Keep saved messages and raw tool arguments intact. Preserve thinking and tool block order when your provider supports them. + +Emit the provider's genuine generation ID as terminal `responseId`, and its reported model as `model`. Omit either field when the provider does not supply it. `chat()` moves these fields into `metadata.tanstack`. + ### 8. Publish and submit a PR Once your adapter is complete: @@ -228,4 +265,4 @@ This includes: - Addressing issues and feedback from users - Updating documentation when features change -If you add new features or breaking changes, open a follow-up PR to keep the docs in sync. \ No newline at end of file +If you add new features or breaking changes, open a follow-up PR to keep the docs in sync. diff --git a/docs/config.json b/docs/config.json index a2d731cf03..1dc8ce868d 100644 --- a/docs/config.json +++ b/docs/config.json @@ -78,7 +78,8 @@ { "label": "Persisted subagents", "to": "tutorials/subagents-persisted", - "addedAt": "2026-09-22" + "addedAt": "2026-09-22", + "updatedAt": "2026-09-30" }, { "label": "Build an MCP Server", @@ -159,25 +160,31 @@ "label": "Stream Events", "to": "chat/stream-events", "addedAt": "2026-08-26", - "updatedAt": "2026-09-27" + "updatedAt": "2026-10-04" }, { "label": "Subagents", "to": "chat/subagents", "addedAt": "2026-09-21", - "updatedAt": "2026-09-25" + "updatedAt": "2026-10-07" }, { "label": "Agentic Cycle", "to": "chat/agentic-cycle", "addedAt": "2026-04-15", - "updatedAt": "2026-07-21" + "updatedAt": "2026-09-30" }, { "label": "Thinking & Reasoning", "to": "chat/thinking-content", "addedAt": "2026-04-15", - "updatedAt": "2026-09-30" + "updatedAt": "2026-10-02" + }, + { + "label": "Reasoning", + "to": "chat/reasoning", + "addedAt": "2026-09-30", + "updatedAt": "2026-10-06" }, { "label": "Message Queue", @@ -236,7 +243,7 @@ "label": "Tools", "to": "tools/tools", "addedAt": "2026-04-15", - "updatedAt": "2026-09-30" + "updatedAt": "2026-10-07" }, { "label": "Server Tools", @@ -301,7 +308,7 @@ "label": "Middleware", "to": "advanced/middleware", "addedAt": "2026-04-15", - "updatedAt": "2026-09-30" + "updatedAt": "2026-10-06" }, { "label": "Built-in Middleware", @@ -325,7 +332,18 @@ "label": "Compaction", "to": "advanced/compaction", "addedAt": "2026-08-24", - "updatedAt": "2026-08-27" + "updatedAt": "2026-10-07" + }, + { + "label": "Prompt Caching", + "to": "advanced/prompt-caching", + "addedAt": "2026-10-02", + "updatedAt": "2026-10-06" + }, + { + "label": "Mid-Conversation Changes", + "to": "advanced/mid-conversation-changes", + "addedAt": "2026-10-02" } ], "tab": "guides" @@ -451,7 +469,7 @@ "label": "Tool Approval", "to": "interrupts/tool-approval", "addedAt": "2026-08-04", - "updatedAt": "2026-08-14" + "updatedAt": "2026-10-05" }, { "label": "Multiple Interrupts", @@ -491,13 +509,13 @@ "label": "MCP Server Tools", "to": "tools/mcp", "addedAt": "2026-06-05", - "updatedAt": "2026-09-25" + "updatedAt": "2026-10-07" }, { "label": "Managed MCP with chat()", "to": "tools/mcp-managed", "addedAt": "2026-06-05", - "updatedAt": "2026-09-25" + "updatedAt": "2026-10-03" }, { "label": "Manual MCP", @@ -514,7 +532,7 @@ "label": "MCP Client Input", "to": "tools/mcp-input", "addedAt": "2026-09-22", - "updatedAt": "2026-09-25" + "updatedAt": "2026-10-04" }, { "label": "MCP Apps", @@ -579,7 +597,7 @@ "label": "Portable Agent Skills", "to": "skills/agent-skills", "addedAt": "2026-08-22", - "updatedAt": "2026-08-23" + "updatedAt": "2026-10-01" }, { "label": "Skill Sources", @@ -660,7 +678,7 @@ "label": "Build Your Own Adapter", "to": "persistence/build-your-own-adapter", "addedAt": "2026-08-04", - "updatedAt": "2026-09-29" + "updatedAt": "2026-09-30" }, { "label": "Migrations", @@ -703,13 +721,13 @@ "label": "Build a Sandbox Adapter", "to": "persistence/build-a-sandbox-adapter", "addedAt": "2026-08-04", - "updatedAt": "2026-08-14" + "updatedAt": "2026-10-05" }, { "label": "Store Reference", "to": "persistence/store-reference", "addedAt": "2026-08-04", - "updatedAt": "2026-09-24" + "updatedAt": "2026-10-07" }, { "label": "How Persistence Works", @@ -867,6 +885,213 @@ ], "tab": "guides" }, + { + "label": "Harness", + "tab": "guides", + "children": [ + { + "label": "Build your first harness", + "to": "harness/overview", + "addedAt": "2026-09-26", + "updatedAt": "2026-10-05" + }, + { + "label": "Connect clients", + "to": "harness/connect", + "addedAt": "2026-09-26", + "updatedAt": "2026-10-07" + }, + { + "label": "Send and show media", + "to": "harness/media", + "addedAt": "2026-09-29" + }, + { + "label": "Run in the terminal", + "to": "harness/cli", + "addedAt": "2026-09-26", + "updatedAt": "2026-10-06" + }, + { + "label": "Durable sessions", + "to": "harness/durable-sessions", + "addedAt": "2026-09-26", + "updatedAt": "2026-10-07" + }, + { + "label": "List, rename, and fork sessions", + "to": "harness/sessions", + "addedAt": "2026-10-06", + "updatedAt": "2026-10-07" + }, + { + "label": "Send inputs safely", + "to": "harness/inputs", + "addedAt": "2026-09-30", + "updatedAt": "2026-10-07" + }, + { + "label": "Run side effects once", + "to": "harness/durable-tools", + "addedAt": "2026-09-30", + "updatedAt": "2026-10-01" + }, + { + "label": "Build on the session log", + "to": "harness/session-log", + "addedAt": "2026-09-30", + "updatedAt": "2026-10-03" + }, + { + "label": "Control how a turn ends", + "to": "harness/turn-control", + "addedAt": "2026-10-01", + "updatedAt": "2026-10-07" + }, + { + "label": "Store settings per thread", + "to": "harness/thread-settings", + "addedAt": "2026-10-05" + }, + { + "label": "Compact a long session", + "to": "harness/compaction", + "addedAt": "2026-10-03", + "updatedAt": "2026-10-07" + }, + { + "label": "Fork and reset a thread", + "to": "harness/fork-and-reset", + "addedAt": "2026-10-05", + "updatedAt": "2026-10-07" + }, + { + "label": "Share one log between sessions", + "to": "harness/shared-logs", + "addedAt": "2026-10-01" + }, + { + "label": "Write a plugin", + "to": "harness/plugins", + "addedAt": "2026-09-26", + "updatedAt": "2026-10-05" + }, + { + "label": "Build a coding agent", + "to": "harness/coding-agent", + "addedAt": "2026-09-26", + "updatedAt": "2026-10-06" + }, + { + "label": "Give your agent coding tools", + "to": "harness/coding-tools", + "addedAt": "2026-10-06", + "updatedAt": "2026-10-07" + }, + { + "label": "Sandbox and extend the coding tools", + "to": "harness/coding-tools-backends", + "addedAt": "2026-10-06" + }, + { + "label": "Ask before risky tool calls", + "to": "harness/permissions", + "addedAt": "2026-10-06", + "updatedAt": "2026-10-07" + }, + { + "label": "Undo a turn", + "to": "harness/snapshots", + "addedAt": "2026-10-06", + "updatedAt": "2026-10-07" + }, + { + "label": "Switch between agent profiles", + "to": "harness/agents", + "addedAt": "2026-10-06", + "updatedAt": "2026-10-07" + }, + { + "label": "Connect model providers", + "to": "harness/provider-keys", + "addedAt": "2026-09-29", + "updatedAt": "2026-10-04" + }, + { + "label": "Auth and connectors", + "to": "harness/auth", + "addedAt": "2026-09-26", + "updatedAt": "2026-10-06" + }, + { + "label": "Share a thread between people", + "to": "harness/shared-threads", + "addedAt": "2026-10-05" + }, + { + "label": "Track usage and cost", + "to": "harness/usage", + "addedAt": "2026-10-05", + "updatedAt": "2026-10-06" + }, + { + "label": "Use MCP servers", + "to": "harness/mcp", + "addedAt": "2026-09-26", + "updatedAt": "2026-10-07" + }, + { + "label": "Use from any MCP client", + "to": "harness/mcp-server", + "addedAt": "2026-09-28", + "updatedAt": "2026-10-06" + }, + { + "label": "Code mode in a harness", + "to": "harness/code-mode", + "addedAt": "2026-09-26", + "updatedAt": "2026-10-06" + }, + { + "label": "Add skills", + "to": "harness/skills", + "addedAt": "2026-10-01" + }, + { + "label": "Delegate to coding agents", + "to": "harness/coding-agents", + "addedAt": "2026-09-28" + }, + { + "label": "Work until a goal is met", + "to": "harness/goal", + "addedAt": "2026-09-28" + }, + { + "label": "Build your own UI", + "to": "harness/custom-ui", + "addedAt": "2026-09-28", + "updatedAt": "2026-10-07" + }, + { + "label": "Run agents from a harness", + "to": "harness/subagents", + "addedAt": "2026-09-26", + "updatedAt": "2026-10-07" + }, + { + "label": "Deploy a harness", + "to": "harness/deploy", + "addedAt": "2026-09-26" + }, + { + "label": "Self-host the dashboard", + "to": "harness/dashboard", + "addedAt": "2026-09-26", + "updatedAt": "2026-10-07" + } + ] + }, { "label": "Sandboxes", "tab": "guides", @@ -910,7 +1135,7 @@ "label": "Tools", "to": "sandbox/tools", "addedAt": "2026-06-29", - "updatedAt": "2026-08-12" + "updatedAt": "2026-10-06" }, { "label": "Policy", @@ -1028,6 +1253,11 @@ "to": "sandbox/cloudflare", "addedAt": "2026-06-29", "updatedAt": "2026-08-20" + }, + { + "label": "Build a Provider", + "to": "sandbox/build-a-provider", + "addedAt": "2026-10-05" } ] }, @@ -1105,7 +1335,7 @@ "label": "Runtime Adapter Switching", "to": "advanced/runtime-adapter-switching", "addedAt": "2026-04-15", - "updatedAt": "2026-08-14" + "updatedAt": "2026-10-05" }, { "label": "Tree-Shaking", @@ -1116,13 +1346,25 @@ { "label": "Extend Adapter", "to": "advanced/extend-adapter", - "addedAt": "2026-04-15" + "addedAt": "2026-04-15", + "updatedAt": "2026-10-02" + }, + { + "label": "Model Catalog", + "to": "models/catalog", + "addedAt": "2026-09-30", + "updatedAt": "2026-10-06" }, { "label": "Typed Pre-Configured Options", "to": "advanced/typed-options", "addedAt": "2026-05-25", - "updatedAt": "2026-09-14" + "updatedAt": "2026-09-30" + }, + { + "label": "Test with a Fake Model", + "to": "advanced/testing", + "addedAt": "2026-09-30" }, { "label": "Approval Flow Processing", @@ -1140,7 +1382,7 @@ "label": "OpenAI", "to": "adapters/openai", "addedAt": "2026-04-15", - "updatedAt": "2026-09-30" + "updatedAt": "2026-10-06" }, { "label": "Anthropic", @@ -1164,7 +1406,7 @@ "label": "Ollama", "to": "adapters/ollama", "addedAt": "2026-04-15", - "updatedAt": "2026-09-03" + "updatedAt": "2026-10-06" }, { "label": "Grok (xAI)", @@ -1176,13 +1418,13 @@ "label": "Groq", "to": "adapters/groq", "addedAt": "2026-04-15", - "updatedAt": "2026-08-07" + "updatedAt": "2026-09-30" }, { "label": "Mistral", "to": "adapters/mistral", "addedAt": "2026-06-30", - "updatedAt": "2026-09-03" + "updatedAt": "2026-10-04" }, { "label": "Cohere", @@ -1251,60 +1493,61 @@ "label": "OpenAI-Compatible", "to": "adapters/openai-compatible", "addedAt": "2026-06-01", - "updatedAt": "2026-09-27" + "updatedAt": "2026-10-04" }, { "label": "LLM Gateway", "to": "adapters/llmgateway", - "addedAt": "2026-07-29" + "addedAt": "2026-07-29", + "updatedAt": "2026-09-30" }, { "label": "Cloudflare", "to": "adapters/cloudflare", "addedAt": "2026-09-03", - "updatedAt": "2026-09-18" + "updatedAt": "2026-10-02" }, { "label": "Claude Code", "to": "adapters/claude-code", "addedAt": "2026-06-12", - "updatedAt": "2026-09-25" + "updatedAt": "2026-10-07" }, { "label": "Codex", "to": "adapters/codex", "addedAt": "2026-06-12", - "updatedAt": "2026-08-18" + "updatedAt": "2026-10-07" }, { "label": "OpenCode", "to": "adapters/opencode", "addedAt": "2026-06-12", - "updatedAt": "2026-08-17" + "updatedAt": "2026-10-07" }, { "label": "Grok Build", "to": "adapters/grok-build", "addedAt": "2026-06-29", - "updatedAt": "2026-08-18" + "updatedAt": "2026-10-07" }, { "label": "ACP-Compatible", "to": "adapters/acp-compatible", "addedAt": "2026-06-30", - "updatedAt": "2026-08-18" + "updatedAt": "2026-10-07" }, { "label": "Amazon Bedrock", "to": "adapters/bedrock", "addedAt": "2026-06-25", - "updatedAt": "2026-09-25" + "updatedAt": "2026-10-07" }, { "label": "BytePlus", "to": "adapters/byteplus", "addedAt": "2026-08-04", - "updatedAt": "2026-09-16" + "updatedAt": "2026-09-30" } ], "tab": "adapters" @@ -1317,7 +1560,7 @@ "label": "Community Adapters Guide", "to": "community-adapters/guide", "addedAt": "2026-04-15", - "updatedAt": "2026-08-21" + "updatedAt": "2026-10-04" }, { "label": "Decart", @@ -1384,6 +1627,11 @@ "to": "migration/sampling-options-to-model-options", "addedAt": "2026-06-03" }, + { + "label": "Reasoning option", + "to": "migration/reasoning-option", + "addedAt": "2026-09-30" + }, { "label": "TextPart markdown", "to": "migration/text-part-markdown", diff --git a/docs/harness/agents.md b/docs/harness/agents.md new file mode 100644 index 0000000000..198b148f2c --- /dev/null +++ b/docs/harness/agents.md @@ -0,0 +1,210 @@ +--- +title: Switch between agent profiles +id: harness-agents +order: 8 +description: "Give one harness several agents, such as a read-only planner and a builder. Each profile has its own prompt, tools, model, rules, and step limit, in code or in Markdown files." +keywords: + - tanstack ai + - harness + - agents + - agent profiles + - plan + - subagents +--- + +Your coding agent has two jobs. First it must plan a change without touching a file. Then it must build the change with every tool. The main model also needs a fast helper that searches the code. `agents()` gives the harness named agent profiles. Each profile has its own system prompt, tools, model, permission rules, and step limit. + +## 1. Add the plugin + +```ts group=harness-agents +import { defineHarness } from '@tanstack/ai-harness' +import { agents, permissions } from '@tanstack/ai-harness/plugins' +import { workspaceTools } from '@tanstack/ai-harness/plugins/coding' +import { runCli } from '@tanstack/ai-harness-cli' +import { anthropicText } from '@tanstack/ai-anthropic' + +const root = process.cwd() +const sonnet = anthropicText('claude-sonnet-5-5') + +const coder = defineHarness({ + name: 'acme/coder', + adapter: sonnet, + plugins: () => [ + permissions({ root }), + workspaceTools({ root }), + agents({ adapter: () => sonnet }), + ], +}) + +process.exitCode = await runCli(coder) +``` + +`adapter` turns a model id into an adapter. Here every agent uses the same model. + +## 2. Plan, then build + +1. Run `npx tsx coder.ts` in your project. +2. Type `/agent plan`. +3. Ask: `how do I add a dark mode?` The agent reads the code and gives numbered steps. It has no tool that changes a file. +4. Type `/agent build`. +5. Ask: `do step 1`. The agent changes the files, after your answer to each [permission question](./permissions). + +- `/agent` alone shows the current agent and the list: `Agent: build. Agents: build, plan.` +- A switch applies at the next turn. + +A client can change the `agent` setting when the harness exposes it. Add `expose: { config: ['agent'] }` to `defineHarness` (see [Choose what clients can change](./connect#choose-what-clients-can-change)). The [session view](./custom-ui) lists the options for a menu: + +```ts group=harness-agents-client +import { createHarnessClient } from '@tanstack/ai-harness/client' +import { createSessionView } from '@tanstack/ai-harness/view' + +const view = createSessionView( + createHarnessClient({ url: '/api/harness', threadId: 'fix-login' }), +) +await view.ready + +const setting = view.store.get().config.find((entry) => entry.key === 'agent') +if (setting?.option.type === 'select') { + console.log(setting.value, setting.option.options) // 'build' ['build', 'plan'] +} + +await view.setConfig('agent', 'plan') +``` + +## The built-in agents + +| Agent | Mode | Tools | What it does | +| --- | --- | --- | --- | +| `build` | `primary`, the default | Every tool | Reads code, changes files, and runs commands. | +| `plan` | `primary` | The read tools and `todo_write` | Reads the code and gives a plan in numbered steps. | +| `general` | `subagent` | Every tool | Does a task in many steps and reports back. | +| `explore` | `subagent` | The read tools | Searches the code and reports the paths that it finds. | + +The read tools are `read_file`, `list_files`, `grep`, `webfetch`, `websearch`, and `question`. An agent gets only the tools that the harness has. + +The mode of a profile says who starts it: + +- `primary`: it answers the user's turns. Pick it with `/agent`. +- `subagent`: the main model starts it as a child. +- `all`: both. + +The `plan` agent and the `plan` mode of [permissions](./permissions#pick-a-mode) are separate. The agent has only the read tools and `todo_write`. The mode denies edits and commands for every agent. To drop the built-in agents, set `builtIns: false`. + +## Add your own agent + +Give `agents()` a profile. This one reviews the changes with a stronger model: + +```ts group=harness-agents +import type { AnyTextAdapter } from '@tanstack/ai' +import type { AgentProfile } from '@tanstack/ai-harness/plugins' + +const review: AgentProfile = { + name: 'review', + description: 'Reviews the changes and lists the risks. Does not change files.', + mode: 'all', + model: 'claude-opus-5-5', + system: 'You review code. Read the diff, then list each risk with its file and line.', + tools: ['read_file', 'list_files', 'grep', 'bash'], + permissions: [{ tool: 'bash', resource: 'git diff*', decision: 'allow' }], + steps: 20, +} + +const models: Record = { + 'claude-opus-5-5': anthropicText('claude-opus-5-5'), +} + +export const withReview = agents({ + adapter: (model) => models[model] ?? sonnet, + agents: [review], +}) +``` + +`name`, `description`, and `mode` are required. The other fields: + +| Field | What it does | +| --- | --- | +| `model` | A model id. `adapter` turns it into an adapter. Without it, a primary agent keeps the harness model, and a subagent uses the model of the main turn. | +| `system` | Text for the system prompt of each turn of this agent. | +| `tools` | The tools that this agent can call, by name. A `*` matches any text, for example `todo_*`. Without it, the agent gets every tool. | +| `permissions` | [Permission rules](./permissions#write-rules) that count while this agent is the primary agent. | +| `steps` | The step limit. See [Limit the steps](#limit-the-steps). | +| `hidden` | `true` leaves the agent out of `/agent` and the `agent` setting. The model can still start it as a subagent. | + +- A profile replaces an earlier profile with the same name. The order is the built-in agents, then `dirs`, then `agents`. So a profile named `plan` replaces the built-in `plan`. +- `default` names the primary agent of a new session. Without it, the first primary agent is the default, which is `build`. +- In the main turn, a model that the thread stored with `session.configure` wins over the profile `model`. See [What wins](./thread-settings#what-wins). +- A `tools` list also filters the subagent tools. To let the agent start subagents, add their tool names, for example `explore`. + +## Load agents from Markdown files + +Keep your agents next to your code as Markdown files. Save this file as `.agents/agents/review.md`: + +```md +--- +description: Reviews the changes and lists the risks. Does not change files. +mode: all +model: claude-opus-5-5 +tools: [read_file, list_files, grep] +steps: 20 +--- + +You review code. Read the diff, then list each risk with its file and line. +``` + +Then give `agents()` the folder: + +```ts group=harness-agents +export const fromFiles = agents({ + adapter: (model) => models[model] ?? sonnet, + dirs: [`${root}/.agents/agents`], +}) +``` + +- The file name is the agent name, here `review`. +- The body is the system prompt. +- The frontmatter sets `description`, `mode`, `model`, `tools`, `steps`, and `hidden: true`. The default `mode` is `all`. +- An unknown `mode`, or a `steps` that is not a whole number, stops `host.open` with an error, for example `Agent file review.md: "mode" is not valid.` +- The plugin reads the folder when a session opens. A folder that it cannot read gives no agents. + +## Limit the steps + +An agent can call tools for a long time. `steps` sets the most model calls with tools in one run. With `steps: 20`: + +1. The agent makes up to 20 model calls with tools. +2. Call 21 gets `toolChoice: 'none'` and a note: answer now, and say what is left to do. +3. Then the run stops. + +The limit counts again in each turn of a primary agent and in each run of a subagent. + +Some providers do not obey `'none'`. On Amazon Bedrock, `'none'` after a tool call goes to the model as `auto`. See [Provider notes](../tools/tools#provider-notes). If the model calls a tool on the last call anyway: + +- The tool does not run. +- The model gets this tool result: `{ "error": "Step limit reached. Answer without calling tools." }` +- The run stops. That turn can end with no text answer. + +## Let the model start a subagent + +A profile with the mode `subagent` or `all` is also a subagent. The main model gets a tool for each one, for example `explore`. Ask the build agent: `use explore to find where the theme is set`. + +A subagent run gets these parts of its profile: + +- The tools of the last main turn, filtered by its `tools`. +- Its `system` prompt. +- Its `model`, or the model of the main turn. +- Its step limit. + +To give the model one `subagent` tool for every agent, see [Agents from plugins](./subagents#agents-from-plugins). The harness also limits the depth and the number of children. See [Stay within limits](./subagents#stay-within-limits). + +## Known limits + +- The `permissions` of a profile count only while it is the primary agent. When the profile runs as a subagent, its own rules do not apply. Put rules that must hold in every run in `permissions({ rules })`. See [Rules in subagents](./permissions#rules-in-subagents). +- A Markdown agent cannot set `permissions`. Put those rules in a profile in code. +- On Amazon Bedrock, the model can still call a tool on the last call. The tool does not run, but the turn can end with no text answer. +- A subagent that starts before the first turn of the session has no tools. It also needs its own `model`. + +## What you have now + +- `/agent plan` for a read-only planner, and `/agent build` to do the work. +- Your own agents in code or in Markdown files, each with its own prompt, tools, model, and rules. +- A step limit that ends a run with a text answer. +- Subagents that the main model starts to search and to research. diff --git a/docs/harness/auth.md b/docs/harness/auth.md new file mode 100644 index 0000000000..2b33a130ba --- /dev/null +++ b/docs/harness/auth.md @@ -0,0 +1,173 @@ +--- +title: Auth and connectors +id: harness-auth +order: 7 +description: "Let users sign in to GitHub and other OAuth services from the harness. Tokens stay in your credential store and never reach the model." +keywords: + - tanstack ai + - harness + - oauth + - connectors + - credentials +--- + +Your agent needs to open a pull request, so it needs the user's GitHub token. The user should sign in once, in the browser, and the model should never see the token. `oauthConnector` does that: it adds `/connect github`, stores the token, and hands it to your tools. + +## 1. Add a connector + +```ts group=harness-auth +import { toolDefinition } from '@tanstack/ai' +import { defineHarness, oauthConnector } from '@tanstack/ai-harness' +import { openaiText } from '@tanstack/ai-openai' + +const github = oauthConnector({ + id: 'github', + label: 'GitHub', + oauth: { + authorizationUrl: 'https://github.com/login/oauth/authorize', + tokenUrl: 'https://github.com/login/oauth/access_token', + deviceUrl: 'https://github.com/login/device/code', + clientId: 'your-github-oauth-app-client-id', + scopes: ['repo'], + }, + tools: (token) => [ + toolDefinition({ name: 'list_issues', description: 'List my open issues' }).server( + async () => { + const response = await fetch('https://api.github.com/issues', { + headers: { authorization: `Bearer ${await token()}` }, + }) + return response.json() + }, + ), + ], +}) + +export const assistant = defineHarness({ + name: 'acme/assistant', + adapter: openaiText('gpt-5.6'), + plugins: () => [github], +}) +``` + +`token()` returns a fresh access token. It refreshes an expired token when the service gave a refresh token. + +## 2. Sign in + +- The user runs `/connect github`. +- The CLI opens the browser. The harness listens on `127.0.0.1` on a random port for one callback, with PKCE and a random `state`. +- The token goes into the credential store. `/disconnect github` deletes it. + +Set `login: 'device'` for SSH sessions, containers, and CI. The CLI then shows a code to enter on the service's page. + +If a tool runs before sign-in, it stops with a `harness.auth_required` event. The CLI shows which service to connect. + +## Wait for the sign-in + +A turn from a background agent or an email runs while nobody watches it. That turn can need a sign-in. Without a sign-in, the tool fails, and the work of that turn stops. Pass `wait: true` to `token()`. Then the turn waits for the sign-in and continues after it. + +```ts group=harness-auth-wait +import { toolDefinition } from '@tanstack/ai' +import { defineHarness, oauthConnector } from '@tanstack/ai-harness' +import { openaiText } from '@tanstack/ai-openai' + +const github = oauthConnector({ + id: 'github', + label: 'GitHub', + oauth: { + authorizationUrl: 'https://github.com/login/oauth/authorize', + tokenUrl: 'https://github.com/login/oauth/access_token', + deviceUrl: 'https://github.com/login/device/code', + clientId: 'your-github-oauth-app-client-id', + scopes: ['repo'], + }, + tools: (token) => [ + toolDefinition({ name: 'list_issues', description: 'List my open issues' }).server( + async () => { + const response = await fetch('https://api.github.com/issues', { + headers: { authorization: `Bearer ${await token({ wait: true })}` }, + }) + return response.json() + }, + ), + ], +}) + +export const assistant = defineHarness({ + name: 'acme/assistant', + adapter: openaiText('gpt-5.6'), + plugins: () => [github], + // A web app starts the sign-in with this command. + expose: { commands: ['connect:github'] }, +}) +``` + +In your own plugin, use `ctx.credentials.require('github', { wait: true })`. + +When the credential is missing, the turn stops with an interrupt: + +- The interrupt has `reason: 'auth_required'`. Its `metadata['tanstack:interruptPayload'].request.connector` is `'github'`. +- The sign-in belongs to the sender of the stopped turn. When that person saves the credential, the same turn continues as that person, and the tool runs again. The save can be `/connect github`, the `connect:github` command, or `ctx.credentials.set`. +- A credential that another person saves never continues the turn. +- To stop the wait, resolve the interrupt with `status: 'cancelled'`. Then the tool fails, the same as without `wait`. +- On a durable host, the wait continues after a restart. +- `wait` is for a tool of the running chat turn. Do not pass it in a command, a background agent, or a hook. The credentials of a command never wait. +- Provider keys (`ctx.keys.require`) do not wait. + +In a web app, the wait shows in `state.signIns` of the session view, also after a reload. When the user presses your connect button, run the connector command. A client can run only the commands in `expose.commands`, so the harness above exposes `connect:github` (see [Choose what clients can change](./connect#choose-what-clients-can-change)): + +```ts group=harness-auth-wait-client +import { createHarnessClient } from '@tanstack/ai-harness/client' +import { createSessionView } from '@tanstack/ai-harness/view' + +const view = createSessionView( + createHarnessClient({ url: '/api/harness', threadId: 'user-1-thread' }), +) + +view.on('signIn', (signIn) => { + console.log(`Connect ${signIn.connector} to continue the turn.`) +}) + +export const connectGithub = () => view.command('connect:github') +``` + +After the sign-in, the turn finishes its work with no new message. + +## Keep credentials + +The host reads credentials from `stores.credentials`, keyed by the sender of the running turn. Without it, credentials live in memory until the process stops. When several people write in one thread, see [Share a thread between people](./shared-threads). + +```ts group=harness-auth +import { defineCredentialStore } from '@tanstack/ai-persistence' +import type { Credential } from '@tanstack/ai-persistence' + +const saved = new Map() +const keyOf = (userId: string | undefined, id: string) => `${userId ?? 'tenant'}:${id}` + +export const credentials = defineCredentialStore({ + get: async (scope, id) => saved.get(keyOf(scope.userId, id)) ?? null, + set: async (scope, id, credential) => { + saved.set(keyOf(scope.userId, id), credential) + }, + delete: async (scope, id) => { + saved.delete(keyOf(scope.userId, id)) + }, + list: async (scope) => + [...saved.entries()] + .filter(([key]) => key.startsWith(`${scope.userId ?? 'tenant'}:`)) + .map(([key, credential]) => ({ id: key.split(':')[1] ?? key, type: credential.type })), +}) +``` + +- Encrypt tokens at rest in a real store. +- `list` returns ids and types only, never the secret values. +- A credential saved without a `userId` belongs to the whole tenant. A user who has no credential of their own uses it. See [Share one credential with the whole organization](./shared-threads#share-one-credential-with-the-whole-organization). + +## Read credentials in your own plugin + +`ctx.credentials.require('github')` returns the credential, or stops with `auth_required` when the user has not signed in. Plugins read only the credentials of the sender of the running turn. Outside a turn, they read the credentials of the person who opened the session. + +## What you have now + +- Browser and device-code sign-in for any OAuth service. +- Tools that get a fresh token without the model seeing it. +- Turns that wait for a sign-in and then finish their work. diff --git a/docs/harness/cli.md b/docs/harness/cli.md new file mode 100644 index 0000000000..648a4fc766 --- /dev/null +++ b/docs/harness/cli.md @@ -0,0 +1,180 @@ +--- +title: Run a harness in the terminal +id: harness-cli +order: 3 +description: "Run your harness in a terminal with line mode or your own screen, a print mode for scripts and CI, NDJSON output, an ACP mode for editors, and an HTTP server." +keywords: + - tanstack ai + - harness + - cli + - terminal + - ink +--- + +You built a harness and want to use it like Claude Code: type in a terminal, watch it work, approve tools. You also want to run it in CI. `runCli` gives one harness all of these modes. + +## Install + + + +react: @tanstack/ai-harness-cli +vue: @tanstack/ai-harness-cli +solid: @tanstack/ai-harness-cli +svelte: @tanstack/ai-harness-cli +preact: @tanstack/ai-harness-cli +angular: @tanstack/ai-harness-cli +vanilla: @tanstack/ai-harness-cli +octane: @tanstack/ai-harness-cli + + + +## 1. Write the entry file + +```ts group=harness-cli +import { defineHarness } from '@tanstack/ai-harness' +import { runCli } from '@tanstack/ai-harness-cli' +import { openaiText } from '@tanstack/ai-openai' + +const assistant = defineHarness({ + name: 'acme/assistant', + adapter: openaiText('gpt-5.6'), +}) + +process.exitCode = await runCli(assistant) +``` + +## 2. Pick a mode + +To use the harness yourself: + +- No flags: line mode. Type a message and press Enter. With a `ui`, your own screen starts in its place (see [Run your own screen](#run-your-own-screen)). +- `-p "prompt"`: run one prompt, print the answer, and exit. +- `-p "prompt" --output ndjson`: print every AG-UI event as one JSON line. + +To use the harness from another program: + +- `--acp`: serve the harness as an ACP v2 agent over stdio, for editors. Needs `@tanstack/ai-acp`. +- `--mcp`: serve the harness as an MCP server over stdio, for Claude Code, Cursor, and other MCP clients. Needs `@tanstack/ai-mcp`. Add `--yes` to approve every tool call. Read [Use a harness from any MCP client](./mcp-server). +- `--serve`: serve the session protocol over HTTP on `127.0.0.1:8787`. Every request needs the bearer token. Pass `--token`, set `HARNESS_TOKEN`, or copy the token the CLI prints. With `@tanstack/ai-mcp`, it also serves MCP at `/mcp`. + +`--mcp` and `--serve` run only the commands in `expose.commands` of the harness. `--serve` also sets only the config keys in `expose.config`. Add the names to `expose` in `defineHarness`, for example `expose: { commands: ['connect:github'] }`. Line mode and `-p` run every command. See [Choose what clients can change](./connect#choose-what-clients-can-change). + +Line mode reads one message or command per line and waits for each turn. It works the same in a terminal and with piped input. In a terminal, it also opens sign-in links in the browser. + +## 3. Use it in CI + +`-p` exits with a code your script can check: + +| Code | Meaning | +|---|---| +| 0 | The turn finished. | +| 1 | The turn failed. | +| 2 | The turn waits for an approval. | +| 130 | The turn was cancelled. | + +## Use your own host and user + +By default, the CLI builds its own host in memory and opens sessions with no user. Your server already has a durable host with leases, recovery, and stores, and each person has their own credentials. Give the CLI that host and the user: + +```ts group=harness-cli-host +import { createHarnessHost, defineHarness } from '@tanstack/ai-harness' +import { runCli } from '@tanstack/ai-harness-cli' +import { memoryLogStore, memoryPersistence } from '@tanstack/ai-persistence' +import { openaiText } from '@tanstack/ai-openai' + +const assistant = defineHarness({ + name: 'acme/assistant', + adapter: openaiText('gpt-5.6'), +}) + +const { runs, metadata, credentials } = memoryPersistence().stores +const host = createHarnessHost({ + persistence: { stores: { log: memoryLogStore(), runs, metadata, credentials } }, +}) + +process.exitCode = await runCli(assistant, { host, principal: { id: 'ada' } }) +await host.close() +``` + +- You own the host, so the CLI does not close it. Close it yourself when `runCli` returns. Without `host`, the CLI builds a host and closes it. +- Give `host` or `persistence`, not both. With both, `runCli` writes `Give runCli a host or persistence, not both.` and returns code 1. +- Line mode, `-p`, and your own `ui` open the session as `principal`, so `/connect` saves keys for that user. +- `--serve` gives the same principal to each request with the token. +- `--acp`, `--mcp`, and `--dashboard` open their sessions without the principal. + +## Commands in line mode + +For the session: + +- `/agents`: list the agents. +- `/agent {"json":"input"}`: run an agent in the background. When it is done, a new turn starts with its result. +- `/config`: show the settings. `/config ` changes one. +- `/connect ` and `/disconnect `: sign in to a connector, or out. With [`providerKeys`](./provider-keys), they also connect a model provider, for example `/connect openai`. + +For the running work: + +- `/cancel`: cancel the running turn. +- `/status`: show what runs and what waits. +- `/exit`: quit. + +Plugin commands (for example `/model` or `/todos`) show up in `/help`. When a turn stops for an approval or a plugin asks a question, type your answer. For yes-or-no questions, `y` approves and `n` refuses. + +## Send files and save media + +To send a file with a message, write `@` and the path of the file: + +```text +what is wrong in @./bug.png +compare @"old logo.png" with @./new-logo.png +``` + +- The path is relative to the working folder. Put quotes around a path with spaces. +- `@path` works in line mode, with `-p`, with piped input, and in your own `ui`. +- A `@word` that is not a file stays in the text, for example `@types/node`. +- If the CLI does not know the type of the file, it does not send the message. To send the path as text, remove the `@`. + +The CLI saves the media that a turn makes in a folder. The default folder is `./-media` in the working folder, for example `./acme-assistant-media`. To use another folder, add `--media-dir `: + +```bash +npx tsx cli.ts -p "Draw a logo for acme" --media-dir ./out +``` + +- Line mode prints `[image saved: ]` for each file. +- `-p` prints the same line on stderr, so stdout keeps only the answer. +- `-p --output ndjson` adds the saved `path` to the value of the `harness.media` event. + +The CLI never replaces a file. If the name is taken, it adds `-1`, `-2`, and so on to the new name. + +## Run your own screen + +Line mode prints plain lines. For a full screen with your own layout, pass `ui` to `runCli`. It works with any TUI library, for example Ink, OpenTUI, or blessed. + +`ui` gets a ready [session view](./custom-ui) and resolves when the user quits. This entry file starts an Ink screen: + +```tsx ignore +import { render } from 'ink' +import { runCli } from '@tanstack/ai-harness-cli' +import { assistant } from './harness' +import { Screen } from './screen' + +process.exitCode = await runCli(assistant, { + ui: async (view) => { + await render().waitUntilExit() + }, +}) +``` + +- `ui` runs only in an interactive terminal. Piped input uses line mode. `-p`, `--acp`, `--mcp`, `--serve`, and `--dashboard` do not use `ui`. +- When `ui` resolves, the CLI disposes the view and `runCli` returns. +- To write `Screen`, read [Build your own UI](./custom-ui). + +For a full Ink screen with approvals, questions, sign-ins, and child agents, copy [`examples/harness-cli/src/tui.tsx`](https://github.com/TanStack/ai/blob/main/examples/harness-cli/src/tui.tsx). + +## What you have now + +- One entry file that runs your harness as a terminal app, a script step, an editor agent, an MCP server, or an HTTP server. +- A CLI that uses your durable host and runs as your user. +- Your own terminal screen on the same session, with any TUI library. +- Files that you send with `@path`, and a folder with the media that your agents make. + +Next: keep long turns alive through crashes with [durable sessions](./durable-sessions). diff --git a/docs/harness/code-mode.md b/docs/harness/code-mode.md new file mode 100644 index 0000000000..a3e54768b4 --- /dev/null +++ b/docs/harness/code-mode.md @@ -0,0 +1,122 @@ +--- +title: Code mode in a harness +id: harness-code-mode +order: 10 +description: "Let the harness model write one TypeScript program that calls many tools, and run it in an isolate. Any TanStack AI isolate driver plugs in." +keywords: + - tanstack ai + - harness + - code mode + - isolate + - quickjs +--- + +Your agent has 100 tools from Notion and Linear, and it calls them one at a time. Each call is a model round trip, and every tool schema goes into every request. With code mode, the model writes one TypeScript program that calls the tools it needs, and the program runs in an isolate. The read-only tools leave the tool list and become functions in that program. + +## 1. Add the plugin + + + +react: @tanstack/ai-code-mode @tanstack/ai-isolate-quickjs +vue: @tanstack/ai-code-mode @tanstack/ai-isolate-quickjs +solid: @tanstack/ai-code-mode @tanstack/ai-isolate-quickjs +svelte: @tanstack/ai-code-mode @tanstack/ai-isolate-quickjs +preact: @tanstack/ai-code-mode @tanstack/ai-isolate-quickjs +angular: @tanstack/ai-code-mode @tanstack/ai-isolate-quickjs +octane: @tanstack/ai-code-mode @tanstack/ai-isolate-quickjs +vanilla: @tanstack/ai-code-mode @tanstack/ai-isolate-quickjs + + + +```ts group=harness-code-mode +import { defineHarness } from '@tanstack/ai-harness' +import { codeMode } from '@tanstack/ai-code-mode/harness' +import { createQuickJSIsolateDriver } from '@tanstack/ai-isolate-quickjs' +import { mcpConnector } from '@tanstack/ai-mcp/connector' +import { openaiText } from '@tanstack/ai-openai' + +const linear = mcpConnector({ + id: 'linear', + label: 'Linear', + url: 'https://mcp.linear.app/mcp', +}) + +export const assistant = defineHarness({ + name: 'acme/assistant', + adapter: openaiText('gpt-5.6'), + plugins: () => [ + linear, + codeMode({ driver: createQuickJSIsolateDriver() }), + ], +}) +``` + +The model now has an `execute_typescript` tool. Inside the program, each moved tool is an `external_*` function, for example `external_linear_list_issues()`. + +## 2. Pick the isolate + +`driver` takes any isolate driver. Change one line to run the programs somewhere else: + +| Driver | Package | Runs in | +| --- | --- | --- | +| `createQuickJSIsolateDriver()` | `@tanstack/ai-isolate-quickjs` | WebAssembly, on any runtime | +| `createNodeIsolateDriver()` | `@tanstack/ai-isolate-node` | A V8 isolate in Node | +| `createCloudflareIsolateDriver()` | `@tanstack/ai-isolate-cloudflare` | A Cloudflare Worker | +| `createDaytonaIsolateDriver()` | `@tanstack/ai-isolate-daytona` | A Daytona sandbox | + +[Code mode isolates](../code-mode/code-mode-isolates) has the options of each driver. + +## 3. Choose which tools move + +A call inside the program does not stop for approval. So by default, a tool moves into code mode only when it is safe to run without a question: + +- The tool runs on the server and does not set `needsApproval`. +- The permission rules allow it in plan mode. File reads move. `write_file` and `bash` stay. +- MCP tools that the server does not mark read-only ask for approval, so they stay too. + +Where the permission rules come from: + +- With [`permissions()`](./permissions) in the plugin list, code mode asks it about each tool. The answer is the one a call gets in plan mode: your `rules` win, and a tool with no rule gets your `default`. So with `default: 'ask'`, a tool with no rule stays. +- Without `permissions()`, code mode reads the rules that tool plugins add. + +Every other tool stays a normal tool call, with its approvals. To pick the tools yourself, pass `include`: + +```ts group=harness-code-mode +export const onlyLinear = codeMode({ + driver: createQuickJSIsolateDriver(), + include: (tool) => tool.name.startsWith('linear_'), +}) +``` + +The other options of `createCodeMode` pass through: `timeout`, `memoryLimit`, `lazyToolsConfig`, and more. + +## 4. Send only the tool names + +The prompt of code mode lists each moved tool with its full TypeScript type. With 100 tools, that is a lot of tokens on every turn, and most turns use a few tools. Set `lazy: true`, and the prompt lists only the names: + +```ts group=harness-code-mode +export const lazyCodeMode = codeMode({ + driver: createQuickJSIsolateDriver(), + lazy: true, + lazyToolsConfig: { includeDescription: 'first-sentence' }, +}) +``` + +The model then works in three steps: + +1. It reads the tool names in the prompt. +2. It calls `discover_tools` with the names it needs, and gets their TypeScript signatures. +3. It calls those tools in `execute_typescript`. + +`includeDescription` sets what the prompt shows with each name: `'none'` (the default, the name only), `'first-sentence'`, or `'full'`. Tools that stay outside code mode keep their full schema. + +## Tools that appear after sign-in + +The plugin chooses the tools again before every turn. When the user runs `/connect linear`, the read-only Linear tools move into code mode on the next turn. A plugin can do the same with its own tool changes: return the new list from `prepareTools`. [Use MCP servers](./mcp) shows `discoverTools`, which adds the tools first. + +## What you have now + +- One `execute_typescript` tool in place of many read-only tools. +- With `lazy: true`, a prompt that lists only the tool names. +- Programs that run in the isolate you choose. +- Approvals kept for every tool that can change something. diff --git a/docs/harness/coding-agent.md b/docs/harness/coding-agent.md new file mode 100644 index 0000000000..03419e6818 --- /dev/null +++ b/docs/harness/coding-agent.md @@ -0,0 +1,151 @@ +--- +title: Build a coding agent +id: harness-coding-agent +order: 6 +description: "Turn a harness into a coding agent with file tools, permission modes, undo, a todo list, a model picker, project instructions, session titles, a cost count, and /compact." +keywords: + - tanstack ai + - harness + - coding agent + - permissions + - workspace tools + - project instructions + - undo +--- + +You want your own coding agent in the terminal. It reads and edits files, and runs commands after you approve them. When a turn goes wrong, you undo it. The first-party plugins in `@tanstack/ai-harness/plugins` and `@tanstack/ai-harness/plugins/coding` give you those parts. You pick the model and the rules. + +## 1. Define the agent + +```ts group=harness-coding-agent +import { homedir } from 'node:os' +import { join } from 'node:path' +import { defineHarness } from '@tanstack/ai-harness' +import { + compact, + fileCommands, + modelPicker, + permissions, + projectInstructions, + title, + todos, + usage, +} from '@tanstack/ai-harness/plugins' +import { snapshots, workspaceTools } from '@tanstack/ai-harness/plugins/coding' +import { runCli } from '@tanstack/ai-harness-cli' +import { openaiText } from '@tanstack/ai-openai' + +const root = process.cwd() +const smart = openaiText('gpt-6.1-sol') +const fast = openaiText('gpt-6-luna') + +const coder = defineHarness({ + name: 'acme/coder', + adapter: smart, + systemPrompts: ['You are a careful coding agent. Read before you edit.'], + plugins: () => [ + permissions({ root }), + workspaceTools({ root }), + snapshots({ root, dataDir: join(homedir(), '.acme-coder', 'snapshots') }), + todos(), + modelPicker({ choices: { smart, fast }, default: 'smart' }), + fileCommands({ dir: `${root}/.claude/commands` }), + compact({ adapter: fast }), + title({ adapter: fast }), + usage(), + // Keep this plugin last. See "Project instructions" below. + projectInstructions({ root }), + ], +}) + +process.exitCode = await runCli(coder) +``` + +Both plugin entries run only in Node. The coding tools use the file system and the shell of the machine. + +## 2. Run it + +1. Run `npx tsx coder.ts` in your project. +2. Ask for a change. The agent reads files freely, and asks before it writes a file or runs a command. +3. Answer `once` to allow the call, `always` to allow it from now on, or `reject` to refuse it. +4. If you do not like the result, run `/undo`. The files go back to how they were before the turn. + +## What each plugin adds + +| Plugin | Adds | +|---|---| +| `permissions({ root })` | Asks before risky tool calls. `/mode` switches between `default`, `plan` (read-only), `acceptEdits` (edits run without asking), and `bypass`. See [Ask before risky tool calls](./permissions). | +| `workspaceTools({ root })` | Tools that read, search, and change the files in `root`, run commands, and read web pages. A path outside `root` is refused, or asks first with `outside: 'ask'`. See [Give your agent coding tools](./coding-tools). | +| `snapshots({ root, dataDir })` | `/undo` puts the files back as they were before the last turn, and `/redo` brings them back. It needs `git`. See [Undo a turn](./snapshots). | +| `todos()` | A `todo_write` tool the model uses for multi-step work, and `/todos`. | +| `modelPicker({ choices })` | `/model ` switches the model at the next turn. | +| `fileCommands({ dir })` | Each `.md` file becomes a slash command. `$ARGUMENTS` is replaced by what you type after it. | +| `compact({ adapter })` | `/compact` replaces a long conversation with a summary. | +| `title({ adapter })` | Names the session from its first message. | +| `usage()` | `/usage` shows the tokens and the cost of the session, with a line for each model. See [Track usage and cost](./usage). | +| `projectInstructions({ root })` | Adds `AGENTS.md`, `CLAUDE.md`, and facts about the environment to the system prompt. | + +To compact on its own before the context limit, see [Compact a harness session](./compaction). + +## Project instructions + +Your repository has an `AGENTS.md` or a `CLAUDE.md` with its rules. `projectInstructions({ root })` gives them to the model: + +- It reads `AGENTS.md` and `CLAUDE.md` in each folder from the repository root (the folder with `.git`) down to `root`. Set other file names with `files`. +- `global` adds files that come first, for example `['~/.config/AGENTS.md']`. +- When `read_file` reads a file in a subfolder, the model also gets the instruction files of the folders on the way, once each. +- An environment block gives the date, the platform, the working folder, and the git repository and branch. `env: false` turns it off. +- When you change an instruction file, the next turn sends the new text as a note. + +Put `projectInstructions()` last in the plugin list. Then a changed file goes into the conversation as a [mid-conversation change](../advanced/mid-conversation-changes), and the prompt cache stays. When the prompt of another plugin comes after it, a change builds the system prompt again, and the prompt cache starts again. + +## Session titles and cost + +`title({ adapter })` and `usage()` write to the session index, one entry per thread: + +- `title` names the session from its first message. The turn does not wait for it. +- When `usage` knows the cost of the session, it writes the cost. A cost that the provider reports is used as it is. +- To price the calls of a provider that reports no cost, pass `usage({ model })`. [Track usage and cost](./usage#show-usage-and-price-the-calls) shows how. + +A [sidebar of sessions](./sessions) reads the title and the cost from the index. The `usage()` state also has `contextTokens`, the size of the lead model's context at its latest call, for a context meter. + +## Add more + +- `agents({ adapter })`: named agents with their own prompt, tools, and model. `/agent plan` switches to a read-only agent. See [Switch between agent profiles](./agents). +- `question()`: a `question` tool. The model asks the user to pick an option, and waits for the answer. +- `goal({ judge })`: `/goal ` keeps the agent working until a judge model says that the goal is met. See [Work until a goal is met](./goal). +- `mcp({ servers })` from `@tanstack/ai-mcp/harness`: tools from MCP servers. See [Use MCP servers](./mcp). +- `formatter({ root })`: formats each file after a write. See [Format each file after a write](./coding-tools-backends#format-each-file-after-a-write). + +## Add your own rules + +Tool plugins add permission rules to the `PermissionRules` extension point. Add your own for any tool: + +```ts group=harness-coding-agent +import { PermissionRules } from '@tanstack/ai-harness/plugins' +import { definePlugin } from '@tanstack/ai-harness' + +export const noDeploys = definePlugin({ + name: 'acme/no-deploys', + setup: () => ({ + contribute: [PermissionRules.item({ tool: 'deploy_*', decision: 'deny' })], + }), +}) +``` + +A trailing `*` matches every tool that starts with the text. The last matching rule wins. [Ask before risky tool calls](./permissions) explains the rules, the modes, and the saved answers. + +The workspace tools run on your machine with your permissions. For code that you do not trust, [run the tools in a sandbox](./coding-tools-backends#run-the-tools-in-a-sandbox). + +## Hand work to Claude Code or Codex + +Your agent can also give tasks to coding agents you already use. [Delegate to coding agents](./coding-agents) shows how. + +## What you have now + +- A terminal coding agent with file tools, approvals, modes, a todo list, and a model picker. +- `/undo` and `/redo` for the file changes of a turn. +- The rules of your repository in every turn. A change to them keeps the prompt cache. +- A title and a cost for each session. + +Next: connect the agent to GitHub and other services with [auth and connectors](./auth). diff --git a/docs/harness/coding-agents.md b/docs/harness/coding-agents.md new file mode 100644 index 0000000000..ce4ed7d33d --- /dev/null +++ b/docs/harness/coding-agents.md @@ -0,0 +1,122 @@ +--- +title: Delegate to coding agents +id: harness-coding-agents +order: 11 +description: "Let a harness hand coding work to Claude Code, Codex, Grok Build, or any ACP agent. Each one works in a sandbox and keeps its own session." +keywords: + - tanstack ai + - harness + - claude code + - codex + - grok build + - sandbox + - subagents +--- + +Your lead agent plans the work, but you want Claude Code or Codex to do the edits, because they are good at code and you already use them. `codingAgents` gives the lead model one tool per coding agent. Each agent works in a sandbox, keeps its own session between calls, and its tool calls show up in your UI. + +## 1. Add the plugin + + + +react: @tanstack/ai-harness @tanstack/ai-sandbox @tanstack/ai-sandbox-local-process @tanstack/ai-claude-code @tanstack/ai-codex +vue: @tanstack/ai-harness @tanstack/ai-sandbox @tanstack/ai-sandbox-local-process @tanstack/ai-claude-code @tanstack/ai-codex +solid: @tanstack/ai-harness @tanstack/ai-sandbox @tanstack/ai-sandbox-local-process @tanstack/ai-claude-code @tanstack/ai-codex +svelte: @tanstack/ai-harness @tanstack/ai-sandbox @tanstack/ai-sandbox-local-process @tanstack/ai-claude-code @tanstack/ai-codex +preact: @tanstack/ai-harness @tanstack/ai-sandbox @tanstack/ai-sandbox-local-process @tanstack/ai-claude-code @tanstack/ai-codex +angular: @tanstack/ai-harness @tanstack/ai-sandbox @tanstack/ai-sandbox-local-process @tanstack/ai-claude-code @tanstack/ai-codex +octane: @tanstack/ai-harness @tanstack/ai-sandbox @tanstack/ai-sandbox-local-process @tanstack/ai-claude-code @tanstack/ai-codex +vanilla: @tanstack/ai-harness @tanstack/ai-sandbox @tanstack/ai-sandbox-local-process @tanstack/ai-claude-code @tanstack/ai-codex + + + +```ts group=harness-coding-agents +import { defineHarness } from '@tanstack/ai-harness' +import { permissions } from '@tanstack/ai-harness/plugins' +import { claudeCodeText } from '@tanstack/ai-claude-code' +import { codexText } from '@tanstack/ai-codex' +import { openaiText } from '@tanstack/ai-openai' +import { defineSandbox, defineWorkspace, localSource } from '@tanstack/ai-sandbox' +import { codingAgents } from '@tanstack/ai-sandbox/harness' +import { localProcessSandbox } from '@tanstack/ai-sandbox-local-process' + +const repo = '/path/to/your/repo' + +export const lead = defineHarness({ + name: 'acme/lead', + adapter: openaiText('gpt-6.1-sol'), + plugins: () => [ + permissions(), + codingAgents({ + sandbox: defineSandbox({ + id: 'repo', + provider: localProcessSandbox({ dir: repo }), + workspace: defineWorkspace({ source: localSource(repo) }), + }), + agents: { + claude_code: { + adapter: claudeCodeText('claude-opus-4-8', { permissionMode: 'acceptEdits' }), + description: 'Larger changes, refactors, and reviews', + }, + codex: { + adapter: codexText('gpt-5.3-codex', { sandboxMode: 'workspace-write' }), + description: 'Quick fixes and tests', + }, + }, + }), + ], +}) +``` + +The lead model now has a `claude_code` tool and a `codex` tool. Each tool takes one `task`: the whole job, in words. The agent sees only that text and the files in its sandbox. + +## 2. Ask for work + +1. Start the harness, for example with the CLI. +2. Ask the lead: `have claude_code add a test for the date parser, then have codex fix what fails`. +3. Watch the child work. The CLI shows each agent's tool calls, then one line with the start of its answer: + +```text +[agent claude_code started] +[claude_code: tool Write] +[agent claude_code finished: Added parse-date.test.ts with three cases.] +``` + +The dashboard shows the same work in a block under the lead's message. An ACP editor shows the child's tool calls next to the lead's own. + +## Sessions + +Each agent keeps its own session per harness thread. The next task for `claude_code` resumes the same Claude Code session, so it still knows the files it read. The session ids live in the plugin state, so they also survive a restart of the host. + +To start over, run `/fresh claude_code`, or `/fresh` for every agent. + +## Workspaces + +`workspace` chooses where the agents work: + +| Value | Where each agent works | At the same time | +| --- | --- | --- | +| `'shared'` (default) | One sandbox per harness thread | One agent at a time | +| `'per-agent'` | A sandbox per agent | Yes | + +Use `'shared'` when the agents build on each other's changes. Use `'per-agent'` when they work on separate copies and you merge the results. + +## Plan mode + +When the `permissions()` plugin is in [`plan` mode](./permissions#pick-a-mode), the agents start read-only: + +- Claude Code gets `permissionMode: 'plan'`. +- Codex gets `sandboxMode: 'read-only'`. +- Any other agent gets its `planModelOptions`, for example `{ permissionMode: 'default' }` for an ACP agent that asks before each edit. + +## Other agents and sandboxes + +- `adapter` takes any coding-agent adapter: `grokBuildText` from `@tanstack/ai-grok-build`, or `acpCompatibleText` from `@tanstack/ai-acp` for any ACP agent. +- `sandbox` takes any sandbox provider. [Sandbox providers](../sandbox/providers) lists them. +- `modelOptions` on an agent is added to every call, for example a fixed `permissionMode`. + +## What you have now + +- A lead model that hands tasks to Claude Code and Codex. +- One sandbox per thread, and one saved session per agent. +- Child tool calls in the CLI, the dashboard, and ACP editors. diff --git a/docs/harness/coding-tools-backends.md b/docs/harness/coding-tools-backends.md new file mode 100644 index 0000000000..8a391df688 --- /dev/null +++ b/docs/harness/coding-tools-backends.md @@ -0,0 +1,263 @@ +--- +title: Sandbox and extend the coding tools +id: harness-coding-tools-backends +order: 7 +description: "Run the coding tools in a sandbox or on your own backend, format each written file, and keep tool output small." +keywords: + - tanstack ai + - harness + - coding agent + - workspace tools + - sandbox + - formatter + - workspace hooks +--- + +Your agent must work on files that are not on this machine, or on code that you do not trust. Or you want code to run after each write, like a formatter. The [coding tools](./coding-tools) send every read, write, and command through a backend, and they run hooks after each read and write. This page shows how to change both. + +## Run the tools in a sandbox + +By default, the tools run on your machine, with the rights of your user. For code that you do not trust, run them in a sandbox. `sandboxWorkspaceBackend()` sends every read, write, and command to the sandbox. + +1. Install the sandbox packages: + + + +react: @tanstack/ai-sandbox @tanstack/ai-sandbox-docker +vue: @tanstack/ai-sandbox @tanstack/ai-sandbox-docker +solid: @tanstack/ai-sandbox @tanstack/ai-sandbox-docker +svelte: @tanstack/ai-sandbox @tanstack/ai-sandbox-docker +preact: @tanstack/ai-sandbox @tanstack/ai-sandbox-docker +angular: @tanstack/ai-sandbox @tanstack/ai-sandbox-docker +octane: @tanstack/ai-sandbox @tanstack/ai-sandbox-docker +vanilla: @tanstack/ai-sandbox @tanstack/ai-sandbox-docker + + + +2. Start a sandbox, and give its backend to `workspaceTools()`: + +```ts group=harness-coding-tools-sandbox +import { defineHarness } from '@tanstack/ai-harness' +import { permissions } from '@tanstack/ai-harness/plugins' +import { workspaceTools } from '@tanstack/ai-harness/plugins/coding' +import { openaiText } from '@tanstack/ai-openai' +import { sandboxWorkspaceBackend } from '@tanstack/ai-sandbox/harness' +import { dockerSandbox } from '@tanstack/ai-sandbox-docker' + +const sandbox = await dockerSandbox({ image: 'node:22' }).create({}) +// Without `dir`, the clone goes to the workspace root, /workspace. +await sandbox.git.clone({ url: 'https://github.com/acme/web-app.git' }) + +const root = '/workspace' + +export const sandboxed = defineHarness({ + name: 'acme/sandboxed-coder', + adapter: openaiText('gpt-6.1-sol'), + plugins: () => [ + permissions({ root }), + workspaceTools({ root, backend: sandboxWorkspaceBackend(sandbox) }), + ], +}) +``` + +3. When you are done, stop the sandbox with `await sandbox.destroy()`. + +Every provider works. Get the handle from `create()`, or from `resume()` for a sandbox that runs already. See [Providers](../sandbox/providers). + +In a sandbox, some things are different: + +- Paths are POSIX paths in the sandbox, like `/workspace/src/a.ts`, also on a Windows host. +- The tools do not check where links lead. The sandbox backend has no `realpath`. +- `background: true` needs a sandbox with `capabilities.backgroundProcesses`. +- A sandbox with `capabilities.killableProcesses` off cannot stop a command after the time limit. The call returns, and the command continues in the sandbox. +- `list_files` and `grep` use `rg` in the sandbox. The `@vscode/ripgrep` binary is a host file, so the sandbox does not use it. + +## Write your own backend + +A `WorkspaceBackend` is where the tools read and write files and run commands. Write one for another place, for example a remote machine. Every path that the tools give it is absolute. + +The backend must have: + +| Method | What it does | +| --- | --- | +| `readFile(path)` | Gives the bytes of a file. Throws when there is no file. | +| `writeFile(path, data)` | Creates or replaces a file. Creates the missing parent folders. | +| `stat(path)` | Gives `{ type, size, mtimeMs }`, or `undefined` when nothing is at `path`. | +| `readdir(path)` | Gives the entries of a folder. A link has the type `'link'`. | +| `exec(command, options)` | Runs a shell command and waits. Gives `{ exitCode, stdout, stderr }`. | + +`exec` resolves with the exit code of a command that fails. It does not throw. After `options.timeoutMs`, it stops the command with every process that it started, and gives exit code 124. + +The backend can also have: + +| Member | What it does | Without it | +| --- | --- | --- | +| `shell` | `'sh'` or `'cmd'`. Sets the path style and how arguments are quoted. | `sh` quotes, and the path style of this machine. | +| `remove(path)` | Removes one file. | `patch` cannot delete or move files. | +| `realpath(path)` | Gives the path with every link resolved. | The tools do not check links. | +| `spawn(command, options)` | Starts a command in the background. | `background: true` gives an error. | + +The fastest start is to wrap `hostBackend`. This backend logs each command before it runs: + +```ts group=harness-coding-tools-backends +import { hostBackend, workspaceTools } from '@tanstack/ai-harness/plugins/coding' +import type { WorkspaceBackend } from '@tanstack/ai-harness/plugins/coding' + +const root = process.cwd() + +const logged: WorkspaceBackend = { + ...hostBackend, + exec: (command, options) => { + console.log(`$ ${command}`) + return hostBackend.exec(command, options) + }, +} + +export const loggedTools = workspaceTools({ root, backend: logged }) +``` + +Only `hostBackend` itself uses the `@vscode/ripgrep` binary. A wrapped backend uses `rg` on the `PATH`. + +## Run code after a read or a write + +Another plugin can add code that runs after the tools read or write a file. Add a `WorkspaceHooks` item: + +```ts group=harness-coding-tools-backends +import { definePlugin } from '@tanstack/ai-harness' +import { WorkspaceHooks } from '@tanstack/ai-harness/plugins/coding' + +export const generatedFiles = definePlugin({ + name: 'acme/generated-files', + setup: () => ({ + contribute: [ + WorkspaceHooks.item({ + afterRead: async (path) => + path.endsWith('.gen.ts') + ? 'This file is generated. Change its source, not this file.' + : undefined, + afterWrite: async (path) => { + console.log(`changed ${path}`) + }, + }), + ], + }), +}) +``` + +- `afterRead` gets the absolute path. The text that it gives is added after the lines that `read_file` shows. +- `afterWrite` runs after `write_file`, `edit_file`, and `patch`. +- A hook runs while the tool holds the lock on that path. A hook that throws fails the tool call. + +Add the plugin next to the tools: `plugins: () => [workspaceTools({ root }), generatedFiles]`. + +## Format each file after a write + +The model writes code that your formatter changes later, and the diff gets noisy. `formatter()` formats each file that the tools write: + +```ts group=harness-coding-tools-backends +import { defineHarness } from '@tanstack/ai-harness' +import { permissions } from '@tanstack/ai-harness/plugins' +import { formatter } from '@tanstack/ai-harness/plugins/coding' +import { openaiText } from '@tanstack/ai-openai' + +export const formatted = defineHarness({ + name: 'acme/formatted-coder', + adapter: openaiText('gpt-6.1-sol'), + plugins: () => [ + permissions({ root }), + workspaceTools({ root }), + formatter({ root }), + ], +}) +``` + +At the first write of a session, `formatter()` looks at the files in `root`. For each file, the first formatter that the project uses for that file type runs. + +For JavaScript, TypeScript, JSON, CSS, and other web files: + +| Formatter | Runs when `root` has | +| --- | --- | +| prettier | `.prettierrc*`, `prettier.config.*`, or a `prettier` field in `package.json` | +| biome | `biome.json` or `biome.jsonc` | +| oxfmt | `.oxfmtrc*`, or `oxfmt` in the `devDependencies` | + +These three run with `npx --no-install`, so the project must have them installed. + +For other languages: + +| Formatter | Files | Runs when `root` has | +| --- | --- | --- | +| ruff | `.py`, `.pyi` | `ruff.toml`, `.ruff.toml`, or `[tool.ruff]` in `pyproject.toml` | +| gofmt | `.go` | `go.mod` | +| rustfmt | `.rs` | `Cargo.toml` | + +Add your own formatter. Your formatters come before the built-in ones: + +```ts group=harness-coding-tools-backends +export const withToml = formatter({ + root, + formatters: [ + { + name: 'taplo', + extensions: ['.toml'], + command: (file) => `taplo fmt ${file}`, + }, + ], +}) +``` + +- `command` gets the absolute path, already quoted for the shell. +- `when(project)` limits a formatter to some projects. Without it, the formatter is used in every project. +- `builtins: false` turns the built-in formatters off. +- `timeoutMs` limits each run. The default is 20 seconds. +- In a sandbox, give `formatter()` the same `backend` as `workspaceTools()`. + +A formatter that fails, stops at the time limit, or is not installed does not fail the tool call. The file stays as the tool wrote it, and the plugin sends `FormatFailed`: + +```ts group=harness-coding-tools-backends +import { FormatFailed } from '@tanstack/ai-harness/plugins/coding' + +export const formatAlerts = definePlugin({ + name: 'acme/format-alerts', + setup: (ctx) => { + ctx.on(FormatFailed, (failure) => + console.warn(`${failure.path}: ${failure.message}`), + ) + }, +}) +``` + +Clients get the same event as a `harness.plugin.event` with the name `tanstack/formatter:failed`. + +## Keep tool output small + +One tool result with a full log or a big JSON file fills the context of the model. `boundToolOutput()` cuts every tool result that is too long: + +```ts group=harness-coding-tools-backends +import { boundToolOutput } from '@tanstack/ai-harness/plugins' + +export const bounded = defineHarness({ + name: 'acme/bounded-coder', + adapter: openaiText('gpt-6.1-sol'), + plugins: () => [ + workspaceTools({ root }), + boundToolOutput({ dir: '.agent/tool-output' }), + ], +}) +``` + +- A result over 2000 lines or 50 KiB keeps its first lines and gets a note. Set other limits with `maxLines` and `maxBytes`. +- With `dir`, the full output goes to a file in `dir`, and the note gives the path. Files older than `retentionDays` (default 7) are removed. +- It cuts the result of every tool, MCP tools too, in the lead turn and in agent runs. +- An error stays an error. Images and other content parts do not change. + +`read_file` and `bash` keep their output inside the default limits, so `boundToolOutput()` does not cut them again. + +## What you have now + +- The coding tools in a sandbox, with one option. +- A backend of your own for another place. +- Code that runs after each read and each write. +- Formatted files after each write, and tool results that fit the context. + +Next: let the user undo the changes of a turn with [snapshots](./snapshots). diff --git a/docs/harness/coding-tools.md b/docs/harness/coding-tools.md new file mode 100644 index 0000000000..2434f13583 --- /dev/null +++ b/docs/harness/coding-tools.md @@ -0,0 +1,259 @@ +--- +title: Give your agent coding tools +id: harness-coding-tools +order: 6 +description: "Let a harness agent read, change, and search files, run commands, and read web pages, with a question before each risky call." +keywords: + - tanstack ai + - harness + - coding agent + - workspace tools + - sandbox + - permissions + - formatter +--- + +Your agent must read and change the files of a project, search them, and run commands like `pnpm test`. It must stay in that project, and it must ask before it changes a file. `workspaceTools()` gives the model these tools for one folder. `permissions()` asks the user before each risky call. + +## 1. Install + + + +react: @tanstack/ai @tanstack/ai-harness @tanstack/ai-persistence @tanstack/ai-harness-cli @tanstack/ai-openai +vue: @tanstack/ai @tanstack/ai-harness @tanstack/ai-persistence @tanstack/ai-harness-cli @tanstack/ai-openai +solid: @tanstack/ai @tanstack/ai-harness @tanstack/ai-persistence @tanstack/ai-harness-cli @tanstack/ai-openai +svelte: @tanstack/ai @tanstack/ai-harness @tanstack/ai-persistence @tanstack/ai-harness-cli @tanstack/ai-openai +preact: @tanstack/ai @tanstack/ai-harness @tanstack/ai-persistence @tanstack/ai-harness-cli @tanstack/ai-openai +angular: @tanstack/ai @tanstack/ai-harness @tanstack/ai-persistence @tanstack/ai-harness-cli @tanstack/ai-openai +octane: @tanstack/ai @tanstack/ai-harness @tanstack/ai-persistence @tanstack/ai-harness-cli @tanstack/ai-openai +vanilla: @tanstack/ai @tanstack/ai-harness @tanstack/ai-persistence @tanstack/ai-harness-cli @tanstack/ai-openai + + + +Two optional packages make the tools better: + + + +react: @vscode/ripgrep turndown +vue: @vscode/ripgrep turndown +solid: @vscode/ripgrep turndown +svelte: @vscode/ripgrep turndown +preact: @vscode/ripgrep turndown +angular: @vscode/ripgrep turndown +octane: @vscode/ripgrep turndown +vanilla: @vscode/ripgrep turndown + + + +- `@vscode/ripgrep`: a fast `rg` binary for `list_files` and `grep`. Without it, the tools use `rg` on the `PATH`. Without `rg`, they use `git ls-files`, then a walk through the folders. +- `turndown`: `webfetch` gives HTML pages as Markdown. Without it, HTML pages come back as plain text. + +## 2. Add the tools + +```ts group=harness-coding-tools +import { defineHarness } from '@tanstack/ai-harness' +import { permissions } from '@tanstack/ai-harness/plugins' +import { workspaceTools } from '@tanstack/ai-harness/plugins/coding' +import { runCli } from '@tanstack/ai-harness-cli' +import { openaiText } from '@tanstack/ai-openai' + +const root = process.cwd() + +const coder = defineHarness({ + name: 'acme/coder', + adapter: openaiText('gpt-6.1-sol'), + plugins: () => [permissions({ root }), workspaceTools({ root })], +}) + +process.exitCode = await runCli(coder) +``` + +`@tanstack/ai-harness/plugins/coding` runs only in Node. The tools use the file system and the shell of the machine. For code that you do not trust, [run the tools in a sandbox](./coding-tools-backends#run-the-tools-in-a-sandbox). + +## 3. Run it + +1. Run `npx tsx coder.ts` in your project. +2. Ask: `find the failing test and fix it`. +3. When the agent wants to change a file or run a command, answer the question in the CLI. + +The agent reads and searches without a question. For a full terminal agent with a todo list and a model picker, see [Build a coding agent](./coding-agent). + +## The tools + +### Read and search + +These tools run without a question. + +| Tool | What it does | +| --- | --- | +| `read_file` | Reads a text file as numbered lines, at most 2000 lines or 50 KiB a page. The model reads more with `offset` and `limit`. | +| `list_files` | Lists the files in a folder. An optional glob, like `src/**/*.ts`, filters them. | +| `grep` | Searches the file contents with a regular expression. Gives `file:line: text`. | + +- `read_file` gives images (PNG, JPEG, GIF, WebP) and PDFs up to 5 MB to the model as media. +- `read_file` refuses a binary file with a short note. For a missing file, it names up to 3 files with a close name. +- `list_files` and `grep` skip the files that `.gitignore` names, the `.git` and `node_modules` folders, and links. +- `list_files` shows at most 1000 files. `grep` shows at most 200 matches. + +### Change files + +These tools ask first. + +| Tool | What it does | +| --- | --- | +| `write_file` | Creates or replaces a file. | +| `edit_file` | Replaces text in a file. The old text must be in the file once, unless the model sets `replaceAll`. | +| `patch` | Applies a patch that adds, deletes, updates, or moves files. | + +A model gets `edit_file` or `patch`, not both. Every model gets `write_file`. The `editStyle` option picks the edit tool: + +- `'auto'` (default): GPT models (ids that start with `gpt-`, or `o` and a digit, like `o3`) get `patch`. Other models get `edit_file`. +- `'edit'` or `'patch'`: every model gets that tool. + +`edit_file` keeps the line endings and the byte order mark of the file. It also finds text that is different only in Unicode form or in the spaces at the end of a line. + +To let the user undo the changes of a turn, add [snapshots](./snapshots). + +### Commands and the web + +| Tool | What it does | Asks first | +| --- | --- | --- | +| `bash` | Runs a shell command in the working folder. | Yes | +| `webfetch` | Reads a web page by its URL. | Yes | +| `websearch` | Searches the web. It is there only with a search provider. | No | + +For `bash`: + +- The model gets the exit code and the last 1000 lines (20 KiB) of the output. +- A command stops after 2 minutes, with exit code 124. Every process that it started stops too. Set another limit with `bashTimeoutMs`. The model can ask for up to 10 minutes. +- With `background: true`, the call returns at once with a job id. When the job ends, the model gets a note with the output, and a new turn starts. +- With `spillDir`, long output goes to a file in that folder, and the model gets the path. `spillDir` is relative to `root`, or absolute. +- When the session closes, the background jobs stop. + +On Windows, the default backend runs commands in `cmd.exe`. A program name like `rg` or `gradlew` comes from the `PATH` only, not from the working folder. So a cloned repo cannot put its own `rg.cmd` in place of `rg`. To run a script in the working folder, give its path: `.\gradlew build`. + +For `webfetch`: + +- It reads `http` and `https` URLs only, and follows at most 5 redirects. +- It refuses `localhost`, private network addresses, and cloud metadata addresses. It checks the DNS answer of each connection. +- HTML comes back as Markdown, or as text without `turndown`. The model can ask for `format: 'html'`. Text and JSON come back as they are. +- It reads at most 5 MB. It stops after 30 seconds, and the model can ask for up to 120 seconds. + +Change the web tools with the `web` option: + +- `web: false`: no web tools. +- `web: { allowPrivateHosts: true }`: `webfetch` can read `localhost` and your private network. Set it only for a model that is allowed to reach your network. +- `web: { search }`: adds `websearch`. + +There is no built-in search engine. Give `websearch` a `SearchProvider`: + +```ts group=harness-coding-tools +import type { SearchProvider } from '@tanstack/ai-harness/plugins/coding' +import { searchTheWeb } from './search' + +const search: SearchProvider = { + search: (query, { limit, signal }) => searchTheWeb(query, { limit, signal }), +} + +export const withSearch = workspaceTools({ root, web: { search } }) +``` + +The provider gives at most `limit` results. Each result is `{ title, url, snippet? }`. The model reads a result with `webfetch`. + +## Permissions and approvals + +Each tool tells `permissions()` what kind of call it makes: + +| Tools | `default` mode | `plan` mode | +| --- | --- | --- | +| `read_file`, `list_files`, `grep`, `websearch` | Runs | Runs | +| `write_file`, `edit_file`, `patch` | Asks | Denied | +| `bash` | Asks | Denied | +| `webfetch` | Asks | Denied | + +- `acceptEdits` mode runs the edits without a question. `bash` still asks. +- `bypass` mode runs every call. +- `read_file` asks for a `.env` file. +- `grep` skips each file that `read_file` asks about or denies, also in `plan` mode. A note at the end gives the number of skipped files. +- `list_files` still shows the names of these files. A name is not the contents. + +Let `webfetch` run without a question with one rule: + +```ts group=harness-coding-tools +export const trustTheWeb = permissions({ + root, + rules: [{ tool: 'webfetch', decision: 'allow' }], +}) +``` + +In `plan` mode, a call that asks is denied. This rule makes `webfetch` an allow, so `webfetch` runs in `plan` mode too. + +Each call also tells `permissions()` the paths or the commands that it touches. So a rule with a `resource` can allow one folder or one command: + +```ts group=harness-coding-tools +export const trustTests = permissions({ + root, + rules: [ + { tool: 'edit_file', resource: 'src/**', decision: 'allow' }, + { tool: 'bash', resource: 'pnpm test*', decision: 'allow' }, + ], +}) +``` + +- Paths match from `root`. Give `permissions()` and `workspaceTools()` the same `root`. +- A rule also matches the real path of a link, and the stricter decision wins. So a link `notes` to `.env` asks, and a `deny` for `secrets/**` also stops a link to `secrets`. +- `bash` splits a command at `&&`, `||`, `;`, `|`, and new lines. Every part must be allowed. +- A command with `$(`, a backtick, or a heredoc never gets an allow. + +[Permissions](./permissions) explains the rules, the modes, and the saved answers. + +## Outside paths and the working folder + +### A path outside `root` + +The tools refuse a path outside `root`. To ask the user instead, set `outside: 'ask'`: + +```ts group=harness-coding-tools +export const askOutside = workspaceTools({ root, outside: 'ask' }) +``` + +- A yes allows that folder, and the folders in it, until the session ends. +- `bypass` mode allows every path without a question. +- A link in `root` that leads out of `root` counts as outside. +- With `permissions()`, `grep` skips the files outside `root`, because `read_file` asks for them. A `read_file` rule for that folder, for example `{ tool: 'read_file', resource: '/shared/**', decision: 'allow' }`, lets `grep` show them. + +### A working folder for each thread + +Each thread can work in its own folder in `root`. Set the `cwd` thread setting: + +```ts group=harness-coding-tools-cwd +import { createHarnessHost, defineHarness } from '@tanstack/ai-harness' +import { workspaceTools } from '@tanstack/ai-harness/plugins/coding' +import { memoryPersistence } from '@tanstack/ai-persistence' +import { openaiText } from '@tanstack/ai-openai' + +const repos = defineHarness({ + name: 'acme/repos', + adapter: openaiText('gpt-6.1-sol'), + plugins: () => [workspaceTools({ root: '/srv/repos' })], +}) + +const host = createHarnessHost({ persistence: memoryPersistence() }) +const session = await host.open(repos, { threadId: 'fix-login' }) +await session.configure({ cwd: 'web-app' }) +``` + +- Relative paths start in `/srv/repos/web-app`, and `bash` runs there. +- The system prompt tells the model the working folder. +- Permission rules still match paths from `root`, for example `web-app/src/**`. +- `cwd` is relative to `root`. A `cwd` outside `root` counts as an outside path. + +[Store settings per thread](./thread-settings) shows how a client sets `cwd`. + +## What you have now + +- An agent that reads, searches, and changes the files of one folder, and runs commands in it. +- A question before each edit, command, and web fetch. +- A working folder for each thread, and a question before a path outside `root`. + +Next: run the same tools in a sandbox, or format each file after a write. See [Sandbox and extend the coding tools](./coding-tools-backends). diff --git a/docs/harness/compaction.md b/docs/harness/compaction.md new file mode 100644 index 0000000000..71fe4b2c80 --- /dev/null +++ b/docs/harness/compaction.md @@ -0,0 +1,150 @@ +--- +title: Compact a harness session +id: harness-compaction +order: 4 +description: "Keep a long harness session under the context limit. The compaction goes into the session log, so a session that opens again sees the same context." +keywords: + - tanstack ai + - harness + - compaction + - context window + - projectCompaction + - compactNext + - session log +--- + +A harness session can run for days. Its transcript grows until a model call passes the context limit. [Compaction](../advanced/compaction) shrinks what the model sees. In a durable harness, the result must also survive a restart. A session that opens again must see the same context as before. + +Set `durable: true`. Compaction then writes a `tanstack.compaction` record into the session log. The host folds that record into the messages with `projectCompaction`, live and when the session opens again. + +## Add compaction to a harness + +1. Create the compaction middleware with `durable: true`. +2. Add it to `middleware` in `defineHarness`. +3. Give the host `projectCompaction` in its `project` option. + +```ts group=harness-compaction +import { createHarnessHost, defineHarness } from '@tanstack/ai-harness' +import { memoryLogStore, memoryPersistence } from '@tanstack/ai-persistence' +import { openaiText } from '@tanstack/ai-openai' +import { + conversationSummarizer, + projectCompaction, + summarizeOldest, + withCompaction, +} from '@tanstack/ai-compaction' + +const compaction = withCompaction({ + maxTokens: 180_000, + contextWindow: 200_000, + countTokens: 'usage', + durable: true, + strategy: summarizeOldest({ + cut: 'turn', + keepRecentTokens: 8_000, + summarize: conversationSummarizer({ adapter: openaiText('gpt-6.1-sol') }), + }), +}) + +const assistant = defineHarness({ + name: 'acme/assistant', + adapter: openaiText('gpt-6.1-sol'), + middleware: [compaction], +}) + +const { runs, metadata } = memoryPersistence().stores +const host = createHarnessHost({ + persistence: { stores: { log: memoryLogStore(), runs, metadata } }, + project: { version: 'v1', record: projectCompaction }, +}) +``` + +- Compaction runs before a model call when the count passes `maxTokens`. If the usage of the last model call of a turn passed `maxTokens` or `contextWindow`, compaction also runs right after that call. This check needs `countTokens: 'usage'`. +- The log only grows, so the original messages stay in the store. The fold that the model gets starts with the summary. +- A host without a log has no place for the record. Compaction then keeps its result in the `metadata` store, the same as in a plain `chat()` call. It also skips the check after the turn. + +## Prepare the summary in the background + +A summary call takes time. When compaction runs at `maxTokens`, the user waits for that call before the model answers. Set `background` to prepare the summary earlier, while the session goes on: + +```ts group=harness-compaction +const earlyCompaction = withCompaction({ + maxTokens: 180_000, + contextWindow: 200_000, + countTokens: 'usage', + durable: true, + background: { atTokens: 150_000 }, + strategy: summarizeOldest({ + cut: 'turn', + keepRecentTokens: 8_000, + summarize: conversationSummarizer({ adapter: openaiText('gpt-6.1-sol') }), + }), +}) +``` + +- `atTokens` must be below `maxTokens`. If it is not, `withCompaction` throws. +- When the count at a model call passes `atTokens`, the summary starts beside the turn. The turn does not wait for it. +- The summary applies at the first model call of the next run, never in the middle of a run. On a durable host, it goes into the log as a `tanstack.compaction` record with `reason: 'background'`. +- A model call over `maxTokens` while the summary runs waits for it and applies it. It compacts inline only when the messages are still over `maxTokens`. +- If you cancel the run during that wait, the wait stops. The summary continues and applies at the next run. + +Sometimes a summary does not apply: + +- A newer compaction cut past the message that the summary keeps. The summary is dropped. `onCompact` and `compaction:ended` report it with `stale: true`. +- The summary call failed. Nothing applies. `onCompact` and `compaction:ended` report it with `error`. The next model call past `atTokens` starts a new summary. + +## Fold your own records too + +If your host already folds records of its own, call `projectCompaction` first: + +```ts group=harness-compaction +import type { ProjectOptions } from '@tanstack/ai-harness' + +const project: ProjectOptions = { + version: 'v2', + record: (input) => { + const compacted = projectCompaction(input) + if (compacted) return compacted + const { record, messages } = input + if (record.type === 'app.signal' && typeof record.text === 'string') { + return [...messages, { role: 'user', content: record.text }] + } + return undefined + }, +} +``` + +Change `version` when you change the function. A fold checkpoint with another version is ignored. + +## Retry after a context overflow + +A model call can still fail with a context overflow. Call `compactNext` in `turn.onModelError`, then return `'retry'`. The retried call compacts first and writes the record. + +```ts group=harness-compaction +import { isContextOverflow } from '@tanstack/ai' + +const recovering = defineHarness({ + name: 'acme/recovering', + adapter: openaiText('gpt-6.1-sol'), + middleware: [compaction], + turn: { + onModelError: ({ session, error, retries }) => { + if (retries > 0 || !isContextOverflow({ error: error.message })) { + return undefined + } + compaction.compactNext(session.threadId) + return 'retry' + }, + }, +}) +``` + +- `retries > 0` stops after one retry. A second overflow fails the turn with its error. +- `compactNext` also works with `auto: false`. With `auto: false`, compaction runs only for this recovery and for the `contextWindow` check. + +## What you have now + +- A session that compacts before it passes the context limit, counted with the real usage of each call. +- The same context after a restart, because the compaction is in the log. +- With `background`, a summary that is ready before the session reaches `maxTokens`. +- One retry after an overflow, with a compaction first. diff --git a/docs/harness/connect.md b/docs/harness/connect.md new file mode 100644 index 0000000000..1728289000 --- /dev/null +++ b/docs/harness/connect.md @@ -0,0 +1,351 @@ +--- +title: Connect clients to a harness +id: harness-connect +order: 2 +description: "Serve a harness session over HTTP, SSE, or WebSocket. Talk to it from a web app, an editor over ACP, or another chat() call." +keywords: + - tanstack ai + - harness + - AG-UI + - ACP + - websocket + - harnessText +--- + +Your harness runs on a server, but the people who use it are in a browser, an editor, or another agent. `createHarnessHandler` serves a session over HTTP. A web app talks to it with `createHarnessClient`, an editor with ACP, and another `chat()` call with `harnessText`. + +## Serve the session over HTTP + +Mount one fetch handler on a route. `authorize` is required, so no endpoint is open by accident. + +```ts group=harness-connect +import { defineHarness, createHarnessHandler, createHarnessHost } from '@tanstack/ai-harness' +import { memoryPersistence } from '@tanstack/ai-persistence' +import { openaiText } from '@tanstack/ai-openai' + +export const assistant = defineHarness({ + name: 'acme/assistant', + adapter: openaiText('gpt-5.6'), +}) + +const host = createHarnessHost({ persistence: memoryPersistence() }) + +export const handler = createHarnessHandler({ + host, + harness: assistant, + authorize: (request) => + request.headers.get('authorization') === `Bearer ${process.env.HARNESS_TOKEN}` + ? { id: 'user-1' } + : null, + canAccess: (principal, threadId) => threadId.startsWith(principal.id), +}) +``` + +The handler answers these paths under your route: + +- `GET capabilities`: the AG-UI capabilities, with the agents in `expose.agents`. +- `POST run`: standard AG-UI. One request runs one prompt and streams it as SSE. The prompt keeps every content part of the last user message, for example an image. `useChat` works with it. See [Use it from useChat](#use-it-from-usechat). +- `GET run?threadId=`: the saved messages, the running turn, and the waiting approvals, for a `useChat` that loads the thread after a reload. +- `GET run?runId=`: the turn with that run id as SSE, from its first event. A reloaded `useChat` joins a running turn with it. +- `GET events?threadId=`: every event of the session as SSE. Each event id is a cursor, so a reconnect with `Last-Event-ID` continues where it stopped. +- `POST control`: send `{ threadId, input }`, for example `{ op: 'prompt', message }`. You get a receipt back. Agents, settings, config keys, and commands need `expose`. See [Choose what clients can change](#choose-what-clients-can-change). +- `GET snapshot?threadId=`: the status, running operations, and waiting approvals. + +## Choose what clients can change + +A client can send a prompt and answer questions. Other inputs can change how the session works, so a client can use only what the harness exposes. Nothing is exposed by default. + +Name what a client can use in `expose`: + +```ts group=harness-connect +import { permissions, todos } from '@tanstack/ai-harness/plugins' + +export const team = defineHarness({ + name: 'acme/team-assistant', + adapter: openaiText('gpt-6.1-sol'), + plugins: () => [permissions(), todos()], + expose: { settings: ['instructions'], commands: ['todos'] }, +}) +``` + +| Field | What a client can do | +| --- | --- | +| `agents` | Run these agents, for example with `client.agents..start(input)`. | +| `settings` | Change these [thread settings](./thread-settings) with `client.configure(settings)`. | +| `config` | Set these plugin config keys with `client.setConfig(key, value)`. | +| `commands` | Run these commands with `client.command(name)`, or with `/name` in a [session view](./custom-ui) on a client. | + +- Any other input of these kinds gets `{ status: 'rejected', reason: 'not_exposed' }`. +- `GET describe` and `client.describe()` list only the exposed commands and config keys. A UI built from them shows only what a client can use. +- Server code that calls the session is not limited. `session.setConfig()` and `session.command()` work for every key and command, and `session.describe()` lists all of them. +- `defineHarness` checks the names in `agents`. Plugins add their config keys and commands when a session opens, so it does not check those names. + +Do not expose the `mode` of `permissions()` to clients that you do not trust. With `bypass`, every tool call runs with no question. See [Keep the mode on the server](./permissions#keep-the-mode-on-the-server). + +## Talk to it from a web app + +`createHarnessClient` wraps those endpoints. Import the harness with `import type`, so no server code reaches the browser. + +```ts group=harness-connect-client +import { createHarnessClient } from '@tanstack/ai-harness/client' +import type { assistant } from './harness' + +const client = createHarnessClient({ + url: '/api/harness', + threadId: 'user-1-thread', + headers: { authorization: 'Bearer my-token' }, +}) + +await client.prompt('Summarize my open tickets.') + +for await (const entry of client.events()) { + if (entry.event.type === 'TEXT_MESSAGE_CONTENT') { + console.log(entry.event.delta) + } +} +``` + +`events()` reconnects after a network error and continues from the last cursor. A second tab, a phone, or a reload all see the same session. + +To upload files and show the media that your agents make, read [Send and show media](./media). + +## Send a message to a running agent + +A background agent runs for minutes, and the user wants to correct it from the page, or ask for one more change after it ends. The browser sends the run a message with `client.sendToAgent`. + +1. On the server, list the agent in `expose.agents`. A client can start and message only those agents: + +```ts group=harness-connect-agents +import { defineAgent } from '@tanstack/ai' +import { createHarnessHandler, createHarnessHost, defineHarness } from '@tanstack/ai-harness' +import { memoryPersistence } from '@tanstack/ai-persistence' +import { openaiText } from '@tanstack/ai-openai' + +const drafter = defineAgent({ + name: 'drafter', + description: 'Drafts a blog post', + run: (ctx) => ctx.chat({ adapter: openaiText('gpt-6.1-sol'), stream: false }), +}) + +export const studio = defineHarness({ + name: 'acme/studio', + adapter: openaiText('gpt-6.1-sol'), + agents: [drafter], + expose: { agents: ['drafter'] }, +}) + +export const handler = createHarnessHandler({ + host: createHarnessHost({ persistence: memoryPersistence() }), + harness: studio, + authorize: (request) => + request.headers.get('authorization') === `Bearer ${process.env.HARNESS_TOKEN}` + ? { id: 'user-1' } + : null, + canAccess: (principal, threadId) => threadId.startsWith(principal.id), +}) +``` + +2. In the browser, send the message to the id of the run: + +```ts group=harness-connect-agents-client +import { createHarnessClient } from '@tanstack/ai-harness/client' +import type { studio } from './harness' + +const client = createHarnessClient({ + url: '/api/harness', + threadId: 'user-1-thread', + headers: { authorization: 'Bearer my-token' }, +}) + +export async function correct(operationId: string) { + const receipt = await client.sendToAgent(operationId, 'Cite papers, not blogs.') + if (receipt.status === 'rejected') { + console.warn(receipt.reason) + } +} + +export function addTitle(operationId: string) { + return client.sendToAgent(operationId, 'Add a title.', { mode: 'followUp' }) +} +``` + +The id of the run is the `operationId` in the receipt of `client.agents.drafter.start()`. A [session view](./custom-ui#3-act-on-the-session) also lists the running agents with their ids. + +- `mode: 'steer'` (the default): the message joins the next model call of the run. +- `mode: 'followUp'`: the agent runs again after the run ends. `receipt.operationId` is the id of the new run. + +`sendToAgent` posts an `agentMessage` input to `POST control`. It runs as the user that `authorize` returned. The input looks like this: + +```json +{ "op": "agentMessage", "operationId": "op-agent-m1x2-3", "message": "Add a title.", "mode": "followUp" } +``` + +The session rejects the message with a `reason` in these cases: + +- `not_running`: the session does not know the run. Another host can run it, or the id is wrong. +- `not_exposed`: the agent of the run is not in `expose.agents`. + +For how steers and follow-ups run on the server, see [Message a running agent](./subagents#message-a-running-agent). + +## Use it from useChat + +Your app already uses `useChat`. You want the same chat screen on a harness, with approvals, tools that run in the browser, and the thread after a reload. Point `useChat` at the `run` path of the handler. + +1. Put the tools in a file that the server and the browser both import: + +```ts group=harness-connect-tools +import { toolDefinition } from '@tanstack/ai' +import { z } from 'zod' + +export const deploy = toolDefinition({ + name: 'deploy', + description: 'Deploy the app. A person must approve it.', + needsApproval: true, + inputSchema: z.object({ env: z.string() }), + outputSchema: z.object({ ok: z.boolean() }), +}) +``` + +2. On the server, give the harness the tool with its implementation: + +```ts group=harness-connect-tools-server +import { defineHarness } from '@tanstack/ai-harness' +import { openaiText } from '@tanstack/ai-openai' +import { deploy } from './tools' + +export const deployer = defineHarness({ + name: 'acme/deployer', + adapter: openaiText('gpt-5.6'), + tools: [ + deploy.server(async () => { + // Your deploy code runs here. + return { ok: true } + }), + ], +}) +``` + +Serve `deployer` with `createHarnessHandler`, as in [Serve the session over HTTP](#serve-the-session-over-http). + +3. In the browser, give `useChat` the same tool definitions: + +```tsx group=harness-connect-tools-client +import { fetchServerSentEvents, useChat } from '@tanstack/ai-react' +import { deploy } from './tools' + +export function Chat() { + const { messages, interrupts, sendMessage } = useChat({ + threadId: 'user-1-thread', + connection: fetchServerSentEvents('/api/harness/run'), + tools: [deploy.client()] as const, + // The input's context: tools and plugins read it. + forwardedProps: { page: '/releases' }, + }) + + return ( + <> +

{messages.length} messages

+ {interrupts.map((interrupt) => + interrupt.kind === 'tool-approval' ? ( + + ) : null, + )} + + + ) +} +``` + +The approval reaches the harness, and the tool runs on the server. What else `POST run` does for `useChat`: + +- Each turn runs as the `runId` of its request. A retry with the same `runId` runs once. The same `runId` with another message gets `409`. +- A tool with only a `.client()` implementation runs in the browser, and the same turn gets its result. +- A question from `ctx.session.ask` arrives on the stream of the turn that waits for it. +- The `forwardedProps` of the request are the input's [context](./inputs#send-context-with-a-message). +- With `persistence: true`, `useChat` loads the thread from `GET run?threadId=`. +- If a turn still runs after a reload, `useChat` joins it with `GET run?runId=`. The turn shows from its first event, and its answer continues to stream. +- With more than one server, send `GET run?runId=` to the server that runs the turn. Only that server has the live events of the turn. + +To keep the thread after a reload, set `persistence: true` in `useChat`: + +```tsx group=harness-connect-tools-client +export function ReloadableChat() { + const { messages } = useChat({ + threadId: 'user-1-thread', + connection: fetchServerSentEvents('/api/harness/run'), + tools: [deploy.client()] as const, + persistence: true, + }) + return

{messages.length} messages

+} +``` + +The join has these rules: + +- `canAccess` decides who can join, the same as for every other path. +- The turn must run on the host that answers the request. On another host, the join gets `404`. +- A turn that ended gives the events that the session still keeps, and then the stream ends. + +## Use a WebSocket + +For one connection that carries events and inputs, authorize the upgrade, then hand the socket to `handleHarnessSocket`: + +```ts group=harness-connect +import { handleHarnessSocket } from '@tanstack/ai-harness' +import type { WebSocketLike } from '@tanstack/ai' + +export function onUpgrade(socket: WebSocketLike) { + handleHarnessSocket({ host, harness: assistant, socket, principal: { id: 'user-1' } }) +} +``` + +The client sends `{ type: 'harness.subscribe', threadId }` first. After that, each `harness.input` frame gets a `harness.receipt`, and every event arrives as a `harness.event` frame with its cursor. + +## Use it from an editor (ACP) + +Editors such as Zed start agents as a process and talk ACP over stdio. `serveAcp` makes the harness an ACP v2 agent: + +```ts group=harness-connect +import { serveAcp } from '@tanstack/ai-acp/agent' + +serveAcp({ host, harness: assistant }) +``` + +Tool approvals become permission requests in the editor. Image blocks, audio blocks, and embedded files in a prompt go to the media store of the session. ACP v2 is still a draft, so this API is experimental. + +The editor does not show harness questions, for example the questions of `permissions()`. A tool call that asks a question waits, and the editor cannot answer it. See [Known limits](./permissions#known-limits). + +## Use it as the model of another chat + +`harnessText` turns a harness into a text adapter. The outer chat sends a message, and the harness runs a full turn with its own tools, plugins, and agents: + +```ts group=harness-connect +import { chat } from '@tanstack/ai' +import { harnessText } from '@tanstack/ai-harness' + +const stream = chat({ + adapter: harnessText(assistant, { host }), + messages: [{ role: 'user', content: 'Fix the failing test.' }], + threadId: 'outer-thread', +}) +``` + +Each outer thread gets its own inner session, so the harness keeps its own history. The inner session gets every content part of the last user message, so images and files go through too. + +If a plugin picks the model of the harness, the harness has no adapter to read its input kinds from. Pass them with `harnessText(assistant, { host, inputModalities: ['text', 'image'] })`. + +## What you have now + +- One handler that serves standard AG-UI and the session stream. +- Clients that can change only what you expose. +- A `useChat` screen on the harness, with approvals, tools that run in the browser, and a reload that joins the running turn. +- A typed web client that reconnects and resumes from a cursor. +- Messages from the browser to a running agent. +- A WebSocket, an ACP agent for editors, and a harness you can call from `chat()`. + +Next: run the same harness in a terminal with the [CLI](./cli). diff --git a/docs/harness/custom-ui.md b/docs/harness/custom-ui.md new file mode 100644 index 0000000000..4b2d81bcde --- /dev/null +++ b/docs/harness/custom-ui.md @@ -0,0 +1,477 @@ +--- +title: Build your own UI +id: harness-custom-ui +order: 13 +description: "Show a harness session in your own terminal screen or web app. One live store holds the messages, tool calls, approvals, and plugin state, and works with any UI library." +keywords: + - tanstack ai + - harness + - session view + - tanstack store + - ink + - custom ui +--- + +You want your own screen for your harness: your own layout in a terminal, or a page in your web app. `session.events()` gives you raw AG-UI events. To show them, you must join text deltas, follow each tool call, and keep a list of open approvals. That is a reducer, and you must keep it correct. + +`createSessionView` does that work. It keeps one live state in a TanStack Store, and it gives you actions and typed events. You show the state with any UI library. + +At the end of this page you have a terminal screen in about 20 lines, and the same view in a web app. + +## Install + +Install the harness and the TanStack Store adapter for your UI library. With no UI library, the harness is enough. + + + +react: @tanstack/ai-harness @tanstack/react-store +vue: @tanstack/ai-harness @tanstack/vue-store +solid: @tanstack/ai-harness @tanstack/solid-store +svelte: @tanstack/ai-harness @tanstack/svelte-store +preact: @tanstack/ai-harness @tanstack/preact-store +angular: @tanstack/ai-harness @tanstack/angular-store +vanilla: @tanstack/ai-harness +octane: @tanstack/ai-harness + + + +The terminal screen on this page uses Ink, which needs Node 22 or later. Install Ink and React too: + + + +react: ink react + + + +## 1. Make a view + +Open a session, then give it to `createSessionView`: + +```ts group=harness-custom-ui +import { createHarnessHost, defineHarness } from '@tanstack/ai-harness' +import { createSessionView } from '@tanstack/ai-harness/view' +import { memoryPersistence } from '@tanstack/ai-persistence' +import { openaiText } from '@tanstack/ai-openai' + +const assistant = defineHarness({ + name: 'acme/assistant', + adapter: openaiText('gpt-5.6'), +}) + +const host = createHarnessHost({ persistence: memoryPersistence() }) +const session = await host.open(assistant, { threadId: 'thread-1' }) + +const view = createSessionView(session) +await view.ready +``` + +`view.ready` resolves when the view has the saved messages, the status, and the commands and settings of the session. If you open a thread again, its history is in the view immediately. + +## 2. Show the state + +`view.store` is a TanStack Store. `useSelector` reads one part of the state, and your component updates only when that part changes. This Ink screen shows the messages and the status: + +```tsx ignore +import { Box, Text, render } from 'ink' +import { useSelector } from '@tanstack/react-store' +import type { SessionView } from '@tanstack/ai-harness/view' + +function Screen({ view }: { view: SessionView }) { + const messages = useSelector(view.store, (state) => state.messages) + const status = useSelector(view.store, (state) => state.status) + return ( + + {messages.map((message) => + message.role === 'assistant' ? ( + + {message.parts.map((part) => (part.type === 'text' ? part.text : '')).join('')} + + ) : ( + + {message.role === 'user' ? `> ${message.text}` : message.text} + + ), + )} + {status} + + ) +} + +render() +await view.send('Summarize the README.') +``` + +Put the code of steps 1 and 2 in one `.tsx` file, then run it. The screen shows your prompt, then the answer while it streams, then the status `idle`. Press Ctrl+C to quit. + +A full Ink screen, with approvals, questions, sign-ins, and child agents, is in [`examples/harness-cli/src/tui.tsx`](https://github.com/TanStack/ai/blob/main/examples/harness-cli/src/tui.tsx). It runs from [`runCli({ ui })`](./cli#run-your-own-screen). + +Each TanStack Store adapter reads the store the same way. Vue, Solid, Svelte, and Preact have `useSelector`. Angular has `injectSelector`. With no UI library, subscribe to the store: + +```ts group=harness-custom-ui +const subscription = view.store.subscribe((state) => { + console.log(`${state.status}: ${state.messages.length} messages`) +}) +``` + +Call `subscription.unsubscribe()` to stop. + +### What the state holds + +The conversation: + +- `messages`: the lines to show. Each message has a `role`: `user`, `assistant`, or `notice`. +- `status`: `idle`, `running`, or `requires_action`. +- `connection`: `open`, `reconnecting`, or `closed`. + +What waits for the user: + +- `approvals`: tool calls that wait for a yes or a no. +- `clientTools`: tool calls that run in your UI and wait for their output. See [Run a tool in the browser](#run-a-tool-in-the-browser). +- `questions`: questions from a command or a plugin. +- `signIns`: connectors that need a sign-in, with a `url` and a `userCode` when the connector gives them. + +The session: + +- `threadId`: the id of the conversation. +- `agents`: the background agents that run now. Each one has an `id`, a `name`, and `send`. +- `queuedTurns`: the number of messages that wait for their turn. +- `commands`, `config`, and `tools`: what the session has, for a help screen or a settings panel. +- `plugins`: the saved state of each plugin, by plugin name. + +An assistant message has `parts`. Each part is one of these: + +- `text` or `reasoning`: the text, which grows while it streams. +- `tool-call`: a tool call with its `name`, `args`, and `status` (`running`, `done`, `failed`, or `needs-approval`). +- `agent`: a child agent, with its own `parts`. +- `media`: a file that an agent made. See [Show media](#show-media). + +A notice has a `kind`: + +- `info`: the session continues a turn that a crash stopped, or a [reset](./fork-and-reset#reset-with-a-handoff-note) started a fresh context. +- `error`: a turn or an action failed. +- `rejected`: the session did not accept a message. +- `command`: the text result of a command. +- `ui`: a line that you added with `view.notice(text)`. + +## 3. Act on the session + +Start and stop work: + +- `view.send(text)`: sends a prompt. While a turn runs, the text steers that turn. Text that starts with `/` runs a command, for example `/goal all tests pass`. +- `view.command(name, input)`: runs a command, for example from a button. +- `view.setConfig(key, value)`: changes a setting, for example `view.setConfig('mode', 'plan')`. +- `view.cancel()`: cancels the running turn. +- `view.dispose()`: stops the view when your screen closes. After that, each action throws an error. + +On a view with a [`HarnessClient`](#a-ui-in-the-browser), two actions need `expose` in the harness: + +- `view.command`, and a `/command` in `view.send`, run only the commands in `expose.commands`. +- `view.setConfig` sets only the keys in `expose.config`. + +For any other command or key, the view adds a `rejected` notice with the text `Not accepted: not_exposed`. It also calls your `'error'` handler with that text. On a client, `commands` and `config` in the state list only the exposed items. + +See [Choose what clients can change](./connect#choose-what-clients-can-change). + +Answer what waits: + +- `approval.approve()` and `approval.reject()`: answer one approval. `view.approve(id)` and `view.reject(id)` do the same by id. +- `view.approveAll()` and `view.rejectAll()`: answer all open approvals. +- `call.resolve(output)` and `call.fail(message)`: answer one client tool call. +- `question.answer(value)`: answers a question. For the answers to a permission question, see [Answer a question](./permissions#answer-a-question). + +If a turn waits for more than one approval or client tool, the turn continues after you answer all of them. + +Message a background agent: + +- `agent.send(text, mode)`: sends a message to one item of `state.agents`. `mode` is `'steer'` (the default) or `'followUp'`. See [Message a running agent](./subagents#message-a-running-agent). + +```ts group=harness-custom-ui +const [agent] = view.store.get().agents +if (agent) { + await agent.send('Add a title.', 'followUp') +} +``` + +`send` needs a source that can message agents. A session and a `HarnessClient` both can. If you give `createSessionView` your own `SessionViewSource` without `sendToAgent`, `send` rejects with an error. + +## 4. Listen for events + +Some items need attention when they arrive. `view.on` calls your handler for each one. When a tool call waits for approval, this handler rings the terminal bell and adds a line to the screen: + +```ts group=harness-custom-ui +view.on('approval', (approval) => { + process.stdout.write('\u0007') + view.notice(`${approval.tool} waits for your answer.`) +}) +``` + +Items that wait for the user: + +- `'approval'`: an approval. Call `approve()` or `reject()` on it. +- `'clientTool'`: a client tool call. Call `resolve(output)` or `fail(message)` on it. +- `'question'`: a question with a `message`, a `schema`, and `answer(value)`. +- `'signIn'`: a connector that needs a sign-in, with `connector`, `url`, and `userCode`. + +Progress: + +- `'toolCall'`: `{ id, name }`, when a tool call starts. +- `'agent'`: `{ id, name, status }`, when a child agent starts, finishes, or fails. +- `'error'`: the error text. The view also adds it to `messages` as a notice. +- `'turnEnd'`: `{ operationId }`, when a chat turn ends. + +The view calls your handler after the item is in the state, so the handler can read `view.store.get()` and find the item. `view.on` returns a function that removes the handler. + +## Show media + +A user sends a screenshot, and an agent draws an image. The view gives you each file as a media part, with a URL that works in ``, `