Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
43 commits
Select commit Hold shift + click to select a range
d7da6b9
feat(ai-harness): add the AG-UI bridge subpath
jherr Sep 26, 2026
c82a5ad
fix(ai-harness): coalesce multi-turn operations for strict AG-UI clients
jherr Sep 26, 2026
de730ec
feat(examples/agent-dashboard): live session view + approval queue
jherr Sep 26, 2026
b6f39e7
feat(examples/agent-dashboard): control plane — history, spend, config
jherr Sep 26, 2026
74e4597
feat(examples/agent-dashboard): meta-chat over live agent state
jherr Sep 26, 2026
ab9923a
docs: add STATUS.md summarizing the agent-dashboard build
jherr Sep 26, 2026
be6e9d0
feat(examples/agent-dashboard): teams reframe (Phase 1) + Alem protoc…
jherr Sep 26, 2026
0041255
docs: update STATUS.md with the teams reframe (Phase 1)
jherr Sep 26, 2026
d410e78
feat(ai-harness): out-of-band tool invocation + tool visibility
jherr Sep 27, 2026
2078d29
feat(examples/agent-dashboard): Phase 2 — tool registry + injection
jherr Sep 27, 2026
cb726c5
docs: update STATUS.md with teams Phase 2 (tool registry + injection)
jherr Sep 27, 2026
0eda69e
feat(ai-harness): systemPreamble on the prompt op + in-band tool thre…
jherr Sep 27, 2026
fcc699b
feat(examples/agent-dashboard): teams Phase 3 — system tools, channel…
jherr Sep 27, 2026
089617a
docs: update STATUS.md with teams Phase 3 (system tools, channels, po…
jherr Sep 27, 2026
7cbb240
feat(examples/agent-dashboard): persist agent + team state on the server
jherr Sep 27, 2026
7b35c79
docs: update STATUS.md with server-side persistence
jherr Sep 27, 2026
51cf847
fix(examples/agent-dashboard): resolved approval no longer resurfaces…
jherr Sep 27, 2026
e8d40f3
feat(examples/agent-dashboard): move demo controls into a devtools panel
jherr Sep 27, 2026
7956eaa
docs: update STATUS.md with demo controls devtools panel
jherr Sep 27, 2026
1036f82
feat(examples/agent-dashboard): Reddit pod — real service + real LLM
jherr Sep 27, 2026
993e28c
docs: update STATUS.md with the Reddit pod (real service + real LLM)
jherr Sep 27, 2026
9b0fd53
feat(examples/agent-dashboard): product team-composition UI + default…
jherr Sep 27, 2026
8262872
docs: update STATUS.md with the product team-composition UI
jherr Sep 27, 2026
830c514
fix(examples/agent-dashboard): DM membership + surface pod memory on …
jherr Sep 27, 2026
b8e5381
docs: update STATUS.md with the DM membership + team-page memory fix
jherr Sep 27, 2026
69ba301
feat(examples/agent-dashboard): collapsible JSON trees + markdown mes…
jherr Sep 27, 2026
1c54522
ci: apply automated fixes
autofix-ci[bot] Sep 28, 2026
97f51c0
feat(examples/agent-dashboard): rehydrate timeline from persisted mes…
jherr Sep 28, 2026
c1d0169
Merge remote-tracking branch 'origin/main' into feat/agent-dashboard
jherr Oct 3, 2026
b2a3589
feat(agent-dashboard): add live trace waterfall
jherr Oct 3, 2026
a21e35c
feat(agent-dashboard): add steer stop and ask-user controls
jherr Oct 3, 2026
a6a610f
feat(agent-dashboard): add product schedule controls
jherr Oct 3, 2026
eebd878
feat(agent-dashboard): add webhook delivery management
jherr Oct 3, 2026
9abd895
feat(agent-dashboard): estimate model spend in dollars
jherr Oct 3, 2026
0f278a0
feat(agent-dashboard): expose dashboard MCP endpoint
jherr Oct 3, 2026
d70fd25
docs(agent-dashboard): inventory dashboard capabilities
jherr Oct 3, 2026
f78e7ce
fix: preserve branch behavior after main update
jherr Oct 3, 2026
c7041c6
Merge remote-tracking branch 'origin/feat/harness-p14-mcp-server' int…
jherr Oct 4, 2026
0d7d46a
Merge remote-tracking branch 'origin/feat/harness-p14-mcp-server' int…
jherr Oct 4, 2026
8e23607
feat(agent-dashboard): restyle the dashboard with the TanStack design…
jherr Oct 4, 2026
146a467
ci: apply automated fixes
autofix-ci[bot] Oct 4, 2026
d3b1942
Merge remote-tracking branch 'origin/feat/harness-p14-mcp-server' int…
jherr Oct 5, 2026
b38787f
chore: keep ai-vertex on the zod 4.3.6 copy of @google/genai
jherr Oct 5, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions .changeset/harness-ag-ui-bridge.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
---
'@tanstack/ai-harness': minor
---

New subpath `@tanstack/ai-harness/ag-ui`: the AG-UI bridge. A harness session already streams AG-UI events, so this is a thin, documented seam for AG-UI clients.

- `sessionEventsToAgUi(events, options?)` normalizes a session's `SessionEvent` stream into a pure AG-UI event stream: run usage is surfaced under `metadata.tanstack.usage` (`normalizeUsage` handles both the AG-UI spec array and the TanStack prompt/completion shapes), interrupt/approval waits stay as `RUN_FINISHED` with `outcome.type === 'interrupt'`, and subagent attribution is preserved. It can optionally drop the harness-native `CUSTOM` control events or emit interim `tanstack.spend` `CUSTOM` ticks for a live spend meter.
- `createAgUiHandler(options)` is a `fetch` handler a bare `@ag-ui/client` `HttpAgent` can point at. `POST` a `RunAgentInput` to run a prompt, or one with `resume` entries to answer the last turn's interrupts. It streams AG-UI events as SSE via `@ag-ui/encoder`'s `EventEncoder` (protobuf framing when the client's `Accept` prefers it). The stream is strict AG-UI by default (harness control events dropped so it begins with `RUN_STARTED`); the harness-native approval path and the relay dashboard are unchanged.
- `operationToAgUiRun(entries, options)` coalesces one harness operation — which may span several model turns (each a `RUN_STARTED`/`RUN_FINISHED` pair) — into a single valid AG-UI run: one `RUN_STARTED` (synthesized when a resumed run leads with a tool result), one terminal `RUN_FINISHED` carrying the operation outcome and summed usage, and `RUN_FINISHED.usage` conformed to the AG-UI `SpecTokenUsage[]` shape. This is what makes the stream consumable by a strict `@ag-ui/client` without "run already finished" / usage-shape errors. `createAgUiHandler` uses it per request.

The AG-UI wire is pinned: built and tested against `@ag-ui/core@1.0.0`, `@ag-ui/encoder@1.0.0`, and `@ag-ui/client@1.0.0`.

Generated with Claude Code.
17 changes: 17 additions & 0 deletions .changeset/harness-system-preamble.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
---
'@tanstack/ai-harness': minor
---

Add an additive `systemPreamble` to the `prompt` input op. When present, its
strings are prepended (ahead of the harness's own `systemPrompts`) as
system/developer messages for that one run — a place for a trigger to attach
per-run context (e.g. operational memory) without the agent author doing
anything.

Also make the server-tool execution `context` always carry the live `threadId`
and `runId` (merged over the harness's static `context`) so a tool invoked
in-band (from a model turn) can resolve the calling thread, matching the flat
`{ threadId, runId, signal }` already passed to out-of-band `{ op: 'tool' }`
invocations.

Generated with Claude Code.
12 changes: 12 additions & 0 deletions .changeset/harness-tool-op.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
---
'@tanstack/ai-harness': minor
---

Out-of-band tool invocation: a new `{ op: 'tool', name, args?, meta? }` session input runs one registered tool with no model turn. This is the provisional harness surface the agent-dashboard "injection" work builds on (schedules, run-now, webhooks), following the `@tanstack/ai-harness/ag-ui` precedent of shipping dashboard-facing seams here.

- `session.tool(name, args?, meta?)` (and the `{ op: 'tool' }` client input via `applyInput`) resolves the tool from the harness + session-plugin tools, validates `args` against its input schema, and invokes its server executor directly — modeled on the existing `command` op. The tool's lifecycle is published into the session feed as a normal AG-UI run (`RUN_STARTED`, `TOOL_CALL_START`/`ARGS`/`END`, `TOOL_CALL_RESULT`, `RUN_FINISHED`), so watchers render and persist it exactly like a tool call the model made. `meta` (e.g. an injection trigger) is echoed onto the run as a `tanstack.injection` `CUSTOM` event. Injected tools are fire-and-forget (no approval interrupt, like commands).
- Tool **visibility**: a new `toolVisibility?: Record<string, 'public' | 'private'>` on `defineHarness`. Tools default to `private` — only tools named `public` may be invoked out-of-band; unknown or private tools are rejected (`unknown_tool` / `not_public`). Visibility is reported per tool in the capabilities document (`capabilitiesOf().tools.items[].visibility`).

Additive and backward compatible: existing ops, tools, and streams are unchanged.

Generated with Claude Code.
155 changes: 155 additions & 0 deletions PODS-PROTOCOL-PROPOSAL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,155 @@
# Pods/Teams — harness protocol proposal (for @AlemTuzlak)

**Status:** proposal · **Date:** 2026-09-26 · **Author:** Jack + Claude Code
**Basis:** `~/Downloads/pods-design-doc.md` (v2) and `pods-implementation-plan.md`

This is a **discussion doc, not a change**. No `packages/ai-harness` or
`packages/ai-dashboard` (relay) code has been touched. Phase 1 of the teams
reframe ships entirely in `examples/agent-dashboard` (dashboard-local
collections). Everything below is the net-new **harness/relay protocol surface**
that Phases 2+ need, grounded in what the current code does and doesn't do, so we
can agree on the shape before writing any of it.

The design doc says "pod"; the dashboard ships the noun **"team"**. Same concept.

## Why this doc exists

Three capabilities the design depends on are not expressible in the harness as it
stands today. Each is core protocol surface, so it needs your review before
implementation. Findings were verified against the current tree (head at the top
of the harness PR stack).

---

## 1. Out-of-band tool invocation (`injectToolCall`)

**Design need.** The dashboard's injection model (design §5.4) calls a single
named tool deterministically, with **no model turn and zero tokens**
(`injectToolCall(pod, tool, args)`), and posts the result to the stream. This is
the _default_ trigger path (timers, webhooks, "run now").

**Current reality.** Every tool call originates from a model `chat()` turn. The
input ops are fixed:

```
INPUT_OPS = ['prompt','steer','followUp','resolve','agent','cancel','command','answer','config']
// packages/ai-harness/src/protocol.ts:37
```

There is no op that runs a registered tool directly. The closest precedent is
plugin **commands** (`session.command(name, input)`, `session.ts:579`), which run
user-initiated actions outside a model turn and return a `Receipt` — this is the
shape to copy.

**Proposal.**

- Add input op `{ op: 'tool', name: string, args: unknown }` to `INPUT_OPS`
(`protocol.ts:37`) and a case in `applyInput()` (`protocol.ts:119`) that
resolves the tool, validates `args` against its schema, invokes
`tool.execute(args)` **without** opening a `chat()` stream, and returns a
`Receipt` (mirror `executeCommand`).
- Publish the result into the feed under a fresh `operationId` so it renders in
the stream like any tool result (reuse `OperationImpl`/`SessionFeed`).

**Open questions for you.** Should this reuse the command machinery outright
(register injectable tools _as_ commands) rather than a parallel op? What runs the
tool's `needsApproval` gate on the injection path — does an injected tool that
needs approval still raise an interrupt?

## 2. Tool registry + visibility (`public` vs `private`)

**Design need.** A pod tool registry the dashboard can query, with visibility:
`public` tools are pod-invocable (dashboard, other agents, schedules); `private`
tools run only in the owning agent's own runs (design §5.2).

**Current reality.** Tools are static per harness (`HarnessConfig.tools`,
`define.ts:36`), collected per-run from harness + plugins (`session.ts:1080`).
`expose` exists but only for **agents**, not tools:

```
expose?: { agents?: ReadonlyArray<...> } // define.ts:58
```

`capabilitiesOf()` lists harness-level tool names only (`protocol.ts:197`), with
no visibility concept and no plugin/discovered tools.

**Proposal.**

- Add `expose?: { tools?: ReadonlyArray<string> }` to `HarnessConfig`
(`define.ts`), defaulting to private (owner-only).
- Add `session.tools()` mirroring `session.commands()` (`session.ts:562`).
- Have `capabilitiesOf()` (`protocol.ts:187`) return the visibility-filtered set,
so the dashboard registry query and the §1 injection path share one source of
truth.

## 3. Injection to offline hosts vs. the relay constraint

**Design need.** "Injected tool calls execute on the owning agent's host; offline
hosts use the relay's existing queued-input behavior rather than silent drops"
(design §5.4).

**Current reality — this is a real conflict, flagging it explicitly.** The
offline queue exists, but it is **private** to the relay:

```
const sendToHost = (host, envelope): 'sent' | 'queued' => {
if (host.streams.size === 0) { host.queue.push({ envelope, expiresAt: ... }); return 'queued' }
...
} // packages/ai-dashboard/src/server.ts:160
```

`sendToHost()` is called only from `/api/sessions/:host/:thread/input` and
`/open`. There is **no external API to enqueue** for an offline host, and the
queue is drained only on the host's `/api/host/stream` reconnect
(`server.ts:253`). So:

- Adding an injection route that reuses the queue **requires modifying the relay**
(add a route that calls `sendToHost()`), which collides with the hard "don't
modify the relay" constraint.
- Duplicating the queue outside the relay **does not work**: the relay flushes
only its own queue on reconnect, so externally-queued frames would never be
delivered.

**Proposal / decision needed from you.** Pick one:

- **(a)** Accept a minimal, additive relay change: one new frame type
(`harness.inject`) enqueued through the existing `sendToHost()` path. Smallest
possible surface; keeps offline semantics correct. (Recommended — the
constraint exists to protect the relay's protocol, and this is protocol we'd be
co-designing with you rather than a unilateral edit.)
- **(b)** Keep injection **online-only** for now (dashboard injects only to hosts
with a live stream; offline injection is deferred). No relay change; weaker
guarantee.

## 4. Per-event causal metadata (loop-TTL bounce protection)

**Design need.** Bounce protection (design §6.3) leads with a **causal depth
TTL**: every event carries its causal chain and the pod caps reaction depth. That
requires per-event provenance.

**Current reality.** Only operation-level lineage exists — `parentRunId` on turns
for agent resume chains (`session.ts:163`), passed to `chat({ parentRunId })`.
`SessionEvent` itself (`feed.ts`) carries `{ cursor, operationId, event }` — no
`parentEventId`, no `causedByInputId`, no `depth`.

**Proposal.**

- Extend the feed/`SessionEvent` with `parentEventId?`, `causedByInputId?`, and a
monotonically-increasing `depth` set when an operation is spawned in reaction to
another event.
- Enforcement (TTL cap, cycle detection, quarantine) stays in the
dashboard/harness layer per the design; this proposal is only about **carrying**
the provenance so enforcement is possible later.

**Open question.** Is `parentRunId` enough to derive depth for the agent-to-agent
case, or do we genuinely need event-level parentage (I believe we do, because a
single run reacts to many upstream events)?

---

## Phasing implication

Phases 2–7 of the implementation plan are all gated on §§1–4. Phase 1 (the
dashboard-local teams reframe) is done and needs none of this. Recommend a short
review pass on §§1–3 first (they unblock the injection + registry work in Phase
2), with §4 reviewed alongside Phase 4 (bounce protection).
Loading
Loading