diff --git a/.changeset/harness-ag-ui-bridge.md b/.changeset/harness-ag-ui-bridge.md new file mode 100644 index 0000000000..d4c9c853bc --- /dev/null +++ b/.changeset/harness-ag-ui-bridge.md @@ -0,0 +1,13 @@ +--- +'@tanstack/ai-harness': minor +--- + +New subpath `@tanstack/ai-harness/ag-ui`: the AG-UI bridge. A harness session already streams AG-UI events, so this is a thin, documented seam for AG-UI clients. + +- `sessionEventsToAgUi(events, options?)` normalizes a session's `SessionEvent` stream into a pure AG-UI event stream: run usage is surfaced under `metadata.tanstack.usage` (`normalizeUsage` handles both the AG-UI spec array and the TanStack prompt/completion shapes), interrupt/approval waits stay as `RUN_FINISHED` with `outcome.type === 'interrupt'`, and subagent attribution is preserved. It can optionally drop the harness-native `CUSTOM` control events or emit interim `tanstack.spend` `CUSTOM` ticks for a live spend meter. +- `createAgUiHandler(options)` is a `fetch` handler a bare `@ag-ui/client` `HttpAgent` can point at. `POST` a `RunAgentInput` to run a prompt, or one with `resume` entries to answer the last turn's interrupts. It streams AG-UI events as SSE via `@ag-ui/encoder`'s `EventEncoder` (protobuf framing when the client's `Accept` prefers it). The stream is strict AG-UI by default (harness control events dropped so it begins with `RUN_STARTED`); the harness-native approval path and the relay dashboard are unchanged. +- `operationToAgUiRun(entries, options)` coalesces one harness operation — which may span several model turns (each a `RUN_STARTED`/`RUN_FINISHED` pair) — into a single valid AG-UI run: one `RUN_STARTED` (synthesized when a resumed run leads with a tool result), one terminal `RUN_FINISHED` carrying the operation outcome and summed usage, and `RUN_FINISHED.usage` conformed to the AG-UI `SpecTokenUsage[]` shape. This is what makes the stream consumable by a strict `@ag-ui/client` without "run already finished" / usage-shape errors. `createAgUiHandler` uses it per request. + +The AG-UI wire is pinned: built and tested against `@ag-ui/core@1.0.0`, `@ag-ui/encoder@1.0.0`, and `@ag-ui/client@1.0.0`. + +Generated with Claude Code. diff --git a/.changeset/harness-system-preamble.md b/.changeset/harness-system-preamble.md new file mode 100644 index 0000000000..728e4c9180 --- /dev/null +++ b/.changeset/harness-system-preamble.md @@ -0,0 +1,17 @@ +--- +'@tanstack/ai-harness': minor +--- + +Add an additive `systemPreamble` to the `prompt` input op. When present, its +strings are prepended (ahead of the harness's own `systemPrompts`) as +system/developer messages for that one run — a place for a trigger to attach +per-run context (e.g. operational memory) without the agent author doing +anything. + +Also make the server-tool execution `context` always carry the live `threadId` +and `runId` (merged over the harness's static `context`) so a tool invoked +in-band (from a model turn) can resolve the calling thread, matching the flat +`{ threadId, runId, signal }` already passed to out-of-band `{ op: 'tool' }` +invocations. + +Generated with Claude Code. diff --git a/.changeset/harness-tool-op.md b/.changeset/harness-tool-op.md new file mode 100644 index 0000000000..2cc4cadbf3 --- /dev/null +++ b/.changeset/harness-tool-op.md @@ -0,0 +1,12 @@ +--- +'@tanstack/ai-harness': minor +--- + +Out-of-band tool invocation: a new `{ op: 'tool', name, args?, meta? }` session input runs one registered tool with no model turn. This is the provisional harness surface the agent-dashboard "injection" work builds on (schedules, run-now, webhooks), following the `@tanstack/ai-harness/ag-ui` precedent of shipping dashboard-facing seams here. + +- `session.tool(name, args?, meta?)` (and the `{ op: 'tool' }` client input via `applyInput`) resolves the tool from the harness + session-plugin tools, validates `args` against its input schema, and invokes its server executor directly — modeled on the existing `command` op. The tool's lifecycle is published into the session feed as a normal AG-UI run (`RUN_STARTED`, `TOOL_CALL_START`/`ARGS`/`END`, `TOOL_CALL_RESULT`, `RUN_FINISHED`), so watchers render and persist it exactly like a tool call the model made. `meta` (e.g. an injection trigger) is echoed onto the run as a `tanstack.injection` `CUSTOM` event. Injected tools are fire-and-forget (no approval interrupt, like commands). +- Tool **visibility**: a new `toolVisibility?: Record` on `defineHarness`. Tools default to `private` — only tools named `public` may be invoked out-of-band; unknown or private tools are rejected (`unknown_tool` / `not_public`). Visibility is reported per tool in the capabilities document (`capabilitiesOf().tools.items[].visibility`). + +Additive and backward compatible: existing ops, tools, and streams are unchanged. + +Generated with Claude Code. diff --git a/PODS-PROTOCOL-PROPOSAL.md b/PODS-PROTOCOL-PROPOSAL.md new file mode 100644 index 0000000000..42aadc4f42 --- /dev/null +++ b/PODS-PROTOCOL-PROPOSAL.md @@ -0,0 +1,155 @@ +# Pods/Teams — harness protocol proposal (for @AlemTuzlak) + +**Status:** proposal · **Date:** 2026-09-26 · **Author:** Jack + Claude Code +**Basis:** `~/Downloads/pods-design-doc.md` (v2) and `pods-implementation-plan.md` + +This is a **discussion doc, not a change**. No `packages/ai-harness` or +`packages/ai-dashboard` (relay) code has been touched. Phase 1 of the teams +reframe ships entirely in `examples/agent-dashboard` (dashboard-local +collections). Everything below is the net-new **harness/relay protocol surface** +that Phases 2+ need, grounded in what the current code does and doesn't do, so we +can agree on the shape before writing any of it. + +The design doc says "pod"; the dashboard ships the noun **"team"**. Same concept. + +## Why this doc exists + +Three capabilities the design depends on are not expressible in the harness as it +stands today. Each is core protocol surface, so it needs your review before +implementation. Findings were verified against the current tree (head at the top +of the harness PR stack). + +--- + +## 1. Out-of-band tool invocation (`injectToolCall`) + +**Design need.** The dashboard's injection model (design §5.4) calls a single +named tool deterministically, with **no model turn and zero tokens** +(`injectToolCall(pod, tool, args)`), and posts the result to the stream. This is +the _default_ trigger path (timers, webhooks, "run now"). + +**Current reality.** Every tool call originates from a model `chat()` turn. The +input ops are fixed: + +``` +INPUT_OPS = ['prompt','steer','followUp','resolve','agent','cancel','command','answer','config'] + // packages/ai-harness/src/protocol.ts:37 +``` + +There is no op that runs a registered tool directly. The closest precedent is +plugin **commands** (`session.command(name, input)`, `session.ts:579`), which run +user-initiated actions outside a model turn and return a `Receipt` — this is the +shape to copy. + +**Proposal.** + +- Add input op `{ op: 'tool', name: string, args: unknown }` to `INPUT_OPS` + (`protocol.ts:37`) and a case in `applyInput()` (`protocol.ts:119`) that + resolves the tool, validates `args` against its schema, invokes + `tool.execute(args)` **without** opening a `chat()` stream, and returns a + `Receipt` (mirror `executeCommand`). +- Publish the result into the feed under a fresh `operationId` so it renders in + the stream like any tool result (reuse `OperationImpl`/`SessionFeed`). + +**Open questions for you.** Should this reuse the command machinery outright +(register injectable tools _as_ commands) rather than a parallel op? What runs the +tool's `needsApproval` gate on the injection path — does an injected tool that +needs approval still raise an interrupt? + +## 2. Tool registry + visibility (`public` vs `private`) + +**Design need.** A pod tool registry the dashboard can query, with visibility: +`public` tools are pod-invocable (dashboard, other agents, schedules); `private` +tools run only in the owning agent's own runs (design §5.2). + +**Current reality.** Tools are static per harness (`HarnessConfig.tools`, +`define.ts:36`), collected per-run from harness + plugins (`session.ts:1080`). +`expose` exists but only for **agents**, not tools: + +``` +expose?: { agents?: ReadonlyArray<...> } // define.ts:58 +``` + +`capabilitiesOf()` lists harness-level tool names only (`protocol.ts:197`), with +no visibility concept and no plugin/discovered tools. + +**Proposal.** + +- Add `expose?: { tools?: ReadonlyArray }` to `HarnessConfig` + (`define.ts`), defaulting to private (owner-only). +- Add `session.tools()` mirroring `session.commands()` (`session.ts:562`). +- Have `capabilitiesOf()` (`protocol.ts:187`) return the visibility-filtered set, + so the dashboard registry query and the §1 injection path share one source of + truth. + +## 3. Injection to offline hosts vs. the relay constraint + +**Design need.** "Injected tool calls execute on the owning agent's host; offline +hosts use the relay's existing queued-input behavior rather than silent drops" +(design §5.4). + +**Current reality — this is a real conflict, flagging it explicitly.** The +offline queue exists, but it is **private** to the relay: + +``` +const sendToHost = (host, envelope): 'sent' | 'queued' => { + if (host.streams.size === 0) { host.queue.push({ envelope, expiresAt: ... }); return 'queued' } + ... +} // packages/ai-dashboard/src/server.ts:160 +``` + +`sendToHost()` is called only from `/api/sessions/:host/:thread/input` and +`/open`. There is **no external API to enqueue** for an offline host, and the +queue is drained only on the host's `/api/host/stream` reconnect +(`server.ts:253`). So: + +- Adding an injection route that reuses the queue **requires modifying the relay** + (add a route that calls `sendToHost()`), which collides with the hard "don't + modify the relay" constraint. +- Duplicating the queue outside the relay **does not work**: the relay flushes + only its own queue on reconnect, so externally-queued frames would never be + delivered. + +**Proposal / decision needed from you.** Pick one: + +- **(a)** Accept a minimal, additive relay change: one new frame type + (`harness.inject`) enqueued through the existing `sendToHost()` path. Smallest + possible surface; keeps offline semantics correct. (Recommended — the + constraint exists to protect the relay's protocol, and this is protocol we'd be + co-designing with you rather than a unilateral edit.) +- **(b)** Keep injection **online-only** for now (dashboard injects only to hosts + with a live stream; offline injection is deferred). No relay change; weaker + guarantee. + +## 4. Per-event causal metadata (loop-TTL bounce protection) + +**Design need.** Bounce protection (design §6.3) leads with a **causal depth +TTL**: every event carries its causal chain and the pod caps reaction depth. That +requires per-event provenance. + +**Current reality.** Only operation-level lineage exists — `parentRunId` on turns +for agent resume chains (`session.ts:163`), passed to `chat({ parentRunId })`. +`SessionEvent` itself (`feed.ts`) carries `{ cursor, operationId, event }` — no +`parentEventId`, no `causedByInputId`, no `depth`. + +**Proposal.** + +- Extend the feed/`SessionEvent` with `parentEventId?`, `causedByInputId?`, and a + monotonically-increasing `depth` set when an operation is spawned in reaction to + another event. +- Enforcement (TTL cap, cycle detection, quarantine) stays in the + dashboard/harness layer per the design; this proposal is only about **carrying** + the provenance so enforcement is possible later. + +**Open question.** Is `parentRunId` enough to derive depth for the agent-to-agent +case, or do we genuinely need event-level parentage (I believe we do, because a +single run reacts to many upstream events)? + +--- + +## Phasing implication + +Phases 2–7 of the implementation plan are all gated on §§1–4. Phase 1 (the +dashboard-local teams reframe) is done and needs none of this. Recommend a short +review pass on §§1–3 first (they unblock the injection + registry work in Phase +2), with §4 reviewed alongside Phase 4 (bounce protection). diff --git a/STATUS.md b/STATUS.md new file mode 100644 index 0000000000..611c21bd91 --- /dev/null +++ b/STATUS.md @@ -0,0 +1,470 @@ +# Agent Dashboard — status + feature inventory + +Spec: `~/Downloads/agent-dashboard-spec.md` (original four phases), then the +**teams reframe** (`~/Downloads/pods-design-doc.md` + `pods-implementation-plan.md`), +then the **Reddit pod** (`~/Downloads/teams-reddit-pod.md`). + +> **Latest work: weekend sprint** — live trace waterfall, steer/stop/ask-user, +> cron schedules with upcoming runs, webhook delivery management, dollar spend, +> and an MCP endpoint that lets external agents manage the dashboard. The branch +> was merged with `origin/main` on 2026-10-02 (`c1d0169d6`), no conflicts. + +## Where this lives + +- **Worktree**: `~/projects/tanstack/ai-workingtrees/ai-dashboard-work` +- **Branch**: `feat/agent-dashboard`, pushed to `origin/feat/agent-dashboard`, + open as [PR #1560](https://github.com/TanStack/ai/pull/1560). +- **Base**: up to date with `origin/main` (merged 2026-10-02). The harness work + it builds on (`@tanstack/ai-harness` session view, AG-UI bridge, out-of-band + tools, `@tanstack/ai-dashboard`) is **branch-only** — none of it is on `main`. +- **Do not touch** the relay in `packages/ai-dashboard` (spec §6). + +--- + +## Feature inventory (for competitive analysis) + +There are **two dashboards** on this branch. Compare them separately — they +target different buyers. + +| Product | Where | What it is | +| ---------------------------- | -------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | +| **Agent Dashboard** | `examples/agent-dashboard` | Mission control for teams of agents: live channels, approvals, automations, pod memory, spend, config, meta-chat. A TanStack Start app. | +| **`@tanstack/ai-dashboard`** | `packages/ai-dashboard` | A small self-hosted remote control for one or more harness hosts. Zero dependencies, `npx`-runnable, installs as a phone PWA. | + +Status legend: ✅ shipped · 🟡 partial / demo-grade · 🧪 simulated · ❌ not built + +### A. Agent Dashboard (`examples/agent-dashboard`) + +#### A1. Live observability + +| Feature | Status | Notes | +| ---------------------------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------- | +| Live session stream | ✅ | AG-UI SSE consumed by a bare `@ag-ui/client` `HttpAgent`. Projected into TanStack DB collections, read with `useLiveQuery`. No polling. | +| Standards-based wire protocol | ✅ | AG-UI 1.0 (`@tanstack/ai-harness/ag-ui`). Any AG-UI client can consume a harness run. | +| Message rendering | ✅ | Assistant text renders as markdown (GFM tables) via `react-markdown` + `remark-gfm`. User text stays plain. | +| Tool-call cards | ✅ | Name, status (running/done), args and result as collapsible JSON trees. Plain-text results stay text. Oversize results show "(result truncated)". | +| Injected-call badge | ✅ | Tool calls started by a timer, run-now, or webhook get a violet border and a `⏵ timer/manual/webhook` badge. | +| Subagent attribution | ✅ | Messages from a child agent carry a `subagent` badge (`subagentRunId` preserved through AG-UI). | +| Per-agent attribution in teams | ✅ | Author tag on every message and tool card once a channel has ≥2 members. | +| Channel status pill | ✅ | Aggregates members: `running` / `requires action` / `idle`. | +| Live token counter per channel | ✅ | Sum of `spend` rows for the channel, in the header. | +| Memory-attached badge | ✅ | "🧠 N memory entries attached" on the channel header after a run. | +| Long trigger prompts collapse | ✅ | Injected subscription prompts clamp to 3 lines with Show more / Show less. Internal `[channel:]` routing tags are stripped. | +| Session detail view | ✅ | `/sessions/$threadId`: one thread's stream, kept for back-compat with the pre-teams model. | +| Timeline survives a server restart | ✅ | `/api/tail` rebuilds AG-UI events from persisted messages when the in-memory feed is empty, routed to the right channel. | +| Resolved approvals stay resolved | ✅ | Tail sends a snapshot frame + `replay: true` on history, so reloading never re-raises an approved interrupt. | +| Trace waterfall per run | ✅ | Live run, text, tool, and approval spans. Pending approvals can be resolved in the trace. | +| OpenTelemetry export | ❌ | No OTLP export or cross-service trace correlation. | +| Search / filter across sessions | ❌ | | + +#### A2. Human-in-the-loop + +| Feature | Status | Notes | +| -------------------------------------- | ------ | -------------------------------------------------------------------------------------------- | +| Approval queue (mid-run) | ✅ | Run pauses on an approval-gated tool (`needsApproval`). Card shows tool name, message, args. | +| Approve | ✅ | Resumes through the AG-UI resume flow; the continuation streams back into the view. | +| Deny | ✅ | Goes through the harness control tier (`/api/harness/control`). | +| Edit, then approve | ✅ | Edit tool args as JSON, then approve the edited call. | +| Chat with an agent | ✅ | Message box sends to the channel's primary agent. Memory is attached to every run. | +| Run a member on demand | ✅ | Roster `▶ run` sends a generic nudge to any agent. | +| Plugin questions (ask-user) | ✅ | Question cards support Boolean and free-text answers. | +| Steer / stop a running agent | ✅ | Running sessions turn the message box into Steer and expose a Stop button. | +| Broadcast to all members | ❌ | Human input targets one primary agent. | +| Approval policies / auto-approve rules | 🟡 | Only a per-agent config flag ("Skip approval for low-risk replies", demo only). | + +#### A3. Multi-agent teams ("pods") + +| Feature | Status | Notes | +| ------------------------------------ | ------ | ------------------------------------------------------------------------------------------------------------------------------ | +| Agents catalog | ✅ | Home page is a table of every registered agent (name, description) with **Add to team ▾** (new team or an existing one). | +| Teams list | ✅ | Home page lists teams; `/teams/$teamId` opens one. | +| Progressive disclosure | ✅ | One agent looks like a plain chat. A second member reveals the roster and attribution. | +| + Add agent | ✅ | Add any available agent to a team's main channel. | +| Shared channel (multiplexed streams) | ✅ | N member threads merge into one timeline with no id collisions. | +| Dynamic channels | ✅ | Agents open per-topic channels (e.g. `#pr-123`) via `pod.channel_create`. Channel sidebar + "📢 channel created" system cards. | +| Direct messages | ✅ | **New DM** with any agent creates a working 1:1 channel. | +| Roster with live status dots | ✅ | Per-member running / requires-action / idle. | +| Operator role | ✅ | A member can be an `operator` (e.g. the meta agent) instead of an `agent`. | +| Event subscriptions | ✅ | `channel_created` (join + trigger) and `tool_result` (trigger on another agent's tool output). | +| Default subscriptions by harness | ✅ | Hand-composed teams react like seeded ones with no wiring. Roster toggle: 🔔 subscribed / 🔕 subscribe. | +| Agent-to-agent messaging | ✅ | `pod.message_post`, routed by channel id. | +| Loop / bounce protection | 🟡 | Chat messages don't match `tool_result` subscriptions and replays don't re-dispatch. No general cycle detector. | +| Mixed harnesses in one team | ✅ | Server learns each thread's harness, so watcher + reviewer + meta can share a team. | +| Team-level permission boundary | ❌ | Design doc concept only; no enforcement. | + +#### A4. Memory + +| Feature | Status | Notes | +| ----------------------------- | ------ | --------------------------------------------------------------------------------------------------------------------------- | +| Pod memory store | ✅ | Server-side, keyed per agent thread. Persisted to disk. | +| Agents write memory | ✅ | `pod.memory_write` (in-band or out-of-band) or the provider-safe `remember` tool. | +| Automatic memory injection | ✅ | Every run (chat, subscription, prompt-mode webhook) gets the thread's memory as a `systemPreamble`. Zero agent-author code. | +| Memory panel on the team page | ✅ | One panel per agent. Key on top, value below (clamped to 3 lines). Add, delete (🗑); re-adding a key updates it. | +| Learn-from-correction loop | ✅ | Demoed end to end: human corrects the reviewer → memory written → next PR reviewed correctly. | +| Shared team memory | ❌ | Memory is per agent, not per team. | +| Vector / semantic memory | ❌ | Key-value only. | + +#### A5. Automations (deterministic, zero-token triggers) + +| Feature | Status | Notes | +| --------------------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Out-of-band tool invocation | ✅ | `{ op: 'tool' }` runs one registered tool with no model turn. Result streams into the channel and persists for replay. | +| Tool visibility | ✅ | Only `public` tools can run out-of-band. Unknown or private tools are rejected, enforced on the server at read time. | +| Run-now | ✅ | Product UI: roster **🔧 tools** → RunToolDialog (pick a public tool, JSON params). Also in the Demo Controls panel. | +| Schedules | ✅ | Server-owned clock (1 s scheduler). `everySeconds` or 5-field cron. Pause / resume / delete. Idempotent jobs. | +| Webhooks | ✅ | Tokenized ingress `POST /api/webhooks/:token`. Two modes: `tool` (run a tool) or `prompt` (start a model run, memory attached). | +| Result size cap | ✅ | 64 KB per injected result. | +| Offline host queue | 🧪 | Simulated in the dashboard (dev toggle + "host offline — N queued" banner, flush on reconnect). The real relay queue lives in `@tanstack/ai-dashboard`. | +| Per-field tool param forms | ❌ | JSON only; `/api/tools` doesn't expose input schemas. | +| Schedule UI for cron | ✅ | Product UI supports intervals and five-field cron, three upcoming runs, last run, pause/resume, and delete. | +| Webhook delivery management | ✅ | Product UI shows endpoint URLs and the latest 20 delivery attempts, with retry and delete. | + +#### A6. Control plane + +| Feature | Status | Notes | +| --------------------------- | ------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Run history | ✅ | `/history`, backed by `HarnessPersistence` (`runs.listByThread`). Includes the meta agent's own runs. | +| Session replay | ✅ | Rehydrates any stored session from persisted events (`/api/replay`). | +| Spend tracking | ✅ | `/spend`: per-session input / output / total tokens, live. SSR-safe SVG bar chart with budget reference lines. | +| Budgets + alerts | 🟡 | Editable per-session token budget (default 2,000), "over budget" status and banner. **Alert only — not enforced**, and budgets are browser-local (not persisted). | +| Cost in dollars | 🟡 | Per-session and per-team estimates for `claude-haiku-4-5`; scripted agents show `$0.00`. Cached-token pricing is not tracked. | +| Agent config | ✅ | `/config`: form generated from typed `ConfigOption` schemas; writes through the harness protocol (`op: 'config'`). | +| Protocol versioning | ✅ | Endpoints stamped with `HARNESS_PROTOCOL_VERSION`. | +| Agent versioning / rollback | ❌ | Deferred (spec §5). | +| Evals / red-teaming | ❌ | Deferred (spec §5). | +| Prompt playground | ❌ | Deferred (spec §5). | + +#### A7. Meta-chat (the dashboard's own agent) + +| Feature | Status | Notes | +| -------------------- | ------ | ---------------------------------------------------------------------------------------------------------- | +| Chat over live state | ✅ | `/chat`, harness `dashboard/meta`. Also usable as an operator inside any team. | +| Tools | ✅ | `list_agents`, `list_sessions`, `query_runs`, `get_agent_config`, `set_agent_config`, `summarize_session`. | +| Dogfooding | ✅ | Its tool calls render inline; its runs appear in History like any other agent. | +| Real LLM | 🟡 | Scripted keyword router in the example (no API key needed). Swap in a real adapter for production. | + +#### A8. Persistence, deployment, security + +| Feature | Status | Notes | +| -------------------------- | ------ | -------------------------------------------------------------------------------------------------------------------------- | +| Durable state | ✅ | One JSON file (`.data/state.json`, override `DASHBOARD_STATE_FILE`). Atomic, debounced write-through; replayed on boot. | +| What persists | ✅ | Messages, runs, interrupts, agent config, threads, pod memory, schedules, webhooks, team roster. | +| What doesn't | — | In-flight injection queue, offline toggle, budgets, inbox/credentials. | +| Multi-node / real database | ❌ | Single process, single file. A DB can swap in behind the same seam. | +| Auth | 🟡 | Local single user (`authorize` returns a fixed principal). No multi-user, RBAC, or SSO. | +| Audit trail | 🟡 | Every action is a message or tool call in the stream (`pod.memory_write` cards stay visible). No dedicated audit log view. | + +#### A9. Developer experience + +| Feature | Status | Notes | +| ---------------------------- | ------ | ---------------------------------------------------------------------------------------------------------- | +| Runs with no API key | ✅ | Scripted demo agents. Only the Reddit sentiment agent needs `ANTHROPIC_API_KEY`. | +| Demo Controls devtools panel | ✅ | Custom TanStack DevTools plugin: seeded-team launchers, triage demo, automations, PR webhook, offline sim. | +| Hermetic e2e | ✅ | `VITE_E2E` swaps in deterministic doubles + a recorded Reddit fixture. No key, no network. | +| TanStack stack showcase | ✅ | Start (routes + server API), DB (`localOnly` collections), Query (server state), DevTools, Router. | + +#### A10. Bundled demo agents + +| Agent | Kind | What it shows | +| ----------------- | --------------------------- | ------------------------------------------------------------------------------------------------ | +| `support/triage` | Scripted | Auto tool (`lookup_ticket`) → approval-gated tool (`send_reply`) → resume. Public `fetch_stats`. | +| `ops/pr-watcher` | Scripted | Webhook → `github.check_pr` → opens a `#pr-…` channel → posts to it. | +| `security/review` | Scripted | Subscribes to `channel_created`, reviews, learns from human correction via pod memory. | +| `reddit/fetcher` | Procedural (no LLM) | Real external I/O: reads Reddit's public RSS. Run-now / schedule / webhook target. | +| `sentiment/react` | Real LLM (Claude Haiku 4.5) | Triggered by the fetcher's `tool_result`. Posts a markdown sentiment digest, writes memory. | +| `dashboard/meta` | Scripted | The meta-chat operator. | + +### B. `@tanstack/ai-dashboard` (self-hosted remote control) + +| Feature | Status | Notes | +| ----------------------------------------- | ------ | ----------------------------------------------------------------------------------------------------------------------------------------------- | +| One-command self-host | ✅ | `npx @tanstack/ai-dashboard --port 8790` or `startDashboard()`. Plain `node:http`, no new dependencies. | +| Owner sign-in | ✅ | Prints a sign-in link with an owner token. `DASHBOARD_OWNER_TOKEN` keeps it stable across restarts. | +| Outbound-only hosts | ✅ | Agents dial out with `connectDashboard()`; they need no open port. | +| Pairing | ✅ | One-time code shown by the host; the owner approves it in the UI; host gets a revocable host token. | +| Host revocation | ✅ | `POST /api/hosts/revoke`. | +| Relay transport | ✅ | SSE + POST. Events batched so a streaming answer is a few requests, not one per token. | +| Event cache | ✅ | Recent events per session (default 2,000) so a late viewer catches up. | +| Offline input queue | ✅ | Inputs for an offline host wait (default 10 min TTL) and are delivered on reconnect. Real, not simulated. | +| Hosts + sessions list | ✅ | Multiple hosts, each with its sessions. | +| Live messages + tool calls | ✅ | | +| Child-agent blocks | ✅ | Work from delegated coding agents (Claude Code, Codex via `codingAgents`) groups into one block per child, with tools, text, and failure state. | +| Approval cards | ✅ | | +| Plugin questions | ✅ | Yes/No or free-text answers. | +| Prompt / steer / stop | ✅ | Prompt when idle, steer while running, stop. | +| Mobile PWA | ✅ | Web manifest + service worker; installs on a phone. | +| CLI integration | ✅ | `@tanstack/ai-harness-cli --dashboard `. | +| Persistence | ❌ | In-memory relay; restart drops hosts' cached events. | +| Teams, memory, automations, spend, config | ❌ | Those live only in the Agent Dashboard example. | + +### C. Platform pieces underneath (in `@tanstack/ai-harness`, branch-only) + +These are what a competitor would need to match the dashboards, not UI. + +- **AG-UI bridge** (`@tanstack/ai-harness/ag-ui`): `sessionEventsToAgUi()`, + `operationToAgUiRun()` (one harness operation → one valid AG-UI run), + `createAgUiHandler()` (SSE). Usage in `metadata.tanstack.usage` (conformed to + `SpecTokenUsage[]`), optional `tanstack.spend` ticks, subagent attribution. +- **Out-of-band tool op** `{ op: 'tool', name, args?, meta? }` + `toolVisibility`. +- **`systemPreamble`** on the prompt op (per-run context injection) + live + `threadId`/`runId` in tool execution context. +- **`createSessionView`**: a live TanStack Store view of a session for any UI, + with a pure reducer and `selectGoal`. +- **Remote reads**: `session.transcript()`, `session.describe()`, plugin state in + snapshots; client `answer` / `command` / `setConfig` + connection state. +- **Goal plugin**: keeps working until a goal is met. +- **`codingAgents` plugin** (`ai-sandbox`): delegate to Claude Code, Codex, and + other coding agents; child work shows in the CLI, ACP, and dashboard. +- **`agentMiddleware`** runs plugin middleware in every agent; `usage()` counts + the model calls of every agent. + +### D. External management + +- **Dashboard MCP endpoint** ✅ `/api/mcp`: lists agents, sessions, runs, teams, + and schedules; reads/writes config; sends/steers/stops messages; answers + questions; resolves approvals; runs public tools; toggles schedules. +- The endpoint is unauthenticated because this is a local single-user demo. + +### E. Not built (consolidated gap list) + +- Cross-service trace correlation and OpenTelemetry export. +- Cross-session search and filtering. +- Cached-token pricing and budget **enforcement** (budgets only alert). +- Evals, red-teaming, prompt playground, agent versioning / rollback. +- Multi-user auth, RBAC, SSO, org / workspace model. +- Multi-node deployment and a real database. +- Broadcast to all team members; team-level permission boundary; shared team memory. +- Per-field tool-param forms. +- Polyglot tool face, bridging, install model, hibernation, pricing (pods design doc Phases 5–8). +- One dashboard: the two products don't share a UI yet. + +--- + +## Commits (current hashes on this branch) + +Hashes changed when the branch was rebuilt onto the harness stack; these are +the current ones. + +| Commit | What | +| ----------- | -------------------------------------------------------------------------------------------- | +| `9409b06db` | `feat(ai-sandbox, ai-harness)`: `codingAgents` plugin; child agent work in CLI/ACP/dashboard | +| `06595208a` | `feat(ai-harness)`: goal plugin | +| `8b46a04b6` | `feat(ai-harness)`: `session.transcript()`, `session.describe()`, plugin state in snapshot | +| `7364ac59f` | `feat(ai-harness)`: remote transcript/describe, client answer/command/setConfig | +| `3bf4c0c81` | `feat(ai-harness)`: `createSessionView` | +| `83825af22` | `feat(ai, ai-harness)`: `agentMiddleware` | +| `47ee22f76` | `feat(ai-harness)`: `usage()` counts every agent's model calls | +| `61ce0041f` | `feat(ai-dashboard)`: self-hosted dashboard, `--dashboard`, runnable example | +| `d7da6b9a0` | `feat(ai-harness)`: the `@tanstack/ai-harness/ag-ui` bridge subpath | +| `c82a5ad45` | `fix(ai-harness)`: coalesce multi-turn operations for strict AG-UI clients | +| `de730ecd0` | `feat(examples/agent-dashboard)`: live session view + approval queue (Phase 2) | +| `b6f39e7dc` | `feat(examples/agent-dashboard)`: control plane — history, spend, config (Phase 3) | +| `74e45973a` | `feat(examples/agent-dashboard)`: meta-chat over live agent state (Phase 4) | +| `be6e9d050` | `feat(examples/agent-dashboard)`: teams reframe (Phase 1) + Alem protocol proposal | +| `d410e788b` | `feat(ai-harness)`: out-of-band tool invocation + tool visibility | +| `2078d29ce` | `feat(examples/agent-dashboard)`: teams Phase 2 — tool registry + injection | +| `0eda69e1d` | `feat(ai-harness)`: `systemPreamble` on the prompt op + in-band tool thread id | +| `fcc699b1f` | `feat(examples/agent-dashboard)`: teams Phase 3 — system tools, channels, pod memory | +| `7cbb2402e` | `feat(examples/agent-dashboard)`: persist agent + team state on the server | +| `51cf847fa` | `fix(examples/agent-dashboard)`: resolved approval no longer resurfaces on reload | +| `e8d40f36e` | `feat(examples/agent-dashboard)`: move demo controls into a devtools panel | +| `1036f8274` | `feat(examples/agent-dashboard)`: Reddit pod — real service + real LLM | +| `9b0fd5328` | `feat(examples/agent-dashboard)`: product team-composition UI + default subscriptions | +| `830c514b9` | `fix(examples/agent-dashboard)`: DM membership + surface pod memory on the team page | +| `69ba30134` | `feat(examples/agent-dashboard)`: collapsible JSON trees + markdown messages | +| `1c54522a6` | `ci`: apply automated fixes | +| `97f51c0e8` | `feat(examples/agent-dashboard)`: rehydrate timeline from persisted messages + UI polish | +| `c1d0169d6` | Merge `origin/main` (2026-10-02) | + +--- + +## Build history (detail) + +### Phase status + +- **Phase 1 — AG-UI bridge** ✅ `packages/ai-harness/src/ag-ui.ts` + - `sessionEventsToAgUi()` (normalizer), `operationToAgUiRun()` (coalesces one + harness operation → one valid AG-UI run), `createAgUiHandler()` (SSE via + `@ag-ui/encoder`). Consumable by a bare `@ag-ui/client` `HttpAgent`. + - Interrupts stay as `RUN_FINISHED.outcome`, usage → `metadata.tanstack.usage` + (+ conformed to `SpecTokenUsage[]`), subagent attribution preserved, optional + `tanstack.spend` ticks. + - 22 unit tests (`tests/ag-ui.test.ts`), changeset, `docs/harness/ag-ui.md` + (registered in `docs/config.json`). AG-UI pinned at `1.0.0`. +- **Phase 2 — Dashboard shell + live session view** ✅ `examples/agent-dashboard` + - TanStack Start; routes hosts → sessions → session detail. + - Session view consumes AG-UI SSE via `@ag-ui/client`, projected into TanStack + DB `localOnly` collections read with `useLiveQuery` (no polling). + - Approval queue: approve → AG-UI resume, deny → harness control tier, edit → + edited tool args. +- **Phase 3 — Control plane** ✅ + - History (`/history`, `runs.listByThread`) + session replay (`/api/replay`). + - Spend (`/spend`): per-session token rollups (live query) + budgets/alerts. + - Config (`/config`): form from `ConfigOption` schemas → writes via the harness + protocol (`op: 'config'`). Endpoints stamped `HARNESS_PROTOCOL_VERSION`. +- **Phase 4 — Meta-chat** ✅ `/chat`, harness `dashboard/meta` + - Tools over live state: `list_agents`, `list_sessions`, `query_runs`, + `get_agent_config`, `set_agent_config`, `summarize_session`. + - Runs on the same host (harness registry); its tool calls are visible inline + and its runs appear in History like any other agent. + +### Teams reframe (Phase 1) ✅ `be6e9d050` + +A product-model reframe on top of the four-phase build: `hosts → sessions` +becomes **teams → channels**. A team is a group of agents sharing a chat, a tool +registry, and a permission boundary (the design doc calls it a "pod"; the UI noun +is **"team"**). One agent looks like a plain chat; a **second member reveals the +team** — roster appears, per-agent attribution shows up, and both agents' AG-UI +streams merge into one channel. **Dashboard-local only — zero harness/relay +changes.** + +- **Key trick:** each member owns its own thread, so `agentId == threadId`. + Namespacing every projected row id by `agentId` is therefore per-thread, which + lets the old threadId-keyed `/sessions/$threadId` and `/chat` routes keep + working **unchanged** while the new `/teams/$teamId` route queries by + `channelId` and gets the multiplex. +- `src/db/collections.ts` — `teams`/`channels`/`memberships` localOnly + collections; `channelId`/`agentId` (+ raw `interruptId`) on the existing rows. +- `src/lib/session-controller.ts` — `project(ctx)` namespaces ids + (`${agentId}:${raw}`) so N members share one channel with no collisions. +- **Part B — `PODS-PROTOCOL-PROPOSAL.md`** (worktree root, for @AlemTuzlak): the + net-new harness surface Phases 2+ need. Proposal only. +- **Known cosmetic:** `/` runs a live query, so SSR falls back to client + rendering (`useLiveQuery` has no `getServerSnapshot`). + +### Teams Phase 2 — tool registry + injection ✅ `d410e788b`, `2078d29ce` + +The dashboard invokes work **deterministically** — scheduled timers, run-now, and +webhooks — with the structured result streaming into the team channel and **zero +LLM tokens** on the trigger path. "The dashboard owns the clock." + +- **Harness (additive):** `{ op: 'tool', name, args?, meta? }` runs one + registered tool with no model turn, publishing `RUN_STARTED`/`TOOL_CALL_*`/ + `RUN_FINISHED` into the feed. `toolVisibility` on `defineHarness` (default + private). **Rebuild the dist after harness edits** (the example imports the + built package). +- **Architecture:** trigger via the control plane, observe via a **live feed + tail** (`/api/tail`, replays then follows live). One `project()`/`ToolCard` + path for interactive runs, injected tools, timers, and webhooks. +- **Server:** `server/scheduler.ts` (1s interval, lazy-boot), + `server/injection.ts` (server-owned schedule + webhook registries; idempotent + jobs; 64KB result cap), `server/cron.ts` (5-field cron + `everySeconds`). +- **Offline** hosts are **simulated** dashboard-side (the example embeds the host). +- **Deviation from doc §4.1:** schedules/webhooks are **server** state because + the server owns the clock. + +### Teams Phase 3 — system tools, channels, pod memory ✅ `0eda69e1d`, `fcc699b1f` + +The design doc's **"the pod learns" loop**: a webhook opens a per-PR channel, a +subscribed security agent reviews it, the human corrects it, the correction is +persisted to **pod memory**, and the _next_ PR is handled better — every step a +message or a tool call in the stream, **no hidden state**. + +- **Harness (additive):** the `prompt` op gains `systemPreamble?: string[]`; + server-tool context always carries the live `threadId`/`runId`. +- **System tools (`pod.*`):** `pod.channel_create` / `pod.message_post` / + `pod.memory_write` / `pod.memory_read` — callable in-band and out-of-band. +- **Memory delivery:** `/api/run` prepends the thread's pod memory as a + `systemPreamble` on every run. +- **Dynamic channels + subscriptions:** `channels` gains `kind`/`topic`/ + `createdBy`; `channelMembers` (per-channel opt-in) vs the team roster. +- **Deviations (deliberate):** the watcher opens a review channel it does **not** + join; pod memory is keyed by `threadId`; `pod.*` are excluded from run-now. + +### Persistence ✅ `7cbb2402e`, `51cf847fa` + +- **`src/server/store.ts`** — one JSON snapshot file, written through on every + mutation (debounced, atomic tmp+rename) and replayed on boot. +- **Team roster** persisted as a server blob (`GET|POST /api/roster`), seeded on + app load. +- **Approval replay fix** (`51cf847fa`): `/api/tail` sends a snapshot frame and + flags history `replay: true`, so a resolved approval doesn't come back on reload. + +### Demo controls — devtools panel ✅ `e8d40f36e` + +Demo-only scaffolding lives in a custom **Demo Controls** TanStack DevTools +panel. It renders outside the route tree and re-derives state from the same live +TanStack DB collections (via a `uiState` row) — a state-management stress test. +`TanStackDevtools` sits inside `QueryClientProvider`; `VITE_E2E=1` opens it on load. + +### Reddit pod — real service, real AI ✅ `1036f8274` + +- **`reddit/fetcher`** (no LLM) — `reddit.search_react_news` reads Reddit's + public **RSS (Atom)** feed. Parser unit-tested against a recorded fixture. +- **`sentiment/react`** — `anthropicText('claude-haiku-4-5')`, gated on + `ANTHROPIC_API_KEY`. Real LLM only (no product mock); deterministic double + under `VITE_E2E`. +- **`tool_result` subscription** — the fetcher's result triggers the sentiment + agent with a capped (≤10-item) batch. +- **Live findings:** (1) real digest + memory writes work; (2) dotted `pod.*` tool + names 400 on Anthropic/OpenAI (`^[a-zA-Z0-9_-]{1,128}$`) — worked around with a + provider-safe `remember` tool; **follow-up:** sanitize names in the adapter or + rename to `pod_*`; (3) Reddit `.json` is IP-blocked from some egress, so RSS. + +### Product team-composition UI ✅ `9b0fd5328`, `830c514b9` + +- Home is an agents table with **Add to team**; team page has **+ Add agent** + and per-agent **🔧 tools** (RunToolDialog, JSON params). +- `DEFAULT_SUBSCRIPTIONS` by harness so hand-composed teams react. +- Fixed: roster ▶ run sent a triage prompt to any agent; subscribe toggle set a + hardcoded sub; DMs created with no members (Zod stripped `members`); pod + memory only visible in the devtools panel. + +### Rendering + rehydration ✅ `69ba30134`, `97f51c0e8` + +- Tool args/results as collapsible JSON trees (native `
`, no new dep); + assistant text as markdown (`react-markdown` + `remark-gfm`). +- Timeline rebuilt from persisted messages when the feed is empty; `project()` + honors an event's `channelId`; long trigger prompts collapse; memory panel + layout; shared `TrashIcon` for deletes. + +## Run it + +```bash +# from the worktree root +pnpm install +pnpm --filter @tanstack/ai-harness build # dashboard imports the built dist + +pnpm --filter agent-dashboard dev # http://localhost:3002 (no API key needed) +pnpm --filter agent-dashboard test:e2e # Playwright + +npx @tanstack/ai-dashboard --port 8790 # the self-hosted remote control +``` + +Demo: on the home page, **Add to team** an agent. Open the **Demo Controls** +devtools panel (bottom-left) → **Start triage demo** → approve the drafted reply +mid-run → **+ Add agent** to reveal the roster. In **Automations**, run +`fetch_stats` now, add a schedule, send a test webhook, or simulate the host +offline. Then try **Meta-chat**, **History**, **Spend**, **Config**. + +PR-watcher: **+ PR-watcher demo** → **Send PR webhook**. A `#pr-…` channel opens +and the security agent flags the PR; reply **"not a security problem — we're +intranet-only here"**, watch it write pod memory, then **Send PR webhook** again. + +Reddit pod: **+ React-news demo** → **🔧 tools** → run +`reddit.search_react_news` (or schedule it). Set `ANTHROPIC_API_KEY` before +`dev` for a live digest. + +## Verification + +- `examples/agent-dashboard`: **20 Playwright e2e** (approval mid-run, config, + spend, history + replay, meta-chat, teams, injection, PR-watcher loop, DM, + memory panel, persistence + approval replay, Reddit pod, trace, controls, + product automations, MCP); **5 vitest** (RSS parser + pricing). +- `@tanstack/ai-dashboard`: **4 vitest** (relay, child-agent blocks). +- After the 2026-10-02 merge with `main`: `test:types` + `test:lib` pass for + `agent-dashboard`, `@tanstack/ai-dashboard`, `@tanstack/ai-harness`. Full + `test:pr` and e2e not re-run yet. + +## Deliberate decisions / open questions (flagged for @AlemTuzlak) + +- **Chart**: SSR-safe SVG, **not** `@tanstack/react-charts` (0.18 auto-sizing + loops a ResizeObserver and freezes the main thread). +- **Location**: `examples/agent-dashboard` (no `apps/` dir in the repo). +- **Auth**: local single-user. Multi-user auth beyond pairing/host-token is a follow-up. +- **Spend**: tokens-first with per-session budgets as data; cost is a pluggable follow-up. + +## Known cosmetic issue + +- `@ag-ui/client` warns when it strips a nonstandard `/toolName` field the + harness core adds to `TOOL_CALL_START`. Warning only. diff --git a/docs/config.json b/docs/config.json index e4a77d62b5..9d7610db3a 100644 --- a/docs/config.json +++ b/docs/config.json @@ -1048,6 +1048,11 @@ "label": "Self-host the dashboard", "to": "harness/dashboard", "addedAt": "2026-09-26" + }, + { + "label": "Stream a harness over AG-UI", + "to": "harness/ag-ui", + "addedAt": "2026-09-26" } ] }, diff --git a/docs/harness/ag-ui.md b/docs/harness/ag-ui.md new file mode 100644 index 0000000000..0239f59800 --- /dev/null +++ b/docs/harness/ag-ui.md @@ -0,0 +1,146 @@ +--- +title: Stream a harness over AG-UI +id: harness-ag-ui +order: 13 +description: "Serve a harness session as a pure AG-UI event stream a bare @ag-ui/client can consume, with approvals as interrupts and token usage as metadata." +keywords: + - tanstack ai + - harness + - AG-UI + - SSE + - interrupts + - usage +--- + +A harness session already speaks AG-UI: `session.events()` yields events whose payloads are AG-UI protocol events. The `@tanstack/ai-harness/ag-ui` subpath is the thin, documented seam an AG-UI client consumes — a normalizer and an SSE handler a bare [`@ag-ui/client`](https://docs.ag-ui.com) `HttpAgent` can point at. + +This does not replace the [harness protocol](./connect.md) or the [relay dashboard](./dashboard.md). AG-UI is the *run stream*; the harness protocol carries the *control plane* (pairing, snapshots, history, config). Use both: AG-UI for the conversation, the harness protocol for everything around it. + +## Serve a session over AG-UI + +`createAgUiHandler` is one `fetch` handler. `authorize` is required, so no endpoint is open by accident. + +```ts group=harness-ag-ui +import { defineHarness, createHarnessHost } from '@tanstack/ai-harness' +import { createAgUiHandler } from '@tanstack/ai-harness/ag-ui' +import { memoryPersistence } from '@tanstack/ai-persistence' +import { openaiText } from '@tanstack/ai-openai' + +const assistant = defineHarness({ + name: 'acme/assistant', + adapter: openaiText('gpt-5.6'), +}) + +const host = createHarnessHost({ persistence: memoryPersistence() }) + +export const handler = createAgUiHandler({ + host, + harness: assistant, + authorize: (request) => + request.headers.get('authorization') === + `Bearer ${process.env.HARNESS_TOKEN}` + ? { id: 'user-1' } + : null, + canAccess: (principal, threadId) => threadId.startsWith(principal.id), +}) +``` + +`POST` a `RunAgentInput` to it. A body with a trailing user message runs a prompt; a body with `resume` entries answers the last turn's interrupts and continues. The response is `text/event-stream` of AG-UI events, encoded with `@ag-ui/encoder` so a protobuf-accepting client gets binary framing for free. The stream closes when the run reaches a terminal state. + +## Consume it from a client + +Point a bare `@ag-ui/client` `HttpAgent` at the handler's URL: + +```ts group=harness-ag-ui-client +import { HttpAgent } from '@ag-ui/client' + +const agent = new HttpAgent({ + url: 'https://example.com/agent', + threadId: 'user-1/main', + headers: { authorization: `Bearer ${process.env.HARNESS_TOKEN}` }, +}) + +agent.addMessage({ id: 'u1', role: 'user', content: 'Summarize incidents' }) +await agent.runAgent() + +console.log(agent.messages.at(-1)) // the assistant's reply +``` + +The stream is strict AG-UI by default: it begins with `RUN_STARTED`, as a bare client requires. (The harness-native `CUSTOM` control events — `harness.operation.*`, `harness.question`, and friends — are dropped from the AG-UI stream; read them over the harness protocol's `/events` tier instead. Pass `stream: { includeHarnessEvents: true }` to keep them for a lenient consumer.) + +## Event mapping + +| Harness concept | AG-UI event | +| --- | --- | +| run start / end | `RUN_STARTED` / `RUN_FINISHED` | +| assistant text | `TEXT_MESSAGE_START` / `TEXT_MESSAGE_CONTENT` / `TEXT_MESSAGE_END` | +| tool call | `TOOL_CALL_START` / `TOOL_CALL_ARGS` / `TOOL_CALL_END` / `TOOL_CALL_RESULT` | +| approval / wait | `RUN_FINISHED` with `outcome.type === 'interrupt'` | +| resume | a new run whose `RunAgentInput.resume` answers the interrupts | +| subagent | `SUBAGENT_STARTED` / `SUBAGENT_FINISHED` (with `subagentRunId`) | +| token usage | `metadata.tanstack.usage` on `RUN_FINISHED` | +| live spend (interim) | a `tanstack.spend` `CUSTOM` event | + +## Approvals as interrupts + +A turn that stops for approval finishes with a `RUN_FINISHED` whose `outcome.type === 'interrupt'`. `@ag-ui/client` collects these on `agent.pendingInterrupts`: + +```ts group=harness-ag-ui-client +await agent.runAgent() + +for (const interrupt of agent.pendingInterrupts) { + console.log(interrupt.id, interrupt.message) // "Approval required to run remove" +} +``` + +Answer by starting a new run whose `resume` entries reference the interrupt ids: + +```ts group=harness-ag-ui-resume +await fetch('https://example.com/agent', { + method: 'POST', + headers: { + 'content-type': 'application/json', + authorization: `Bearer ${process.env.HARNESS_TOKEN}`, + }, + body: JSON.stringify({ + threadId: 'user-1/main', + runId: 'run-2', + messages: [], + tools: [], + context: [], + resume: [{ interruptId: 'approval_call_1', status: 'resolved', payload: true }], + }), +}) +``` + +The harness-native approval path (`session.resolve`, the CLI, and the relay dashboard) keeps working; AG-UI interrupt/resume is the AG-UI-shaped view of the same wait. + +## Token usage and spend + +`RUN_FINISHED` carries usage under `metadata.tanstack.usage` as a `NormalizedUsage` (`inputTokens`, `outputTokens`, `totalTokens`). Read it directly, or normalize any provider's usage with `normalizeUsage`: + +```ts group=harness-ag-ui-usage +import { normalizeUsage } from '@tanstack/ai-harness/ag-ui' + +normalizeUsage([{ inputTokens: 10, outputTokens: 5 }]) +// { inputTokens: 10, outputTokens: 5, totalTokens: 15 } +normalizeUsage({ promptTokens: 8, completionTokens: 2, totalTokens: 10 }) +// { inputTokens: 8, outputTokens: 2, totalTokens: 10 } +``` + +For a live spend meter, opt into interim ticks. After each run that reports usage, the stream carries a `tanstack.spend` `CUSTOM` event with that run's usage and a running cumulative total: + +```ts group=harness-ag-ui +const meteredHandler = createAgUiHandler({ + host, + harness: assistant, + authorize: () => ({ id: 'user-1' }), + stream: { emitSpendEvents: true }, +}) +``` + +Spend rides `CUSTOM` because the AG-UI spec has no first-class spend event yet — track that convention deliberately. + +## Pin the AG-UI version + +The AG-UI spec is still evolving. This bridge is built and tested against `@ag-ui/core@1.0.0`, `@ag-ui/encoder@1.0.0`, and `@ag-ui/client@1.0.0`. Pin those versions and move them on purpose, not by floating range. diff --git a/examples/agent-dashboard/.gitignore b/examples/agent-dashboard/.gitignore new file mode 100644 index 0000000000..771f390df3 --- /dev/null +++ b/examples/agent-dashboard/.gitignore @@ -0,0 +1,15 @@ +node_modules +.DS_Store +dist +dist-ssr +*.local +.env +.env.local +.nitro +.tanstack +.output +.vinxi +test-results +playwright-report +*.log +.data diff --git a/examples/agent-dashboard/README.md b/examples/agent-dashboard/README.md new file mode 100644 index 0000000000..4225f60198 --- /dev/null +++ b/examples/agent-dashboard/README.md @@ -0,0 +1,194 @@ +# Agent Dashboard + +Mission control for TanStack AI agents — a TanStack Start app that watches live +agent sessions, approves tool calls mid-run, and tracks spend, all as a live +projection of the [AG-UI](../../docs/harness/ag-ui.md) event stream. + +It embeds a deterministic **support-triage** agent, so it runs with **no API +key**: the agent looks up a ticket (an auto tool), drafts a customer reply (an +approval-gated tool that pauses the run), and sends it once you approve. (One +team is the exception — the **Reddit pod** uses a real service and a real LLM; +see below.) + +```bash +pnpm --filter agent-dashboard dev # http://localhost:3002 +``` + +## What it exercises + +- **TanStack Start** — app shell + routing: hosts → sessions → session detail + (`src/routes`), and server API routes that host the agent (`src/routes/api.*`). +- **`@tanstack/ai-harness/ag-ui`** — the session view consumes the AG-UI SSE + stream with a bare `@ag-ui/client` `HttpAgent` (`src/lib/session-controller.ts`); + no bespoke protocol code. +- **TanStack DB** — run state (messages, tool calls, approvals, spend) lives in + `localOnly` collections written from the stream and read with `useLiveQuery` + (`src/db/collections.ts`). The UI is a projection of the stream, not a poller. +- **TanStack Query** — server state: host/session/run lists and agent config. +- **Composing teams (product UI)** — the home page is an **agents table**; each + row's **Add to team** starts a new team with that agent or drops it into an + existing one. On a team, **Add agent** adds any available agent, and each + roster agent has a **Run a tool** (wrench) button to run one of its public tools with + JSON parameters. Agents carry **default subscriptions** by harness (what they + react to), so a hand-composed team behaves like a seeded one. +- **Styling** — Tailwind v4 with the TanStack AI design tokens (colors, type, + radii) in `src/styles.css`. Dark is the default; the rail's theme toggle + switches to light. Icons are Phosphor (`@phosphor-icons/react`). +- **TanStack DevTools** — the demo-only scaffolding (the seeded-team launchers, + the triage demo, automations, pod memory) lives in a custom **Demo Controls** + panel, kept out of the product UI so it's clear what's scaffolding vs. the real + experience. The panel renders from the devtools root (outside the route tree) + and drives the app purely by reading the same live TanStack DB state the UI + does — so it doubles as a state-management stress test. + +## Reddit pod (real service, real AI) + +The **+ React-news demo** team is the first pod wired to a _real_ external +service and a _real_ LLM — the graduation from the scripted demo agents: + +- **`reddit/fetcher`** — a procedural agent (no LLM) carrying one real tool, + `reddit.search_react_news`, which reads Reddit's public **RSS (Atom)** feed + (read-only, no auth, no key). Run it from the roster's **Run a tool** button, a + 30-min schedule, or the Demo Controls panel. +- **`sentiment/react`** — a **real LLM** agent (Anthropic). Its harness default + subscription is the fetcher's tool _result_ + (`{ event: 'tool_result', tool: 'reddit.search_react_news', action: 'trigger' }`), + so whether you spin up the seeded demo or compose the team by hand, a news + batch triggers it automatically — nobody runs it — and it posts a sentiment + digest, persisting standout signals to pod memory. + +The loop is **timer → tool result → subscription → LLM digest**. The trigger +path spends zero tokens; the only cost is the digest itself. Chat messages don't +match the `tool_result` subscription, so the loop doesn't feed itself (a Phase 4 +scoping case study — verified by letting it run several cycles). + +### The one manual step: an API key + +`sentiment/react` needs a real provider. Set `ANTHROPIC_API_KEY` before starting +the dev server to get a live digest: + +```bash +ANTHROPIC_API_KEY=sk-ant-... pnpm --filter agent-dashboard dev +``` + +Without a key the agent posts a "set the key" message instead of a digest — the +rest of the dashboard still runs key-free. The e2e suite uses a deterministic +double and a recorded Reddit fixture, so it needs neither a key nor the network. + +> **Why RSS, not `.json`:** Reddit's public JSON (`/r/x/new.json`) returns `403` +> for many datacenter/VPN egress IPs regardless of `User-Agent`, while the Atom +> feed (`/r/x/new.rss`) is served — so the tool reads RSS. RSS carries title, +> link, author, timestamp and body (enough for a digest) but not score/comment +> counts. Reddit still rate-limits bursts (a rapid retry can `429`); the 30-min +> schedule stays well clear. The e2e suite uses the recorded fixture, so it +> needs neither network nor key. + +## Control plane + +- **History** (`/history`) — past runs backed by `HarnessPersistence` + (`runs.listByThread`). Opening a session replays it from stored events + (`/api/replay`), so a session you didn't run in this tab rehydrates from the + persisted feed. +- **Spend** (`/spend`) — per-session token rollups (a live query over the + `spend` collection) with per-session budgets and over-budget alerts. +- **Config** (`/config`) — a form generated from the agent's typed + `ConfigOption` schemas; edits write through the harness protocol + (`op: 'config'`). Endpoints are versioned with `HARNESS_PROTOCOL_VERSION`. + +## The approval queue + +When a run pauses on an approval, an approval card appears. It resolves the +interrupt through **both** paths the interrupt supports: + +- **Approve** → the AG-UI resume flow (`runAgent({ resume })` over `/api/agent`), + so the continuation streams back into the view. +- **Deny** → the harness-native control endpoint (`/api/harness/control`). +- **Edit** → approve with edited tool arguments. + +## Operate a live run + +Open a team and start a run. The team page gives you these controls: + +1. Select **Trace** to see run, model, tool, and approval spans in a live waterfall. +2. Select **approve** or **deny** on an approval span to continue the run. +3. Send a message during a run to steer it. +4. Select **Stop** to cancel the run. +5. Answer an agent question with its question card. + +The **Spend** page shows token use and estimated dollar cost. Scripted demo +agents cost `$0.00`. The `sentiment/react` agent uses the price for +`claude-haiku-4-5` from the Anthropic model metadata. + +## Manage automations + +Use the **Automations** section on a team page to run work on a schedule or from +a webhook. + +1. Select an agent and one of its public tools. +2. Add an interval or a five-field cron schedule. +3. Use the schedule list to see the next three runs, pause the schedule, or delete it. +4. Add a webhook and copy its endpoint. +5. Use the delivery list to inspect a result or retry the payload. + +## Manage the dashboard through MCP + +The dashboard serves its management tools at `/api/mcp`. Connect Claude Code: + +```bash +claude mcp add --transport http dashboard http://localhost:3002/api/mcp +``` + +The MCP server can list dashboard state, send and steer messages, stop runs, +answer questions, resolve approvals, run public tools, and control schedules. + +This local demo does not require authentication. Add an authentication gate +before you expose the endpoint outside localhost. + +## Meta-chat (the demo) + +`/chat` is the dashboard's **own** agent (`dashboard/meta`), a tool-using chat +over live dashboard state. It runs on the same host as every other agent, so its +runs and tool calls show up in History and the trace view — the dashboard +dogfooding itself. Its tools: `list_agents`, `list_sessions`, `query_runs`, +`get_agent_config`, `set_agent_config`, `summarize_session`. + +Demo script: + +1. Open `/chat`. +2. Ask **"List the agents on this host"** — it calls `list_agents` (visible + inline) and answers with the agents registered on the host. +3. Run a triage session (open the **Demo Controls** devtools panel → **+ New + team** → **Start triage demo**). +4. Back in `/chat`, ask **"How many runs so far?"** and **"Summarize the latest + session"** — it queries live run history and the session snapshot. +5. Open **History** — the meta-chat's own runs are listed alongside the agents', + and each replays. + +## Server wiring + +- `POST /api/agent` — the AG-UI run stream (`createAgUiHandler`, spend ticks on). +- `GET|POST /api/harness/*` — the harness control tier (`createHarnessHandler`): + `snapshot`, `events`, `control`. +- `GET /api/hosts`, `GET /api/sessions` — host and session lists. +- `GET|POST /api/config` — read/write agent config (versioned). +- `GET /api/runs` — run history; `GET /api/replay` — a session's stored events. +- `POST /api/meta` — the meta-chat's AG-UI run stream. + +## State + +Server state is durable, so you can close the tab, restart the server, and find +each team where you left it. It lives in one JSON file (`.data/state.json`, +gitignored; override with `DASHBOARD_STATE_FILE`), written through on every +mutation and replayed on boot (`src/server/store.ts`). Persisted: chat run state +(messages, runs, interrupts, agent config) plus the dashboard's side tables +(threads, pod memory, schedules, webhooks). Not persisted: the in-flight +injection queue and the dev offline toggle. It's a single-process file store — a +multi-node dashboard would swap in a real database behind the same seam. + +## Test + +```bash +pnpm --filter agent-dashboard test:e2e # Playwright: stream + approve mid-run +``` + +Generated with Claude Code. diff --git a/examples/agent-dashboard/e2e/automations.spec.ts b/examples/agent-dashboard/e2e/automations.spec.ts new file mode 100644 index 0000000000..ea7014796e --- /dev/null +++ b/examples/agent-dashboard/e2e/automations.spec.ts @@ -0,0 +1,29 @@ +import { expect, test } from '@playwright/test' +import { closeDemo, openDemo } from './devtools' + +test('manages cron schedules and webhook deliveries from the team page', async ({ + page, +}) => { + await page.goto('/') + await openDemo(page) + await page.getByRole('button', { name: '+ New team' }).click() + await closeDemo(page) + + await page.getByLabel('schedule cadence').selectOption('cron') + await page.getByLabel('cron expression').fill('*/30 * * * *') + await page.getByRole('button', { name: '+ Add schedule' }).click() + await expect(page.getByText(/upcoming:/)).toContainText('|') + await page.getByRole('button', { name: 'pause' }).click() + await expect(page.getByRole('button', { name: 'resume' })).toBeVisible() + + await page.getByRole('button', { name: '+ Add webhook' }).click() + const endpoint = await page + .locator('code') + .filter({ hasText: '/api/webhooks/' }) + .textContent() + expect(endpoint).toBeTruthy() + await page.request.post(endpoint!, { data: { queue: 'e2e' } }) + await expect(page.getByText('accepted')).toBeVisible() + await page.getByRole('button', { name: 'retry' }).click() + await expect(page.getByText('accepted')).toHaveCount(2) +}) diff --git a/examples/agent-dashboard/e2e/control-plane.spec.ts b/examples/agent-dashboard/e2e/control-plane.spec.ts new file mode 100644 index 0000000000..a47f9a0cec --- /dev/null +++ b/examples/agent-dashboard/e2e/control-plane.spec.ts @@ -0,0 +1,55 @@ +import { expect, test } from '@playwright/test' +import { answerAgentQuestion, closeDemo } from './devtools' + +test('config form reads and writes ConfigOption schemas', async ({ page }) => { + await page.goto('/config') + const tone = page.getByLabel('tone') + await expect(tone).toBeVisible() + await tone.selectOption('formal') + + // Written through the harness protocol and persisted. + await page.waitForTimeout(500) + const res = await page.request.get('/api/config?threadId=settings') + const body = await res.json() + const toneValue = body.options.find( + (o: { key: string }) => o.key === 'tone', + ).value + expect(toneValue).toBe('formal') +}) + +test('spend dashboard shows live token usage after a run', async ({ page }) => { + const threadId = `spend-${Date.now()}` + await page.goto(`/sessions/${threadId}`) + await closeDemo(page) + await page.getByRole('button', { name: 'Start triage demo' }).click() + await answerAgentQuestion(page) + await expect( + page.getByText('Approval required', { exact: true }), + ).toBeVisible() + + // Client-side nav (no reload) so the in-memory TanStack DB projection survives. + await page.getByRole('link', { name: 'Spend' }).click() + await expect(page.getByText(/tokens across/)).toBeVisible() + await expect(page.getByText(threadId).first()).toBeVisible() +}) + +test('run history lists a run and replays it into the session view', async ({ + page, +}) => { + const threadId = `replay-${Date.now()}` + // Run server-side so the browser has no local state for this thread. + await page.request.post('/api/run', { + data: { + threadId, + harness: 'dashboard/meta', + message: 'List agents', + }, + }) + + await page.goto('/history') + await expect(page.getByText(threadId).first()).toBeVisible() + + // Open the session fresh — it rehydrates from stored events (replay). + await page.goto(`/sessions/${threadId}`) + await expect(page.getByText('list_agents').first()).toBeVisible() +}) diff --git a/examples/agent-dashboard/e2e/control.spec.ts b/examples/agent-dashboard/e2e/control.spec.ts new file mode 100644 index 0000000000..90904456b3 --- /dev/null +++ b/examples/agent-dashboard/e2e/control.spec.ts @@ -0,0 +1,17 @@ +import { expect, test } from '@playwright/test' +import { closeDemo, openDemo } from './devtools' + +test('answers an agent question, steers, and stops a run', async ({ page }) => { + await page.goto('/') + await openDemo(page) + await page.getByRole('button', { name: '+ New team' }).click() + await page.getByRole('button', { name: 'Start triage demo' }).click() + await closeDemo(page) + + await expect(page.getByText('Agent question')).toBeVisible() + await expect(page.getByRole('button', { name: 'Steer' })).toBeVisible() + await page.getByPlaceholder('Send a message…').fill('Focus on the customer.') + await page.getByRole('button', { name: 'Steer' }).click() + await page.getByRole('button', { name: 'Stop' }).click() + await expect(page.getByRole('button', { name: 'Send' })).toBeVisible() +}) diff --git a/examples/agent-dashboard/e2e/dashboard.spec.ts b/examples/agent-dashboard/e2e/dashboard.spec.ts new file mode 100644 index 0000000000..ff9ee67ed9 --- /dev/null +++ b/examples/agent-dashboard/e2e/dashboard.spec.ts @@ -0,0 +1,39 @@ +import { expect, test } from '@playwright/test' +import { answerAgentQuestion, closeDemo } from './devtools' + +// The dashboard embeds a deterministic support-triage agent, so this runs with +// no API key: send a prompt, watch the AG-UI stream, approve a tool call +// mid-run, and see the run resume — the north-star approval demo. +test('streams a session and approves a tool call mid-run', async ({ page }) => { + const threadId = `e2e-${Date.now()}` + await page.goto(`/sessions/${threadId}`) + await closeDemo(page) + + const startButton = page.getByRole('button', { name: 'Start triage demo' }) + await expect(startButton).toBeVisible() + await startButton.click() + await answerAgentQuestion(page) + + // The agent looks up the ticket (auto tool) then drafts a reply for approval. + await expect(page.getByText('lookup_ticket').first()).toBeVisible() + await expect( + page.getByText("Here's a draft reply for your approval."), + ).toBeVisible() + + // The approval card appears (the run paused on the interrupt). + await expect( + page.getByText('Approval required', { exact: true }), + ).toBeVisible() + await expect(page.getByText('send_reply').first()).toBeVisible() + + // Spend meter is live (tokens accrued from the stream). + await expect(page.getByText(/[1-9][0-9,]* tokens/)).toBeVisible() + + // Approve via the AG-UI resume flow; the run continues and finishes. + await page.getByRole('button', { name: 'Approve', exact: true }).click() + + await expect(page.getByText(/Sent ✅/)).toBeVisible() + await expect( + page.getByText('Approval required', { exact: true }), + ).toHaveCount(0) +}) diff --git a/examples/agent-dashboard/e2e/devtools.ts b/examples/agent-dashboard/e2e/devtools.ts new file mode 100644 index 0000000000..55e251f001 --- /dev/null +++ b/examples/agent-dashboard/e2e/devtools.ts @@ -0,0 +1,45 @@ +import type { Page } from '@playwright/test' + +// The demo controls now live in a custom TanStack DevTools panel (see +// src/components/demo-controls.tsx). Under E2E the panel opens on load +// (VITE_E2E, see playwright.config.ts), so demo controls are reachable by +// default. The open panel is a fixed overlay along the bottom, so it can cover +// bottom-of-page product controls (the message box, roster "run" buttons) — +// close it before clicking those, reopen before the next demo control. + +const trigger = (page: Page) => + page.getByRole('button', { name: 'Open TanStack Devtools' }) + +/** Open the demo-controls panel if it's currently closed. */ +export async function openDemo(page: Page) { + // When the panel is open the trigger is hidden; if it's visible, we're closed. + if ( + await trigger(page) + .isVisible() + .catch(() => false) + ) { + await trigger(page).click() + } +} + +/** Close the panel so bottom-of-page product controls are clickable. */ +export async function closeDemo(page: Page) { + if ( + await trigger(page) + .isVisible() + .catch(() => false) + ) + return // already closed + await page.locator('button.close').first().click() +} + +export async function answerAgentQuestion(page: Page) { + // The question card sits at the bottom of the stream, under the open panel. + await closeDemo(page) + await page.getByText('Agent question').first().waitFor() + await page.getByPlaceholder('Your answer').first().fill('yes') + await page + .getByRole('button', { name: 'Answer', exact: true }) + .first() + .click() +} diff --git a/examples/agent-dashboard/e2e/injection.spec.ts b/examples/agent-dashboard/e2e/injection.spec.ts new file mode 100644 index 0000000000..f70caa349b --- /dev/null +++ b/examples/agent-dashboard/e2e/injection.spec.ts @@ -0,0 +1,85 @@ +import { expect, test } from '@playwright/test' +import { closeDemo, openDemo } from './devtools' +import type { Page } from '@playwright/test' + +// Phase 2: the dashboard invokes work deterministically — run-now, timers, and +// webhooks — with the structured result streaming into the channel. These share +// server-side state (the scheduler + the offline flag), so run them serially. +test.describe.configure({ mode: 'serial' }) + +async function newTeam(page: Page) { + await page.goto('/') + // The team-creation demos now live in the Demo Controls devtools panel. + await openDemo(page) + await page.getByRole('button', { name: '+ New team' }).click() + await expect(page).toHaveURL(/\/teams\//) + // The automations panel is present once the channel view mounts. + await expect(page.getByText('Automations').first()).toBeVisible() +} + +test('run-now executes a public tool; private tools are never listed', async ({ + page, +}) => { + await newTeam(page) + + // Only the public tool (fetch_stats) is offered — the reply tools stay private. + await expect( + page.getByRole('button', { name: 'run fetch_stats' }), + ).toBeVisible() + await expect( + page.getByRole('button', { name: /run send_reply/ }), + ).toHaveCount(0) + await expect( + page.getByRole('button', { name: /run lookup_ticket/ }), + ).toHaveCount(0) + + // Run it now — the result streams back as a structured, injected card. + await page.getByRole('button', { name: 'run fetch_stats' }).click() + await expect(page.getByText('manual trigger')).toBeVisible() + await expect(page.getByText('fetch_stats').first()).toBeVisible() +}) + +test('a scheduled timer fires a public tool into the channel', async ({ + page, +}) => { + await newTeam(page) + + // Fire every second so the test doesn't wait on a slow interval. + await closeDemo(page) + await page.getByLabel('schedule interval seconds').fill('1') + await page.getByRole('button', { name: '+ Add schedule' }).click() + // The timer (dashboard-owned clock) fires within a couple of seconds. + await expect(page.getByText('timer trigger').first()).toBeVisible({ + timeout: 10000, + }) + + // Pause it so it doesn't keep firing after the test. + await page.getByRole('button', { name: 'pause' }).first().click() +}) + +test('a webhook produces the same injected result as the other triggers', async ({ + page, +}) => { + await newTeam(page) + + await page.getByRole('button', { name: 'Send test webhook' }).click() + await expect(page.getByText('webhook trigger')).toBeVisible() +}) + +test('an injected job queues while the host is offline and flushes on reconnect', async ({ + page, +}) => { + await newTeam(page) + + await page.getByRole('button', { name: 'Simulate host offline' }).click() + await expect(page.getByText(/host offline/)).toBeVisible() + + // Injecting while offline queues instead of running — no card appears yet. + await page.getByRole('button', { name: 'run fetch_stats' }).click() + await expect(page.getByText(/1 queued/)).toBeVisible() + await expect(page.getByText('manual trigger')).toHaveCount(0) + + // Reconnect → the queue flushes → the result appears. + await page.getByRole('button', { name: 'Bring host online' }).click() + await expect(page.getByText('manual trigger')).toBeVisible({ timeout: 8000 }) +}) diff --git a/examples/agent-dashboard/e2e/mcp.spec.ts b/examples/agent-dashboard/e2e/mcp.spec.ts new file mode 100644 index 0000000000..dddb4f233c --- /dev/null +++ b/examples/agent-dashboard/e2e/mcp.spec.ts @@ -0,0 +1,35 @@ +import { expect, test } from '@playwright/test' +import { createMCPClient } from '@tanstack/ai-mcp' +import { closeDemo, openDemo } from './devtools' + +test('exposes dashboard management tools over MCP', async ({ page }) => { + await page.goto('/') + await openDemo(page) + await page.getByRole('button', { name: '+ New team' }).click() + await closeDemo(page) + await page.waitForTimeout(250) + + const roster = await page.request + .get('/api/roster') + .then((response) => response.json()) + const teamId = new URL(page.url()).pathname.split('/').pop() + const threadId = roster.memberships.find( + (membership: { teamId?: string }) => membership.teamId === teamId, + ).threadId as string + const client = await createMCPClient({ + transport: { type: 'http', url: 'http://localhost:3002/api/mcp' }, + }) + try { + const names = (await client.tools()).map((tool) => tool.name) + expect(names).toContain('list_teams') + expect(names).toContain('send_message') + expect(names).toContain('resolve_approval') + await client.callTool('send_message', { + threadId, + message: 'Please handle ticket T-1042.', + }) + await expect(page.getByText('Agent question')).toBeVisible() + } finally { + await client.close() + } +}) diff --git a/examples/agent-dashboard/e2e/meta-chat.spec.ts b/examples/agent-dashboard/e2e/meta-chat.spec.ts new file mode 100644 index 0000000000..565f036914 --- /dev/null +++ b/examples/agent-dashboard/e2e/meta-chat.spec.ts @@ -0,0 +1,26 @@ +import { expect, test } from '@playwright/test' +import { closeDemo } from './devtools' + +// The meta-chat is the dashboard's own tool-using agent. A quick prompt makes +// it call a tool over live state and summarize the result, with the tool call +// visible inline (dogfooding the trace view). +test('meta-chat calls a tool over live state and summarizes', async ({ + page, +}) => { + await page.goto('/chat') + await closeDemo(page) + + await page + .getByRole('button', { name: 'List the agents on this host' }) + .click() + + // The tool call is visible in the trace… + await expect(page.getByText('list_agents').first()).toBeVisible() + // …and the agent summarizes the real result: triage, meta, the two PR-watcher + // demo agents, and the two Reddit-pod agents are registered on this host. + await expect(page.getByText(/host runs 6 agent/)).toBeVisible() + + // Its run shows up in history like any other agent. + await page.getByRole('link', { name: 'History' }).click() + await expect(page.getByText(/meta-/).first()).toBeVisible() +}) diff --git a/examples/agent-dashboard/e2e/persistence.spec.ts b/examples/agent-dashboard/e2e/persistence.spec.ts new file mode 100644 index 0000000000..d859d35700 --- /dev/null +++ b/examples/agent-dashboard/e2e/persistence.spec.ts @@ -0,0 +1,37 @@ +import { expect, test } from '@playwright/test' +import { answerAgentQuestion, openDemo } from './devtools' + +// Server-side persistence: a team, its runs, and its resolved approvals survive a +// full page reload (the roster rehydrates from the server; the session replays +// from stored events). Regression guard for two bugs: +// 1. the team vanished on reload (roster was client-only), and +// 2. an already-approved tool call resurfaced as "Approval required" because the +// tail replayed the historical interrupt as if it were live. +test('a team and its approved run survive a reload', async ({ page }) => { + await page.goto('/') + + // Create a triage team and run it to the approval (demos live in the panel). + await openDemo(page) + await page.getByRole('button', { name: '+ New team' }).click() + await expect(page).toHaveURL(/\/teams\//) + await page.getByRole('button', { name: 'Start triage demo' }).click() + await answerAgentQuestion(page) + await expect( + page.getByText('Approval required', { exact: true }).first(), + ).toBeVisible() + + // Approve via the AG-UI resume flow; the run continues and finishes. + await page.getByRole('button', { name: 'Approve', exact: true }).click() + await expect(page.getByText(/Sent ✅/)).toBeVisible() + await expect( + page.getByText('Approval required', { exact: true }), + ).toHaveCount(0) + + // Reload: the team must reappear (roster is persisted) and the finished run must + // replay — WITHOUT the resolved approval coming back. + await page.reload() + await expect(page.getByText(/Sent ✅/)).toBeVisible() + await expect( + page.getByText('Approval required', { exact: true }), + ).toHaveCount(0) +}) diff --git a/examples/agent-dashboard/e2e/pr-watcher.spec.ts b/examples/agent-dashboard/e2e/pr-watcher.spec.ts new file mode 100644 index 0000000000..41df6d7df4 --- /dev/null +++ b/examples/agent-dashboard/e2e/pr-watcher.spec.ts @@ -0,0 +1,118 @@ +import { expect, test } from '@playwright/test' +import { closeDemo, openDemo } from './devtools' +import type { Page } from '@playwright/test' + +// Phase 3: the PR-watcher "the pod learns" loop. A webhook opens a per-PR channel, +// a subscribed security agent reviews it, the human corrects it, the correction is +// persisted to pod memory, and the next PR is handled better. These share +// server-side state (the seeded PR sequence, pod memory), so run serially. +test.describe.configure({ mode: 'serial' }) + +async function newPrWatcherTeam(page: Page) { + await page.goto('/') + // The seeded demos live in the Demo Controls devtools panel now. + await openDemo(page) + await page.getByRole('button', { name: '+ PR-watcher demo' }).click() + await expect(page).toHaveURL(/\/teams\//) + await expect(page.getByText('Members · 2')).toBeVisible() + await expect(page.getByText('Automations').first()).toBeVisible() +} + +test('the PR-watcher loop: review, correct, remember, handle the next PR better', async ({ + page, +}) => { + await newPrWatcherTeam(page) + + // 1–2. Webhook in → the watcher opens a per-PR channel and posts a summary. + await page.getByRole('button', { name: 'Send PR webhook' }).click() + await expect(page.getByText(/Channel #pr-\d+ created/)).toBeVisible() + // The new channel shows up in the sidebar; open it. + const firstChannel = page.getByRole('button', { name: /# pr-\d+/ }).first() + await expect(firstChannel).toBeVisible() + const firstChannelName = ((await firstChannel.textContent()) ?? '').trim() + await firstChannel.click() + + // 3. The watcher's summary and the security agent's finding are both in-channel, + // and the finding flags the public-internet exposure (no memory yet). + await expect(page.getByText(/Requesting a security review/)).toBeVisible() + await expect(page.getByText(/Security finding/)).toBeVisible() + // Match the finding message specifically (the PR risk also mentions "public + // internet" but lives in the collapsed tool-result JSON tree). + await expect( + page.getByText(/exposes an endpoint to the public internet/), + ).toBeVisible() + + // 4. The human corrects it. The agent persists the standing instruction. + // Close the demo panel so it doesn't cover the message box / Send button. + await closeDemo(page) + await page + .getByPlaceholder('Send a message…') + .fill("not a security problem — we're intranet-only here") + await page.getByRole('button', { name: 'Send', exact: true }).click() + + // The write is a visible tool call; its args render as a collapsed JSON tree, + // so the key is present in the DOM but not visible until expanded. + await expect(page.getByText('pod.memory_write').first()).toBeVisible() + await expect(page.locator('body')).toContainText('intranet-policy') + + // 5. The next PR: a second webhook opens a new channel; this time the review is + // clean and the run carries the "memory attached" badge. + await page.getByRole('button', { name: /# main/ }).click() + // Reopen the demo panel to reach the webhook button again. + await openDemo(page) + await page.getByRole('button', { name: 'Send PR webhook' }).click() + + // A second, different PR channel appears; open it. + const secondChannel = page + .getByRole('button', { name: /# pr-\d+/ }) + .filter({ hasNotText: firstChannelName.replace(/^#\s*/, '') }) + await expect(secondChannel.first()).toBeVisible({ timeout: 15000 }) + await secondChannel.first().click() + + // The pod learned: no intranet finding this time, and memory was attached. + await expect(page.getByText(/intranet-only per policy/)).toBeVisible() + await expect(page.getByText(/memory entr(y|ies) attached/)).toBeVisible() + await expect(page.getByText(/Security finding/)).toHaveCount(0) +}) + +test('a DM with an agent is created and routes messages to it', async ({ + page, +}) => { + await newPrWatcherTeam(page) + + // Create a DM with the security member from the roster (a page control, so + // close the demo panel that would otherwise cover it). + await closeDemo(page) + await page.getByRole('button', { name: /New DM with security/ }).click() + + // A dm channel appears in the sidebar; open it. + const dm = page.getByRole('button', { name: /@ dm-/ }) + await expect(dm).toBeVisible({ timeout: 15000 }) + await dm.click() + + // The DM seats the security agent as its member, so a message reaches it and + // it replies (regression: DMs used to be created with no members → no reply). + await page.getByPlaceholder('Send a message…').fill('please review this') + await page.getByRole('button', { name: 'Send', exact: true }).click() + await expect(page.getByText(/Security finding/).first()).toBeVisible({ + timeout: 15000, + }) +}) + +test('the memory panel adds and removes entries', async ({ page }) => { + await newPrWatcherTeam(page) + // Memory now lives on the team page (one panel per agent); close the demo + // overlay and scope to the watcher's panel. + await closeDemo(page) + const mem = page.getByRole('group', { name: 'memory pr-watcher' }) + await expect(mem.getByRole('heading', { name: /Memory ·/ })).toBeVisible() + await mem.getByLabel('memory key').fill('watchlist') + await mem.getByLabel('memory value').fill('repo:tanstack/ai') + await mem.getByRole('button', { name: '+ Add entry' }).click() + + await expect(mem.getByText('watchlist')).toBeVisible() + await expect(mem.getByText('repo:tanstack/ai')).toBeVisible() + + await mem.getByRole('button', { name: 'delete memory watchlist' }).click() + await expect(mem.getByText('repo:tanstack/ai')).toHaveCount(0) +}) diff --git a/examples/agent-dashboard/e2e/react-news.spec.ts b/examples/agent-dashboard/e2e/react-news.spec.ts new file mode 100644 index 0000000000..6c0db5ee1a --- /dev/null +++ b/examples/agent-dashboard/e2e/react-news.spec.ts @@ -0,0 +1,86 @@ +import { expect, test } from '@playwright/test' +import { closeDemo, openDemo } from './devtools' + +// The Reddit pod — the first real-service + real-AI team. A procedural fetcher +// tool pulls React news (served from a recorded fixture under VITE_E2E, so this +// runs offline and deterministically), and a sentiment agent — subscribed to +// the tool's *result* — posts a digest with nobody clicking run on it. +// +// This also exercises the product controls: the demo team is seeded from the +// Demo Controls devtools panel, and the fetch is driven from the roster's +// per-agent "run tool" dialog. It asserts the LOOP structurally (fetch → result +// card → sentiment message), not specific live headlines (see design doc §10). +test('the Reddit pod loop: a news batch triggers an unprompted sentiment digest', async ({ + page, +}) => { + await page.goto('/') + // The seeded demos live in the Demo Controls devtools panel now. + await openDemo(page) + await page.getByRole('button', { name: '+ React-news demo' }).click() + await expect(page).toHaveURL(/\/teams\//) + + // Two members: the fetcher and the sentiment agent. + await expect(page.getByText('Members · 2')).toBeVisible() + + // Close the panel and drive the fetch from the roster's run-tool dialog (the + // product control): open it on the fetcher, then Run its default tool. + await closeDemo(page) + await page.getByRole('button', { name: 'Run a tool on fetcher' }).click() + await expect(page.getByRole('heading', { name: /Run a tool/ })).toBeVisible() + await page.getByRole('button', { name: 'Run', exact: true }).click() + // Dismiss the dialog so it doesn't cover the channel. + await page.getByRole('button', { name: 'Close' }).click() + + // 1. The batch lands as a result card with (fixtured) real headlines. + // The result JSON renders as a collapsed tree, so the headline is present in + // the DOM but not visible until expanded — assert content, not visibility. + await expect(page.locator('body')).toContainText( + /React Compiler is now stable/, + { timeout: 15000 }, + ) + + // 2. The sentiment agent posts a digest — unprompted, driven by the + // tool_result subscription. + await expect(page.getByText(/Sentiment digest/).first()).toBeVisible({ + timeout: 15000, + }) +}) + +// Regression: a team COMPOSED BY HAND (agents table → New team, then + Add +// agent) must react the same as the seeded demo — the sentiment agent carries +// its tool_result subscription by harness default, so nobody has to wire it. +test('a hand-composed Reddit team auto-reacts to a fetch', async ({ page }) => { + await page.goto('/') + // Close the demo panel so it doesn't overlay the lower agents-table rows. + await closeDemo(page) + + // New team from the agents table with reddit/fetcher. + const fetcherRow = page.getByRole('row').filter({ hasText: 'reddit/fetcher' }) + await fetcherRow.getByRole('button', { name: /Add to team/ }).click() + await page.getByRole('button', { name: /New team with this agent/ }).click() + await expect(page).toHaveURL(/\/teams\//) + + // Add the sentiment agent from the team-page product control. + await page.getByRole('button', { name: /Add agent/ }).click() + await page + .getByRole('button', { name: 'sentiment/react', exact: true }) + .click() + await expect(page.getByText('Members · 2')).toBeVisible() + + // Run the fetch from the roster — the sentiment agent reacts with no manual + // subscription wiring. + await page.getByRole('button', { name: 'Run a tool on fetcher' }).click() + await expect(page.getByRole('heading', { name: /Run a tool/ })).toBeVisible() + await page.getByRole('button', { name: 'Run', exact: true }).click() + await page.getByRole('button', { name: 'Close' }).click() + + // The result JSON renders as a collapsed tree, so the headline is present in + // the DOM but not visible until expanded — assert content, not visibility. + await expect(page.locator('body')).toContainText( + /React Compiler is now stable/, + { timeout: 15000 }, + ) + await expect(page.getByText(/Sentiment digest/).first()).toBeVisible({ + timeout: 15000, + }) +}) diff --git a/examples/agent-dashboard/e2e/team.spec.ts b/examples/agent-dashboard/e2e/team.spec.ts new file mode 100644 index 0000000000..af0412c090 --- /dev/null +++ b/examples/agent-dashboard/e2e/team.spec.ts @@ -0,0 +1,50 @@ +import { expect, test } from '@playwright/test' +import { answerAgentQuestion, closeDemo, openDemo } from './devtools' + +// The teams reframe: one agent looks like a plain chat; a second member reveals +// the team (roster appears), and both members' AG-UI streams merge into ONE +// channel view. Exercises the product controls: the home agents table's +// "Add to team" and the team page's "+ Add agent" picker. +test('a second agent reveals the team and both members share one channel', async ({ + page, +}) => { + await page.goto('/') + + // Start a team from the agents table: support/triage → New team. + const triageRow = page.getByRole('row').filter({ hasText: 'support/triage' }) + await triageRow.getByRole('button', { name: /Add to team/ }).click() + await page.getByRole('button', { name: /New team with this agent/ }).click() + await expect(page).toHaveURL(/\/teams\//) + + // One member: it looks like a normal chat — no team roster. + await expect(page.getByText(/Members ·/)).toHaveCount(0) + + // Run the first agent from the Demo Controls panel: it looks up the ticket and + // pauses for approval. + await openDemo(page) + await page.getByRole('button', { name: 'Start triage demo' }).click() + await answerAgentQuestion(page) + await expect(page.getByText('lookup_ticket').first()).toBeVisible() + await expect( + page.getByText('Approval required', { exact: true }).first(), + ).toBeVisible() + + // Add a second agent from the team page's product control → the roster appears. + await closeDemo(page) + await page.getByRole('button', { name: /Add agent/ }).click() + await page + .getByRole('button', { name: 'support/triage', exact: true }) + .click() + await expect(page.getByText('Members · 2')).toBeVisible() + + // Run the second member from the roster; its stream joins the SAME channel. + await page.getByRole('button', { name: 'Run triage 2' }).click() + await answerAgentQuestion(page) + + // Both members' runs are visible in one timeline — two lookup_ticket cards and + // two approvals. If the two threads' row ids collided, we'd see only one each. + await expect(page.getByText('lookup_ticket')).toHaveCount(2) + await expect( + page.getByText('Approval required', { exact: true }), + ).toHaveCount(2) +}) diff --git a/examples/agent-dashboard/e2e/trace.spec.ts b/examples/agent-dashboard/e2e/trace.spec.ts new file mode 100644 index 0000000000..65c0852eff --- /dev/null +++ b/examples/agent-dashboard/e2e/trace.spec.ts @@ -0,0 +1,22 @@ +import { expect, test } from '@playwright/test' +import { answerAgentQuestion, openDemo } from './devtools' + +test('shows a live waterfall and resolves an approval from it', async ({ + page, +}) => { + await page.goto('/') + await openDemo(page) + await page.getByRole('button', { name: '+ New team' }).click() + await page.getByRole('button', { name: 'Start triage demo' }).click() + await answerAgentQuestion(page) + await expect( + page.getByText('Approval required', { exact: true }), + ).toBeVisible() + + await page.getByRole('button', { name: 'trace', exact: true }).click() + await expect(page.getByText(/tool · lookup_ticket/)).toBeVisible() + await expect(page.getByText(/approval ·/)).toBeVisible() + await page.getByRole('button', { name: 'approve', exact: true }).click() + await page.getByRole('button', { name: 'timeline', exact: true }).click() + await expect(page.getByText(/Sent ✅/)).toBeVisible() +}) diff --git a/examples/agent-dashboard/package.json b/examples/agent-dashboard/package.json new file mode 100644 index 0000000000..3cfeeb0560 --- /dev/null +++ b/examples/agent-dashboard/package.json @@ -0,0 +1,50 @@ +{ + "name": "agent-dashboard", + "private": true, + "type": "module", + "scripts": { + "dev": "vite dev --port 3002", + "build": "vite build", + "serve": "vite preview", + "start": "node .output/server/index.mjs", + "test": "vitest run", + "test:lib": "vitest run", + "test:e2e": "playwright test", + "test:types": "tsc --noEmit" + }, + "dependencies": { + "@ag-ui/client": "1.0.0", + "@phosphor-icons/react": "^2.1.10", + "@tailwindcss/vite": "^4.1.18", + "@tanstack/ai": "workspace:*", + "@tanstack/ai-anthropic": "workspace:*", + "@tanstack/ai-harness": "workspace:*", + "@tanstack/ai-mcp": "workspace:*", + "@tanstack/ai-persistence": "workspace:*", + "@tanstack/react-db": "^0.1.55", + "@tanstack/react-devtools": "^0.9.10", + "@tanstack/react-query": "^5.90.12", + "@tanstack/react-router": "^1.170.41", + "@tanstack/react-router-devtools": "^1.158.4", + "@tanstack/react-start": "^1.168.60", + "@tanstack/router-plugin": "^1.168.42", + "nitro": "3.0.260610-beta", + "react": "^19.2.3", + "react-dom": "^19.2.3", + "react-markdown": "^10.1.0", + "remark-gfm": "^4.0.1", + "tailwindcss": "^4.1.18", + "zod": "^4.2.0" + }, + "devDependencies": { + "@playwright/test": "^1.57.0", + "@tanstack/devtools-vite": "^0.5.3", + "@types/node": "^24.10.1", + "@types/react": "^19.2.7", + "@types/react-dom": "^19.2.3", + "@vitejs/plugin-react": "^5.2.0", + "typescript": "5.9.3", + "vite": "^8.2.1", + "vitest": "^4.1.10" + } +} diff --git a/examples/agent-dashboard/playwright.config.ts b/examples/agent-dashboard/playwright.config.ts new file mode 100644 index 0000000000..60dfaf7af8 --- /dev/null +++ b/examples/agent-dashboard/playwright.config.ts @@ -0,0 +1,29 @@ +import { defineConfig, devices } from '@playwright/test' + +export default defineConfig({ + testDir: './e2e', + fullyParallel: false, + forbidOnly: !!process.env.CI, + retries: 0, + workers: 1, + reporter: [['list']], + timeout: 30_000, + expect: { timeout: 15_000 }, + use: { + baseURL: 'http://localhost:3002', + screenshot: 'only-on-failure', + trace: 'on-first-retry', + }, + projects: [{ name: 'chromium', use: { ...devices['Desktop Chrome'] } }], + webServer: { + // Isolate persisted state to a throwaway file, wiped on start, so runs never + // inherit a previous run's runs/memory/automations (see src/server/store.ts). + // VITE_E2E opens the devtools panel on load so specs can reach the demo + // controls that now live in the "Demo Controls" panel. + command: + 'rm -f .data/e2e-state.json && VITE_E2E=1 DASHBOARD_STATE_FILE=.data/e2e-state.json pnpm run dev', + url: 'http://localhost:3002', + reuseExistingServer: !process.env.CI, + timeout: 120_000, + }, +}) diff --git a/examples/agent-dashboard/src/components/automations-panel.tsx b/examples/agent-dashboard/src/components/automations-panel.tsx new file mode 100644 index 0000000000..98ee28f6a5 --- /dev/null +++ b/examples/agent-dashboard/src/components/automations-panel.tsx @@ -0,0 +1,163 @@ +/** + * The channel's automations: the tool registry (public tools only, with run-now), + * the schedule table (the dashboard's clock), and a webhook tester. All three + * produce the same thing — an injected tool call whose structured result streams + * into the channel via the feed tail. + */ +import { useMutation, useQuery, useQueryClient } from '@tanstack/react-query' +import { runInjection } from '@/lib/session-controller' +import type { MembershipRow } from '@/db/collections' + +interface ToolInfo { + name: string + description: string +} +export function AutomationsPanel({ + channelId, + primary, +}: { + channelId: string + primary: MembershipRow +}) { + const qc = useQueryClient() + const invalidate = () => { + void qc.invalidateQueries({ queryKey: ['offline'] }) + } + + const tools = useQuery<{ tools: Array }>({ + // Pass the harness so the registry is resolved directly — otherwise the + // thread may not yet be noted with its harness and we'd get triage's tools. + queryKey: ['tools', primary.threadId, primary.harness], + queryFn: () => + fetch( + `/api/tools?threadId=${primary.threadId}&harness=${encodeURIComponent(primary.harness)}`, + ).then((r) => r.json()), + }) + const offline = useQuery<{ offline: boolean; queued: number }>({ + queryKey: ['offline'], + queryFn: () => fetch('/api/dev/offline').then((r) => r.json()), + refetchInterval: 1500, + }) + + const toolNames = tools.data?.tools ?? [] + const selectedTool = toolNames[0]?.name || '' + + const setOffline = useMutation({ + mutationFn: (value: boolean) => + fetch('/api/dev/offline', { + method: 'POST', + headers: { 'content-type': 'application/json' }, + body: JSON.stringify({ offline: value }), + }).then((r) => r.json()), + onSuccess: invalidate, + }) + const sendWebhook = useMutation({ + mutationFn: async () => { + const created = await fetch('/api/webhooks', { + method: 'POST', + headers: { 'content-type': 'application/json' }, + body: JSON.stringify({ + threadId: primary.threadId, + channelId, + tool: selectedTool, + argMapping: { queue: 'queue' }, + }), + }).then((r) => r.json()) + const token = created.webhook.token as string + return fetch(`/api/webhooks/${token}`, { + method: 'POST', + headers: { 'content-type': 'application/json' }, + body: JSON.stringify({ queue: 'from-webhook' }), + }).then((r) => r.json()) + }, + }) + // The PR-watcher demo: a prompt-mode webhook drives the watcher's full run + // (check_pr → channel_create → message_post). Each send is the next seeded PR. + const sendPrWebhook = useMutation({ + mutationFn: async () => { + const created = await fetch('/api/webhooks', { + method: 'POST', + headers: { 'content-type': 'application/json' }, + body: JSON.stringify({ + threadId: primary.threadId, + channelId, + harness: primary.harness, + mode: 'prompt', + message: + 'A PR webhook arrived. Check for a new PR and open a review channel.', + }), + }).then((r) => r.json()) + const token = created.webhook.token as string + return fetch(`/api/webhooks/${token}`, { + method: 'POST', + headers: { 'content-type': 'application/json' }, + body: JSON.stringify({ event: 'pull_request' }), + }).then((r) => r.json()) + }, + }) + + return ( +
+
+

Automations

+ + deterministic tool runs — no tokens + + {offline.data?.offline && ( + + host offline — {offline.data.queued} queued + + )} +
+ + {/* Tool registry + run-now (public tools only) */} +
+
Public tools
+
+ {toolNames.length === 0 && ( + none + )} + {toolNames.map((t) => ( + + ))} +
+
+ + {/* Webhook + offline simulation */} +
+ + {primary.harness === 'ops/pr-watcher' && ( + + )} + +
+
+ ) +} diff --git a/examples/agent-dashboard/src/components/channel-view.tsx b/examples/agent-dashboard/src/components/channel-view.tsx new file mode 100644 index 0000000000..100f06403c --- /dev/null +++ b/examples/agent-dashboard/src/components/channel-view.tsx @@ -0,0 +1,468 @@ +/** + * The shared channel view. It unions every member thread of a channel into one + * timeline (a live query over the collections keyed by `channelId`). With one + * member it looks exactly like a single-agent chat; when a second member joins, + * team chrome (the member list, per-agent attribution) appears — the "second + * agent reveals the team" moment. + * + * A channel's members are its team roster for the `main` channel, or the opt-in + * `channelMembers` for a `dynamic`/`dm` channel. Attribution is resolved against + * the whole team roster either way. + */ +import { eq, useLiveQuery } from '@tanstack/react-db' +import { useQuery } from '@tanstack/react-query' +import { useEffect, useState } from 'react' +import { CoinsIcon, DatabaseIcon, UserPlusIcon } from '@phosphor-icons/react' +import { + DEFAULT_BUDGET, + approvals, + budgets, + channelMembers, + channels, + memberships, + messages, + questions, + runMeta, + sessions, + spend, + toolCalls, + uiState, + upsert, +} from '@/db/collections' +import { + addAgentToChannel, + channelSendPrompt, + controlInput, + createDm, + defaultSubscriptions, + openChannelMember, +} from '@/lib/session-controller' +import { MemberList } from '@/components/member-list' +import { MemoryPanel } from '@/components/memory-panel' +import { + Composer, + StatusPill, + Stream, + buildTimeline, +} from '@/components/stream' +import { compact } from '@/components/ui' +import { TraceWaterfall } from '@/components/trace-waterfall' +import { TeamAutomations } from '@/components/team-automations' +import { costUsd } from '@/lib/pricing' +import type { + ApprovalRow, + BudgetRow, + ChannelMemberRow, + ChannelRow, + MembershipRow, + MessageRow, + QuestionRow, + RunMetaRow, + SessionRow, + SpendRow, + ToolCallRow, + UiStateRow, +} from '@/db/collections' + +// `pod.channel_create` / `pod.message_post` are realized as a channel and a +// message respectively, so their raw tool cards are hidden (they'd duplicate). +// `pod.memory_write` stays visible (the write is auditable mechanics). +const HIDDEN_TOOL_CARDS = new Set([ + 'pod.channel_create', + 'pod.message_post', + 'pod.memory_read', +]) + +export function ChannelView({ + channelId, + teamId, +}: { + channelId: string + teamId?: string +}) { + const [input, setInput] = useState('') + const [view, setView] = useState<'timeline' | 'trace'>('timeline') + + const { data: chanRows = [] } = useLiveQuery( + (q) => q.from({ c: channels }).where(({ c }) => eq(c.id, channelId)), + [channelId], + ) + const channel = (chanRows as Array)[0] + const resolvedTeamId = teamId ?? channel?.teamId + + // The whole team roster (for attribution + main-channel membership). + const { data: roster = [] } = useLiveQuery( + (q) => + q + .from({ m: memberships }) + .where(({ m }) => eq(m.teamId, resolvedTeamId ?? '')), + [resolvedTeamId], + ) + const rosterRows = roster as Array + // Opt-in members for a non-main channel. + const { data: chanMembers = [] } = useLiveQuery( + (q) => + q + .from({ cm: channelMembers }) + .where(({ cm }) => eq(cm.channelId, channelId)), + [channelId], + ) + const chanMemberRows = chanMembers as Array + + const isMain = channel?.kind !== 'dynamic' && channel?.kind !== 'dm' + const memberRows: Array = isMain + ? rosterRows + : chanMemberRows + .map((cm) => rosterRows.find((m) => m.agentId === cm.agentId)) + .filter((m): m is MembershipRow => Boolean(m)) + + // Open a live tail for every team member so background runs (a watcher firing, + // a subscribed agent reviewing) project even when we're not looking at them. + const rosterKey = rosterRows.map((m) => m.id).join(',') + useEffect(() => { + for (const m of rosterRows) { + openChannelMember({ + channelId: m.channelId, + agentId: m.agentId, + threadId: m.threadId, + teamId: m.teamId, + harness: m.harness, + role: m.role, + displayName: m.displayName, + }) + } + // eslint-disable-next-line react-hooks/exhaustive-deps + }, [rosterKey]) + + // Publish the active channel so the demo-controls devtools panel (rendered + // out of the route tree) knows which channel to drive. Clear on unmount + // unless another channel already took over. + useEffect(() => { + upsert(uiState, { id: 'active' }, (d) => { + d.channelId = channelId + d.teamId = resolvedTeamId + }) + return () => { + if (uiState.get('active')?.channelId === channelId) { + upsert(uiState, { id: 'active' }, (d) => { + d.channelId = undefined + d.teamId = undefined + }) + } + } + }, [channelId, resolvedTeamId]) + + const { data: msgs = [] } = useLiveQuery( + (q) => q.from({ m: messages }).where(({ m }) => eq(m.channelId, channelId)), + [channelId], + ) + const { data: tools = [] } = useLiveQuery( + (q) => + q.from({ t: toolCalls }).where(({ t }) => eq(t.channelId, channelId)), + [channelId], + ) + const { data: apprs = [] } = useLiveQuery( + (q) => + q.from({ a: approvals }).where(({ a }) => eq(a.channelId, channelId)), + [channelId], + ) + const { data: questionRows = [] } = useLiveQuery( + (query) => + query + .from({ question: questions }) + .where(({ question }) => eq(question.channelId, channelId)), + [channelId], + ) + const { data: spendRows = [] } = useLiveQuery( + (q) => q.from({ s: spend }).where(({ s }) => eq(s.channelId, channelId)), + [channelId], + ) + const { data: sess = [] } = useLiveQuery( + (q) => q.from({ s: sessions }).where(({ s }) => eq(s.channelId, channelId)), + [channelId], + ) + const { data: runMetaRows = [] } = useLiveQuery((q) => q.from({ r: runMeta })) + const { data: budgetRows = [] } = useLiveQuery((q) => q.from({ b: budgets })) + + const isTeam = memberRows.length > 1 + const nameByAgent = new Map(rosterRows.map((m) => [m.agentId, m.displayName])) + const statusByThread: Record = {} + for (const s of sess as Array) + statusByThread[s.threadId] = s.status + + const sessionRows = sess as Array + const status = sessionRows.some((s) => s.status === 'requires_action') + ? 'requires_action' + : sessionRows.some((s) => s.status === 'running') + ? 'running' + : 'idle' + const tokens = (spendRows as Array<{ totalTokens: number }>).reduce( + (sum, s) => sum + (s.totalTokens ?? 0), + 0, + ) + const dollars = (spendRows as Array).reduce((sum, row) => { + const member = rosterRows.find((item) => item.threadId === row.threadId) + return sum + costUsd(member?.harness, row.inputTokens, row.outputTokens) + }, 0) + const timeline = buildTimeline({ + msgs: msgs as Array, + tools: (tools as Array).filter( + (t) => !HIDDEN_TOOL_CARDS.has(t.name), + ), + approvals: apprs as Array, + questions: questionRows as Array, + }) + + // Human input targets the primary agent member (broadcast is a later phase). + const primary = + memberRows.find((m) => m.role === 'agent') ?? memberRows[0] ?? undefined + + // How much pod memory the platform attached to this channel's agents' last run. + const attachedById = new Map( + (runMetaRows as Array).map((r) => [r.threadId, r.attached]), + ) + const attached = primary ? (attachedById.get(primary.threadId) ?? 0) : 0 + + const send = async (text: string) => { + const t = text.trim() + if (!t || !primary) return + setInput('') + if (status === 'running') { + await controlInput(primary.threadId, { op: 'steer', message: t }) + return + } + await channelSendPrompt(primary, t, channelId) + } + + // A generic nudge to run a specific member. (The triage-specific demo prompt + // lives in the Demo Controls panel, not here — this must work for any agent.) + const runMember = (member: MembershipRow) => + channelSendPrompt(member, 'Please proceed.', channelId) + + const createDmWith = (member: MembershipRow) => { + if (!primary || member.agentId === primary.agentId) return + void createDm( + { + channelId: primary.channelId, + agentId: primary.agentId, + threadId: primary.threadId, + teamId: primary.teamId, + harness: primary.harness, + role: primary.role, + displayName: primary.displayName, + }, + member.agentId, + member.displayName, + ) + } + + // Toggle a member's subscription on/off. "On" restores the harness's default + // triggers (what it reacts to) rather than a hardcoded channel_created one. + const toggleSubscription = (member: MembershipRow) => { + const on = (member.subscriptions ?? []).length > 0 + memberships.update(member.id, (draft) => { + draft.subscriptions = on + ? [] + : (defaultSubscriptions(member.harness) ?? []) + }) + } + + const budget = memberRows.reduce( + (sum, m) => + sum + + ((budgetRows as Array).find((b) => b.threadId === m.threadId) + ?.maxTokens ?? DEFAULT_BUDGET), + 0, + ) + + const sigil = channel?.kind === 'dm' ? '@' : '#' + const title = + isMain && !isTeam + ? (primary?.displayName ?? channel?.name ?? channelId) + : isMain + ? 'main' + : (channel?.name ?? channelId) + const column = isTeam ? 'px-6' : 'mx-auto w-full max-w-[720px] px-6' + + return ( +
+
+
+
+
+

+ {(isTeam || !isMain) && ( + {sigil} + )} + {title} +

+ {!isMain && ( + + {channel?.kind} + + )} + +
+ {channel?.topic && ( +
+ {channel.topic} +
+ )} +
+
+ {attached > 0 && ( + + + {attached} memory {attached === 1 ? 'entry' : 'entries'}{' '} + attached + + )} + + ${dollars.toFixed(4)} ·{' '} + {tokens.toLocaleString()} tokens + + + {(['timeline', 'trace'] as const).map((item) => ( + + ))} + + {isMain && } +
+
+ + {/* column-reverse keeps the stream bottom-anchored, like chat. */} +
+
+ {view === 'trace' ? ( + + ) : timeline.length === 0 ? ( +

+ No activity yet. Send a message below, or drive it from the Demo + controls devtools panel. +

+ ) : ( + + )} +
+
+ +
+ send(input)} + running={status === 'running' && Boolean(primary)} + onStop={() => + primary && controlInput(primary.threadId, { op: 'cancel' }) + } + solo={!isTeam} + /> +
+
+ + +
+ ) +} + +const CHIP = + 'inline-flex items-center gap-1.5 rounded-full bg-ui px-2.5 py-1 text-xs font-normal text-ink-2' + +/** Add any available agent to this team (product control, main channel only). */ +function AddAgentControl({ channelId }: { channelId: string }) { + const [open, setOpen] = useState(false) + const hosts = useQuery< + Array<{ agents: Array<{ name: string; description: string }> }> + >({ + queryKey: ['hosts'], + queryFn: () => fetch('/api/hosts').then((r) => r.json()), + }) + const agents = (hosts.data ?? []).flatMap((h) => h.agents) + + return ( +
+ + {open && ( +
+ {agents.map((a) => ( + + ))} +
+ )} +
+ ) +} diff --git a/examples/agent-dashboard/src/components/demo-controls.tsx b/examples/agent-dashboard/src/components/demo-controls.tsx new file mode 100644 index 0000000000..95041b01a0 --- /dev/null +++ b/examples/agent-dashboard/src/components/demo-controls.tsx @@ -0,0 +1,151 @@ +/** + * The demo-only controls, rendered inside a custom TanStack DevTools panel + * (see `__root.tsx`) rather than inline in the channel — so it's obvious what + * drives the demo vs. what an operator would actually use. + * + * The panel lives in the devtools render root, outside the route tree, so it + * can't take props from the current route. It reads the active channel from the + * `uiState` DB row (published by `ChannelView`) and re-derives the primary + * member from the same live collections the channel view uses — one shared + * source of truth, two independent readers. + */ +import { eq, useLiveQuery } from '@tanstack/react-db' +import { useNavigate } from '@tanstack/react-router' +import { + channelMembers, + channels, + memberships, + uiState, +} from '@/db/collections' +import { + channelSendPrompt, + createPrWatcherTeam, + createReactNewsTeam, + createTeam, +} from '@/lib/session-controller' +import { AutomationsPanel } from '@/components/automations-panel' +import type { + ChannelMemberRow, + ChannelRow, + MembershipRow, + UiStateRow, +} from '@/db/collections' + +/** The devtools plugin body: seeded-demo launchers, then the active channel's controls. */ +export function DemoControlsPanel() { + const navigate = useNavigate() + const { data: rows = [] } = useLiveQuery((q) => q.from({ u: uiState })) + const active = (rows as Array).find((r) => r.id === 'active') + + const go = (teamId: string) => + navigate({ to: '/teams/$teamId', params: { teamId } }) + + return ( +
+
+

Demo controls

+

+ Scaffolding to drive the demo — not part of the product UX. Reads live + app state (TanStack DB) from the devtools render root. +

+
+ +
+
Seeded demos
+
+ + + +
+
+ + {active?.channelId ? ( + + ) : ( +

+ Open a team channel to see its per-channel demo controls. +

+ )} +
+ ) +} + +function DemoControls({ channelId }: { channelId: string }) { + const { data: chanRows = [] } = useLiveQuery( + (q) => q.from({ c: channels }).where(({ c }) => eq(c.id, channelId)), + [channelId], + ) + const channel = (chanRows as Array)[0] + const teamId = channel?.teamId + + const { data: roster = [] } = useLiveQuery( + (q) => + q.from({ m: memberships }).where(({ m }) => eq(m.teamId, teamId ?? '')), + [teamId], + ) + const rosterRows = roster as Array + const { data: chanMembers = [] } = useLiveQuery( + (q) => + q + .from({ cm: channelMembers }) + .where(({ cm }) => eq(cm.channelId, channelId)), + [channelId], + ) + const chanMemberRows = chanMembers as Array + + // Mirror ChannelView's member/primary derivation so the controls target the + // same agent the channel does. + const isMain = channel?.kind !== 'dynamic' && channel?.kind !== 'dm' + const memberRows: Array = isMain + ? rosterRows + : chanMemberRows + .map((cm) => rosterRows.find((m) => m.agentId === cm.agentId)) + .filter((m): m is MembershipRow => Boolean(m)) + const primary = + memberRows.find((m) => m.role === 'agent') ?? memberRows[0] ?? undefined + + if (!channel) { + return

This channel isn't loaded.

+ } + + return ( +
+ {isMain && primary?.harness === 'support/triage' && ( + + )} + {primary && isMain && ( + + )} +
+ ) +} diff --git a/examples/agent-dashboard/src/components/json-tree.tsx b/examples/agent-dashboard/src/components/json-tree.tsx new file mode 100644 index 0000000000..1bb477e63a --- /dev/null +++ b/examples/agent-dashboard/src/components/json-tree.tsx @@ -0,0 +1,69 @@ +/** + * A minimal collapsible JSON tree. Objects and arrays render as native + *
nodes, initially collapsed; primitives render inline and coloured. + * `tryParse` returns null for non-JSON strings so callers can fall back to raw + * text (a plain sentence isn't valid JSON and shouldn't be a tree). + * + * ponytail: native
recursion, no dependency. Reach for a json-view + * lib only if we need search / copy-path / big-array virtualization. + */ + +/** Parse a string as JSON, but only surface object/array roots as trees. */ +export function tryParse(s: string | undefined): unknown { + if (!s) return undefined + try { + const v = JSON.parse(s) + return v && typeof v === 'object' ? v : undefined + } catch { + return undefined + } +} + +export function JsonTree({ value }: { value: unknown }) { + return +} + +function Node({ label, value }: { label: string | null; value: unknown }) { + if (value !== null && typeof value === 'object') { + const entries = Array.isArray(value) + ? value.map((v, i) => [String(i), v] as const) + : Object.entries(value) + const count = entries.length + const kind = Array.isArray(value) + ? `[ ] ${count} item${count === 1 ? '' : 's'}` + : `{ } ${count} key${count === 1 ? '' : 's'}` + return ( +
+ + {label !== null && } + {kind} + +
+ {entries.map(([k, v]) => ( + + ))} +
+
+ ) + } + return ( +
+ {label !== null && } + +
+ ) +} + +function Key({ label }: { label: string }) { + return {label}: +} + +function Leaf({ value }: { value: unknown }) { + if (typeof value === 'string') + return "{value}" + if (typeof value === 'number') + return {value} + if (typeof value === 'boolean') + return {String(value)} + return null +} diff --git a/examples/agent-dashboard/src/components/markdown.tsx b/examples/agent-dashboard/src/components/markdown.tsx new file mode 100644 index 0000000000..5bdf679935 --- /dev/null +++ b/examples/agent-dashboard/src/components/markdown.tsx @@ -0,0 +1,19 @@ +/** + * Message text is LLM output (the sentiment digest is a markdown table + prose), + * so render it as markdown. react-markdown doesn't emit raw HTML by default, so + * no sanitize plugin is needed for this untrusted content; remark-gfm adds the + * table/strikethrough/autolink support the digests use. + * + * ponytail: reuse the sibling examples' markdown stack; `prose` (Tailwind + * typography defaults) is enough styling without a custom component map. + */ +import ReactMarkdown from 'react-markdown' +import remarkGfm from 'remark-gfm' + +export function Markdown({ children }: { children: string }) { + return ( +
+ {children} +
+ ) +} diff --git a/examples/agent-dashboard/src/components/member-list.tsx b/examples/agent-dashboard/src/components/member-list.tsx new file mode 100644 index 0000000000..23728d898c --- /dev/null +++ b/examples/agent-dashboard/src/components/member-list.tsx @@ -0,0 +1,130 @@ +import { useState } from 'react' +import { + BellIcon, + BellSlashIcon, + ChatCircleIcon, + PlayIcon, + WrenchIcon, +} from '@phosphor-icons/react' +import { RunToolDialog } from '@/components/run-tool-dialog' +import { Avatar } from '@/components/ui' +import { defaultSubscriptions } from '@/lib/session-controller' +import type { MembershipRow, SessionRow } from '@/db/collections' + +const STATUS: Record = { + running: { dot: 'bg-ok', label: 'Running' }, + requires_action: { dot: 'bg-warn', label: 'Waiting for approval' }, + idle: { dot: 'bg-ink-3', label: 'Idle' }, +} + +function subscribed(member: MembershipRow): boolean { + return (member.subscriptions ?? []).length > 0 +} + +/** Only agents that actually react to events (have default triggers) can subscribe. */ +function reactive(member: MembershipRow): boolean { + return Boolean(defaultSubscriptions(member.harness)) +} + +/** The team roster. Renders only when a channel has more than one member. */ +export function MemberList({ + members, + statusByThread, + onRun, + onCreateDm, + onToggleSubscription, +}: { + members: Array + statusByThread: Record + onRun: (member: MembershipRow) => void + onCreateDm?: (member: MembershipRow) => void + onToggleSubscription?: (member: MembershipRow) => void +}) { + // The member whose run-tool dialog is open (product control, not the demo run-now). + const [toolMember, setToolMember] = useState() + + return ( +
+

Members · {members.length}

+
    + {members.map((m) => { + const operator = m.role === 'operator' + const status = STATUS[statusByThread[m.threadId] ?? 'idle'] + return ( +
  • + + + + +
    +
    + {m.displayName} +
    +
    + {operator + ? 'Operator' + : subscribed(m) + ? `${status.label} · subscribed` + : status.label} +
    +
    + {!operator && ( +
    + + + {onToggleSubscription && reactive(m) && ( + + )} + {onCreateDm && ( + + )} +
    + )} +
  • + ) + })} +
+ + {toolMember && ( + setToolMember(undefined)} + /> + )} +
+ ) +} diff --git a/examples/agent-dashboard/src/components/memory-panel.tsx b/examples/agent-dashboard/src/components/memory-panel.tsx new file mode 100644 index 0000000000..f47c5b7318 --- /dev/null +++ b/examples/agent-dashboard/src/components/memory-panel.tsx @@ -0,0 +1,111 @@ +/** + * Pod memory for one member. The human operator views the standing instructions + * the agent has accumulated (and can add/remove them). This is the operational + * memory the platform attaches to every run — visible mechanics, not hidden state. + */ +import { useMutation, useQuery, useQueryClient } from '@tanstack/react-query' +import { useState } from 'react' +import { FileTextIcon, TrashIcon } from '@phosphor-icons/react' + +export function MemoryPanel({ + threadId, + name, +}: { + threadId: string + name: string +}) { + const qc = useQueryClient() + const [key, setKey] = useState('') + const [value, setValue] = useState('') + + const memory = useQuery<{ entries: Record }>({ + queryKey: ['memory', threadId], + queryFn: () => + fetch(`/api/memory?threadId=${encodeURIComponent(threadId)}`).then((r) => + r.json(), + ), + refetchInterval: 1500, + }) + const invalidate = () => + void qc.invalidateQueries({ queryKey: ['memory', threadId] }) + + const add = useMutation({ + mutationFn: () => + fetch('/api/memory', { + method: 'POST', + headers: { 'content-type': 'application/json' }, + body: JSON.stringify({ threadId, key, value }), + }).then((r) => r.json()), + onSuccess: () => { + setKey('') + setValue('') + invalidate() + }, + }) + const remove = useMutation({ + mutationFn: (k: string) => + fetch( + `/api/memory?threadId=${encodeURIComponent(threadId)}&key=${encodeURIComponent(k)}`, + { method: 'DELETE' }, + ).then((r) => r.json()), + onSuccess: invalidate, + }) + + const entries = Object.entries(memory.data?.entries ?? {}) + + return ( +
+

Memory · {name}

+
    + {entries.length === 0 && ( +
  • no entries yet
  • + )} + {entries.map(([k, v]) => ( +
  • + +
    +
    {k}
    +
    + {v} +
    +
    + +
  • + ))} +
+
+ setKey(e.target.value)} + aria-label="memory key" + placeholder="key" + className="input w-20 min-w-0 py-1 text-xs" + /> + setValue(e.target.value)} + aria-label="memory value" + placeholder="value" + className="input min-w-0 flex-1 py-1 text-xs" + /> + +
+
+ ) +} diff --git a/examples/agent-dashboard/src/components/run-tool-dialog.tsx b/examples/agent-dashboard/src/components/run-tool-dialog.tsx new file mode 100644 index 0000000000..a9b9eab1f6 --- /dev/null +++ b/examples/agent-dashboard/src/components/run-tool-dialog.tsx @@ -0,0 +1,136 @@ +/** + * Run a public tool on a team member, out-of-band. Pick a tool from the + * member's harness (`/api/tools`), supply JSON parameters, and inject it via + * `runInjection` — the result streams back into the channel like any other tool + * call. A product control (the roster's per-agent "🔧 tools" button), distinct + * from the demo panel's run-now. + */ +import { useQuery } from '@tanstack/react-query' +import { useState } from 'react' +import { XIcon } from '@phosphor-icons/react' +import { runInjection } from '@/lib/session-controller' +import type { MembershipRow } from '@/db/collections' + +interface ToolInfo { + name: string + description: string +} + +export function RunToolDialog({ + member, + onClose, +}: { + member: MembershipRow + onClose: () => void +}) { + const tools = useQuery<{ tools: Array }>({ + queryKey: ['tools', member.threadId, member.harness], + queryFn: () => + fetch( + `/api/tools?threadId=${member.threadId}&harness=${encodeURIComponent(member.harness)}`, + ).then((r) => r.json()), + }) + const toolNames = tools.data?.tools ?? [] + const [tool, setTool] = useState('') + const [argsText, setArgsText] = useState('{}') + const [error, setError] = useState() + const [result, setResult] = useState() + const selected = tool || toolNames[0]?.name || '' + const selectedInfo = toolNames.find((t) => t.name === selected) + + const run = async () => { + setError(undefined) + setResult(undefined) + let args: Record + try { + args = argsText.trim() ? JSON.parse(argsText) : {} + } catch { + setError('Parameters must be valid JSON.') + return + } + if (!selected) return + const res = await runInjection(member, selected, args) + if (res.status === 'rejected') { + setError(res.reason ?? 'Tool was rejected.') + return + } + setResult(`${selected} ${res.status} — result streaming into the channel.`) + } + + return ( +
+
e.stopPropagation()} + > +
+

+ Run a tool · {member.displayName} +

+ +
+ + {toolNames.length === 0 ? ( +

+ {tools.isLoading + ? 'Loading tools…' + : 'This agent has no public tools to run.'} +

+ ) : ( + <> + + {selectedInfo?.description && ( +

+ {selectedInfo.description} +

+ )} +