Skip to content

About

React viewer for agent conversations, tool execution, topology, and recorded usage.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Agent Record

A React component for reading what an agent team did: conversations, tool inputs and outputs, topology, timelines, and recorded token usage.

Live example · Research using the viewer

Agent Record with a real research run

Install

pnpm add https://github.com/drewstone/agent-record/releases/download/v0.5.1/drewstone-agent-record-0.5.1.tgz react react-dom

The release is an ESM package with TypeScript declarations and CSS. It is distributed through GitHub Releases; this version is not published to the npm registry. React 18.3 and 19 are supported peer versions; release verification uses React 19.

import { AgentRecord, parseRecord } from '@drewstone/agent-record'
import '@drewstone/agent-record/styles.css'

const record = parseRecord(await response.json())

export function RunPage() {
  return <AgentRecord records={[record]} />
}

Your application loads the data and decides which content may be displayed. The component makes no network requests, stores no data, starts no agents, and does not change your URL. It can render on the server; hydration enables the interactive controls. Use a client component when embedding it in a React Server Components application.

Research reports

Use ResearchReport to read authored claims, limitations, checks, and source references alongside each play’s recorded events. A compact play navigation opens Results, Activity, or Sources in one reading area. Search retained activity across agents, inspect readable tool inputs/results, and open exact source citations. Missing conversation capture remains explicit. The offline agent-record-report REPORT.json OUTPUT.html command creates a self-contained interactive report and LaTeX document. Reports remain separate from immutable execution records. Source search can optionally use one commissioned same-origin endpoint; offline reports make no network requests. See the report contract and exports.

Input

agent-record.v1 contains a run ID, title, nodes, and timestamped events. parseRecord(unknown) validates the record and preserves additional metadata. It rejects duplicate node/event IDs, invalid timestamps, parent cycles, invalid session joins, and duplicate published call IDs within an event.

Nodes distinguish agents, native sessions, and findings. An importer can join a session to an agent with agentId; the viewer never infers that relationship from names or timing. Roles and assignments are independent fields. A parent missing from a partial capture remains unresolved.

Events carry messages, tool inputs/results, usage, or lifecycle information. Source paths, hashes, original IDs, timestamps, publication notes, and unknown metadata remain in the downloadable record. Missing measurements remain unknown. An explicit zero remains zero.

See the record format for the complete input rules and an example.

Existing blog records

import { fromResearchPublication } from '@drewstone/agent-record/adapters/research-publication'

const record = fromResearchPublication(reviewedPublication)

This adapter accepts the blog's research-publication.events.v1 files. It preserves their events and source references while translating node identities into the reusable format. It does not review, redact, or sanitize private logs.

Agent-runtime run directories

node tools/ingest.mjs <dir> --out record.json [--run-id <id>] [--manifest <snapshot.json>] [--native <dir>] converts an agent-runtime run directory, including Claude Code, OpenCode, pi, Codex and Kimi conversations where the run retained them, or a bundle of native sessions (bundle.json). Event IDs are anchored to the stored bytes they came from (anchor.v1, see the record format). Each node's capture channel and every missing transcript are recorded in the output. See the record format.

Shared native sessions

Use @tangle-network/harness-sessions to read native Claude Code, Codex, OpenCode, Pi, Kimi and Factory sessions. Pass its normalized sessions to fromHarnessSessions for a browser-renderable record; pass recorded relationship facts from @tangle-network/traces when a workflow should draw parent edges. This projection does not parse a second native format or infer joins from names or time.

import { readerFor } from '@tangle-network/harness-sessions'
import { fromHarnessSessions } from '@drewstone/agent-record/adapters/harness-sessions'

const session = await readerFor('pi').read(ref)
const record = fromHarnessSessions([session], { recordId: 'review-1', title: 'Review' })

Session projection event IDs use shared-session.v1, not publication anchor.v1. Keep the existing anchored converter for published research records until the shared readers expose equivalent source locations and publication verification passes. Review raw text and paths before serving the projection publicly.

Normalized trace spans

fromTraceSpans projects OpenInference/OTLP spans already emitted by @tangle-network/traces or the ChatGPT fleet into the same record model. fromStoryboardSpans projects the separate @tangle-network/agent-eval Span[] shape used by @tangle-network/run-capsule. Both retain recorded trace/run, span and parent-span IDs on events. Each trace or run is one session node; span containment is not treated as agent delegation. The caller can supply exact retained source locations and claim/verdict/publication references. Missing parents, conflicting session IDs and untimed spans become coverage gaps.

import { fromTraceSpans } from '@drewstone/agent-record/adapters/trace-spans'

const record = fromTraceSpans(spans, { recordId: 'fleet-review', title: 'Fleet review' })

OTLP projection uses trace-span.v1 IDs; the storyboard projection uses storyboard-span.v1 IDs. Neither claims original capture completeness. Review content and source paths before public serving. They do not replace the anchored research-publication converter. The two upstream span formats remain different; this package does not convert ChatGPT OTLP spans into run-capsule's storyboard input.

The Overview, run progress, and play charts use the live React renderers from @tangle-network/charts. Agent-record supplies the data, state colors, labels, links, and theme tokens; the shared package owns chart geometry, interaction, and CSS. The distributed styles.css includes the package's live chart stylesheet for hosts that load one renderer CSS file.

RecordWorkGraph draws a validated RunRecord as one interactive entity and evidence graph. It shows the run, agents, sessions, artifacts, precise claims, verdicts, publications, and profile versions. A dashed container edge attaches each root node to its record without claiming execution parentage. Other edges use only recorded parent IDs, page SHA-256 and claim IDs, reviewer session IDs, publication citations, and profile digests; absent targets remain unlinked. Events stay in the record viewer rather than becoming thousands of graph cards. The live run Work graph and this bundle graph share the profile canvas's pan, zoom, pinch, and keyboard engine.

import { RecordWorkGraph, type RunRecord } from '@drewstone/agent-record'
import '@drewstone/agent-record/styles.css'

export function BundleGraph({ record }: { record: RunRecord }) {
  return <RecordWorkGraph record={record} />
}

Run workspace

dist/workspace-viewer.js renders a play page and a run page from a same-origin run workspace API, for example the Tangle Discovery wall:

<link rel="stylesheet" href="/assets/agent-record/styles.css">
<div id="agent-workspace" data-api="/api/discovery" data-mode="run" data-id="research-math-20261004system3"></div>
<script src="/assets/agent-record/workspace-viewer.js" defer></script>

The play page opens on Versions: the play's version graph, left to right. What the versions supersede comes first. Each run follows as a version: its registered profile, its readout judges as bars on the absolute 0–100 vs world-class scale (filled calibrated, outlined advisory; a judge on the retired relative 0–4 scale draws no bar and is named retired) and what changed from the version before it. Last come the profiles the selected version's agents wrote at runtime, as a tree under their authors. The Lab records no edge between two registered profiles, so the viewer compares their content field by field (model, harness, tools, skills, instructions, system prompt, files and budget) and labels the result a comparison. An authored profile is compared with the profile of the same name in the version before it. Selecting a node opens it in the inspector beside the graph: its facts, the comparison, its absolute 0–100 judge scores with their median on the world-class axis ("no absolute score yet" when the judges used the retired relative scale), its verdicts, and for an authored profile its runs, parents and content (plays/<id>/profiles, the discovery-lab.profile-graph document). The other tabs hold the runs table, the play input, spend and the assessment matrix. The run page opens on its overview, which answers five questions before any detail. Its header states the goal in one sentence and four facts: status and time (time used and left against the deadline a fork inherits, and on track or at risk with the reason, at risk when no better version came in a quarter of the window), the best version against its bar (8 of 18 must-pass checks · 2 of 6 held-out · judge 71, and whether later versions beat it), what is still missing, and the cost so far: API-equivalent dollars when the model is priced, else the tokens Runtime metered, and the billed figure apart. When the run has gone without a better version, a Decide line says for how long and how many tokens it used since. Below come the deliverable (the best tag's receipt: open buttons for its report, deck and workbook, the executive summary it quotes with machine keys removed, and the reader test against its bar), what's left (every check the best version fails, must-pass first, with its pass and fail across the scored versions and its evidence without hashes), progress (score per version and the lineage's cumulative tokens on one time axis each, then every tag with its score delta), and the team (one card per agent: role, live state such as working · 47 min · 23 tool calls · active 2 min ago, model, tokens, its latest words and a link to its conversation). History (the lineage, with what each run contributed, and each role's profile versions with the reason its commit states, the score before and after, and what changed: instructions added and removed), the brief and one data-quality line stay collapsed. The page draws all of it from runs/<id>'s story, lineage and spend.meter and the record's node list; it is computed in src/workspace/run-story.ts. Every detail is one click from the section it supports, as a detail view with a bar back to the overview (?section=): Conversations (a rail of the run's agents beside the selected agent's conversation, timeline and usage), Versions (the commit graph), Profiles (the profile graph), Outputs, Readout, Spend ledger, Findings, Graph, Assessments, Capture and Input. Findings lead with what the run found (agent-workspace.findings.v1): each agent's stated answer, then its results and claims; the overview links them only for a research run with no release bar. Graph lays each agent's results, claims and checks on its own lane in writing order, the papers they cite above, with an edge wherever one page names another (findings.links) or cites a paper (citedBy). The outputs view lists each deliverable the readout names (its bar, present or missing, a published copy) with the files under it, then the declared folder's other files, and draws a selected file by its kind: Markdown as a document (a link to another delivered file opens it), images inline, CSV and TSV as a table, JSON pretty-printed, and code, HTML and anything else as numbered text. Nothing a run wrote is executed: HTML is shown as source, and a model is "not runnable here" with the reason, so no editable inputs are offered. A run that ended without its deliverable points to its brief as its readable result. On the play page the inspector also shows what the output changed from the version before: deliverable by deliverable (matched by the readout's id), which files were added, removed or changed by content hash, each changed text one click from its line diff and each changed image drawn before and after. It is a comparison of the delivered files, not a measured effect. When the brief says what an agent is doing, that line sits under the agent's row, labelled as the observer's words. Every score sits on the readout's absolute 0–100 vs world-class scale, side by side per category: the readout's AI judges (advisory unless calibrated, and "retired scale" for a judge that used the old relative 0–4 scale), the AI persona panel (finalOutput.panel: median, range and count) and each person's latest grade (finalOutput.grades). The table appears in the readout drawer and the version inspector, and the answer strip shows the AI judges' median, or the personas' median when no judge used the absolute scale. The readout's charts (readout.charts) are drawn in the readout drawer and the outputs view, and the version inspector pairs two versions' charts before and after: the same chart where both drew it, else charts of the same unit. The readout drawer also shows the run's latest trace review (finalOutput.traceReview, discovery-lab runner/trace-review.mjs): the requester's goal, whether it is a live review or the final one written after the run settled, how many agent sessions the reviewing model read and how many it could not, and one row per question in plain words with its 0–100 score (higher is better), verdict, or the failure of a question it could not answer; each row's quotes from the agents and suggestions fold under it. A run without a review shows no section. When the selected version ran an optimizer search (the profile index's searches, with proposed versions), the play page shows it below the graph: the claim it made (shipped, with the held-out test effect and its interval), the best selection score against evaluation spend, a replay over the search's clock, and every proposed version as a tree with its operator, decision, selection score, the paired effect on its parent drawn as an interval and its proposal and evaluation cost. Selecting a version shows the optimizer's quoted reason, its train, selection and test scores, its decision, the diff from its parent and its prompt blame (which version introduced each line). A page served by a writer (me says {writer: true}, from the host's tailnet identity check) adds a one-tap grade: tap the 0–100 track, add an optional comment, and Save posts runs/<id>/grades with the target (the run, a deliverable or a file). Elsewhere no write control appears. Under Conversations, the agents are a tree on one time axis (the rail and the conversation both stay in the window as the page scrolls): outcome, when each agent ran, time, tokens, list price, capture and adverse assessment flags. Selecting a row opens that agent in the drawer, with its conversation, timeline, usage, spend, assessments, input and coverage, and a replay over its own events. The page draws from the record's node list (runs/<id>/record?part=nodes: every node, no events, and nodeStats per agent with its event count, time span, token usage, list price, tool calls, latest words and profile digest) and loads one agent's events when it is opened (runs/<id>/record?node=<agent>). A run with a 5 MB record opens as fast as a small one. A readout not yet written reads "Readout pending"; a missing number reads "unknown", never 0, and only absolute http(s) links are followed. The run page's Versions section draws runs/<id>/versions (agent-workspace.version-graph): every version the run's agents wrote, newest first, one lane per branch (main, the shared store, and one per worker), merges joining a worker's write into main, and the profiles the run spawned in the same graph. Release candidates list their scores (exact checks, held-out checks, open blockers, judge), the tag the registered rule picks, and regression and suspected judge-gaming flags; selecting a version shows its files, message and checks, and opens the conversation of the worker that wrote it. The play page draws its versions with the same graph.

It reads plays/<id>, plays/<id>/profiles, runs/<id>, runs/<id>/record, runs/<id>/versions, runs/<id>/profile-change/<commit>, runs/<id>/assessments, runs/<id>/source/<sha256>, runs/<id>/page/<sha256>, runs/<id>/final/<path> and dimensions under data-api; the shapes are the schemas in @drewstone/agent-record/workspace and /assessment. The Overview's 30-day lost agent-hours KPI covers all catalog runs started in the stated UTC window. Beside the daily runs chart, it stacks lost hours by the five largest causes plus other and shows lost hours divided by all runs started each day. Days without a start have no per-run rate. The 7/14/30-day selector changes these charts while the KPI retains its stated 30-day scope. Page state lives in the URL (?tab=&v=&profile=&run=&show= on a play, ?q=&program=&sort=&show= on the plays index, ?node=&view=&file=&drawer=&event=&tab=&t=&dim= on a run, where section picks a detail view and no section is the overview; older view=readout|outputs, tab= and node= links still land on the matching view), and a running run polls every 3 seconds. The plays page also shows the last seven days' run reliability from reliability (discovery-lab tools/reliability.mjs, discovery-lab.reliability.v1): the share of settled runs that ran clean (below 80% in the failure colour), each day's share and run count, and the top three causes still happening (seen in the last 48 hours), most runs lost first, each with its layer, owner and when it was last seen. A cause that topped the week but stopped days ago is not listed as live. Tests and smoke runs, failed runs and archived runs are hidden until show=all ("Show them"), and the page counts what it hides, failures first; the host decides each run's and play's hidden reason, the viewer never guesses it. A document the host serves as last composed while it recomposes it (X-Workspace-Stale: 1) is fetched again after 2.5 s, backing off, for about a minute. An agent of a live run without a terminal state reads working; one whose status says done with no recorded usage reads done · no usage recorded (and ; capture missing when its conversation was not captured); an agent with no recorded event says so in place of an empty conversation. Theme tokens come from the host: --ar-background, --ar-surface, --ar-raised, --ar-line, --ar-foreground, --ar-muted, --ar-faint, --ar-accent, --ar-frame, --ar-ok, --ar-warn, --ar-crit, --ar-font and --ar-mono.

The conversation list renders only the rows near its viewport, so records with tens of thousands of events stay responsive.

Interaction and embedding

The viewer includes a recursive agent tree, scrollable activity timeline, conversation, source inspector, and usage plots. Tool inputs and results are paired by the original node ID and call ID, within one run. Repeated call identities are marked ambiguous; a missing result is not treated as success. A result returned before its call's recorded timestamp stays visible, with its timing discrepancy labeled.

Replay supports adjustable speed, recorded time, event steps, and reduced motion. The time cursor hides future content in conversation, source details, and tooltips. Usage plots distinguish input, output, cache read, and cache write counters. Tool return time is the observed call/result interval, including queue and tool time; it is not model latency. Recorded order follows the supplied event array.

import { useState } from 'react'
import { AgentRecord, type RecordSelection, type RunRecord } from '@drewstone/agent-record'

function ControlledViewer({ records }: { records: readonly RunRecord[] }) {
  const [selection, setSelection] = useState<RecordSelection>({
    runId: records[0].runId,
    eventId: 'an-original-event-id',
    view: 'source',
  })
  return <AgentRecord records={records} selection={selection}
    onSelectionChange={setSelection} theme="dark" />
}
Prop Meaning
records Readonly array of validated RunRecord objects; run IDs must be unique.
selection Controlled run, node, event, time cursor, and tab.
defaultSelection Initial selection when the component owns its state.
onSelectionChange User selection callback; use it to implement URLs or persistence.
theme auto, light, or dark; defaults to the system theme.
className Additional class for host styles.

RecordSelection has runId, optional nodeId, eventId, at, and view (chat, source, or usage). Omit at to show the complete record. Pass new immutable record objects when evidence changes. Several viewers can share a page without sharing state or DOM IDs. The example includes a second-instance toggle.

Import CSS once. All selectors are scoped to .agent-record; it does not restyle the surrounding page. Override --ar-background, --ar-foreground, --ar-font, and --ar-mono, or the semantic tokens in styles.css. A host with a sticky header of its own sets --ar-sticky-top (default 8px) so the topology column and the profile detail stick below it. Text uses one scale of three sizes: --ar-text-s (15px: labels, chips, axes, table headers), --ar-text-m (17px: body, tabs, table cells) and --ar-text-l (26px: titles); override them together to scale the viewer.

Run the example

pnpm install --frozen-lockfile
pnpm build
pnpm dev

The example uses reviewed records from the research site. Redactions and excerpts are labeled; the original discovery capture is incomplete. Open a record file to inspect your own data entirely in the browser.

Release checks

pnpm install --frozen-lockfile
pnpm build
pnpm typecheck
pnpm test
pnpm build:example
pnpm pack

This repository is public. Keep evidence screenshots and recordings in private storage and describe them in the pull request; never commit them under docs/evidence. Never commit a tailnet address (a Tailscale 100.x IP or a ts.net name). tests/public-repo.test.mjs refuses both.

Consumer evidence covers the built package, browser interactions, source preservation, server rendering, and the blog integration. The conversation renders only the rows near its viewport.

Extracted from Drew Stone's research site at 95e0aaf. Code is MIT licensed. Example records retain their source references and publication notes; upstream papers and excerpts retain their original rights.

About

React viewer for agent conversations, tool execution, topology, and recorded usage.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages