You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Part of the docs & demos tracking issue: #140. The keystone for "docs that coding agents can actually follow": don't guess whether agents can complete tasks from the docs — measure it.
Concept
A recurring eval harness where a fresh agent session (no prior context, no repo access) gets only the published docs (via llms.txt / reactonrails.com) plus a bare environment, and must complete each canonical task end-to-end. Every failure or stumble becomes a docs issue with the exact page and step where the agent went wrong. Run per release (and optionally weekly), so docs regressions are caught the way tests catch code regressions.
This is the empirical extension of the existing /prompts page: the prompts promise agent-driven setup works; the evals prove it and keep it true.
Canonical task list (v1)
Create a new app (npx create-react-on-rails-app) and render a component with props from Rails
Add React on Rails to an existing bare Rails app
Enable server-side rendering for one component
Enable streaming SSR (Pro)
Enable React Server Components via the react_on_rails:rsc generator (Pro)
Implement a mutation with useRailsForm + a Rails controller (per mutations.md)
Migrate one page of a small Inertia app using the coexistence setup
Harness sketch
Runner: scripted agent invocations (Claude Code headless / Codex CLI) with a pinned system prompt: "use only the docs at reactonrails.com; do not use prior knowledge of the framework where docs conflict"
Each task: clean workspace (Docker or throwaway dir), success command (e.g., page renders, test passes, curl returns SSR HTML), transcript capture
Scoring: pass/fail + time + number of doc lookups + where the agent got stuck (page + step)
Output: a scoreboard (task × model × date) and auto-drafted issues for failures
Cadence: on release tags of react_on_rails / react_on_rails_pro; manual dispatch for docs PRs that touch journey pages
Why this is worth it
Catches doc rot mechanically instead of by user complaint
Produces a defensible marketing claim no competing framework makes: the docs are agent-tested every release
The stumble logs are the highest-signal input for the Importance × Quality triage (see the inventory pipeline issue)
Notes
The harness likely graduates to its own repo (react-on-rails-docs-evals) once it exists; tracked here because this is the docs-quality program home
Tasks 1, 2, 3, and 6 runnable end-to-end by a scripted fresh agent with pass/fail output and captured transcripts; failures produce actionable stumble reports naming the doc page and step.
Part of the docs & demos tracking issue: #140. The keystone for "docs that coding agents can actually follow": don't guess whether agents can complete tasks from the docs — measure it.
Concept
A recurring eval harness where a fresh agent session (no prior context, no repo access) gets only the published docs (via llms.txt / reactonrails.com) plus a bare environment, and must complete each canonical task end-to-end. Every failure or stumble becomes a docs issue with the exact page and step where the agent went wrong. Run per release (and optionally weekly), so docs regressions are caught the way tests catch code regressions.
This is the empirical extension of the existing /prompts page: the prompts promise agent-driven setup works; the evals prove it and keep it true.
Canonical task list (v1)
npx create-react-on-rails-app) and render a component with props from Railsreact_on_rails:rscgenerator (Pro)useRailsForm+ a Rails controller (permutations.md)Harness sketch
Why this is worth it
Notes
react-on-rails-docs-evals) once it exists; tracked here because this is the docs-quality program homellms.txtAcceptance criteria (v1)
🤖 Generated with Claude Code