Skip to content

Docs-as-evals: harness that measures whether coding agents can complete core tasks from the docs alone #148

Description

@justin808

Part of the docs & demos tracking issue: #140. The keystone for "docs that coding agents can actually follow": don't guess whether agents can complete tasks from the docs — measure it.

Concept

A recurring eval harness where a fresh agent session (no prior context, no repo access) gets only the published docs (via llms.txt / reactonrails.com) plus a bare environment, and must complete each canonical task end-to-end. Every failure or stumble becomes a docs issue with the exact page and step where the agent went wrong. Run per release (and optionally weekly), so docs regressions are caught the way tests catch code regressions.

This is the empirical extension of the existing /prompts page: the prompts promise agent-driven setup works; the evals prove it and keep it true.

Canonical task list (v1)

  1. Create a new app (npx create-react-on-rails-app) and render a component with props from Rails
  2. Add React on Rails to an existing bare Rails app
  3. Enable server-side rendering for one component
  4. Enable streaming SSR (Pro)
  5. Enable React Server Components via the react_on_rails:rsc generator (Pro)
  6. Implement a mutation with useRailsForm + a Rails controller (per mutations.md)
  7. Deploy a demo to a fresh host (pairs with Add one-click deploy buttons to all demo repos + 'Deploy a demo in 5 minutes' docs page #142)
  8. Migrate one page of a small Inertia app using the coexistence setup

Harness sketch

  • Runner: scripted agent invocations (Claude Code headless / Codex CLI) with a pinned system prompt: "use only the docs at reactonrails.com; do not use prior knowledge of the framework where docs conflict"
  • Each task: clean workspace (Docker or throwaway dir), success command (e.g., page renders, test passes, curl returns SSR HTML), transcript capture
  • Scoring: pass/fail + time + number of doc lookups + where the agent got stuck (page + step)
  • Output: a scoreboard (task × model × date) and auto-drafted issues for failures
  • Cadence: on release tags of react_on_rails / react_on_rails_pro; manual dispatch for docs PRs that touch journey pages

Why this is worth it

  • Catches doc rot mechanically instead of by user complaint
  • Produces a defensible marketing claim no competing framework makes: the docs are agent-tested every release
  • The stumble logs are the highest-signal input for the Importance × Quality triage (see the inventory pipeline issue)

Notes

Acceptance criteria (v1)

  • Tasks 1, 2, 3, and 6 runnable end-to-end by a scripted fresh agent with pass/fail output and captured transcripts; failures produce actionable stumble reports naming the doc page and step.

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentationenhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions