I'm Zain Dana Harper. I build tools that let anyone recheck what an AI system did, and I publish investigations into who knew about AI incidents first and who pays the people who check. I work independently as a sole proprietor in Kent, Washington, and I take scoped evaluation work.
A result is worth trusting when an outside skeptic can rerun the check on their own machine and reach the same verdict. That holds whoever ran the model: any lab, open or closed, any company, any nation. My own verdicts get the same treatment.
| Pick a door | Go |
|---|---|
| Run an AI task with any model and keep a record you can recheck | Flywheel and the flagships |
| Read the investigations | Who Knew First and the series |
| See where the work is heading | Watching the trace, checking before the action |
| Hire me for evaluation work | Work with me |
| Get in touch | Reach me |
Flywheel is a self-hostable,
model-agnostic AI workstation and coding harness. It runs any model, frontier
or local, behind one OpenAI-compatible surface, with your keys and data kept on
your machine. Rowan, the desktop assistant, turns a plain request into a
recorded run. Relay runs a permission-gated coding agent over your own folders.
Lanes add research intake, workspace maps, memory, agent routing and writing.
flywheel check-output checks answers against finance, medicine and law packs
and can emit Lean 4 proofs. Every accepted run leaves a sealed receipt that an
independent witness reruns offline, with no learned model deciding the verdict.
pip install flywheel-verify
flywheel up
flywheel lanes --probe
On the shipped benchmark the verified loop shows no measured accuracy gain over a single pass; the interval includes zero. The value is the workstation and a record you can check yourself. Flywheel is source-available under FSL-1.1-MIT.
Each flagship below also works on its own and plugs into Flywheel.
Current release of every flagship (refreshed daily from GitHub)
| Tool | What it does | Release |
|---|---|---|
| Flywheel | AI workstation and coding harness: any model, gated agent, rerunnable receipts | v1.1.2 (2026-09-29) |
| Articulate | Local writing checker and editor with content-free receipts | v0.7.0 (2026-10-01) |
| Telos | Accountable actuation: senses, actions and hardware control in permission tiers | v0.6.0 (2026-10-01) |
| Accountable Surface | Gates agent actions on explicit grants, with a hash-chained journal | v0.2.0 (2026-09-18) |
| Forum | Coordinates agent teams with a replayable ledger | v1.16.0 (2026-10-01) |
| Relay | Permission-checked coding agent for any model endpoint | v0.6.0 (2026-10-01) |
| Gather | Research intake from the web, papers, video, scans and audio, with provenance | v2.1.0 (2026-10-01) |
| Index | Offline repository and workspace maps with file and line evidence | v2.15.0 (2026-10-01) |
| Mneme | Agent memory where every recall can be rechecked | v0.6.0 (2026-10-01) |
| Canon | One memory and personality record shared across models and tools | v0.6.0 (2026-10-01) |
| Crucible | Tests falsifiable claims and records MATCH, DRIFT or UNVERIFIABLE | v1.4.0 (2026-10-01) |
| EMET | Checks that bytes reaching a model still match their source | v1.3.0 (2026-09-13) |
| Learn | Turns your own material into a course that never takes the test for you | v2.1.0 (2026-10-01) |
| Plexus | Finds and wires compatible tools in an agent toolchain | v0.3.0 (2026-10-01) |
| Phantom | Reversible hardware-identity privacy for owned Windows and Linux machines | v1.1.1 (2026-09-10) |
Work accepted upstream
Changes that survived another maintainer's review and merged:
- AgentFence PR 261: a Go engine optimization with deterministic rule selection and allocation coverage.
- Free Law Project PR 820: documentation for the fast database-free test suite.
- Mergewarden PR 107: replay fixtures for reusable-workflow pinning, revised after owner review.
The portfolio lists merged, open and closed contributions separately.
Who Knew First is a record of nine 2026 incidents in which an AI agent crossed a boundary. In six of them, someone outside the organization that ran the model told the public first. The argument is simple: whoever holds an incident's logs gets to name it, and the name decides how fast anyone else hears about it.
The series tests the questions that argument raises. Each piece stands alone, lists its sources and the confidence of each claim, and says what each claim does not prove.
| Piece | The question | Status |
|---|---|---|
| Who Pays the Referees | The people who check AI models depend on the labs they check. Which of those terms are public? | Published 1 October 2026 |
| The Terms for Telling | The party that holds the records also writes the contracts of the people who could tell. Who got heard? | Published 1 October 2026 |
| Who Kept the Books | In money cases from 1514 to Iran-Contra, what made the first account move? | Planned |
| The Maker Is Part of the Story | Three famous stories, read for who funds the work and who edits the record. | Planned |
| A Check It Cannot Predict | Does a check that is certain and outside the actor's control work on AI models too? | Planned, no result yet |
An Anthropic-built model helped compile these pieces, and Anthropic appears in the record, so each piece marks where Anthropic is a party and invites an outside check of those items.
Latest writing (refreshed daily from the site feed)
| Date | Piece |
|---|---|
| 2026-10-01 | Who Pays the Referees |
| 2026-10-01 | The Terms for Telling |
| 2026-10-01 | The Number Has a Vintage |
| 2026-09-28 | What the Formula Counts |
| 2026-09-28 | The Timestamp Is Not the Order |
| 2026-09-28 | The Scene the Song Did Not Tell You |
Everything else, essays, briefings and papers, is on the writing page.
A receipt tells you what happened after the fact. The next step is to watch an agent's trace as it runs and check each consequential action before it happens. If the check fails, the action halts on its own, with no person needing to step in, and the record shows why.
This is work in progress, and the pieces exist at different stages:
- Accountable Surface lets an agent take only the action a person approved, then verifies the outcome and rolls back what it can.
- Telos places sensing, actions and workstation hardware control in explicit permission tiers, each with confirmation points and receipts.
- Rowan's monitor, in Flywheel 1.2.0 on PyPI, refuses to trust a check that rewrote its own grading files.
Trace observation uses what providers document: reasoning summaries, token counts, effort settings and ordinary outputs. It never tries to pull hidden reasoning out of a model through jailbreaks or prompt injection, and it never bypasses an access control.
How the pieces connect
flowchart LR
T[Agent trace] --> M{Check before the action}
M -- passes --> A[Action runs]
M -- fails --> H[Halted, with the reason recorded]
A --> R[Sealed receipt]
H --> R
R --> W[Independent rerun: MATCH, DRIFT or UNVERIFIABLE]
W --> P[Published finding with its limits]
I take scoped work on evaluation design review, harness integration, agent safety review before an audit, incident review, and conflict-of-interest review. Each engagement gets a quote built from the labor, time, compute and tooling it needs. I have no paid client today and no current sponsors.
Independence comes first, so the rules are public:
- Income from this work is published with its exact source, API credits included.
- When one source passes 15 percent of income over twelve months, I disclose it and give my findings about that party a second review.
- At 50 percent, I decline new work evaluating that party. A first contract is most of the income by arithmetic, so it is disclosed in full and the decline rule waits for the second.
- One standard for every lab. I build with Anthropic and OpenAI models and publish investigations that name both.
Roles and paths
I'm also open to technical and nontechnical roles in AI governance and evaluation, and willing to relocate to London or travel to San Francisco.
| Path | Where I fit | Start here |
|---|---|---|
| Technical and evaluation | Agent and model evaluation, developer tooling, CI, security testing, technical support and documentation. | Engineering path |
| Public, union, and field | Public service, facilities, parks and grounds, arboriculture, scheduling and safety judgment, from eleven years of field work. | Public-service and field path |
| Education and research | Fellowships, research operations and evidence-centered technical writing. | Research |
- Email: zaindharper@gmail.com
- Site: harperz9.github.io
- LinkedIn: zaindanaharper
- Writing feed: feed.xml
- Hiring page: hire
The art on this page is generated by scripts/profile_art.py. Motion stops when your system asks for reduced motion.



