feat(iris_doc): run the named test after a compiled write, and say whether anything compiled - #371
Merged
Merged
Conversation
This was referenced Sep 23, 2026
…ether anything compiled
Items 2 and 4 of issue 327, measured over 31,009 parsed `tool_use` blocks from 364 transcripts.
`iris_doc(put)` is immediately followed by `iris_test` 579 times. A boolean would cover only
the 346 where the class written IS the class tested; a STRING covers all 579, because in the
other 233 the suite lives beside the class (`Hospital.BO.PatientDb` ->
`Hospital.Tests.BO.PatientDbTest`).
The decision is a pure `TestGate` with four states, not a boolean, because the ways a test does
NOT run need different fixes from the caller:
* `Run(pattern)` — the payload reports a compiled class.
* `NotRequested` — nothing added to the payload at all. A `test_skipped` on every ordinary put
would be noise on the 2,979 measured puts that want none of this.
* `SkippedCallFailed` — the write did not succeed. Its own envelope is returned UNTOUCHED: it
already carries the compile error, SCM elicitation or refusal the caller must act on, and
rebuilding it to add a note would risk dropping the isError flag and hints that are the answer.
* `SkippedNotCompiled { compile_requested }` — two different sentences, because "you did not ask
for a compile" and "the compile produced nothing" are different states and need different
fixes.
Keyed off the payload's `compiled`, not off the mode, so it needs no list of modes to keep in
step with `DocMode`.
The write's verdict and the test's verdict stay separate (#310): `success` stays the WRITE's —
the document landed — and the test arrives as `test` with `test_ok` beside it. A red test must
not read as a failed write, and a failed write must not read as a red test.
Sequenced in the `iris_doc` tool method rather than inside `doc::handle_iris_doc`, because
running a suite is `iris_test`'s own 600-line job (polling, the result query, the per-case
shaping of #233/#273). Calling the tool means the two paths cannot disagree about what a red
test looks like.
89 of the measured `put` -> `iris_compile` pairs are a put that did not pass `compile`, followed
by a compile of the SAME class. Issue 327 records that as a 3% tail and proposes the cheap
option rather than flipping the default: have the put say so.
The no-compile branch carried NO `compiled` key at all, so "was it compiled?" was answerable
only by noticing an absence — and an absence reads as "does not apply here", not as "no". All
three write outcomes now report `compiled` AND `compile_requested`, so the states are distinct
from those two fields alone: (false,false) written and nobody asked, (false,true) the compile ran
and failed, (true,true) compiled clean.
fmt, clippy `-D warnings`, `cargo build --workspace`, and the gate with IRIS env unset. 15 new
tests. 10 mutations applied one at a time — including "a blank test counts as a request", "a red
test rewrites the write's own verdict", "iris_doc stops handing its result to the sequencer" and
"the sequencer stops calling iris_test" — each red on its named tests, restored green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Ran this PR's test-after-put against a real instance for the first time. The behaviour is right in all five cases — a green suite gives test_ok: true; a RED suite gives test_ok: false while the write's own success stays true, which is the #310 separation working live; no compile gives test_skipped; a compile failure gives COMPILE_ERROR with the write envelope untouched and no test block; and no `test` parameter adds nothing at all. What was wrong is the text. The skip message came back as: not run: `test` needs a COMPILED class and this call did not compile. Pass compile=true Four messages, all in this PR: the three TestGate::skipped_reason arms and the unparseable-result note. Same collapsed line-continuation as the seven fixed in #376/#377 — I fixed those on the branches I happened to be working in and did not check the others. ## The scan that found the rest, and the one that did not Running the earlier detector over every branch was useless: master-descended branches report 15-40 hits each, because most matches are whitespace test fixtures, embedded ObjectScript whose indentation is deliberate, and aligned doc-comment tables. No conclusion is available from those numbers. Scoping it to the lines a branch ADDS (`git diff origin/master...origin/<br>`, `^+` only) makes the population precise. Across all fifteen open PR branches that gives four real leaks, all four here. Everything else the scan flags on those branches is intentional: measurement tables in doc comments, the deliberate `"ERROR #5002: "` truncation fixture, the real IRIS SQLCODE text in #368's fixtures (which genuinely contains double spaces), and the leaked string the #380 detector test asserts on. Verified at runtime after the fix by driving the built binary: the skip message contains no double space. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
PYDuquesnoy
force-pushed
the
feat/doc-put-runs-the-test
branch
from
September 28, 2026 09:34
0b514cc to
028e176
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Items 2 and 4 of issue 327. Items 1 and 3 are not here — see the last section.
Every number below is from the issue's corpus: 31,009 parsed
tool_useblocks across 364transcripts (parsed, not grepped — the issue records that a raw grep overcounts by 10×).
Item 2 —
test: "<%UnitTest class>"on a write (579 calls)iris_doc(put)is immediately followed byiris_test579 times. A boolean would cover onlythe 346 where the class written IS the class tested; a string covers all 579, because in the
other 233 the suite lives beside the class (
Hospital.BO.PatientDb→Hospital.Tests.BO.PatientDbTest).iris_doc { mode: "put", name: "Hospital.BO.PatientDb.cls", content: "...", compile: true, test: "Hospital.Tests.BO.PatientDbTest" }The decision is four states, not a boolean
The ways a test does not run need different fixes from the caller, so each is its own state with
its own sentence:
Run(pattern)NotRequestedtest_skippedon every ordinary put would be noise on the 2,979 measured puts that want none of thisSkippedCallFailedSkippedNotCompiled { compile_requested }The last row is the one that matters: "you did not ask for a compile" and "a compile ran and
produced nothing" are different states, and one sentence for both tells half the callers the wrong
fix.
not_asking_for_a_compile_and_a_compile_that_produced_nothing_are_different_answersassertsthe two sentences differ and what each must name.
A failed write is returned exactly as it came back. Its envelope already carries what the
caller must act on — a compile error, an SCM elicitation, a refusal — and rebuilding it to add a
note about a test that did not run risks dropping the
isErrorflag and the hints that are theactual answer. Running the suite there would report a red for a class that is not present.
Gated on the payload's
compiled, not on a list of modes, so nothing has to be kept in step withDocMode: every mode that compiles reports it the same way, and a mode that never compiles has nokey to be true. A result carrying nothing parseable is treated as a failed call, never as a
compiled class — guessing the other way would run a suite against a state this code cannot see.
The two verdicts stay apart
successremains the write's: the document landed. The test arrives astest(its wholepayload — error code, per-case failures and all) with
test_okbeside it. A red test must not readas a failed write, and a failed write must not read as a red test — the issue names #310 for
exactly this, and
a_red_test_does_not_rewrite_the_writes_own_verdictpins it, witha_green_test_is_reported_as_greenas the control that "always false" would not satisfy.Why the sequencing lives in the tool method
Running a suite is
iris_test's own ~600-line job: the poll loop, the result query, the per-caseshaping from #233/#273. A copy of that beside the writer would drift, so
iris_doccalls the tool.The two paths then cannot disagree about what a red test looks like.
Item 4 — every write says whether it compiled (the 89-call tail)
89 of the measured
put→iris_compilepairs are a put that did not passcompile, followed bya compile of the same class. The issue calls that a 3% tail rather than a feature gap and
proposes the cheap option over flipping the default: have the put say so.
The no-compile branch carried no
compiledkey at all, so "was it compiled?" was answerableonly by noticing an absence — and an absence reads as "does not apply here", not as "no". All three
write outcomes now report
compiledandcompile_requested:compiledcompile_requestedCOMPILE_ERROR)Distinguishable from those two fields alone, which is the point:
compiled: falseby itself cannotsay why.
Verification
cargo fmt --all -- --checkclean,cargo clippy --workspace --all-targets -- -D warningsclean,
cargo build --workspacebefore the run.Gate with IRIS env unset: 4 result lines, 0 FAILED, 1071 lib tests. Each new test confirmed
by name; a fabricated name returns 0 as the control.
mcp_handshakeandput_refuses_namesarein the gate because this changes the advertised description.
15 new tests.
12 mutations, each red on its named tests and restored green:
testcounts as a requesta_blank_test_is_not_a_requesta_failed_write_does_not_run_the_testan_unreadable_result_is_not_treated_as_a_compiled_classnot_asking_for_a_compile_…_different_answersa_red_test_does_not_rewrite_the_writes_own_verdicttest_okis always truecompile_requestedall_three_write_outcomes_report_both_compile_keysiris_docstops handing its result to the sequenceriris_doc_hands_its_result_to_the_test_sequenceriris_testthe_description_says_the_test_needs_a_compiletest_skippedOne mutation survived first, and the oracle was the bug
"The description drops (needs compile=true)" passed against the first version of that test,
which asserted
description.contains("compile"). The word appears eight other times in thatdescription (
compile=trueon the compile flag,compile_errors,compile_console, …), so theassertion could not fail for the reason it was written. It now reads only the sentence that
introduces
test— bounded, with a length check and a control phrase from elsewhere in thedescription that must not fall inside the window. Both description mutations are red against the
repaired version. Recorded in the test's own doc comment so it is not loosened again.
Not in this change
per-document write gate (file-on-disk, storage guard) applies per element, the response has to
carry per-document status rather than one boolean, and the point of doing it in the tool at all is
issuing one compile after every write lands. That deserves its own PR.
iris_production_itemaccepting a list onget_settings/add, 174 calls) isindependent of both items here.
Issue 327 stays open for items 1 and 3; its checklist boxes 2 and 4 are what this covers.