Skip to content

feat(iris_doc): run the named test after a compiled write, and say whether anything compiled - #371

Merged
PYDuquesnoy merged 2 commits into
masterfrom
feat/doc-put-runs-the-test
Sep 28, 2026
Merged

PYDuquesnoy merged 2 commits into
masterfrom
feat/doc-put-runs-the-test

Conversation

@PYDuquesnoy

Copy link
Copy Markdown
Contributor

Items 2 and 4 of issue 327. Items 1 and 3 are not here — see the last section.

Every number below is from the issue's corpus: 31,009 parsed tool_use blocks across 364
transcripts (parsed, not grepped — the issue records that a raw grep overcounts by 10×).

Item 2 — test: "<%UnitTest class>" on a write (579 calls)

iris_doc(put) is immediately followed by iris_test 579 times. A boolean would cover only
the 346 where the class written IS the class tested; a string covers all 579, because in the
other 233 the suite lives beside the class (Hospital.BO.PatientDb →
Hospital.Tests.BO.PatientDbTest).

iris_doc { mode: "put", name: "Hospital.BO.PatientDb.cls", content: "...",
           compile: true, test: "Hospital.Tests.BO.PatientDbTest" }

The decision is four states, not a boolean

The ways a test does not run need different fixes from the caller, so each is its own state with
its own sentence:

state what the caller gets
Run(pattern) the payload reports a compiled class → the suite runs
NotRequested nothing added to the payload. A test_skipped on every ordinary put would be noise on the 2,979 measured puts that want none of this
SkippedCallFailed the write's own envelope, returned untouched
SkippedNotCompiled { compile_requested } two different sentences — "pass compile=true" vs "read compile_errors"

The last row is the one that matters: "you did not ask for a compile" and "a compile ran and
produced nothing" are different states, and one sentence for both tells half the callers the wrong
fix. not_asking_for_a_compile_and_a_compile_that_produced_nothing_are_different_answers asserts
the two sentences differ and what each must name.

A failed write is returned exactly as it came back. Its envelope already carries what the
caller must act on — a compile error, an SCM elicitation, a refusal — and rebuilding it to add a
note about a test that did not run risks dropping the isError flag and the hints that are the
actual answer. Running the suite there would report a red for a class that is not present.

Gated on the payload's compiled, not on a list of modes, so nothing has to be kept in step with
DocMode: every mode that compiles reports it the same way, and a mode that never compiles has no
key to be true. A result carrying nothing parseable is treated as a failed call, never as a
compiled class — guessing the other way would run a suite against a state this code cannot see.

The two verdicts stay apart

success remains the write's: the document landed. The test arrives as test (its whole
payload — error code, per-case failures and all) with test_ok beside it. A red test must not read
as a failed write, and a failed write must not read as a red test — the issue names #310 for
exactly this, and a_red_test_does_not_rewrite_the_writes_own_verdict pins it, with
a_green_test_is_reported_as_green as the control that "always false" would not satisfy.

Why the sequencing lives in the tool method

Running a suite is iris_test's own ~600-line job: the poll loop, the result query, the per-case
shaping from #233/#273. A copy of that beside the writer would drift, so iris_doc calls the tool.
The two paths then cannot disagree about what a red test looks like.

Item 4 — every write says whether it compiled (the 89-call tail)

89 of the measured put → iris_compile pairs are a put that did not pass compile, followed by
a compile of the same class. The issue calls that a 3% tail rather than a feature gap and
proposes the cheap option over flipping the default: have the put say so.

The no-compile branch carried no compiled key at all, so "was it compiled?" was answerable
only by noticing an absence — and an absence reads as "does not apply here", not as "no". All three
write outcomes now report compiled and compile_requested:

compiled compile_requested state
false false written; nobody asked for a compile
false true the compile ran and failed (with COMPILE_ERROR)
true true compiled clean

Distinguishable from those two fields alone, which is the point: compiled: false by itself cannot
say why.

Verification

  • cargo fmt --all -- --check clean, cargo clippy --workspace --all-targets -- -D warnings
    clean, cargo build --workspace before the run.

  • Gate with IRIS env unset: 4 result lines, 0 FAILED, 1071 lib tests. Each new test confirmed
    by name; a fabricated name returns 0 as the control. mcp_handshake and put_refuses_names are
    in the gate because this changes the advertised description.

  • 15 new tests.

  • 12 mutations, each red on its named tests and restored green:

    mutation red
    a blank test counts as a request a_blank_test_is_not_a_request
    the test runs even when the write failed a_failed_write_does_not_run_the_test
    an unreadable result is treated as a compiled class an_unreadable_result_is_not_treated_as_a_compiled_class
    one sentence for both not-compiled states not_asking_for_a_compile_…_different_answers
    a red test rewrites the write's own verdict a_red_test_does_not_rewrite_the_writes_own_verdict
    test_ok is always true same, with the green test still passing
    the no-compile branch stops reporting compile_requested all_three_write_outcomes_report_both_compile_keys
    iris_doc stops handing its result to the sequencer iris_doc_hands_its_result_to_the_test_sequencer
    the sequencer stops calling iris_test same
    the description drops "(needs compile=true)" the_description_says_the_test_needs_a_compile
    the description stops naming test_skipped same

One mutation survived first, and the oracle was the bug

"The description drops (needs compile=true)" passed against the first version of that test,
which asserted description.contains("compile"). The word appears eight other times in that
description (compile=true on the compile flag, compile_errors, compile_console, …), so the
assertion could not fail for the reason it was written. It now reads only the sentence that
introduces test
— bounded, with a length check and a control phrase from elsewhere in the
description that must not fall inside the window. Both description mutations are red against the
repaired version. Recorded in the test's own doc comment so it is not loosened again.

Not in this change

  • Item 1 (a put accepting a document list, 724 calls) is a bigger change than it looks: the
    per-document write gate (file-on-disk, storage guard) applies per element, the response has to
    carry per-document status rather than one boolean, and the point of doing it in the tool at all is
    issuing one compile after every write lands. That deserves its own PR.
  • Item 3 (iris_production_item accepting a list on get_settings / add, 174 calls) is
    independent of both items here.

Issue 327 stays open for items 1 and 3; its checklist boxes 2 and 4 are what this covers.

PYDuquesnoy and others added 2 commits September 28, 2026 11:30
…ether anything compiled

Items 2 and 4 of issue 327, measured over 31,009 parsed `tool_use` blocks from 364 transcripts.

`iris_doc(put)` is immediately followed by `iris_test` 579 times. A boolean would cover only
the 346 where the class written IS the class tested; a STRING covers all 579, because in the
other 233 the suite lives beside the class (`Hospital.BO.PatientDb` ->
`Hospital.Tests.BO.PatientDbTest`).

The decision is a pure `TestGate` with four states, not a boolean, because the ways a test does
NOT run need different fixes from the caller:

* `Run(pattern)` — the payload reports a compiled class.
* `NotRequested` — nothing added to the payload at all. A `test_skipped` on every ordinary put
  would be noise on the 2,979 measured puts that want none of this.
* `SkippedCallFailed` — the write did not succeed. Its own envelope is returned UNTOUCHED: it
  already carries the compile error, SCM elicitation or refusal the caller must act on, and
  rebuilding it to add a note would risk dropping the isError flag and hints that are the answer.
* `SkippedNotCompiled { compile_requested }` — two different sentences, because "you did not ask
  for a compile" and "the compile produced nothing" are different states and need different
  fixes.

Keyed off the payload's `compiled`, not off the mode, so it needs no list of modes to keep in
step with `DocMode`.

The write's verdict and the test's verdict stay separate (#310): `success` stays the WRITE's —
the document landed — and the test arrives as `test` with `test_ok` beside it. A red test must
not read as a failed write, and a failed write must not read as a red test.

Sequenced in the `iris_doc` tool method rather than inside `doc::handle_iris_doc`, because
running a suite is `iris_test`'s own 600-line job (polling, the result query, the per-case
shaping of #233/#273). Calling the tool means the two paths cannot disagree about what a red
test looks like.

89 of the measured `put` -> `iris_compile` pairs are a put that did not pass `compile`, followed
by a compile of the SAME class. Issue 327 records that as a 3% tail and proposes the cheap
option rather than flipping the default: have the put say so.

The no-compile branch carried NO `compiled` key at all, so "was it compiled?" was answerable
only by noticing an absence — and an absence reads as "does not apply here", not as "no". All
three write outcomes now report `compiled` AND `compile_requested`, so the states are distinct
from those two fields alone: (false,false) written and nobody asked, (false,true) the compile ran
and failed, (true,true) compiled clean.

fmt, clippy `-D warnings`, `cargo build --workspace`, and the gate with IRIS env unset. 15 new
tests. 10 mutations applied one at a time — including "a blank test counts as a request", "a red
test rewrites the write's own verdict", "iris_doc stops handing its result to the sequencer" and
"the sequencer stops calling iris_test" — each red on its named tests, restored green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Ran this PR's test-after-put against a real instance for the first time. The behaviour is right in all
five cases — a green suite gives test_ok: true; a RED suite gives test_ok: false while the write's own
success stays true, which is the #310 separation working live; no compile gives test_skipped; a compile
failure gives COMPILE_ERROR with the write envelope untouched and no test block; and no `test`
parameter adds nothing at all.

What was wrong is the text. The skip message came back as:

  not run: `test` needs a COMPILED class and this call did not compile. Pass                  compile=true

Four messages, all in this PR: the three TestGate::skipped_reason arms and the unparseable-result
note. Same collapsed line-continuation as the seven fixed in #376/#377 — I fixed those on the branches
I happened to be working in and did not check the others.

## The scan that found the rest, and the one that did not

Running the earlier detector over every branch was useless: master-descended branches report 15-40 hits
each, because most matches are whitespace test fixtures, embedded ObjectScript whose indentation is
deliberate, and aligned doc-comment tables. No conclusion is available from those numbers.

Scoping it to the lines a branch ADDS (`git diff origin/master...origin/<br>`, `^+` only) makes the
population precise. Across all fifteen open PR branches that gives four real leaks, all four here.
Everything else the scan flags on those branches is intentional: measurement tables in doc comments,
the deliberate `"ERROR #5002:   "` truncation fixture, the real IRIS SQLCODE text in #368's fixtures
(which genuinely contains double spaces), and the leaked string the #380 detector test asserts on.

Verified at runtime after the fix by driving the built binary: the skip message contains no double
space.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@PYDuquesnoy
PYDuquesnoy force-pushed the feat/doc-put-runs-the-test branch from 0b514cc to 028e176 Compare September 28, 2026 09:34
@PYDuquesnoy
PYDuquesnoy merged commit fe5c17a into master Sep 28, 2026
5 checks passed
@PYDuquesnoy
PYDuquesnoy deleted the feat/doc-put-runs-the-test branch September 28, 2026 09:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant