Skip to content

Fail when the mirror is more permissive than the system it mirrors - #16

Merged
rajatrnaura merged 4 commits into
rajatfrom
sumer/drift-and-labels
Aug 20, 2026
Merged

rajatrnaura merged 4 commits into
rajatfrom
sumer/drift-and-labels

Conversation

@sumerjohal

Copy link
Copy Markdown
Contributor

Fail when the mirror is more permissive than the system it mirrors

hub.users.phone was relaxed to nullable and non-unique in migration/models.py,
and phone was removed from the constraint guard, while ar2-hub still enforced
NOT NULL UNIQUE on that column. The import consequently reported 0 rejected
accounts where a real run against the hub would still have rejected 85 -- 8 with
no phone, 77 sharing a number. The constraint was removed from the simulation,
not from the system.

That is the worse failure shape available: not an error, but a clean bill of
health for an import that cannot happen. The existing drift guard compared
column names only, so it passed throughout -- the column was still called
phone.

  • test_mirrored_constraints_match_the_real_schema compares nullability and
    uniqueness against the real models, and fails only on permissiveness. A
    mirror that is stricter than reality predicts rejections that will not
    happen, which is a false alarm; a mirror that is looser predicts acceptance
    for rows the system rejects, which reads as success.

  • The docstring on test_hub_constraints_the_import_depends_on_are_still_there
    records why phone legitimately left both lists, and that the order it
    happened in was backwards. Relaxing the guard is the last step, never the
    first.

  • The real relaxation is ar2-hub#7, with the DDL create_all will not apply.

Also, three reporting defects. Every number below is one a reader would act on,
and each was wrong in the direction of looking worse or looking better than the
run actually was.

  • 268 fields that resolved correctly on an identical content hash were counted
    as quarantined and then listed under DECISIONS REQUIRED. Every one had
    aliased to its canonical row and still resolves. A run that rejected 7 fields
    read as a run with 275 open problems. They now count as resolutions, split
    into exact-geometry and above-threshold matches, because certainty and
    judgement should not share a number.

  • Every imported point was counted as canonicalization-altered geometry,
    taking that figure from 3,597 to 18,184 and converting a real signal about
    boundary repair into a headcount of pins. A point has no boundary to repair.

  • The point count was not reported at all, though a registry that is half pins
    is a different thing to plan around than one that is all boundaries. It was
    the most consequential fact about the run and appeared nowhere in it.

  • "quarantined" is now "REJECTED, not imported", with the share of the total,
    and the decisions section says that a reason naming a geometry type is
    usually a missing code path rather than a policy question.

Finally, the harness submodule moves to f7ac8b0. The pinned revision counted
ruff's "All checks passed!" banner as one violation, so CI reported 1 where this
tree has 36 -- the ratchet has been reading green while blind. The gating
workflow also installs ruff unpinned, which the earlier pin missed because it
only covered ci.yml; patches/pin-ruff-in-stomata-workflow.patch has that, since
my token cannot write .github/workflows.

sumerjohal and others added 4 commits August 19, 2026 11:28
hub.users.phone was relaxed to nullable and non-unique in migration/models.py,
and phone was removed from the constraint guard, while ar2-hub still enforced
NOT NULL UNIQUE on that column. The import consequently reported 0 rejected
accounts where a real run against the hub would still have rejected 85 -- 8 with
no phone, 77 sharing a number. The constraint was removed from the simulation,
not from the system.

That is the worse failure shape available: not an error, but a clean bill of
health for an import that cannot happen. The existing drift guard compared
column *names* only, so it passed throughout -- the column was still called
phone.

  * test_mirrored_constraints_match_the_real_schema compares nullability and
    uniqueness against the real models, and fails only on permissiveness. A
    mirror that is stricter than reality predicts rejections that will not
    happen, which is a false alarm; a mirror that is looser predicts acceptance
    for rows the system rejects, which reads as success.

  * The docstring on test_hub_constraints_the_import_depends_on_are_still_there
    records why phone legitimately left both lists, and that the order it
    happened in was backwards. Relaxing the guard is the last step, never the
    first.

  * The real relaxation is ar2-hub#7, with the DDL create_all will not apply.

Also, three reporting defects. Every number below is one a reader would act on,
and each was wrong in the direction of looking worse or looking better than the
run actually was.

  * 268 fields that resolved correctly on an identical content hash were counted
    as quarantined and then listed under DECISIONS REQUIRED. Every one had
    aliased to its canonical row and still resolves. A run that rejected 7 fields
    read as a run with 275 open problems. They now count as resolutions, split
    into exact-geometry and above-threshold matches, because certainty and
    judgement should not share a number.

  * Every imported point was counted as canonicalization-altered geometry,
    taking that figure from 3,597 to 18,184 and converting a real signal about
    boundary repair into a headcount of pins. A point has no boundary to repair.

  * The point count was not reported at all, though a registry that is half pins
    is a different thing to plan around than one that is all boundaries. It was
    the most consequential fact about the run and appeared nowhere in it.

  * "quarantined" is now "REJECTED, not imported", with the share of the total,
    and the decisions section says that a reason naming a geometry type is
    usually a missing code path rather than a policy question.

Finally, the harness submodule moves to f7ac8b0. The pinned revision counted
ruff's "All checks passed!" banner as one violation, so CI reported 1 where this
tree has 36 -- the ratchet has been reading green while blind. The gating
workflow also installs ruff unpinned, which the earlier pin missed because it
only covered ci.yml; patches/pin-ruff-in-stomata-workflow.patch has that, since
my token cannot write .github/workflows.

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
… adds up

Seeded .stomata/lessons.json from the review record of 2026-07-24 to 2026-08-19.
Seven recurrence keys, every occurrence traceable to a dated document. Nothing
speculative: a lesson with no incident behind it becomes noise, and noise is what
discredits the checks that were earned.

What the corpus shows, which was not visible one review at a time:

  7x  a gate reported green while blind
  6x  one fact in two places, one of them updated
  5x  a label or count that described a run that did not happen
  4x  a number quoted that was not in the evidence
  3x  a failure mode written down as prose instead of a check
  2x  a guard narrowed to let a change through

The last two are the interesting ones. The prose entry is about this review
process rather than the code: three times a correct warning was written in a
document and did not bind, including one that predicted the guard-weakening
incident five days before it happened. That is the argument for checks J and K
existing in code.

test_report_arithmetic.py mechanizes the label lesson. Every individual number in
the import report was computed correctly and the totals still misdescribed the
run: 268 successful aliases reported as quarantined, canonicalisation overstated
5x by points on another code path. What was missing was an assertion about the
relationship between categories, which is checkable.

It found a live one immediately: resolved_child_of was presented as a peer of
imported_new when it is a subset, so 74 fields summed to 75 and read as one field
counted twice. The display was already correctly nested; the invariant was
unstated, and is now stated where it can fail.

Two guards declared. Verified by removing an assertion from the drift guard and
confirming J fails, quoting the incident, then restoring it. 11/11 checks pass
with mutations, 229 tests, 6/6 recurring lessons mechanized.

Co-authored-by: Cursor <cursoragent@cursor.com>
Declares AGENTS.md as the briefing target and generates its lessons section from
.stomata/lessons.json. AGENTS.md loads automatically, which is the point: the
ledger was being verified after every run and read before none.

The hand-written part covers the two things worth knowing before a first failure.
Do not narrow a check to get past it -- if a guard blocks a change you believe is
correct, one of the two is wrong and it is worth working out which, often the
guard. And if a check fires on something genuinely fine, that is a bug in the
check, to be fixed rather than worked around, because a check that cries wolf
becomes noise and noise discredits the checks that were earned.

Verified by hand-editing the generated block and confirming check L fails naming
the drift, then restoring it. The first attempt at that verification proved nothing
-- it replaced text that was not in the block -- which is the claim-not-in-the-
evidence lesson catching me in the act of testing a check against nothing.

12/12 checks with mutations, 229 tests.

Co-authored-by: Cursor <cursoragent@cursor.com>
@rajatrnaura
rajatrnaura merged commit da675e0 into rajat Aug 20, 2026
6 of 9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants