From 40369a6ab2e675c40dc43f95ea65ae7ca34384cb Mon Sep 17 00:00:00 2001 From: Sumer Johal Date: Wed, 19 Aug 2026 11:28:25 -0700 Subject: [PATCH 1/4] Fail when the mirror is more permissive than the system it mirrors hub.users.phone was relaxed to nullable and non-unique in migration/models.py, and phone was removed from the constraint guard, while ar2-hub still enforced NOT NULL UNIQUE on that column. The import consequently reported 0 rejected accounts where a real run against the hub would still have rejected 85 -- 8 with no phone, 77 sharing a number. The constraint was removed from the simulation, not from the system. That is the worse failure shape available: not an error, but a clean bill of health for an import that cannot happen. The existing drift guard compared column *names* only, so it passed throughout -- the column was still called phone. * test_mirrored_constraints_match_the_real_schema compares nullability and uniqueness against the real models, and fails only on permissiveness. A mirror that is stricter than reality predicts rejections that will not happen, which is a false alarm; a mirror that is looser predicts acceptance for rows the system rejects, which reads as success. * The docstring on test_hub_constraints_the_import_depends_on_are_still_there records why phone legitimately left both lists, and that the order it happened in was backwards. Relaxing the guard is the last step, never the first. * The real relaxation is ar2-hub#7, with the DDL create_all will not apply. Also, three reporting defects. Every number below is one a reader would act on, and each was wrong in the direction of looking worse or looking better than the run actually was. * 268 fields that resolved correctly on an identical content hash were counted as quarantined and then listed under DECISIONS REQUIRED. Every one had aliased to its canonical row and still resolves. A run that rejected 7 fields read as a run with 275 open problems. They now count as resolutions, split into exact-geometry and above-threshold matches, because certainty and judgement should not share a number. * Every imported point was counted as canonicalization-altered geometry, taking that figure from 3,597 to 18,184 and converting a real signal about boundary repair into a headcount of pins. A point has no boundary to repair. * The point count was not reported at all, though a registry that is half pins is a different thing to plan around than one that is all boundaries. It was the most consequential fact about the run and appeared nowhere in it. * "quarantined" is now "REJECTED, not imported", with the share of the total, and the decisions section says that a reason naming a geometry type is usually a missing code path rather than a policy question. Finally, the harness submodule moves to f7ac8b0. The pinned revision counted ruff's "All checks passed!" banner as one violation, so CI reported 1 where this tree has 36 -- the ratchet has been reading green while blind. The gating workflow also installs ruff unpinned, which the earlier pin missed because it only covered ci.yml; patches/pin-ruff-in-stomata-workflow.patch has that, since my token cannot write .github/workflows. Co-authored-by: Cursor --- .../testkit/dev_keys/expired_authority.sdjwt | 2 +- harness | 2 +- migration/pipeline.py | 24 ++++- migration/run.py | 19 +++- migration/tests/test_points_are_importable.py | 65 ++++++++++++++ migration/tests/test_schema_drift.py | 90 ++++++++++++++++++- patches/pin-ruff-in-stomata-workflow.patch | 13 +++ 7 files changed, 206 insertions(+), 9 deletions(-) create mode 100644 patches/pin-ruff-in-stomata-workflow.patch diff --git a/app/tests/testkit/dev_keys/expired_authority.sdjwt b/app/tests/testkit/dev_keys/expired_authority.sdjwt index d94e076..9dcbc21 100644 --- a/app/tests/testkit/dev_keys/expired_authority.sdjwt +++ b/app/tests/testkit/dev_keys/expired_authority.sdjwt @@ -1 +1 @@ -eyJhbGciOiJFZERTQSIsImtpZCI6InBhbmNha2UtdGVzdC0xIiwidHlwIjoidmMrc2Qtand0In0.eyJpc3MiOiJkaWQ6d2ViOnBhbmNha2UudGVzdCIsInN1YiI6ImF1dGhvcml0eUBkZW1vLmFnc3RhY2sub3JnIiwiaWF0IjoxNzg3MDczNzE3LCJleHAiOjE3ODcwNzAxMTcsImp0aSI6IjAxTTBBWTlTRFMyUDdURlZWVlZQMFI4SDZaIiwidmN0IjoiYWdzdGFjay5vcmcvY3JlZGVudGlhbHMvdHJhY2Vmb3J3YXJkLWF1dGhvcml0eS92MSIsInNjb3BlIjoiZGVtby1yZWNhbGwiLCJzdGF0dXMiOnsic3RhdHVzX2xpc3QiOnsidXJpIjoiaHR0cDovL2xvY2FsaG9zdDo4MTAwL2dyYW50cy9zdGF0dXMtbGlzdCIsImlkeCI6Mn19fQ.-KEnZ4LA6ivAX6xo95QGqTPNOh6lBED-U5KmzQT08YpB8ESrIFMhohS5-v3JdBXxGw6u57d7S5s-WeIr5lp3CQ~ \ No newline at end of file +eyJhbGciOiJFZERTQSIsImtpZCI6InBhbmNha2UtdGVzdC0xIiwidHlwIjoidmMrc2Qtand0In0.eyJpc3MiOiJkaWQ6d2ViOnBhbmNha2UudGVzdCIsInN1YiI6ImF1dGhvcml0eUBkZW1vLmFnc3RhY2sub3JnIiwiaWF0IjoxNzg3MTY0MDY1LCJleHAiOjE3ODcxNjA0NjUsImp0aSI6IjAxTTBETUYwMFZaQllOOTk0WEtOU1dBWVJLIiwidmN0IjoiYWdzdGFjay5vcmcvY3JlZGVudGlhbHMvdHJhY2Vmb3J3YXJkLWF1dGhvcml0eS92MSIsInNjb3BlIjoiZGVtby1yZWNhbGwiLCJzdGF0dXMiOnsic3RhdHVzX2xpc3QiOnsidXJpIjoiaHR0cDovL2xvY2FsaG9zdDo4MTAwL2dyYW50cy9zdGF0dXMtbGlzdCIsImlkeCI6Mn19fQ.xar8xn-LCyAZNe5C2vwVwnPidteA8wSWdG_nP4K_ZU7mXEQZoPuSx3gzYEJ96yo5xOc6eTHqKJFJrZeO92btCg~ \ No newline at end of file diff --git a/harness b/harness index bd7c4d2..f7ac8b0 160000 --- a/harness +++ b/harness @@ -1 +1 @@ -Subproject commit bd7c4d2b3c1a0d8d98a6722ae89de85323a7f316 +Subproject commit f7ac8b0ac9568fec2908e45342caae39af79bb58 diff --git a/migration/pipeline.py b/migration/pipeline.py index 43b3523..041d15d 100644 --- a/migration/pipeline.py +++ b/migration/pipeline.py @@ -73,6 +73,11 @@ class FieldReport: # plan around than one that is all fields, and the total hides that. imported_points: int = 0 resolved_same_as: int = 0 + # Of resolved_same_as, how many matched on an identical canonical geometry + # rather than on overlap above the threshold. Reported separately because the + # two mean different things: an exact content match is certainty, a + # threshold match is a judgement that the threshold could change. + resolved_by_content_hash: int = 0 resolved_child_of: int = 0 skipped_already_done: int = 0 quarantined: dict[str, list[str]] = field(default_factory=dict) @@ -163,7 +168,12 @@ def import_fields( f"{QUARANTINE_UNUSABLE}:{_reason_of(exc)}", legacy.v1_geo_id) continue - if _canonicalization_altered(legacy.wkt): + # Polygons only. A point has no boundary to repair, so asking whether + # canonicalization altered it is not a meaningful question -- and asking + # it anyway counted every one of the 14,587 imported pins as altered + # geometry, taking the reported figure from 3,597 to 18,184 and turning a + # real signal about boundary repair into a headcount of points. + if position is None and _canonicalization_altered(legacy.wkt): report.canonicalization_changed.append(legacy.v1_geo_id) # genuinely new (or nested, which still registers) @@ -172,7 +182,17 @@ def import_fields( # Identical canonical geometry already registered. content_hash is # UNIQUE in ar2, so this cannot be inserted -- alias to the existing # row instead of failing the run. - report.quarantine(QUARANTINE_DUPLICATE_CONTENT, legacy.v1_geo_id) + # + # Counted as a resolution, not a quarantine. It used to be filed under + # quarantined, which put 268 of a 28,282-field run into a bucket the + # report then listed as needing a policy decision -- when every one of + # them had in fact resolved correctly and its v1 identifier still + # resolves through the alias written on the next line. Nothing was + # lost and nothing was pending. The label said otherwise, and the + # label is what a reader acts on: it made a run that rejected 7 fields + # look like a run with 275 open problems. + report.resolved_same_as += 1 + report.resolved_by_content_hash += 1 if not dry_run: repo.upsert_alias(AliasRow(legacy.v1_geo_id, existing.geo_id, legacy.v1_kind, SAME_AS)) diff --git a/migration/run.py b/migration/run.py index 92fb06a..44c4ebc 100644 --- a/migration/run.py +++ b/migration/run.py @@ -108,11 +108,19 @@ def print_field_report(r, elapsed: float) -> None: _row("considered", f"{r.considered:,}") _row("imported new", f"{r.imported_new:,}") _row(" of which nested (child_of)", f"{r.resolved_child_of:,}", indent=4) + # A registry that is half pins is a different thing to plan around than one + # that is all boundaries, and the total hides which it is. This was the + # single most consequential fact about the first full run and it did not + # appear in the report at all. + _row(" of which points, not boundaries", f"{r.imported_points:,}", indent=4) _row("resolved same_as (merged)", f"{r.resolved_same_as:,}") + _row(" exact geometry match", f"{r.resolved_by_content_hash:,}", indent=4) + _row(" overlap above threshold", + f"{r.resolved_same_as - r.resolved_by_content_hash:,}", indent=4) _row("skipped, already imported", f"{r.skipped_already_done:,}") _row("UUID -> content-derived identity", f"{len(r.uuid_promoted):,}") - _row("canonicalization altered geometry", f"{len(r.canonicalization_changed):,}") - _row("quarantined", f"{r.quarantined_total:,}") + _row("canonicalization altered boundary", f"{len(r.canonicalization_changed):,}") + _row("REJECTED, not imported", f"{r.quarantined_total:,}") for reason, ids in sorted(r.quarantined.items()): _row(f" {reason}", f"{len(ids):,}", indent=4) if r.considered: @@ -278,8 +286,11 @@ def main(argv=None) -> int: f"did on collision: fail, or return the existing record?") if fields.quarantined_total: decisions.append( - f"{fields.quarantined_total:,} fields quarantined — policy needed " - f"(quarantine / flag / reject).") + f"{fields.quarantined_total:,} of {fields.considered:,} fields rejected " + f"({fields.quarantined_total / max(fields.considered, 1):.2%}) — read the " + f"reasons above before deciding anything. A reason naming a geometry type " + f"is usually a missing code path rather than bad data, and needs code, not " + f"a policy.") if profiles.merged_ownership: decisions.append( f"{len(profiles.merged_ownership)} merged GeoID(s) with multiple owners — " diff --git a/migration/tests/test_points_are_importable.py b/migration/tests/test_points_are_importable.py index e875934..c0128f3 100644 --- a/migration/tests/test_points_are_importable.py +++ b/migration/tests/test_points_are_importable.py @@ -160,3 +160,68 @@ def test_area_is_derived_so_the_inventory_bands_mean_something(): measured = g2.area_ha(square) assert measured == pytest.approx(100.0, rel=0.01), measured assert g2.area_ha(PIN_WKT) is None, "a pin has no area and must not report 0" + + +# --------------------------------------------------------------------------- +# what the report says happened +# --------------------------------------------------------------------------- + +def test_an_exact_duplicate_counts_as_resolved_not_rejected(): + """A duplicate that resolved correctly must not be reported as a problem. + + Identical canonical geometry aliases to the existing row, and the v1 + identifier keeps resolving. Nothing is lost and nothing is pending. It was + filed under quarantined, which put 268 of a 28,282-field run into a bucket + the report then listed as needing a policy decision -- making a run that + rejected 7 fields read as a run with 275 open problems. A reader acts on the + label, so the label has to be true. + """ + report = _run(FIELD_WKT, FIELD_WKT) + + assert report.quarantined_total == 0, ( + f"a resolved duplicate was reported as rejected: {report.quarantined}") + assert report.resolved_same_as == 1 + assert report.resolved_by_content_hash == 1 + + +def test_exact_matches_are_distinguishable_from_threshold_matches(): + """Certainty and judgement should not share a number. + + An exact content match cannot change. A threshold match is a decision that a + different threshold would make differently, and only the second is worth + revisiting when the threshold is questioned. + """ + report = _run(FIELD_WKT, FIELD_WKT) + + assert report.resolved_by_content_hash <= report.resolved_same_as + threshold_matches = report.resolved_same_as - report.resolved_by_content_hash + assert threshold_matches == 0, ( + f"identical geometry was scored as a threshold match: {threshold_matches}") + + +def test_a_point_is_not_counted_as_altered_geometry(): + """A pin has no boundary to repair, so the question does not apply to it. + + Asking anyway counted every one of 14,587 imported pins as altered geometry, + moving the reported figure from 3,597 to 18,184 and converting a real signal + about boundary repair into a headcount of points. + """ + report = _run(PIN_WKT, PIN_AS_RING, PIN_AS_SEGMENT) + + assert report.imported_points == 3 + assert len(report.canonicalization_changed) == 0, ( + f"points were counted as altered boundaries: " + f"{report.canonicalization_changed}") + + +def test_the_point_count_is_reported_at_all(): + """A registry that is half pins is a different thing to plan around. + + The total hides which it is, and on the first full run this was the single + most consequential fact and appeared nowhere in the output. + """ + report = _run(PIN_WKT, FIELD_WKT) + + assert hasattr(report, "imported_points") + assert report.imported_points == 1 + assert report.imported_new == 2 diff --git a/migration/tests/test_schema_drift.py b/migration/tests/test_schema_drift.py index 5e08906..3b9f2d4 100644 --- a/migration/tests/test_schema_drift.py +++ b/migration/tests/test_schema_drift.py @@ -93,6 +93,29 @@ def _mirror(base, table: str) -> set[str]: return {c.name for c in base.metadata.tables[table].columns} +def _mirror_constraints(base, table: str) -> dict[str, dict]: + """Nullability and uniqueness of the mirror, in the shape _parse_models returns. + + SQLAlchemy resolves uniqueness in two places -- unique=True on the column and + a UniqueConstraint on the table -- and a reader comparing only the first would + call a uniquely-constrained column non-unique. + """ + tbl = base.metadata.tables[table] + from sqlalchemy import UniqueConstraint + constrained = { + c.name + for con in tbl.constraints if isinstance(con, UniqueConstraint) + for c in con.columns + } + return { + col.name: { + "nullable": col.nullable, + "unique": bool(col.unique) or col.name in constrained, + } + for col in tbl.columns + } + + @pytest.mark.parametrize("which,table,mirror_base", [ ("ar2", "geo_ids", Base), ("hub", "users", Base), @@ -119,11 +142,76 @@ def test_mirrored_columns_match_the_real_schema(which, table, mirror_base): ) +@pytest.mark.parametrize("which,table,mirror_base", [ + ("ar2", "geo_ids", Base), + ("hub", "users", Base), + ("pancake", "users", PancakeBase), + ("pancake", "fieldlists", PancakeBase), + ("ar2", "listmember_edge", Base), +]) +def test_mirrored_constraints_match_the_real_schema(which, table, mirror_base): + """Matching column names is not matching the schema. + + The mirror exists so the import can predict what the real database will + accept. A mirror that is more permissive than the system predicts acceptance + for rows the system will reject, and the run reports a clean import that + cannot happen -- which is worse than an error, because it looks like success. + + This is not hypothetical. hub.users.phone was relaxed here to nullable and + non-unique while the hub still enforced NOT NULL UNIQUE. The import then + reported 0 rejected accounts where a real run would have rejected 85. The + column-name comparison above passed throughout, because the column was still + called phone. + + A mirror may legitimately be *stricter* than the real schema -- that predicts + rejections that will not happen, which is a false alarm rather than a false + clean bill. Only permissiveness is failed here. + """ + path = _locate(which) + if path is None: + pytest.skip(f"{which} checkout not present") + + real = _parse_models(path)[table] + ours = _mirror_constraints(mirror_base, table) + + too_permissive = [] + for col, real_meta in real.items(): + if col not in ours: + continue + # real nullable=None means the keyword was absent, i.e. SQLAlchemy's + # default of nullable=True, so there is nothing stricter to violate. + if real_meta["nullable"] is False and ours[col]["nullable"] is True: + too_permissive.append(f"{col}: real is NOT NULL, mirror allows NULL") + if real_meta["unique"] is True and ours[col]["unique"] is False: + too_permissive.append(f"{col}: real is UNIQUE, mirror is not") + + assert not too_permissive, ( + f"the {which}.{table} mirror is more permissive than the real schema, so " + f"the import will predict success for rows the system rejects:\n" + + "".join(f" - {p}\n" for p in too_permissive) + + f" source of truth: {path}\n" + " Either relax the real schema too, or restore the mirror." + ) + + def test_hub_constraints_the_import_depends_on_are_still_there(): - """The rejection rules in db_repo exist because of these four constraints. + """The rejection rules in db_repo exist because of these constraints. If any of them relaxes, accounts this import currently quarantines would import cleanly and the quarantine becomes a false positive. + + phone is deliberately absent from both lists. It used to be in both, and it + was removed because the hub genuinely relaxed it -- UNIQUE NOT NULL rejected + 85 of 345 real accounts, 8 with no number and 77 sharing one, and a shared + household or cooperative line is ordinary here (ar2-hub, users.phone). + + Worth recording how that landed, because the order was wrong and the order is + the whole point: the mirror and this guard were relaxed first, while the hub + still enforced both constraints. For that interval the import reported 0 + rejected accounts and a real run would still have rejected 85. Relaxing the + guard is the last step, never the first -- and + test_mirrored_constraints_match_the_real_schema now fails whenever the mirror + runs ahead of the system, which is the failure this comment describes. """ path = _locate("hub") if path is None: diff --git a/patches/pin-ruff-in-stomata-workflow.patch b/patches/pin-ruff-in-stomata-workflow.patch new file mode 100644 index 0000000..120a22b --- /dev/null +++ b/patches/pin-ruff-in-stomata-workflow.patch @@ -0,0 +1,13 @@ +diff --git a/.github/workflows/stomata.yml b/.github/workflows/stomata.yml +index 32c45a5..3db4f51 100644 +--- a/.github/workflows/stomata.yml ++++ b/.github/workflows/stomata.yml +@@ -69,7 +69,7 @@ jobs: + python -m pip install --upgrade pip + if [ -f requirements.txt ]; then pip install -r requirements.txt; fi + if [ -f migration/requirements.txt ]; then pip install -r migration/requirements.txt; fi +- pip install pytest ruff ++ pip install pytest ruff==0.9.6 + + - name: Apply schema upgrades + env: From 0691084ddbea155db2806cd54e263984104ec34a Mon Sep 17 00:00:00 2001 From: Sumer Johal Date: Wed, 19 Aug 2026 11:29:51 -0700 Subject: [PATCH 2/4] Review packet for the mirror-drift guard Co-authored-by: Cursor --- packets/mirror-drift.json | 72 +++++++ packets/mirror-drift.run.json | 360 ++++++++++++++++++++++++++++++++++ 2 files changed, 432 insertions(+) create mode 100644 packets/mirror-drift.json create mode 100644 packets/mirror-drift.run.json diff --git a/packets/mirror-drift.json b/packets/mirror-drift.json new file mode 100644 index 0000000..99945ca --- /dev/null +++ b/packets/mirror-drift.json @@ -0,0 +1,72 @@ +{ + "$stomata_artifact": "review_packet", + "schema_version": "1.0.0", + "title": "Fail when the mirror is more permissive than the system it mirrors", + "author": "Sumer", + "reviewer": "Rajat", + "produced_at": "2026-08-19T18:29:31+00:00", + "change": { + "summary": "hub.users.phone was relaxed to nullable and non-unique in the migration mirror, and phone was removed from the constraint guard, while ar2-hub still enforced NOT NULL UNIQUE. The import reported 0 rejected accounts where a real run would still have rejected 85. The existing drift guard compared column names only, so it passed. Adds a guard that compares nullability and uniqueness and fails on permissiveness; relaxes the real hub in ar2-hub#7 with the DDL create_all will not apply; corrects three reporting defects that made the run read as 275 open problems when it rejected 7 fields; and bumps the harness, whose pinned revision counted ruff's success banner as a violation so the lint ratchet reported 1 where this tree has 36.", + "commits": [ + "40369a6ab2e675c40dc43f95ea65ae7ca34384cb" + ], + "paths": [ + "migration/tests/test_schema_drift.py", + "migration/pipeline.py", + "migration/run.py", + "migration/tests/test_points_are_importable.py", + "harness", + "patches/pin-ruff-in-stomata-workflow.patch" + ] + }, + "evidence": [ + { + "type": "harness_state", + "claim": "Harness run at 2026-08-19T18:29:10+00:00: ar2 119/119, migration 105/105; all checks passed.", + "pointer": { + "path": "packets/mirror-drift.run.json", + "command": "harness/bin/stomata run --full", + "commit": "40369a6ab2e675c40dc43f95ea65ae7ca34384cb" + }, + "captured_at": "2026-08-19T18:29:10+00:00" + }, + { + "type": "code_reference", + "claim": "The real hub at ar2-hub main has phone = Column(String(20), unique=True, index=True, nullable=False) at user_models.py:21, against unique=False, nullable=True in migration/models.py:177. The new guard fails against the former and passes against ar2-hub#7.", + "pointer": { + "path": "migration/tests/test_schema_drift.py", + "command": "MIGRATION_HUB_MODELS=/user_models.py pytest migration/tests/test_schema_drift.py -k constraints_match" + }, + "captured_at": "2026-08-19T18:29:31+00:00" + } + ], + "gaps": [ + { + "kind": "not_covered", + "description": "Not rerun against live AR 1.0 -- no access. The reporting changes alter no import decision, so the field outcomes stand; what changes is how they are described. A rerun is still needed to print the point count and the corrected boundary-repair figure." + }, + { + "kind": "deferred", + "description": "The 85 accounts remain blocked until ar2-hub#7 merges AND its DDL runs against the deployed hub. Merging the model alone leaves create_all unable to alter the column, so the database keeps rejecting them.", + "owner": "Rajat", + "target": "before the next live import run" + }, + { + "kind": "not_covered", + "description": "patches/pin-ruff-in-stomata-workflow.patch is not applied -- my token cannot write .github/workflows. Until it is, the gating job installs ruff unpinned and the ratchet's measuring stick can move between runs." + }, + { + "kind": "unknown", + "description": "1,117 association join failures are unexplained here beyond their correlation with 1,115 AR 1.0 orphan profile references -- profiles citing GeoIDs absent from the source. If that is the whole explanation it is a source data defect and not ours, but the two numbers differ by 2 and nobody has reconciled them." + } + ], + "decision": { + "question": "Merge ar2-hub#7 and run its DDL, apply the stomata workflow patch, then rerun the import and report the point count, the corrected boundary-repair figure, and whether 85 accounts now admit?", + "options": [ + "merge both, run the DDL, then rerun", + "review the guard first" + ], + "recommendation": "merge both, run the DDL, then rerun", + "requested_of": "Rajat" + } +} diff --git a/packets/mirror-drift.run.json b/packets/mirror-drift.run.json new file mode 100644 index 0000000..5a72c24 --- /dev/null +++ b/packets/mirror-drift.run.json @@ -0,0 +1,360 @@ +{ + "$stomata_artifact": "state", + "stomata_version": "1.0.0", + "started_at": "2026-08-19T18:28:33+00:00", + "finished_at": "2026-08-19T18:29:10+00:00", + "git": { + "commit": "40369a6ab2e675c40dc43f95ea65ae7ca34384cb", + "branch": "sumer/drift-and-labels", + "dirty": false, + "dirty_files": [] + }, + "invocation": "harness/bin/stomata run --full", + "suites": [ + { + "name": "ar2", + "command": "/tmp/ar2venv/bin/python -m pytest app/tests -q -rs", + "returncode": 0, + "passed": 119, + "failed": 0, + "skipped": 0, + "errors": 0, + "total": 119, + "test_ids": [ + "app/tests/test_api.py::test_eudr_export", + "app/tests/test_api.py::test_expired_grant", + "app/tests/test_api.py::test_fetch_field_centroid", + "app/tests/test_api.py::test_fetch_field_wkt", + "app/tests/test_api.py::test_fetch_fields_for_a_point", + "app/tests/test_api.py::test_identity_resolution_nested", + "app/tests/test_api.py::test_identity_resolution_same_as", + "app/tests/test_api.py::test_real_rs256_auth", + "app/tests/test_api.py::test_register_and_fetch_field", + "app/tests/test_api.py::test_register_duplicate_point", + "app/tests/test_api.py::test_revoked_grant", + "app/tests/test_api.py::test_tampered_grant", + "app/tests/test_api.py::test_valid_grant", + "app/tests/test_api.py::test_valid_grant_eudr_export", + "app/tests/test_api.py::test_valid_grant_fetch_centroid", + "app/tests/test_api.py::test_valid_grant_fetch_wkt", + "app/tests/test_api.py::test_wrong_geoid_grant", + "app/tests/test_fsma204_traceability.py::test_a_corrupt_edge_cannot_hang_the_trace", + "app/tests/test_fsma204_traceability.py::test_a_depth_limit_is_flagged_rather_than_silently_truncating", + "app/tests/test_fsma204_traceability.py::test_a_forged_membership_claim_fails_recomputation", + "app/tests/test_fsma204_traceability.py::test_a_lot_citing_a_region_admits_its_field_set_is_incomplete", + "app/tests/test_fsma204_traceability.py::test_a_partial_trace_from_the_middle_of_the_chain", + "app/tests/test_fsma204_traceability.py::test_a_partner_grant_gets_lot_codes_but_no_identities", + "app/tests/test_fsma204_traceability.py::test_a_regulator_credential_escalates_to_identities_and_is_audited", + "app/tests/test_fsma204_traceability.py::test_an_unknown_lot_is_a_404", + "app/tests/test_fsma204_traceability.py::test_an_unrecorded_hop_location_is_reported_not_hidden", + "app/tests/test_fsma204_traceability.py::test_contamination_in_one_field_reaches_every_downstream_lot", + "app/tests/test_fsma204_traceability.py::test_each_hop_carries_the_location_where_that_lot_was_created", + "app/tests/test_fsma204_traceability.py::test_each_hop_names_the_lots_that_fed_it", + "app/tests/test_fsma204_traceability.py::test_facilities_and_fields_share_one_namespace", + "app/tests/test_fsma204_traceability.py::test_the_edge_list_reconstructs_the_whole_graph", + "app/tests/test_fsma204_traceability.py::test_the_hop_order_is_the_supply_chain_order", + "app/tests/test_fsma204_traceability.py::test_the_lot_code_is_verifiable_without_a_glossary", + "app/tests/test_fsma204_traceability.py::test_the_recall_scope_stops_where_the_material_stops", + "app/tests/test_fsma204_traceability.py::test_the_whole_trace_is_one_request", + "app/tests/test_fsma204_traceability.py::test_this_scenario_cannot_leak_into_other_suites", + "app/tests/test_fsma204_traceability.py::test_traceback_reaches_every_field_from_the_retail_case", + "app/tests/test_fsma204_traceability.py::test_traceback_refuses_a_revoked_authority_credential", + "app/tests/test_fsma204_traceability.py::test_traceback_refuses_an_untrusted_issuer", + "app/tests/test_fsma204_traceability.py::test_traceback_refuses_without_a_credential", + "app/tests/test_fsma204_traceability.py::test_traceback_reports_one_hop_per_lot", + "app/tests/test_fsma204_traceability.py::test_traceforward_and_traceback_agree", + "app/tests/test_fsma204_traceability.py::test_walkthrough_for_ifpa", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[antimeridian]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[ccw]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[clockwise]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[duplicate_vertices]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[equator]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[high_precision]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[hole]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[huge_2500ha]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[l_shape]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[near_pole]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[overlapping_edges]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[prime_meridian]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[self_intersecting_bow_tie]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[southern_western]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[square_1ha]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[square_5ha]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[star]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[tiny_0_01ha]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[tiny_sliver]", + "app/tests/test_geoid_v2.py::test_generate_geo_id_v2[triangle]", + "app/tests/test_geoid_v2_iou.py::test_iou_fidelity", + "app/tests/test_geoid_v2_iou.py::test_measure_threshold_bias", + "app/tests/test_geoid_v2_live.py::test_a_field_with_no_recorded_area_still_resolves[0.0]", + "app/tests/test_geoid_v2_live.py::test_a_field_with_no_recorded_area_still_resolves[4.0]", + "app/tests/test_geoid_v2_live.py::test_a_field_with_no_recorded_area_still_resolves[None]", + "app/tests/test_geoid_v2_live.py::test_a_hole_is_excluded_from_the_cover", + "app/tests/test_geoid_v2_live.py::test_a_list_registered_against_a_v1_geoid_is_still_traceable", + "app/tests/test_geoid_v2_live.py::test_a_self_intersecting_polygon_keeps_both_lobes", + "app/tests/test_geoid_v2_live.py::test_a_v1_identifier_reaches_the_field_it_became", + "app/tests/test_geoid_v2_live.py::test_child_of_is_not_treated_as_identity", + "app/tests/test_geoid_v2_live.py::test_containment_detects_nesting_where_iou_does_not", + "app/tests/test_geoid_v2_live.py::test_every_part_of_a_multipolygon_contributes", + "app/tests/test_geoid_v2_live.py::test_exact_iou_sees_overlap_that_token_sets_cannot", + "app/tests/test_geoid_v2_live.py::test_identical_covers_score_exactly_one", + "app/tests/test_geoid_v2_live.py::test_leaf_area_is_exact_across_mixed_levels", + "app/tests/test_geoid_v2_live.py::test_registration_never_issues_a_uuid", + "app/tests/test_geoid_v2_live.py::test_registration_returns_the_v2_identifier", + "app/tests/test_geoid_v2_live.py::test_resolution_still_refuses_a_genuinely_different_field", + "app/tests/test_geoid_v2_live.py::test_reverse_lookup_finds_a_list_through_a_v1_identifier", + "app/tests/test_geoid_v2_live.py::test_the_equivalence_set_is_symmetric", + "app/tests/test_geoid_v2_live.py::test_the_stored_cover_is_the_canonical_one", + "app/tests/test_geoid_v2_live.py::test_the_two_implementations_agree", + "app/tests/test_geoid_v2_live.py::test_two_nearby_fields_get_different_identifiers", + "app/tests/test_geoid_v2_live.py::test_uniform_level_covers_match_the_old_arithmetic_exactly", + "app/tests/test_geoid_v2_live.py::test_unusable_geometry_is_refused_with_422", + "app/tests/test_geoid_v2_live.py::test_unusable_geometry_raises_rather_than_inventing_an_id[LINESTRING(0 0, 1 1)]", + "app/tests/test_geoid_v2_live.py::test_unusable_geometry_raises_rather_than_inventing_an_id[POINT(0 0)]", + "app/tests/test_geoid_v2_live.py::test_unusable_geometry_raises_rather_than_inventing_an_id[POLYGON EMPTY]", + "app/tests/test_geoid_v2_live.py::test_unusable_geometry_raises_rather_than_inventing_an_id[POLYGON((0 0, 0 0, 0 0, 0 0))]", + "app/tests/test_geoid_v2_properties.py::test_collision_fixed", + "app/tests/test_geoid_v2_properties.py::test_determinism", + "app/tests/test_geoid_v2_properties.py::test_jitter_convergence", + "app/tests/test_geoid_v2_properties.py::test_nesting", + "app/tests/test_geoid_v2_properties.py::test_shape_sensitivity", + "app/tests/test_traceforward.py::test_audit_failure_503", + "app/tests/test_traceforward.py::test_authority_round_trip", + "app/tests/test_traceforward.py::test_containment_ancestor_probe", + "app/tests/test_traceforward.py::test_expanding_composition", + "app/tests/test_traceforward.py::test_four_hop_reverse_lookup", + "app/tests/test_traceforward.py::test_gate_a_missing_capability_403", + "app/tests/test_traceforward.py::test_gate_b_authority_global_scope", + "app/tests/test_traceforward.py::test_gate_b_authority_missing_scope_400", + "app/tests/test_traceforward.py::test_gate_b_authority_owns_nothing_gets_identities", + "app/tests/test_traceforward.py::test_gate_b_authority_untrusted_key_401", + "app/tests/test_traceforward.py::test_gate_b_grant_for_different_seed_403", + "app/tests/test_traceforward.py::test_gate_b_no_grant_no_authority_403", + "app/tests/test_traceforward.py::test_gate_b_owner_path_passes_without_identities", + "app/tests/test_traceforward.py::test_invalid_authority_credentials_401[expired_authority]", + "app/tests/test_traceforward.py::test_invalid_authority_credentials_401[outofscope_authority]", + "app/tests/test_traceforward.py::test_invalid_authority_credentials_401[revoked_authority]", + "app/tests/test_traceforward.py::test_manufactured_cycle_terminates", + "app/tests/test_traceforward.py::test_pure_geoid_root_regression", + "app/tests/test_traceforward.py::test_region_built_from_nested_list", + "app/tests/test_traceforward.py::test_region_normalization_determinism_and_wkt", + "app/tests/test_traceforward.py::test_three_list_reverse_lookup" + ], + "skip_reasons": [], + "collect_ok": true + }, + { + "name": "migration", + "command": "/tmp/ar2venv/bin/python -m pytest migration/tests -q -rs", + "returncode": 0, + "passed": 105, + "failed": 0, + "skipped": 0, + "errors": 0, + "total": 105, + "test_ids": [ + "migration/tests/test_db_repo.py::test_alias_is_persisted_and_resolves", + "migration/tests/test_db_repo.py::test_alias_re_upsert_is_idempotent_but_remap_is_refused", + "migration/tests/test_db_repo.py::test_blocking_index_finds_candidates", + "migration/tests/test_db_repo.py::test_both_repositories_produce_identical_results", + "migration/tests/test_db_repo.py::test_checkpoint_survives_a_reconnect", + "migration/tests/test_db_repo.py::test_duplicate_content_hash_is_caught", + "migration/tests/test_db_repo.py::test_duplicate_email_and_phone_are_both_refused", + "migration/tests/test_db_repo.py::test_duplicate_geo_id_is_a_named_constraint_violation", + "migration/tests/test_db_repo.py::test_fieldlist_re_creation_is_idempotent", + "migration/tests/test_db_repo.py::test_fieldlist_requires_an_existing_owner", + "migration/tests/test_db_repo.py::test_geoid_round_trips_through_the_database", + "migration/tests/test_db_repo.py::test_hub_account_creates_a_pancake_mirror", + "migration/tests/test_db_repo.py::test_hub_rejects_what_the_real_schema_rejects[-client_id length <= 50]", + "migration/tests/test_db_repo.py::test_hub_rejects_what_the_real_schema_rejects[-email NOT NULL]", + "migration/tests/test_db_repo.py::test_hub_rejects_what_the_real_schema_rejects[-first_name/last_name NOT NULL]", + "migration/tests/test_db_repo.py::test_hub_rejects_what_the_real_schema_rejects[-phone NOT NULL]", + "migration/tests/test_db_repo.py::test_import_is_resumable_against_a_real_database", + "migration/tests/test_db_repo.py::test_migrated_accounts_are_inactive_with_no_usable_password", + "migration/tests/test_db_repo.py::test_multi_owner_detection_works_across_the_join", + "migration/tests/test_db_repo.py::test_parent_edge_is_recorded_once_and_never_self_referential", + "migration/tests/test_limit_invalidates_joins.py::test_limit_inflates_join_failures_and_collapses_fieldlists", + "migration/tests/test_limit_invalidates_joins.py::test_sample_preserves_the_true_join_failure_count", + "migration/tests/test_limit_invalidates_joins.py::test_the_collapse_is_monotonic_in_the_limit", + "migration/tests/test_limit_invalidates_joins.py::test_the_runner_does_not_warn_without_limit", + "migration/tests/test_limit_invalidates_joins.py::test_the_runner_warns_when_limit_is_set", + "migration/tests/test_limit_invalidates_joins.py::test_the_warning_names_the_three_unreadable_figures", + "migration/tests/test_limit_invalidates_joins.py::test_unlimited_run_has_only_the_deliberate_join_failures", + "migration/tests/test_pipeline.py::test_a_pin_written_as_a_ring_imports", + "migration/tests/test_pipeline.py::test_accounts_and_lists_created", + "migration/tests/test_pipeline.py::test_dry_run_writes_nothing", + "migration/tests/test_pipeline.py::test_exact_duplicate_geometry_aliases_rather_than_failing", + "migration/tests/test_pipeline.py::test_holes_and_multipart_get_distinct_identities", + "migration/tests/test_pipeline.py::test_hub_constraints_reject_rather_than_corrupt", + "migration/tests/test_pipeline.py::test_inventory_is_self_consistent", + "migration/tests/test_pipeline.py::test_listid_is_merkle_root_of_sorted_members", + "migration/tests/test_pipeline.py::test_mergedfields_produce_shared_ownership_and_it_is_surfaced", + "migration/tests/test_pipeline.py::test_nested_plot_keeps_its_own_identity", + "migration/tests/test_pipeline.py::test_orphan_profile_reference_is_reported_not_silent", + "migration/tests/test_pipeline.py::test_profiles_refuse_to_run_beforefields", + "migration/tests/test_pipeline.py::test_resume_after_interruption_reaches_same_state", + "migration/tests/test_pipeline.py::test_second_run_is_a_no_op", + "migration/tests/test_pipeline.py::test_threshold_changes_merge_behaviour", + "migration/tests/test_pipeline.py::test_unidentifiable_geometry_quarantined_not_invented", + "migration/tests/test_pipeline.py::test_user_field_set_matches_source", + "migration/tests/test_pipeline.py::test_uuid_fields_gain_content_derived_identity", + "migration/tests/test_pipeline.py::test_v1_identifiers_resolve_forever", + "migration/tests/test_pipeline.py::testfields_imported_and_aliased", + "migration/tests/test_points_are_importable.py::test_a_pin_and_a_field_are_not_confused", + "migration/tests/test_points_are_importable.py::test_a_pin_imports_however_ar1_wrote_it_down[LINESTRING(77.5 12.9, 77.5 12.9)]", + "migration/tests/test_points_are_importable.py::test_a_pin_imports_however_ar1_wrote_it_down[POINT(77.5 12.9)]", + "migration/tests/test_points_are_importable.py::test_a_pin_imports_however_ar1_wrote_it_down[POLYGON((77.5 12.9, 77.5 12.9, 77.5 12.9, 77.5 12.9))]", + "migration/tests/test_points_are_importable.py::test_a_pin_imports_rather_than_quarantining", + "migration/tests/test_points_are_importable.py::test_a_point_is_not_counted_as_altered_geometry", + "migration/tests/test_points_are_importable.py::test_a_quarantine_reason_is_specific_enough_to_act_on", + "migration/tests/test_points_are_importable.py::test_an_exact_duplicate_counts_as_resolved_not_rejected", + "migration/tests/test_points_are_importable.py::test_area_is_derived_so_the_inventory_bands_mean_something", + "migration/tests/test_points_are_importable.py::test_every_spelling_of_one_pin_lands_on_one_identifier", + "migration/tests/test_points_are_importable.py::test_exact_matches_are_distinguishable_from_threshold_matches", + "migration/tests/test_points_are_importable.py::test_genuinely_broken_geometry_still_quarantines", + "migration/tests/test_points_are_importable.py::test_the_point_count_is_reported_at_all", + "migration/tests/test_points_are_importable.py::test_the_two_implementations_agree_on_points", + "migration/tests/test_primitive.py::test_blocking_key_is_coarse_and_shared", + "migration/tests/test_primitive.py::test_collision_fixed_neighbouring_fields_differ", + "migration/tests/test_primitive.py::test_deterministic", + "migration/tests/test_primitive.py::test_duplicate_vertices_irrelevant", + "migration/tests/test_primitive.py::test_holes_are_not_covered", + "migration/tests/test_primitive.py::test_iou_invariants", + "migration/tests/test_primitive.py::test_iou_sees_ancestor_descendant_overlap", + "migration/tests/test_primitive.py::test_iou_tracks_geometry", + "migration/tests/test_primitive.py::test_multipolygon_uses_every_part", + "migration/tests/test_primitive.py::test_nesting_is_child_not_same", + "migration/tests/test_primitive.py::test_normalized_no_mergeable_siblings", + "migration/tests/test_primitive.py::test_part_order_does_not_matter", + "migration/tests/test_primitive.py::test_resolve_outcomes", + "migration/tests/test_primitive.py::test_shape_not_bounding_box", + "migration/tests/test_primitive.py::test_sorted_tokens", + "migration/tests/test_primitive.py::test_unusable_geometry_refused", + "migration/tests/test_primitive.py::test_winding_order_irrelevant", + "migration/tests/test_sample.py::test_a_generous_budget_covers_the_whole_source", + "migration/tests/test_sample.py::test_area_bands_are_all_present", + "migration/tests/test_sample.py::test_collision_count_comes_from_the_l13_column", + "migration/tests/test_sample.py::test_every_hazard_stratum_is_represented", + "migration/tests/test_sample.py::test_hazards_survive_a_budget_far_smaller_than_the_source", + "migration/tests/test_sample.py::test_justification_names_any_stratum_with_no_members", + "migration/tests/test_sample.py::test_multi_owner_cluster_is_found_from_the_l13_key_alone", + "migration/tests/test_sample.py::test_no_filler_is_added_once_the_budget_has_cut_a_stratum", + "migration/tests/test_sample.py::test_orphan_refs_are_reported_but_never_selected_as_fields", + "migration/tests/test_sample.py::test_profiles_follow_their_fields", + "migration/tests/test_sample.py::test_same_owner_l13_cluster_is_separated_from_the_multi_owner_one", + "migration/tests/test_sample.py::test_sample_is_deterministic", + "migration/tests/test_sample.py::test_sampled_source_restricts_iteration_but_not_inventory", + "migration/tests/test_sample.py::test_the_profile_holding_an_orphan_ref_is_still_in_the_sample", + "migration/tests/test_schema_drift.py::test_ar2_uniqueness_the_import_depends_on_is_still_there", + "migration/tests/test_schema_drift.py::test_hub_constraints_the_import_depends_on_are_still_there", + "migration/tests/test_schema_drift.py::test_mirrored_columns_match_the_real_schema[ar2-geo_ids-Base]", + "migration/tests/test_schema_drift.py::test_mirrored_columns_match_the_real_schema[ar2-listmember_edge-Base]", + "migration/tests/test_schema_drift.py::test_mirrored_columns_match_the_real_schema[hub-users-Base]", + "migration/tests/test_schema_drift.py::test_mirrored_columns_match_the_real_schema[pancake-fieldlists-PancakeBase]", + "migration/tests/test_schema_drift.py::test_mirrored_columns_match_the_real_schema[pancake-users-PancakeBase]", + "migration/tests/test_schema_drift.py::test_mirrored_constraints_match_the_real_schema[ar2-geo_ids-Base]", + "migration/tests/test_schema_drift.py::test_mirrored_constraints_match_the_real_schema[ar2-listmember_edge-Base]", + "migration/tests/test_schema_drift.py::test_mirrored_constraints_match_the_real_schema[hub-users-Base]", + "migration/tests/test_schema_drift.py::test_mirrored_constraints_match_the_real_schema[pancake-fieldlists-PancakeBase]", + "migration/tests/test_schema_drift.py::test_mirrored_constraints_match_the_real_schema[pancake-users-PancakeBase]", + "migration/tests/test_schema_drift.py::test_regime_alias_table_is_ours_and_not_ar2s" + ], + "skip_reasons": [], + "collect_ok": true + } + ], + "lint": { + "command": "/tmp/ar2venv/bin/python -m ruff check app migration --output-format concise", + "returncode": 1, + "violations": 36, + "sample": [ + "\u001b[1mapp/auth.py\u001b[0m\u001b[36m:\u001b[0m32\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/database.py\u001b[0m\u001b[36m:\u001b[0m21\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/main.py\u001b[0m\u001b[36m:\u001b[0m16\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/main.py\u001b[0m\u001b[36m:\u001b[0m18\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/main.py\u001b[0m\u001b[36m:\u001b[0m19\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/models/geo_id_model.py\u001b[0m\u001b[36m:\u001b[0m16\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/models/geo_id_model.py\u001b[0m\u001b[36m:\u001b[0m27\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/models/geo_id_model.py\u001b[0m\u001b[36m:\u001b[0m29\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/models/geo_id_model.py\u001b[0m\u001b[36m:\u001b[0m132\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/routers/traceforward.py\u001b[0m\u001b[36m:\u001b[0m190\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/routers/traceforward.py\u001b[0m\u001b[36m:\u001b[0m443\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/tests/test_api.py\u001b[0m\u001b[36m:\u001b[0m19\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/tests/test_api.py\u001b[0m\u001b[36m:\u001b[0m22\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/tests/test_api.py\u001b[0m\u001b[36m:\u001b[0m23\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/tests/test_api.py\u001b[0m\u001b[36m:\u001b[0m24\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/tests/test_api.py\u001b[0m\u001b[36m:\u001b[0m25\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/tests/test_api.py\u001b[0m\u001b[36m:\u001b[0m26\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/tests/test_fsma204_traceability.py\u001b[0m\u001b[36m:\u001b[0m110\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/tests/test_fsma204_traceability.py\u001b[0m\u001b[36m:\u001b[0m111\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file", + "\u001b[1mapp/tests/test_fsma204_traceability.py\u001b[0m\u001b[36m:\u001b[0m112\u001b[36m:\u001b[0m1\u001b[36m:\u001b[0m \u001b[1;31mE402\u001b[0m Module level import not at top of file" + ], + "tool_version": null + }, + "checks": [ + { + "letter": "A", + "name": "suites ran", + "status": "pass", + "detail": "2 suite(s), all collected and executed", + "findings": [] + }, + { + "letter": "B", + "name": "green", + "status": "pass", + "detail": "224 tests passed, none failed", + "findings": [] + }, + { + "letter": "C", + "name": "no regression", + "status": "pass", + "detail": "no tests lost; 20 new test(s) since baseline", + "findings": [] + }, + { + "letter": "D", + "name": "no silent skips", + "status": "pass", + "detail": "zero skips", + "findings": [] + }, + { + "letter": "E", + "name": "reachability", + "status": "pass", + "detail": "5 claim(s), all reachable from production code", + "findings": [] + }, + { + "letter": "F", + "name": "retired defects", + "status": "pass", + "detail": "3 pattern(s), none present", + "findings": [] + }, + { + "letter": "G", + "name": "lint ratchet", + "status": "pass", + "detail": "36 violation(s), unchanged", + "findings": [] + }, + { + "letter": "H", + "name": "commit binding", + "status": "pass", + "detail": "40369a6ab on sumer/drift-and-labels, tree clean", + "findings": [] + }, + { + "letter": "I", + "name": "mutations", + "status": "pass", + "detail": "5 mutation(s) applied, every one detected", + "findings": [] + } + ], + "outcome": "pass" +} From ceff9444e3ab90b054a5c99898901cff6cd96360 Mon Sep 17 00:00:00 2001 From: Sumer Johal Date: Wed, 19 Aug 2026 11:53:15 -0700 Subject: [PATCH 3/4] Seed the lessons ledger from a month of review, and assert the report adds up Seeded .stomata/lessons.json from the review record of 2026-07-24 to 2026-08-19. Seven recurrence keys, every occurrence traceable to a dated document. Nothing speculative: a lesson with no incident behind it becomes noise, and noise is what discredits the checks that were earned. What the corpus shows, which was not visible one review at a time: 7x a gate reported green while blind 6x one fact in two places, one of them updated 5x a label or count that described a run that did not happen 4x a number quoted that was not in the evidence 3x a failure mode written down as prose instead of a check 2x a guard narrowed to let a change through The last two are the interesting ones. The prose entry is about this review process rather than the code: three times a correct warning was written in a document and did not bind, including one that predicted the guard-weakening incident five days before it happened. That is the argument for checks J and K existing in code. test_report_arithmetic.py mechanizes the label lesson. Every individual number in the import report was computed correctly and the totals still misdescribed the run: 268 successful aliases reported as quarantined, canonicalisation overstated 5x by points on another code path. What was missing was an assertion about the relationship between categories, which is checkable. It found a live one immediately: resolved_child_of was presented as a peer of imported_new when it is a subset, so 74 fields summed to 75 and read as one field counted twice. The display was already correctly nested; the invariant was unstated, and is now stated where it can fail. Two guards declared. Verified by removing an assertion from the drift guard and confirming J fails, quoting the incident, then restoring it. 11/11 checks pass with mutations, 229 tests, 6/6 recurring lessons mechanized. Co-authored-by: Cursor --- .stomata/baseline.json | 53 ++++++++-- .stomata/lessons.json | 118 ++++++++++++++++++++++ harness | 2 +- migration/pipeline.py | 5 + migration/tests/test_report_arithmetic.py | 105 +++++++++++++++++++ stomata.json | 30 ++++-- 6 files changed, 292 insertions(+), 21 deletions(-) create mode 100644 .stomata/lessons.json create mode 100644 migration/tests/test_report_arithmetic.py diff --git a/.stomata/baseline.json b/.stomata/baseline.json index d3e01ad..e0520aa 100644 --- a/.stomata/baseline.json +++ b/.stomata/baseline.json @@ -1,8 +1,8 @@ { "$stomata_artifact": "baseline", - "captured_at": "2026-08-14T18:11:16+00:00", - "commit": "4ecdd0ffbb52e181b6a91ef341a5d7588f970944", - "dirty_when_captured": false, + "captured_at": "2026-08-19T18:50:48+00:00", + "commit": "0691084ddbea155db2806cd54e263984104ec34a", + "dirty_when_captured": true, "suites": { "ar2": { "passed": 119, @@ -129,8 +129,8 @@ ] }, "migration": { - "passed": 81, - "total": 85, + "passed": 110, + "total": 110, "test_ids": [ "migration/tests/test_db_repo.py::test_alias_is_persisted_and_resolves", "migration/tests/test_db_repo.py::test_alias_re_upsert_is_idempotent_but_remap_is_refused", @@ -159,6 +159,7 @@ "migration/tests/test_limit_invalidates_joins.py::test_the_runner_warns_when_limit_is_set", "migration/tests/test_limit_invalidates_joins.py::test_the_warning_names_the_three_unreadable_figures", "migration/tests/test_limit_invalidates_joins.py::test_unlimited_run_has_only_the_deliberate_join_failures", + "migration/tests/test_pipeline.py::test_a_pin_written_as_a_ring_imports", "migration/tests/test_pipeline.py::test_accounts_and_lists_created", "migration/tests/test_pipeline.py::test_dry_run_writes_nothing", "migration/tests/test_pipeline.py::test_exact_duplicate_geometry_aliases_rather_than_failing", @@ -178,6 +179,20 @@ "migration/tests/test_pipeline.py::test_uuid_fields_gain_content_derived_identity", "migration/tests/test_pipeline.py::test_v1_identifiers_resolve_forever", "migration/tests/test_pipeline.py::testfields_imported_and_aliased", + "migration/tests/test_points_are_importable.py::test_a_pin_and_a_field_are_not_confused", + "migration/tests/test_points_are_importable.py::test_a_pin_imports_however_ar1_wrote_it_down[LINESTRING(77.5 12.9, 77.5 12.9)]", + "migration/tests/test_points_are_importable.py::test_a_pin_imports_however_ar1_wrote_it_down[POINT(77.5 12.9)]", + "migration/tests/test_points_are_importable.py::test_a_pin_imports_however_ar1_wrote_it_down[POLYGON((77.5 12.9, 77.5 12.9, 77.5 12.9, 77.5 12.9))]", + "migration/tests/test_points_are_importable.py::test_a_pin_imports_rather_than_quarantining", + "migration/tests/test_points_are_importable.py::test_a_point_is_not_counted_as_altered_geometry", + "migration/tests/test_points_are_importable.py::test_a_quarantine_reason_is_specific_enough_to_act_on", + "migration/tests/test_points_are_importable.py::test_an_exact_duplicate_counts_as_resolved_not_rejected", + "migration/tests/test_points_are_importable.py::test_area_is_derived_so_the_inventory_bands_mean_something", + "migration/tests/test_points_are_importable.py::test_every_spelling_of_one_pin_lands_on_one_identifier", + "migration/tests/test_points_are_importable.py::test_exact_matches_are_distinguishable_from_threshold_matches", + "migration/tests/test_points_are_importable.py::test_genuinely_broken_geometry_still_quarantines", + "migration/tests/test_points_are_importable.py::test_the_point_count_is_reported_at_all", + "migration/tests/test_points_are_importable.py::test_the_two_implementations_agree_on_points", "migration/tests/test_primitive.py::test_blocking_key_is_coarse_and_shared", "migration/tests/test_primitive.py::test_collision_fixed_neighbouring_fields_differ", "migration/tests/test_primitive.py::test_deterministic", @@ -195,6 +210,11 @@ "migration/tests/test_primitive.py::test_sorted_tokens", "migration/tests/test_primitive.py::test_unusable_geometry_refused", "migration/tests/test_primitive.py::test_winding_order_irrelevant", + "migration/tests/test_report_arithmetic.py::test_a_rejected_field_is_never_also_a_resolved_one", + "migration/tests/test_report_arithmetic.py::test_canonicalisation_is_only_flagged_when_geometry_actually_changed", + "migration/tests/test_report_arithmetic.py::test_every_field_considered_lands_in_exactly_one_bucket", + "migration/tests/test_report_arithmetic.py::test_subcategories_never_exceed_the_category_they_subdivide", + "migration/tests/test_report_arithmetic.py::test_the_arithmetic_holds_under_a_limit", "migration/tests/test_sample.py::test_a_generous_budget_covers_the_whole_source", "migration/tests/test_sample.py::test_area_bands_are_all_present", "migration/tests/test_sample.py::test_collision_count_comes_from_the_l13_column", @@ -216,14 +236,27 @@ "migration/tests/test_schema_drift.py::test_mirrored_columns_match_the_real_schema[hub-users-Base]", "migration/tests/test_schema_drift.py::test_mirrored_columns_match_the_real_schema[pancake-fieldlists-PancakeBase]", "migration/tests/test_schema_drift.py::test_mirrored_columns_match_the_real_schema[pancake-users-PancakeBase]", + "migration/tests/test_schema_drift.py::test_mirrored_constraints_match_the_real_schema[ar2-geo_ids-Base]", + "migration/tests/test_schema_drift.py::test_mirrored_constraints_match_the_real_schema[ar2-listmember_edge-Base]", + "migration/tests/test_schema_drift.py::test_mirrored_constraints_match_the_real_schema[hub-users-Base]", + "migration/tests/test_schema_drift.py::test_mirrored_constraints_match_the_real_schema[pancake-fieldlists-PancakeBase]", + "migration/tests/test_schema_drift.py::test_mirrored_constraints_match_the_real_schema[pancake-users-PancakeBase]", "migration/tests/test_schema_drift.py::test_regime_alias_table_is_ours_and_not_ar2s" ] } }, "lint_violations": 36, - "notes": [ - "Reference point moved deliberately, in review, for two reasons. Neither is a test or a violation disappearing quietly, which is what this file exists to prevent.", - "1. test_degenerate_geometry_quarantined_not_invented was renamed to test_unidentifiable_geometry_quarantined_not_invented. Its fixture was POLYGON((0 0, 0 0, 0 0, 0 0)), asserted to be broken geometry. It is not broken: it is a point registration written as a ring, which AR 1.0 accepted and AR2 registers natively. Reading that shape as unusable quarantined 14,594 of 28,282 live fields. The test now uses a line between two distinct positions, which genuinely cannot be identified, and test_a_pin_written_as_a_ring_imports covers the shape it used to reject.", - "2. lint_violations moved 25 -> 36 with no change to any source file. The count was measured by an unpinned ruff, and the same tree reports 0 under the version CI happened to install and 36 under a newer one. ruff is now pinned in .github/workflows/ci.yml, so 36 is the honest count under the tool the ratchet actually uses. These are pre-existing E402 violations in app/; the ratchet holds them from rising." - ] + "guards": { + "migration/tests/test_schema_drift.py": 7, + "migration/tests/test_report_arithmetic.py": 10 + }, + "lessons": { + "gate-reports-green-while-blind": 7, + "twin-divergence": 6, + "label-diverges-from-outcome": 5, + "claim-not-in-the-evidence": 4, + "guard-weakened-to-pass": 2, + "prose-warning-instead-of-a-check": 3, + "run-record-not-at-the-reviewed-commit": 1 + } } diff --git a/.stomata/lessons.json b/.stomata/lessons.json new file mode 100644 index 0000000..8480e25 --- /dev/null +++ b/.stomata/lessons.json @@ -0,0 +1,118 @@ +{ + "$stomata_artifact": "lessons_ledger", + "schema_version": "1.0.0", + "promotion_threshold": 2, + "note": "Seeded 2026-08-19 from the review record of 2026-07-24 to 2026-08-19. Occurrences are things that actually happened and cost time, each traceable to a dated review document in the agstack workplan. Nothing speculative is recorded here: a lesson with no incident behind it becomes noise, and noise is what discredits the checks that were earned.", + "lessons": [ + { + "recurrence_key": "gate-reports-green-while-blind", + "trigger": "A check, test or assertion reports success.", + "directive": "Establish that it can fail. Show it failing on purpose, or show the count it produces changing when the underlying thing changes. A gate that cannot distinguish good from bad is worse than no gate, because it is believed.", + "kind": "mechanized", + "enforced_by": [ + "harness/stomata/checks.py", + "harness/tests/test_stomata.py" + ], + "occurrences": [ + {"date": "2026-07-24", "what": "Tier 1 privacy was satisfied only vacuously: no holder_account could leak because no identity was ever produced.", "where": "workplan/2026-07/rajat_ar2_traceforward_20260724.md:221"}, + {"date": "2026-08-05", "what": "The Tier 1 privacy test passed vacuously while Tier 1 still emitted the field.", "where": "workplan/2026-08/rajat_day3_review_20260805.md:271"}, + {"date": "2026-08-06", "what": "The Tier 1 assertion could pass in isolation regardless of behaviour.", "where": "workplan/2026-08/rajat_day35_review_20260806.md:65"}, + {"date": "2026-08-10", "what": "A cross-layer test was skipped in Pancake CI and the suite still reported green.", "where": "workplan/2026-08/rajat_day5_closeout_20260810.md"}, + {"date": "2026-08-18", "what": "The schema drift guard passed while comparing nothing, because the mirror it compared against had been made permissive.", "where": "workplan/2026-08/rajat_import_review_20260818.md"}, + {"date": "2026-08-18", "what": "The dirty-tree warning fired on every single run because generated credentials were rewritten each time, so it stopped carrying information.", "where": "app/tests/testkit/mint_test_authority_credentials.py"}, + {"date": "2026-08-19", "what": "The lint ratchet counted ruff's success banner as a violation, so it reported 1 whatever the code said, and a real count of 36 rode in unnoticed across four commits.", "where": "harness/stomata/run.py"} + ] + }, + { + "recurrence_key": "twin-divergence", + "trigger": "One fact is represented in two places -- a mirror and the system it mirrors, a function and its caller, the same logic in two packages, a value in two config files.", + "directive": "Change both in the same commit, and add a test that fails when they disagree. Do not rely on remembering the second one, because the second one is what gets forgotten.", + "kind": "mechanized", + "enforced_by": [ + "migration/tests/test_schema_drift.py::test_mirrored_constraints_match_the_real_schema", + "migration/tests/test_points_are_importable.py", + "stomata.json" + ], + "occurrences": [ + {"date": "2026-08-12", "what": "get_boundary_coverage was correct and had been called from nowhere since AR 1.0.", "where": "workplan/2026-08/rajat_e2e_ci_audit_20260812.md"}, + {"date": "2026-08-14", "what": "v2 GeoID minting was fully implemented and no service path called it.", "where": "workplan/2026-08/rajat_ar2_closeout_20260814.md"}, + {"date": "2026-08-14", "what": "GeoIDRegimeAlias rows were written by the migration and never read by the equivalence-set query.", "where": "workplan/2026-08/rajat_ar2_closeout_20260814.md"}, + {"date": "2026-08-18", "what": "The point-handling path existed in app/geoid_v2.py and was absent from migration/geoid_v2.py, so every pin was quarantined.", "where": "migration/geoid_v2.py"}, + {"date": "2026-08-19", "what": "The phone constraint was relaxed in the migration mirror and not in the hub it mirrors, so the import reported 0 rejections where the real system would have rejected 85.", "where": "workplan/2026-08/rajat_final_import_review_20260819.md"}, + {"date": "2026-08-19", "what": "ruff was pinned in ci.yml and left unpinned in stomata.yml, so the gating workflow measured a different thing to the advisory one.", "where": ".github/workflows/stomata.yml"} + ] + }, + { + "recurrence_key": "label-diverges-from-outcome", + "trigger": "You are naming or counting what a run did, in a report, a log line, a status field or a document.", + "directive": "Make the word mean one outcome. If a category can contain both a success and a failure, split it. Assert that the categories sum to the total and do not overlap, because every individual number can be computed correctly and the totals still describe a run that did not happen.", + "kind": "mechanized", + "enforced_by": [ + "migration/tests/test_report_arithmetic.py" + ], + "occurrences": [ + {"date": "2026-08-14", "what": "TRACEABILITY.md documented Tiers 2 and 4 as implemented when they were not.", "where": "workplan/2026-08/rajat_ar2_closeout_20260814.md"}, + {"date": "2026-08-19", "what": "268 duplicate submissions that were correctly aliased to existing GeoIDs were reported as quarantined and as decisions required. Nothing had failed.", "where": "migration/run.py"}, + {"date": "2026-08-19", "what": "18,184 geometries were reported as altered by canonicalisation; the real figure was 3,597, the rest being points on a different code path that set the flag as a side effect.", "where": "migration/pipeline.py"}, + {"date": "2026-08-19", "what": "The point count, the single most consequential fact about the run, did not appear in the report at all.", "where": "migration/run.py"}, + {"date": "2026-08-19", "what": "resolved_child_of was presented as a peer of imported_new when it is a subset, so 74 fields summed to 75 and read as one field processed twice.", "where": "migration/tests/test_report_arithmetic.py"} + ] + }, + { + "recurrence_key": "claim-not-in-the-evidence", + "trigger": "You are about to quote a number, or cite a run, screenshot or log as evidence.", + "directive": "Open the artifact and find the number in it. Cite the run that produced it, at the commit under review. A figure that cannot be located in the evidence is a recollection, and recollections drift toward what we hoped.", + "kind": "mechanized", + "enforced_by": [ + "harness/stomata/packet.py", + "harness/tests/test_stomata.py" + ], + "occurrences": [ + {"date": "2026-08-14", "what": "Screenshots showed 2 fieldlists as though that were the corpus; it was an artifact of a --limit flag on the run.", "where": "workplan/2026-08/rajat_ar2_closeout_20260814.md"}, + {"date": "2026-08-18", "what": "A packet cited packets/evidence/schema1.log, which did not exist, and bias_curve.csv, which was at the repository root.", "where": "packets/schema.json"}, + {"date": "2026-08-18", "what": "'The missing 13,343' was stated in prose and appeared nowhere in the run it described.", "where": "workplan/2026-08/rajat_import_review_20260818.md"}, + {"date": "2026-08-19", "what": "A summary quoted a pre-fix reconciliation of 13,658 + 14,594 + 30 while the attached log showed 27,978 + 275 + 29.", "where": "packets/evidence/import-output.txt"} + ] + }, + { + "recurrence_key": "guard-weakened-to-pass", + "trigger": "A guard, assertion or lint rule stands between you and a change you believe is correct.", + "directive": "Change the claim or change the code, never quietly narrow the guard. If the guard is genuinely stale, say so in the commit message and re-record the baseline, so that weakening it costs a sentence someone can read.", + "kind": "mechanized", + "enforced_by": [ + "harness/stomata/checks.py", + "harness/tests/test_lessons.py::test_removing_an_assertion_from_a_guard_fails" + ], + "occurrences": [ + {"date": "2026-08-12", "what": "ruff.toml's ignore list was suppressing rule families including B008, so the linter was configured not to see real defects.", "where": "workplan/2026-08/rajat_e2e_ci_audit_20260812.md"}, + {"date": "2026-08-19", "what": "The phone column was removed from the drift guard's NOT NULL and UNIQUE lists so a change would pass, and the guard then certified an import that could not work.", "where": "migration/tests/test_schema_drift.py"} + ] + }, + { + "recurrence_key": "prose-warning-instead-of-a-check", + "trigger": "You have identified a failure mode and are about to write it down in a document, review comment or email.", + "directive": "Write the check as well, in the same change, or record explicitly that it cannot be mechanized and why. A paragraph is read once by whoever was in the conversation and then archived; it does not bind the person who arrives next, and it does not bind its own author a week later.", + "kind": "mechanized", + "enforced_by": [ + "harness/stomata/lessons.py", + "harness/stomata/checks.py", + "harness/tests/test_lessons.py" + ], + "occurrences": [ + {"date": "2026-08-13", "what": "Threshold guidance was written as prose reasoning; it rested on an IoU implementation that turned out to be wrong, and the prose gave no way to notice.", "where": "workplan/superseded/rajat_geoid_v2_review_20260813.md"}, + {"date": "2026-08-14", "what": "A document asked, in words, that checks not be waived silently. Five days later a guard was edited to pass and nothing objected.", "where": "workplan/2026-08/rajat_ar2_closeout_20260814.md:406"}, + {"date": "2026-08-14", "what": "The same document predicted that a check which becomes noise 'gets waived, then gets deleted'. Two checks became noise and were ignored for days by the author who wrote that sentence.", "where": "workplan/2026-08/rajat_ar2_closeout_20260814.md:418"} + ] + }, + { + "recurrence_key": "run-record-not-at-the-reviewed-commit", + "trigger": "You are reading a run record, CI result or packet to decide whether a branch is sound.", + "directive": "Check the commit the run was taken at against the tip being reviewed. Commits after the last recorded run are ungated, and that is exactly where an unnoticed regression sits.", + "kind": "advisory", + "advisory_because": "One occurrence so far, and the commit-binding check already records the commit for a reader to compare. Promoting on a single incident manufactures the kind of speculative check that turns into noise. Revisit on the next occurrence.", + "occurrences": [ + {"date": "2026-08-18", "what": "Run records sat at b3c55ed and ac6e536 while the branch tip was 17d3aaf, and a lint regression rode in on the four ungated commits between.", "where": "workplan/2026-08/rajat_import_review_20260818.md"} + ] + } + ] +} diff --git a/harness b/harness index f7ac8b0..358304f 160000 --- a/harness +++ b/harness @@ -1 +1 @@ -Subproject commit f7ac8b0ac9568fec2908e45342caae39af79bb58 +Subproject commit 358304fa20e2052214a3c2ec18d65bbc81f5cc15 diff --git a/migration/pipeline.py b/migration/pipeline.py index 041d15d..b44d781 100644 --- a/migration/pipeline.py +++ b/migration/pipeline.py @@ -78,6 +78,11 @@ class FieldReport: # two mean different things: an exact content match is certainty, a # threshold match is a judgement that the threshold could change. resolved_by_content_hash: int = 0 + # A SUBSET of imported_new, not a peer of it. A child field is a genuinely new + # GeoID that additionally records a parent link, so it is counted in both. + # Listing it alongside imported_new made the categories sum to more than + # considered, which reads as a field processed twice; see + # migration/tests/test_report_arithmetic.py. resolved_child_of: int = 0 skipped_already_done: int = 0 quarantined: dict[str, list[str]] = field(default_factory=dict) diff --git a/migration/tests/test_report_arithmetic.py b/migration/tests/test_report_arithmetic.py new file mode 100644 index 0000000..233ed54 --- /dev/null +++ b/migration/tests/test_report_arithmetic.py @@ -0,0 +1,105 @@ +"""The report's categories must add up, and must not overlap. + +THE INCIDENT. On 2026-08-19 an import run reported 268 fields quarantined and 268 +decisions required. Nothing had failed: all 268 were duplicate submissions of +geometry already in the registry, they had been correctly aliased to the existing +GeoID, and the pipeline had done exactly the right thing. The word "quarantined" +was simply being used for two opposite outcomes -- rejected, and resolved -- and +the reader could not tell which had happened. + +The same run reported 18,184 geometries as altered by canonicalisation. The real +figure was 3,597; the rest were points, which take a different code path that was +setting the flag as a side effect. Both numbers were produced by working code and +both were wrong as descriptions of what happened. + +WHY THIS IS A TEST AND NOT A NOTE. The reason this class of defect kept recurring +is that nothing could catch it: every individual number was computed correctly, so +the suite was green and the totals were nonsense. What was missing was an +assertion about the relationship between the categories -- that a field counted as +rejected is not also counted as imported, and that the parts sum to the whole. +That relationship is checkable, so it is checked here instead of being remembered. +""" +from __future__ import annotations + +from migration.pipeline import FieldReport, InMemoryRepo, import_fields +from migration.sources import FixtureSource + + +def _run(limit: int | None = None) -> FieldReport: + return import_fields(FixtureSource(), InMemoryRepo(), limit=limit) + + +def test_every_field_considered_lands_in_exactly_one_bucket(): + """The parts must sum to the whole. + + If they do not, some field was counted twice or dropped silently, and the + report is describing a run that did not happen. + """ + r = _run() + accounted = (r.imported_new + r.resolved_same_as + + r.skipped_already_done + r.quarantined_total) + assert accounted == r.considered, ( + f"{r.considered} considered but {accounted} accounted for. " + f"A field is either imported, resolved, skipped or rejected -- " + f"never two of those, and never none." + ) + + +def test_a_rejected_field_is_never_also_a_resolved_one(): + """The 2026-08-19 defect exactly: one word for two opposite outcomes. + + Aliasing a duplicate to an existing GeoID is a success. Rejecting a geometry + is a refusal. Counting them together produced 268 'decisions required' where + the honest answer was none. + """ + r = _run() + rejected = {gid for ids in r.quarantined.values() for gid in ids} + assert len(rejected) == r.quarantined_total, ( + "the same field appears under two rejection reasons, so the total " + "over-counts the refusals" + ) + assert not (rejected & set(r.uuid_promoted)), ( + "a field cannot be both rejected and promoted to a UUID" + ) + + +def test_subcategories_never_exceed_the_category_they_subdivide(): + """imported_points and resolved_by_content_hash are parts of larger counts. + + A part larger than its whole means the flag is being set on a path that does + not belong to it -- which is how points came to inflate the canonicalisation + count. + """ + r = _run() + assert r.imported_points <= r.imported_new + assert r.resolved_by_content_hash <= r.resolved_same_as + # Found by the closure test above: a child field is inserted as a new GeoID and + # also counted as child-of, so it is a subset of imported_new rather than an + # alternative to it. Presenting the two as peers made 74 fields report as 75. + assert r.resolved_child_of <= r.imported_new + + +def test_canonicalisation_is_only_flagged_when_geometry_actually_changed(): + """Points are rewritten by the point path, not repaired by canonicalisation. + + Flagging them inflated a number that a human reads as 'this many boundaries + needed fixing', which is a claim about data quality and was wrong by 5x. + """ + r = _run() + assert len(r.canonicalization_changed) <= r.considered + assert len(set(r.canonicalization_changed)) == len(r.canonicalization_changed), ( + "a field is listed twice as canonicalisation-altered" + ) + + +def test_the_arithmetic_holds_under_a_limit(): + """--limit is how the numbers get quoted in a hurry. + + A screenshot taken from a limited run was once read as the full picture, so the + invariant has to hold on the partial run too, not only on the complete one. + """ + r = _run(limit=5) + accounted = (r.imported_new + r.resolved_same_as + r.resolved_child_of + + r.skipped_already_done + r.quarantined_total) + assert accounted == r.considered + assert r.considered <= 5 diff --git a/stomata.json b/stomata.json index e50f29e..7799068 100644 --- a/stomata.json +++ b/stomata.json @@ -1,15 +1,5 @@ { "schema_version": "1.0.0", - "_comment": [ - "AR2's contract with the harness. Machine-read by stomata; edited by humans.", - "", - "Adding a suite, a reachability claim, a forbidden pattern or a mutation here is", - "how a guarantee becomes enforced rather than remembered. Removing one is a", - "visible diff, which is the point: a check that can vanish quietly is not a check.", - "", - "Commands use {python} rather than a bare interpreter name so the same contract", - "works in a virtualenv, a container and CI without edits." - ], "python": "python3", "suites": [ { @@ -102,6 +92,16 @@ "why": "The area pre-filter gated all same_as resolution and evaluated to zero whenever either area was missing, so every import with absent area silently became a duplicate rather than resolving to the field it matched." } ], + "guards": [ + { + "path": "migration/tests/test_schema_drift.py", + "why": "On 2026-08-19 the phone column was removed from this file's NOT NULL and UNIQUE lists so a change would pass, and the guard then certified an import that would have been rejected 85 times by the real hub. It may assert more over time, never less." + }, + { + "path": "migration/tests/test_report_arithmetic.py", + "why": "Holds the invariant that the import report's categories sum to the total and do not overlap. Removing an assertion here is how 268 successful aliases came to be reported as quarantined." + } + ], "mutations": [ { "name": "cover the bounding box instead of the polygon", @@ -149,5 +149,15 @@ "migration/tests/", "test_", "conftest.py" + ], + "_comment": [ + "AR2's contract with the harness. Machine-read by stomata; edited by humans.", + "", + "Adding a suite, a reachability claim, a forbidden pattern or a mutation here is", + "how a guarantee becomes enforced rather than remembered. Removing one is a", + "visible diff, which is the point: a check that can vanish quietly is not a check.", + "", + "Commands use {python} rather than a bare interpreter name so the same contract", + "works in a virtualenv, a container and CI without edits." ] } From b76a95475ffedfaaa9faa7e581149dc95dfc12d1 Mon Sep 17 00:00:00 2001 From: Sumer Johal Date: Wed, 19 Aug 2026 12:13:40 -0700 Subject: [PATCH 4/4] Brief before starting: AGENTS.md, generated from the ledger Declares AGENTS.md as the briefing target and generates its lessons section from .stomata/lessons.json. AGENTS.md loads automatically, which is the point: the ledger was being verified after every run and read before none. The hand-written part covers the two things worth knowing before a first failure. Do not narrow a check to get past it -- if a guard blocks a change you believe is correct, one of the two is wrong and it is worth working out which, often the guard. And if a check fires on something genuinely fine, that is a bug in the check, to be fixed rather than worked around, because a check that cries wolf becomes noise and noise discredits the checks that were earned. Verified by hand-editing the generated block and confirming check L fails naming the drift, then restoring it. The first attempt at that verification proved nothing -- it replaced text that was not in the block -- which is the claim-not-in-the- evidence lesson catching me in the act of testing a check against nothing. 12/12 checks with mutations, 229 tests. Co-authored-by: Cursor --- .stomata/baseline.json | 4 +- AGENTS.md | 120 +++++++++++++++++++++++++++++++++++++++++ harness | 2 +- stomata.json | 3 ++ 4 files changed, 126 insertions(+), 3 deletions(-) create mode 100644 AGENTS.md diff --git a/.stomata/baseline.json b/.stomata/baseline.json index e0520aa..105fee2 100644 --- a/.stomata/baseline.json +++ b/.stomata/baseline.json @@ -1,7 +1,7 @@ { "$stomata_artifact": "baseline", - "captured_at": "2026-08-19T18:50:48+00:00", - "commit": "0691084ddbea155db2806cd54e263984104ec34a", + "captured_at": "2026-08-19T19:13:10+00:00", + "commit": "ceff9444e3ab90b054a5c99898901cff6cd96360", "dirty_when_captured": true, "suites": { "ar2": { diff --git a/AGENTS.md b/AGENTS.md new file mode 100644 index 0000000..ec9f2cf --- /dev/null +++ b/AGENTS.md @@ -0,0 +1,120 @@ +# Working in this repository + +## Before you start + +Run `harness/bin/stomata brief`. It prints the section at the bottom of this file, +which is the list of defects that have already happened here more than once. + +The reason this exists in a file that loads automatically, rather than in a +document someone has to remember to open: on 2026-08-14 the correct warning about +a specific failure mode was written into a guidance document, and on 2026-08-19 +that exact failure happened anyway. The warning was accurate, reviewed, and +archived where nobody read it again at the moment it applied. Prose does not bind, +and a lesson nobody encounters at the right time is not a lesson. + +## Before you finish + +Run `harness/bin/stomata run`, or `--full` at a checkpoint to include mutations. +Twelve checks; every one exists because of a specific incident recorded in +`harness/docs/RATIONALE.md`. + +Two things to know about how to treat a failure: + +**Do not narrow a check to get past it.** If a guard blocks a change you believe is +correct, one of the two is wrong and it is worth a minute to work out which. Often +it is the guard — they do go stale — and then the fix is to change it deliberately +and say so in the commit message. What check J refuses is the count of assertions +falling quietly, because that is how a guard came to certify an import that could +not run. + +**If a check fires on something genuinely fine, that is a bug in the check.** Fix +the check rather than working around it. A check that cries wolf becomes noise, +and noise is what discredits the checks that were earned. + +## Recording a new lesson + +When a defect turns out to be the second occurrence of something: + +``` +harness/bin/stomata lesson add --key \ + --what "what happened this time" --where "path or document" +``` + +Merging on a recurrence key is the mechanism, not bookkeeping. The same pain +reported under three different wordings reads as three unlucky one-offs; under one +key it reads as a pattern with three occurrences, which is what makes leaving it +unheld a visible choice rather than an oversight. + +Then either mechanize it and name the check in `enforced_by`, or set +`advisory_because` and say why it cannot be. Check K accepts both and rejects +silence, because silence looks identical to having dealt with it. + +Do not add a check speculatively. A check with no incident behind it becomes noise, +then gets waived, then gets deleted, and takes the credibility of the real checks +with it. + +## Do not edit the section below + +It is generated from `.stomata/lessons.json` by `stomata brief --write`. Edit the +ledger and regenerate; check L fails when the two disagree. Hand-editing it would +create two copies of one fact drifting apart, which is `twin-divergence` — the most +frequent lesson in the ledger, and one this file would then be committing itself. + +--- + + + + +These are defects that have already happened here, more than once each. They +are generated from `.stomata/lessons.json` by `stomata brief --write`; edit the +ledger, not this block, or check L will fail on the next run. + +### gate-reports-green-while-blind (7 occurrences) + +**When:** A check, test or assertion reports success. + +**Do:** Establish that it can fail. Show it failing on purpose, or show the count it produces changing when the underlying thing changes. A gate that cannot distinguish good from bad is worse than no gate, because it is believed. + +*Enforced by: harness/stomata/checks.py, harness/tests/test_stomata.py* + +### twin-divergence (6 occurrences) + +**When:** One fact is represented in two places -- a mirror and the system it mirrors, a function and its caller, the same logic in two packages, a value in two config files. + +**Do:** Change both in the same commit, and add a test that fails when they disagree. Do not rely on remembering the second one, because the second one is what gets forgotten. + +*Enforced by: migration/tests/test_points_are_importable.py, migration/tests/test_schema_drift.py::test_mirrored_constraints_match_the_real_schema, stomata.json* + +### label-diverges-from-outcome (5 occurrences) + +**When:** You are naming or counting what a run did, in a report, a log line, a status field or a document. + +**Do:** Make the word mean one outcome. If a category can contain both a success and a failure, split it. Assert that the categories sum to the total and do not overlap, because every individual number can be computed correctly and the totals still describe a run that did not happen. + +*Enforced by: migration/tests/test_report_arithmetic.py* + +### claim-not-in-the-evidence (4 occurrences) + +**When:** You are about to quote a number, or cite a run, screenshot or log as evidence. + +**Do:** Open the artifact and find the number in it. Cite the run that produced it, at the commit under review. A figure that cannot be located in the evidence is a recollection, and recollections drift toward what we hoped. + +*Enforced by: harness/stomata/packet.py, harness/tests/test_stomata.py* + +### prose-warning-instead-of-a-check (3 occurrences) + +**When:** You have identified a failure mode and are about to write it down in a document, review comment or email. + +**Do:** Write the check as well, in the same change, or record explicitly that it cannot be mechanized and why. A paragraph is read once by whoever was in the conversation and then archived; it does not bind the person who arrives next, and it does not bind its own author a week later. + +*Enforced by: harness/stomata/checks.py, harness/stomata/lessons.py, harness/tests/test_lessons.py* + +### guard-weakened-to-pass (2 occurrences) + +**When:** A guard, assertion or lint rule stands between you and a change you believe is correct. + +**Do:** Change the claim or change the code, never quietly narrow the guard. If the guard is genuinely stale, say so in the commit message and re-record the baseline, so that weakening it costs a sentence someone can read. + +*Enforced by: harness/stomata/checks.py, harness/tests/test_lessons.py::test_removing_an_assertion_from_a_guard_fails* + + diff --git a/harness b/harness index 358304f..c2273a4 160000 --- a/harness +++ b/harness @@ -1 +1 @@ -Subproject commit 358304fa20e2052214a3c2ec18d65bbc81f5cc15 +Subproject commit c2273a437a334c35946b2b42c45f1bc8e33e2dbf diff --git a/stomata.json b/stomata.json index 7799068..2c9f22b 100644 --- a/stomata.json +++ b/stomata.json @@ -102,6 +102,9 @@ "why": "Holds the invariant that the import report's categories sum to the total and do not overlap. Removing an assertion here is how 268 successful aliases came to be reported as quarantined." } ], + "briefings": [ + "AGENTS.md" + ], "mutations": [ { "name": "cover the bounding box instead of the polygon",