Repository navigation
AR2 Closeout tasks & pipeline updates - #14
Merged
Merged
Conversation
…nes so signatures match
…rmance tests, and properties
AR2's registry is empty, so the v1 -> v2 namespace change is an import from
two legacy sources rather than an in-place migration: AR 1.0 holds the
polygons and TerraPipe holds the user-to-polygon association. The join key
between them is a v1 -> v2 identifier mapping that has nowhere to live today,
since GeoIDAlias maps content hashes rather than regimes.
Adds the pipeline end to end against a pluggable source, so the only thing
still to write is the adapter over the two legacy schemas.
geoid_v2 corrected primitive: preserves holes and every part of a
MultiPolygon, both of which were being silently discarded
resolve area-exact IoU via cell-union intersection and leaf weighting;
token-set arithmetic cannot see ancestor/descendant overlap and
is wrong on the multi-level covers normalisation produces
sample adversarial stratified selection, using AR 1.0's own L13 GeoID
as a free blocking key for near-duplicate clusters
models mirrors of the target tables plus the new geo_id_regime_alias
db_repo TargetRepo over the three real databases
pipeline two enforced phases; child_of registers rather than aliasing,
so a nested plot keeps its identity and its owner keeps scope
78 tests, including that the in-memory and database repositories produce
identical reports, and a drift guard that parses the shipped model files.
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
…rst stated Co-authored-by: Cursor <cursoragent@cursor.com>
…xact Three changes that together retire the v1 identity regime. THE IDENTIFIER CASCADE IS GONE. Registration hashed the L13 cover, fell back to hashing L20 on collision, and fell back to uuid4() on a second collision. An L13 cell is roughly 1.2 km across, so distinct fields collided routinely, and the escape hatch produced an identifier derived from nothing -- indistinguishable downstream from a real one. Registration now hashes the polygon's own normalized cover, and a repeat means identical canonicalized geometry, which resolution already reports as same_as at IoU 1.0. Unusable geometry is refused with 422 rather than assigned an invented ID. COVER COMPARISON IS NOW AREA-EXACT. check_percentage_match compared covers by token-set intersection, which cannot see ancestor/descendant overlap: one L16 cell and its own four L17 children cover the identical region and share no token. Measured on two 200 m squares offset 10 m, true IoU 0.905, the old arithmetic returned 0.283 and the corrected form returns 0.912. On uniform-level v1 covers the two agree to within 1e-12, so no existing v1 decision changes -- that equivalence is asserted, not assumed. THE AREA PRE-FILTER IS GONE. It gated the entire same_as branch behind area_ratio >= threshold, and area_ratio is 0 whenever either area is missing, so an imported field with no recorded area could never resolve to an existing one however exactly its geometry matched. Every such import silently became a duplicate. Also: the v2 primitive itself was discarding holes and keeping only the first part of a MultiPolygon, so a field with a pond hashed identically to the solid shape and a multi-part field took the identity of one fragment. Both matter because make_valid() *produces* MultiPolygons from self-intersecting input. The implementation now lives in app/geoid_v2.py, pinned against migration/geoid_v2.py by a cross-implementation agreement test, and correctness is asserted by area rather than by stored hashes -- a conformance vector can always be regenerated to match whatever the code does. Two of the twenty conformance vectors were regenerated for exactly that reason; both carry a note recording why, and the area assertions that justify the new values. GeoIDRegimeAlias is added to the models and read by _equivalence_set in both directions, so an identifier AR 1.0 issued years ago still reaches its field. child_of is deliberately not followed: it is containment, not identity, and collapsing it would widen every trace to the parent. Co-authored-by: Cursor <cursoragent@cursor.com>
…04 scenario
A hop is now one lot, not one level of the traversal. 21 CFR 1.1320(a) assigns a
new traceability lot code at exactly three events -- initial packing, first
land-based receiving, transformation -- and a ListArtifact is created at exactly
those moments, so one list is one lot is one hop. Reporting per depth level
grouped unrelated lots together, which is a fact about how we walk the graph
rather than about the supply chain.
Each hop now carries where its lot was created, as a GeoID. That is the lot code
source the rule asks for at 1.1330(a)(14) and 1.1350(a)(2)(ii), and making it a
GeoID means a packhouse is identified exactly as a field is, with no second
registry to keep in sync. It is nullable: a lot whose creation site was never
recorded is a common real state, and the trace reports the gap in
hops_missing_location rather than refusing to answer.
Regions are reported but deliberately not expanded. Region membership is spatial
-- /traceforward matches a field by probing its S2 cells and their ancestors
against RegionCoverCell -- and there is nothing to read back, so inverting it is
a prefix query rather than a lookup. regions_unexpanded says the field set is a
lower bound, because a partial answer that looks total is what makes an
investigator under-scope a recall.
Also fixes the reverse lookup to expand the seed through its equivalence set, so
GET /list-artifact/reverse/{geoid} and /traceforward no longer disagree about
what counts as the same field.
The scenario suite is both a regression harness and the IFPA demonstration: three
fields, two growers, one blending processor, traced in both directions, with
every assertion written against a clause of the rule. Verified by mutation --
dropping the hop location, removing the recursion, collapsing hops to one per
level, hiding truncation, hiding the region gap, and unlinking inputs are each
caught by at least one test.
The two new columns need scripts/schema_upgrade_traceback.sql on any existing
database: AR2 has no Alembic, and create_all adds missing tables but never
missing columns, so this is a change that passes CI on an empty database and
fails on a deployed one.
Co-authored-by: Cursor <cursoragent@cursor.com>
The harness is a pinned submodule at harness/, driven by stomata.json. One
command runs both suites and nine checks:
harness/bin/stomata the gate
harness/bin/stomata run --full plus mutations, at a checkpoint
The contract declares what must stay true, with a reason for each entry:
- four reachability claims, so a tested module cannot ship unwired. GeoID v2
was fully tested and called by nothing for two days; that is the shape of
defect this catches.
- three forbidden patterns -- the L13 cover hash, the uuid4 fallback, the area
pre-filter -- so fixed defects cannot return in a merge.
- five mutations, so a green suite has to prove it can fail.
Three findings from turning it on, all real:
- test_geoid_v2_live.py was not re-runnable. Its fixture counter restarted each
process, so a second run re-registered identical geometry and hit the geo_id
unique constraint: nine failures on the second run of a suite that passed on
the first. Each run now takes its own latitude band.
- the migration schema drift guard skips without the hub and pancake checkouts,
so locally it passes without comparing anything. Waived here with its reason,
and the CI job refuses the waiver, where both are present.
- one mutation was undetectable, because S2CellUnion.Init already returns ids
in order, so deleting the sort changed nothing. It tested whether a line
existed rather than whether order mattered. Reversing the order tests the
real invariant.
Also consolidates the two schema upgrade scripts into
scripts/schema_upgrade_20260814.sql, so one command brings a database up to date.
Baseline: 119 AR2 tests, 74 migration tests, 25 lint violations held.
The first full run against live AR 1.0 quarantined 14,594 of 28,282 fields --
51.6% -- under one reason, "unusable_geometry", and 14,229 of 26,265
user-to-field associations then failed to join, because a quarantined field
never enters the alias table the join reads. Better than half the registry, and
better than half of every user's holdings, did not arrive.
The cause was not the data. AR 1.0 accepted point registrations; a pin has no
area, so the polygon coverer cannot describe one. AR2's own registration has
handled points since the v2 wiring -- app/geoid_v2.point_geo_id_with_tokens --
but migration/geoid_v2.py was ported as a polygon-only primitive and the
pipeline had no point path at all, so every pin reached a bare except and was
filed as broken geometry. It was reported as a question about policy when it
was a question about code.
* migration/geoid_v2.py gains point_geo_id_with_tokens, byte-for-byte the same
rule as the live path, so a pin registered through the API and the same pin
imported land on one identifier rather than two. point_coords accepts a pin
however AR 1.0 wrote it down, including a ring whose vertices coincide,
which is a pin recorded as a boundary rather than a broken boundary.
* Quarantine reasons now carry the specific fault. One bucket of 14,594 says
how many rows failed and nothing about what to do with them; it cannot
distinguish a missing code path from a corrupt row, and those need opposite
responses.
* area_ha is derived from the boundary rather than read from AR 1.0, which
exposes no area this adapter can trust. Every row reporting None collapsed
the inventory bands into a single "unknown" bucket describing nothing.
* The test fixture POLYGON((0 0, 0 0, 0 0, 0 0)) was named bad_geom and
asserted to be unusable. It is a pin at the origin. The fixture now uses a
line between two distinct positions, which genuinely cannot be identified.
Also, three things that made the gate quieter than it should have been:
* conftest re-minted tracked credentials on every test run, so any run left
the tree modified and every result was reported as coming from a dirty tree.
A warning that fires every time is one people learn to skip past. The mint
is now idempotent and a run leaves the tree clean.
* ruff is pinned, in patches/pin-ruff-in-ci.patch rather than in this commit:
my token cannot write .github/workflows. Unpinned, the same tree measured
0 violations under the
version CI installed and 36 under a newer one, and a ratchet whose measuring
stick moves reports upgrades as regressions and hides regressions behind
downgrades. The baseline moves to 36, the honest count under the pinned
tool, with no source change; .stomata/baseline.json records why.
* A reachability claim on point_coords. The primitive existing was never the
point -- the live path had a working point implementation the whole time and
the importer never called it.
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
hub.users.phone was relaxed to nullable and non-unique in migration/models.py,
and phone was removed from the constraint guard, while ar2-hub still enforced
NOT NULL UNIQUE on that column. The import consequently reported 0 rejected
accounts where a real run against the hub would still have rejected 85 -- 8 with
no phone, 77 sharing a number. The constraint was removed from the simulation,
not from the system.
That is the worse failure shape available: not an error, but a clean bill of
health for an import that cannot happen. The existing drift guard compared
column *names* only, so it passed throughout -- the column was still called
phone.
* test_mirrored_constraints_match_the_real_schema compares nullability and
uniqueness against the real models, and fails only on permissiveness. A
mirror that is stricter than reality predicts rejections that will not
happen, which is a false alarm; a mirror that is looser predicts acceptance
for rows the system rejects, which reads as success.
* The docstring on test_hub_constraints_the_import_depends_on_are_still_there
records why phone legitimately left both lists, and that the order it
happened in was backwards. Relaxing the guard is the last step, never the
first.
* The real relaxation is ar2-hub#7, with the DDL create_all will not apply.
Also, three reporting defects. Every number below is one a reader would act on,
and each was wrong in the direction of looking worse or looking better than the
run actually was.
* 268 fields that resolved correctly on an identical content hash were counted
as quarantined and then listed under DECISIONS REQUIRED. Every one had
aliased to its canonical row and still resolves. A run that rejected 7 fields
read as a run with 275 open problems. They now count as resolutions, split
into exact-geometry and above-threshold matches, because certainty and
judgement should not share a number.
* Every imported point was counted as canonicalization-altered geometry,
taking that figure from 3,597 to 18,184 and converting a real signal about
boundary repair into a headcount of pins. A point has no boundary to repair.
* The point count was not reported at all, though a registry that is half pins
is a different thing to plan around than one that is all boundaries. It was
the most consequential fact about the run and appeared nowhere in it.
* "quarantined" is now "REJECTED, not imported", with the share of the total,
and the decisions section says that a reason naming a geometry type is
usually a missing code path rather than a policy question.
Finally, the harness submodule moves to f7ac8b0. The pinned revision counted
ruff's "All checks passed!" banner as one violation, so CI reported 1 where this
tree has 36 -- the ratchet has been reading green while blind. The gating
workflow also installs ruff unpinned, which the earlier pin missed because it
only covered ci.yml; patches/pin-ruff-in-stomata-workflow.patch has that, since
my token cannot write .github/workflows.
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
… adds up Seeded .stomata/lessons.json from the review record of 2026-07-24 to 2026-08-19. Seven recurrence keys, every occurrence traceable to a dated document. Nothing speculative: a lesson with no incident behind it becomes noise, and noise is what discredits the checks that were earned. What the corpus shows, which was not visible one review at a time: 7x a gate reported green while blind 6x one fact in two places, one of them updated 5x a label or count that described a run that did not happen 4x a number quoted that was not in the evidence 3x a failure mode written down as prose instead of a check 2x a guard narrowed to let a change through The last two are the interesting ones. The prose entry is about this review process rather than the code: three times a correct warning was written in a document and did not bind, including one that predicted the guard-weakening incident five days before it happened. That is the argument for checks J and K existing in code. test_report_arithmetic.py mechanizes the label lesson. Every individual number in the import report was computed correctly and the totals still misdescribed the run: 268 successful aliases reported as quarantined, canonicalisation overstated 5x by points on another code path. What was missing was an assertion about the relationship between categories, which is checkable. It found a live one immediately: resolved_child_of was presented as a peer of imported_new when it is a subset, so 74 fields summed to 75 and read as one field counted twice. The display was already correctly nested; the invariant was unstated, and is now stated where it can fail. Two guards declared. Verified by removing an assertion from the drift guard and confirming J fails, quoting the incident, then restoring it. 11/11 checks pass with mutations, 229 tests, 6/6 recurring lessons mechanized. Co-authored-by: Cursor <cursoragent@cursor.com>
Declares AGENTS.md as the briefing target and generates its lessons section from .stomata/lessons.json. AGENTS.md loads automatically, which is the point: the ledger was being verified after every run and read before none. The hand-written part covers the two things worth knowing before a first failure. Do not narrow a check to get past it -- if a guard blocks a change you believe is correct, one of the two is wrong and it is worth working out which, often the guard. And if a check fires on something genuinely fine, that is a bug in the check, to be fixed rather than worked around, because a check that cries wolf becomes noise and noise discredits the checks that were earned. Verified by hand-editing the generated block and confirming check L fails naming the drift, then restoring it. The first attempt at that verification proved nothing -- it replaced text that was not in the block -- which is the claim-not-in-the- evidence lesson catching me in the act of testing a check against nothing. 12/12 checks with mutations, 229 tests. Co-authored-by: Cursor <cursoragent@cursor.com>
Fail when the mirror is more permissive than the system it mirrors
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Submission of AR2 closeout tasks, import pipeline evidence, and CI configuration.