diff --git a/changelog.d/smoke-gapfill-ordering-fix.fixed.md b/changelog.d/smoke-gapfill-ordering-fix.fixed.md new file mode 100644 index 00000000..57ccd68c --- /dev/null +++ b/changelog.d/smoke-gapfill-ordering-fix.fixed.md @@ -0,0 +1,5 @@ +Stage stacked-spine transfer targets by producer availability, route early housing rent from its actual ASEC producer to ACS, and preserve exact assembly-bound ACS group-quarters rent absence without zero synthesis. Channel-aware precedence guards prove every direction's producer runs before activation and reject unknown execution scopes, while complete native ACS support/raw/household-parent/classification lineage, lifecycle-exact receipted clone roles, count-coherent terminal authority checks, and evaluator-sealed evidence snapshots keep relabeled, substituted, grafted, incomplete, or noncanonical receipts out of publication. Reject every nonstructural early gap-fill residual before clone attachment. + +Match certified ASEC produced-frame earnings semantics by applying named `acs_2024_pums_wagp_age_15_plus` and `acs_2024_pums_semp_age_15_plus` universe zeros only to mapped under-15 ACS leaves, with exact person/unit counts and digests; preserve raw WAGP/SEMP blanks, retain all-child tax units with receipted zero predictors, and restrict person allocation to age 15+. Eligible mapped or raw nulls still fail at the original pre-coercion source boundary with the rule ID. Reject ambiguous predictor grains and reject post-aggregation NaN, positive infinity, and negative infinity by named predictor before QRF fitting. Bind the universe receipt and live feature digest through every primary-QRF checkpoint and finalization. + +Authenticate the exact QBI mutation receipt and deterministic live transition at generation, simulated checkpoint emission, durable checkpoint write/load, simulated resume, legacy and stacked manifest construction, and both publication entry points. Generation proves exact preservation of every undeclared person column, non-person entity, link, weight, stratum, mass-log entry, and frame-metadata value, then binds the validated receipt SHA and its preimage/output digests into immutable frame metadata and a separately carried, checkpoint-content-bound transition authority. Later boundaries require that independent authority plus exact schemas, types, nonnegative integer counts, the recomputed live universe receipt, declared-output and QBI-driver SHA-256 digests, preservation-count equations, canonical seed-only person additions, and the kernel fixed point. Missing, wrong-route, ambiguous, forged, rebound, inventory-laundered, or freshly reissued alternate-fixed-point receipts fail closed. Preserve mapped ACS under-15 self-employment universe zeros in every clone role while continuing to reconcile every derived QBI leaf. Version the primary-QRF chain schema, outer stacked checkpoint materializer, and canonical stacked authority at 6, bind the ACS-universe and QBI-mutation contracts into base identity, and reject every stale v1--v5 payload. diff --git a/docs/us-multispine-operator-ordering.md b/docs/us-multispine-operator-ordering.md index d0ad508c..645328c4 100644 --- a/docs/us-multispine-operator-ordering.md +++ b/docs/us-multispine-operator-ordering.md @@ -110,40 +110,192 @@ are allowed only when named by the ACS native-input receipt. both survey arms with the single `sample_fraction` and `sample_seed`, restores each sample to its full-source design-weight mass, and assembles one origin-labeled frame. Standard rungs are `f001`, `f010`, and `f100`; - the manifest binds the fraction, seed, exact realized ASEC/ACS counts, and - selected-lineage digests. The full PUF remains a donor and is never sampled. -2. The spine-blind source-preparation chain derives the native predictors - needed by the declared cross-origin fills. No population operator selects + the manifest binds the fraction, seed, exact realized ASEC/ACS counts, + selected-lineage digests, the complete ordered native ACS household and + person support/raw/household-parent/classification mappings, and the sampled + native ACS TYPEHUGQ 2/3 household and person lineage digests. The full PUF + remains a donor and is never sampled. +2. The spine-blind source-preparation chain derives the native predictors and + pre-clone operator outputs needed by the early declared cross-origin fills. + Historical kernels run on the raw-`PERIDNUM` CPS/ASEC availability + projection and merge only their declared outputs back into the stack. In + particular, the pinned ACS rent artifact trains `with_us_housing_inputs`, + which materializes `pre_subsidy_rent` on ASEC; native ACS `RNTP`/`GRNTP` + remain predictors and are not relabeled as that model input. + Targets produced only by the later PUF pass or source-completion chain are + excluded from this early authority surface. No population operator selects behavior from the source-channel labels. 3. `gap_fill_stacked_spine(...)` runs the two immutable directions over the - same frame: ASEC survey fields fill ACS nulls and ACS housing fills ASEC - nulls. Activation authority is source-and-role exact, observed zero is not - absence, native donor cells must remain byte-identical, and the #608 - per-target banks sit beneath the stack-bound checkpoint identity. + same frame: ASEC survey fields fill ACS nulls, then ASEC-produced housing + rent fills ACS housing-unit nulls in a separately banked direction. + Activation authority is source-and-role exact, observed zero is not + absence, and native donor cells must remain byte-identical. Every declared + target resolves through the operator-output registry to a producer channel + and stage that must strictly precede its direction's check. Only the + explicit `cps_source` and `whole_pool` execution scopes carry authority; + unknown scopes fail rather than inheriting whole-pool authority. ACS TYPEHUGQ + 2/3 people remain null only when their live mask matches the assembly-bound + native group-quarters lineage in every clone role and the declared + structural-absence rule; relabeling a housing unit later cannot create + absence authority. Filling those rows with zero or a donor housing value + would synthesize an unobserved housing unit. Every other early target must + finish with zero `unmodeled_rows` and zero residual nulls: transfer + accounting alone is not terminal absence authority. The #608 per-target + banks sit beneath the stack-bound checkpoint identity. 4. `run_stacked_puf_pass(...)` attaches the separately controlled PUF clone arm (`clone_attachment_fraction`, default `1.0`) and runs one primary QRF - pass across both survey origins. PUF donors stay full. The clone-2 - capital-gains-tail operator runs inside this pass; exact tail-owned and - QRF-owned cells are checked after source completion and every later phase. -5. The transferred checkpoint records the gap-fill banks, primary-QRF bank, - tail manifest, weights audit, stack-manifest digest, fraction/seed, and - clone controls. The same identity regime governs cold and resumed builds. -6. Schedule-D preparation, deterministic derivation, seeded inputs, and - batched simulation run on the transferred stack. The tool retains the - existing `assembled`, `transferred`, and `simulated` #599 boundaries. -7. A fresh `us_stacked_completeness` gate proves every declared input is + pass across both survey origins. The strict recipient surface applies the + [2024 ACS PUMS Data Dictionary](https://www2.census.gov/programs-surveys/acs/tech_docs/pums/data_dict/PUMS_Data_Dictionary_2024.pdf) + universes for `WAGP` and `SEMP`. The ASEC evidence is the established + producer, not a new convention: [`derive_us_cps_carried_inputs`](../packages/microcosm-build/src/microcosm/build/us_runtime/cps_carried.py#L150-L155) + maps `WSAL_VAL` and `SEMP_VAL`, while its [`_source`](../packages/microcosm-build/src/microcosm/build/us_runtime/cps_carried.py#L355-L362) + coercion materializes source blanks as numeric zero. A direct audit of the + certified Build J artifact (`populace-us-2024-buildj-sparse-rmloss100-75d5add-20260710T094201Z`, + built with PolicyEngine US 1.764.6) found all 15,509 under-15 ASEC people + numeric zero on raw and mapped earnings and retained all seven all-child + ASEC tax units with zero earnings sums. The PUF-detail ASEC clone likewise + has 21,737 under-15 zeros and 22 retained all-child units. Cross-arm + comparison therefore requires the ACS produced frame to use the same + semantics: a + named pre-QRF operator materializes `0.0` only on declared under-15 mapped + leaves, records the exact person/unit counts and the + `acs_2024_pums_wagp_age_15_plus` and + `acs_2024_pums_semp_age_15_plus` rule IDs in its receipt, and leaves raw + `WAGP`/`SEMP` blanks unchanged. Eligible mapped or raw nulls still fail in + [`_require_complete_recipient_predictor_sources`](../packages/microcosm-build/src/microcosm/build/us_runtime/puf_support.py#L3036-L3167) + with the original greppable `missing values before coercion` diagnostic and + the responsible rule ID. All-child units remain + recipients with receipted zero earnings predictors; they are not generic + `fillna` results. PUF earnings allocation is limited to age-15-plus people, + so an under-15 first person and every member of an all-child unit remain + zero even when the unit prediction is positive. The mapped leaves' exact + universe zeros must agree with the raw `WAGP`/`SEMP` authority columns, + which are mandatory on scoped ACS rows. A + same-named tax-unit/person source + collision is ambiguous and fails instead of letting receipt and feature + construction choose different grains. + The full ACS 2024 archive audit found exactly 510,098 under-15 people, equal + to the complete WAGP/SEMP blank set. At the 1% rung the original failure set + was exactly 2,998 recipient tax units touching 5,294 under-15 people: 2,980 + mixed-age units and 18 units with no eligible member. These populations are + counted by the universe receipt; the latter 18 remain zero-basis recipients + to match the certified ASEC arm. + Exact structural-person, affected-unit, mixed-unit, and empty-unit counts + and lineage/feature digests bind the root QRF manifest, both immutable + banks, every target checkpoint, live finalization, and the outer stacked + receipt. PUF donors stay full. The clone-2 capital-gains-tail operator runs + inside this pass; exact tail-owned and QRF-owned cells are checked after + source completion and every later phase. + This semantic change is authority-gated: the primary-QRF root and target + checkpoint schema, outer stacked checkpoint materializer, and canonical + stacked authority are all version 6. The outer base identity binds the + primary-QRF schema plus the ACS universe and QBI reconciliation contract + identities. Every v1--v5 payload is stale and refused, including the former + strict v5 two-control payload. +5. The post-clone source-completion chain runs, then the declared post-PUF + transfer fills the targets first materialized by that chain or the PUF pass. + Its complete model donor is the ASEC-origin PUF-detail role. Authority is + target-specific: every live positive-index clone must already observe a + PUF-produced target, every ASEC-origin clone must already observe a + source-produced target, and dual-produced targets require the union. A null + on any such producer row is terminal; only the complementary recipient + rows may be filled from QRF predictions. No blanket null-to-zero synthesis + occurs, every producer cell stays byte-identical, and zero residual nulls + are required. +6. The transferred checkpoint records the early gap-fill banks, post-PUF + transfer bank, primary-QRF bank, tail manifest, weights audit, + stack-manifest digest, fraction/seed, clone controls, and the channel-aware + producer-precedence schedule. The same identity regime governs cold and + resumed builds. Checkpoint emission, resume, and final publication each + reject the post-PUF receipt unless it carries the exact canonical stacked + authority; NON-CANONICAL test receipts cannot ship. +7. Schedule-D preparation, deterministic derivation, seeded inputs, and + batched simulation run on the transferred stack. QBI reconciliation uses + the same source declaration: it fails on any in-universe self-employment + null and preserves the receipted ACS under-15 base self-employment zero in + every clone role. That narrow source exception does not suppress independent QBI + identities: every QBI detail row is still reconciled and checked. The + operator declares the base self-employment rewrite alongside its QBI + outputs and receipts a whole-person input-table digest, a declared-output + digest, changed-row counts, exact preservation of every undeclared person + column, entity, link, weight, stratum, mass-log, and metadata surface, an + exact structural-source exclusion count, and zero forbidden + structural-source mutations. Receipt generation recomputes the complete + deterministic before/after transition and validates exact keys, types, + nonnegative counts, nested universe receipts, preservation claims, and + every SHA-256. Only after that succeeds does the derive boundary bind the + canonical receipt SHA and its preimage/output digests into deeply immutable + frame metadata and a separately carried transition-authority field. + Checkpoint H5 metadata and its sidecar content-bind that independent field. + Simulated checkpoint emission, durable checkpoint write and load, simulated + resume, legacy and stacked manifest construction, and both publication + entry points require the live receipt SHA to match the carried authority. + They also recompute the declared-output and QBI-driver digests, exact live + universe receipt, kernel fixed point, preservation-count equations, and + exact person-column inventory. The only later person additions permitted by + that inventory are outputs named by the checked-in take-up contract and the + seed-stage receipt; an arbitrary receipt cannot whitelist a column. Legacy + receipts must use + `derive.qbi_input_reconciliation`; stacked receipts must use + `derive.pool_derivation.qbi_input_reconciliation`; missing, wrong-route, or + ambiguous receipts fail. A forged envelope, a laundered input inventory, + and even a self-consistent alternate fixed point paired with a freshly + generated receipt therefore fail rather than becoming publication + authority. The tool + retains the existing `assembled`, `transferred`, and `simulated` #599 + boundaries. +8. A fresh `us_stacked_completeness` gate proves every declared input is observed or has exact source-by-role absence authority. The terminal `us_by_origin_battery` then evaluates all 131 declared targets (114 person, 9 tax-unit, 8 SPM-unit), plus joint immigration structure, using an immutable live-digested per-column metric registry. Metric choice never - dispatches from physical dtype. At small rungs, comparisons outside the - validity domain receipt `insufficient_support`; tolerances do not widen. -8. Only after both gates run does publication write the nullable H5, + dispatches from physical dtype. A digest-bound structural-absence rule may + remove only its exact proven cells from a comparison's applicability scope; + any additional null or filled structural cell is terminal. Manifest + emission revalidates the exact structural-rule schema, row arithmetic, + per-role proofs, and battery exclusion count from the immutable gate + snapshot, so authority metadata cannot be grafted onto invented absence. + At small rungs, comparisons outside the validity domain receipt + `insufficient_support`; tolerances do not widen. +9. Only after both gates run does publication write the nullable H5, diagnostics, and readiness manifest. Success, failed gate, and exception paths each append a durable Logbook spool row beside the output, with the fraction token, seed, code/input/identity pins, phases, gate-receipt pointers, wall time, artifact location, and disposition. +### Downstream hard-completeness audit + +This table makes the stacked 1% supplier and starvation behavior explicit at +every remaining boundary. An early `unmodeled_rows` receipt is merely an +accounting result; `insufficient_support` is a later battery status reached +only after a comparison surface is complete and valid. + +| Boundary | Hard requirement | Stacked 1% supplier | Can an upstream insufficient-support/unmodeled state starve it? | +|---|---|---|---| +| Early gap-fill handoff | Donors observe every declared target; every recipient null is filled except the exact ACS group-quarters rent rule. Nonstructural `unmodeled_rows` and residual nulls are forbidden. | The pre-clone ASEC source operators supply 48 early targets to the two ASEC-to-ACS directions. | No accepted starvation remains. A nonstructural residual fails before cloning; literal `insufficient_support` is not an early-transfer outcome. | +| Clone attachment | Input rows are all clone 0; the seeded whole-household selection, lineage, pair weights, fraction, and seed agree exactly. | The completed gap-filled stack and the attachment sampler. | A permitted structural rent null is copied with its authority. Any other early residual has already failed. | +| PUF raw predictor sources | Every filing-status, count, and income component is observed in its declared source universe. Raw WAGP/SEMP authority is present and agrees with mapped leaves; a cross-grain source collision is rejected. A null on any eligible member fails before coercion. | Structure supplies status/count; ACS-native or ASEC-carried earnings supply earnings; early transfer supplies interest, dividends, and gains. | No. ACS under-15 WAGP/SEMP blanks are an exact source-universe state, not transfer starvation; all other source nulls fail. | +| PUF tax-unit features | Every clone-1 recipient has a finite feature vector. Post-aggregation NaN, `+inf`, and `-inf` are counted by named predictor and rejected before fitting; none is coerced or snapped to zero. | Universe-aware person sums plus tax-unit structural inputs. | No. Eligible member values must be complete; the only special case is an all-child unit whose numeric-zero predictor is explicitly owned and counted by the named universe-zero rule. | +| Primary QRF banks and chain | Donor/recipient banks are immutable; target order and RNG prefix are contiguous; all targets complete; live recipient identity, source-universe receipt, and feature digest match before finalization. | The processed full PUF donor and strict recipient checkpoint initialized above. | No. Mutation or missing receipt invalidates the bank; it cannot resume under legacy semantics. | +| Outer pool checkpoint identity and resume | Schema/materializer/authority v6 plus the ACS-universe and QBI-mutation contract identities must match exactly before any cached stage is discovered. | Fresh input pins, live stack receipt, scale controls, code identity, and both semantic contract identities. | No. Every v1--v5 root, target, materializer, or authority payload is stale; a self-consistent old receipt cannot reopen a checkpoint. | +| Clone-2 capital-gains tail | Candidate recipients have the required filing-status/AGI support, positive donor mass, unique household lineage, and sufficient weight capacity; every selected donor is assigned once. | Completed clone-1 QRF output and full PUF tail donors. | No early residual is accepted. Universe-aware PUF recipients remain eligible, including explicitly receipted empty-universe tax units. | +| Post-clone source completion | Each source operator preserves structure and emits its declared ASEC-evidenced outputs; unavailable peer cells remain null only until late transfer. | ASEC evidence rows plus completed PUF clone outputs. | Temporarily: peer nulls are intentional here, but the next zero-residual transfer must consume them. | +| Post-PUF transfer | Every declared PUF-clone or ASEC source-producer cell is nonnull; all complementary recipients are filled; the allowed count for both unmodeled and residual rows is zero. | Forty-three PUF and 30 source targets, with three overlaps, supply the 70-target late surface. | No. A missing producer or recipient value is terminal at this boundary. | +| Fit-weight audit | Every primary and post-PUF QRF fit receipts its resolved entity weight kind, and the collected fit records pass the weights audit before a transferred checkpoint can exist. | Calibrated household weights mapped by the frame to each modeled entity. | No. A missing, inconsistent, or manually substituted weight declaration fails before checkpoint emission. | +| Tail preservation | Tail manifest, descendants, IDs, weights, provenance, joint vector, and non-tail QRF cells remain exact after completion, transfer, derive, seed, and simulation. | The tail manifest bound during the PUF pass. | Completeness receipts cannot authorize a mutation; any byte or identity change fails. | +| Schedule-D derive | Both transferred parent columns are finite for every person and align to every tax unit. | Completed post-PUF transfer plus tail replacements. | No. A residual would fail late transfer first and derive again by name. | +| QBI derive | All QBI detail outputs are finite; self-employment is finite wherever its source applies; every independent archived QBI identity holds. The declared surface includes the base self-employment rewrite and binds pre/post digests. Its exact receipt is recomputed and authenticated at every persisted and publication boundary. | PUF/source detail plus ACS/ASEC native self-employment. Raw under-15 ACS `SEMP` remains structurally blank; mapped `self_employment_income_before_lsr` is a named, receipted universe zero. | No silent starvation. Every mapped ACS under-15 base value is held at its receipted universe zero across clone roles; all derived QBI cells remain in scope, and an in-universe null, forged receipt, or non-kernel output fails. | +| Take-up seed | Every administratively seeded variable completes; transfer-owned take-up cannot use a default; only explicitly non-transfer-owned inputs may use receipted engine defaults. | Seed kernels, the complete transfer surface, and declared defaults. | Transfer-owned residuals fail. A declared default is a separate modeled state, not an insufficient-support receipt. | +| SSI simulation projection | Every nullable engine input has a declared default on the disposable projection; the engine returns exactly one SSI value per person. | The persistent derived/seeded pool plus separately receipted ephemeral defaults. | A projection default can enable simulation but cannot cure the persistent pool; terminal evaluation returns to the original inputs plus SSI. | +| Simulated checkpoint pair and resume | The persistent input-only frame and temporary evaluation frame must share exact assembly provenance; SSI exists only on the evaluation half. The live QBI receipt must authenticate the persistent frame at emission, durable write/load, and resume. | Derived/seeded persistent inputs plus the separately materialized SSI evaluation output. | No. A forged QBI receipt, altered persistent value, invalid SSI binding, or mismatched pair invalidates the simulated checkpoint and falls back only to an independently valid earlier stage. | +| Terminal completeness | All 131 registered targets exist; every positive-weight value is metric-valid; a null needs exact source/role authority, and post-PUF targets forbid absence authority. | The 48 early targets, 70 late targets, derived leaves, take-up inputs, and SSI output. | No. Only the canonical group-quarters rent rule reaches this gate as null; base WAGP/SEMP leaves are outside the 131-target terminal surface. | +| By-origin battery | All 131 clone-0 comparison surfaces are complete and valid before support is measured. | The terminal simulation frame, comparing ASEC and ACS native origins. | No. `insufficient_support` is assigned only after null and validity checks, so it cannot hide an upstream missing value. | +| Manifest construction and canonical publication closure | Legacy and stacked builders reauthenticate QBI live output, canonical stacked authority, terminal-gate receipts, H5/diagnostics run IDs, and artifact digests before readiness can be asserted. | The validated persistent pool, immutable stage receipts, terminal gate snapshot, and atomically staged publication files. | No. Construction rejects forged or wrong-route receipts; publication begins with a non-ready tombstone, and only one fully authenticated run can replace it with a ready manifest. | + +The audit leaves no generic “receipted but null” path into a hard consumer. +Structural absence is target- and universe-exact; sample-size support affects +only whether an otherwise complete terminal comparison is testable. + ### Retiring `--legacy-two-spine` sequence The explicit compatibility flag preserves the previous assemble-first pool @@ -286,6 +438,7 @@ source ingestion and faithful schema harmonization -> banked cross-origin gap-fill -> one PUF QRF pass plus clone-2 capital-gains tail -> source completion + -> banked post-PUF transfer of newly materialized targets -> derive -> seed take-up and other stochastic inputs -> simulate @@ -396,8 +549,10 @@ The completeness gate and by-origin battery run after simulation and before publication or calibration. Their authority bundle freezes the complete declared surface, gap-fill plan, per-column metric registry, joint metrics, and support profile; live content digests are recomputed at evaluation and -manifest emission. Production entrypoints accept no caller-supplied surface, -metric, or tolerance authority. +manifest emission. Each canonical result is also sealed to the exact evidence +snapshot minted by its evaluator, so even another internally valid canonical +receipt cannot be grafted onto it. Production entrypoints accept no +caller-supplied surface, metric, or tolerance authority. Each declared comparison uses positive record weights and exactly one named metric: boolean/rare incidence, monetary sign-separated incidence plus @@ -408,6 +563,21 @@ the immutable effective-support floor are explicitly untestable and receipt `insufficient_support`; they do not silently pass as tested and do not alter the production tolerances. +The one canonical applicability exception is ACS group-quarters rent. Its +rule is part of the gap-fill-plan digest. Assembly separately binds the native +sampled TYPEHUGQ 2/3 household and person lineage counts and SHA-256 digests, +plus complete ordered native mappings of household support/raw/classification +and person support/raw/household-parent/classification identity. Each later +boundary proves the live native set and full native mapping, then expands those +mappings to every clone and proves exact person coverage of each clone's +receipted household selection. Clone roles are lifecycle-exact in every entity: +unattached `{0}`, full or partial PUF detail `{0, 1}`, or tail-descendant +`{0, 1, 2}` backed by its fully validated attachment receipt. The +`pre_subsidy_rent` null mask must equal those linked people across every clone +role and receipts those cells as recipient-exact structural absence. +Completeness proves the exact mask; the battery excludes only the same clone-0 +cells before applying the unchanged support floor and tolerances. + Both gates return batched `GateResult` failures. Calibration cannot consume a frame whose completeness or by-origin gate failed, and no source-specific target, loss term, seed, inferred dtype metric, or caller threshold may shape diff --git a/packages/microcosm-build/src/microcosm/build/gates.py b/packages/microcosm-build/src/microcosm/build/gates.py index 4860a82a..c181d21f 100644 --- a/packages/microcosm-build/src/microcosm/build/gates.py +++ b/packages/microcosm-build/src/microcosm/build/gates.py @@ -126,6 +126,12 @@ class GateResult: passed: bool failures: tuple[str, ...] = () details: Mapping[str, object] = field(default_factory=dict) + _stacked_authority_seal: bytes | None = field( + init=False, + default=None, + repr=False, + compare=False, + ) _details_protocol_5_snapshot: bytes | None = field( init=False, repr=False, @@ -148,6 +154,29 @@ def __post_init__(self) -> None: ) +def _sealed_stacked_gate_result( + *, + name: str, + passed: bool, + failures: tuple[str, ...], + details: Mapping[str, object], +) -> GateResult: + """Mint a stacked-gate result whose evaluated evidence cannot be replaced.""" + + seal = pickle.dumps( + (name, passed, failures, dict(details)), + protocol=5, + ) + result = GateResult( + name=name, + passed=passed, + failures=failures, + details=details, + ) + object.__setattr__(result, "_stacked_authority_seal", seal) + return result + + @dataclass(frozen=True) class GateReport: """The full acceptance suite: every gate, one verdict. @@ -248,7 +277,25 @@ def to_manifest(self) -> dict[str, object]: _validate_stacked_gate_manifest_details, ) - _validate_stacked_gate_manifest_details(result.name, result.details) + _validate_stacked_gate_manifest_details( + result.name, + result.details, + passed=result.passed, + ) + expected_seal = pickle.dumps( + ( + result.name, + result.passed, + result.failures, + dict(result.details), + ), + protocol=5, + ) + if result._stacked_authority_seal != expected_seal: + raise ValueError( + f"Gate {result.name!r} was not sealed by its evaluator; " + "production manifest emission is forbidden." + ) return { "passed": self.passed, "gates": { diff --git a/packages/microcosm-build/src/microcosm/build/us_runtime/acs_income_universe.py b/packages/microcosm-build/src/microcosm/build/us_runtime/acs_income_universe.py new file mode 100644 index 00000000..44b5730d --- /dev/null +++ b/packages/microcosm-build/src/microcosm/build/us_runtime/acs_income_universe.py @@ -0,0 +1,466 @@ +"""Declared ACS PUMS earnings-source universes for stacked builds. + +The 2024 ACS PUMS ``WAGP`` and ``SEMP`` fields apply to people age 15 and +older. Census encodes younger people as not applicable (blank), while the +certified ASEC arm produces explicit zero earnings for the same age universe. +The stacked ACS arm therefore materializes mapped universe zeros under named, +digest-bound rules while preserving the raw PUMS blanks as source authority. +""" + +from __future__ import annotations + +import hashlib +import json +from collections.abc import Mapping, Sequence +from dataclasses import dataclass +from types import MappingProxyType + +import numpy as np +import pandas as pd + +from microcosm.build.us_runtime.support_provenance import ( + support_channel_column, + support_clone_index_column, + support_source_id_column, +) +from microcosm.frame import US_SCHEMA, Frame + +__all__ = [ + "ACS_PUMS_2024_DATA_DICTIONARY_URL", + "ACS_PUMS_EARNINGS_MINIMUM_AGE", + "ACS_PUMS_EARNINGS_SOURCE_COLUMNS", + "AcsPumsEarningsUniverse", + "AcsPumsEarningsUniverseApplication", + "acs_pums_earnings_universe_contract_identity", + "apply_acs_pums_earnings_universe_zeros", + "resolve_acs_pums_earnings_universe", +] + +ACS_PUMS_2024_DATA_DICTIONARY_URL = ( + "https://www2.census.gov/programs-surveys/acs/tech_docs/pums/" + "data_dict/PUMS_Data_Dictionary_2024.pdf" +) +ACS_PUMS_EARNINGS_MINIMUM_AGE = 15 +ACS_PUMS_EARNINGS_SOURCE_COLUMNS: Mapping[str, str] = MappingProxyType( + { + "employment_income_before_lsr": "WAGP", + "self_employment_income_before_lsr": "SEMP", + } +) + +_ACS_SUPPORT_CHANNEL = "acs" +_RULE_VERSION = 1 +_UNIVERSE_DESCRIPTION = "ACS persons age 15 and older" +_AGGREGATION_SEMANTICS = ( + "sum explicit mapped universe-zero values with eligible person values; " + "an empty eligible set remains the receipted numeric total 0.0" +) +_PRODUCED_FRAME_SEMANTICS = ( + "explicit zero below age 15, matching the certified ASEC produced frame" +) + + +@dataclass(frozen=True) +class AcsPumsEarningsUniverse: + """Exact raw-source absence masks plus their JSON-ready receipt.""" + + structural_absence_masks: Mapping[str, pd.Series] + receipt: Mapping[str, object] + + +@dataclass(frozen=True) +class AcsPumsEarningsUniverseApplication: + """A frame with exact mapped universe zeros plus its application receipt.""" + + frame: Frame + receipt: Mapping[str, object] + + +def acs_pums_earnings_universe_contract_identity() -> dict[str, object]: + """Return the complete cache identity for ACS earnings-universe semantics.""" + + rules = [ + { + "rule_id": f"acs_2024_pums_{source.lower()}_age_15_plus", + "source_column": source, + "mapped_column": mapped, + } + for mapped, source in ACS_PUMS_EARNINGS_SOURCE_COLUMNS.items() + ] + body: dict[str, object] = { + "version": _RULE_VERSION, + "source_dataset": "2024 ACS 1-year PUMS", + "source_channel": _ACS_SUPPORT_CHANNEL, + "minimum_age": ACS_PUMS_EARNINGS_MINIMUM_AGE, + "rules": rules, + "produced_frame_semantics": _PRODUCED_FRAME_SEMANTICS, + "out_of_universe_person_policy": ( + "materialize mapped zero only through the named universe rule; " + "preserve the raw PUMS blank" + ), + "eligible_person_null_policy": ( + "fail at recipient predictor source preflight before coercion" + ), + "empty_universe_tax_unit_policy": ( + "retain the recipient with receipted zero earnings predictors" + ), + "person_allocation_policy": ( + "allocate earnings only to age-eligible people; never use an " + "out-of-universe first-person fallback" + ), + "aggregation": _AGGREGATION_SEMANTICS, + } + return {**body, "sha256": _mapping_sha256(body)} + + +def apply_acs_pums_earnings_universe_zeros( + frame: Frame, + *, + columns: Sequence[str] = tuple(ACS_PUMS_EARNINGS_SOURCE_COLUMNS), + person_scope: Sequence[bool] | pd.Series | np.ndarray | None = None, + boundary: str, +) -> AcsPumsEarningsUniverseApplication: + """Materialize mapped ACS earnings zeros only below the declared age floor. + + Raw ``WAGP``/``SEMP`` blanks remain untouched. Every mapped cell in the + structural scope must still be null at this producer boundary: even a + pre-existing zero is refused because it lacks this operator's receipt. + Eligible nulls deliberately remain null so the primary-QRF source preflight + retains its original diagnostic. + """ + + if frame.schema != US_SCHEMA: + raise ValueError(f"{boundary}: ACS earnings universes require the US schema.") + requested = tuple(dict.fromkeys(columns)) + unsupported = sorted(set(requested) - set(ACS_PUMS_EARNINGS_SOURCE_COLUMNS)) + if unsupported: + raise ValueError( + f"{boundary}: unsupported ACS earnings-universe column(s) {unsupported}." + ) + + person = frame.table("person") + required = { + "age", + support_channel_column("person"), + *requested, + *(ACS_PUMS_EARNINGS_SOURCE_COLUMNS[column] for column in requested), + } + missing = sorted(required - set(person.columns)) + if missing: + raise ValueError( + f"{boundary}: ACS earnings universe cannot materialize required person " + f"column(s) {missing}." + ) + scope = _aligned_scope(person, person_scope, boundary=boundary) + channel = person[support_channel_column("person")].astype(str) + acs_scope = scope & channel.eq(_ACS_SUPPORT_CHANNEL) + age = pd.to_numeric(person["age"], errors="coerce") + invalid_age = acs_scope & (~np.isfinite(age.to_numpy(dtype=np.float64))) + if invalid_age.any(): + raise ValueError( + f"{boundary}: ACS earnings universe has {int(invalid_age.sum())} " + "missing or nonfinite age value(s)." + ) + structural = acs_scope & age.lt(ACS_PUMS_EARNINGS_MINIMUM_AGE) + + failures: list[str] = [] + for column in requested: + source_column = ACS_PUMS_EARNINGS_SOURCE_COLUMNS[column] + rule_id = f"acs_2024_pums_{source_column.lower()}_age_15_plus" + mapped = person[column] + mapped_preexisting = structural & mapped.notna() + raw_nonblank = structural & person[source_column].notna() + if mapped_preexisting.any() or raw_nonblank.any(): + failures.append( + f"{rule_id}: unreceipted_preexisting_mapped_rows=" + f"{int(mapped_preexisting.sum())}, raw_source_nonblank_rows=" + f"{int(raw_nonblank.sum())}" + ) + if failures: + raise ValueError( + f"{boundary}: ACS earnings universe-zero application failed:\n " + + "\n ".join(failures) + ) + + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + output_person = tables["person"] + for column in requested: + output_person.loc[structural, column] = 0.0 + applied = Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + metadata=frame.metadata, + ) + resolved = resolve_acs_pums_earnings_universe( + applied, + columns=requested, + person_scope=scope, + boundary=boundary, + ) + return AcsPumsEarningsUniverseApplication( + frame=applied, + receipt=resolved.receipt, + ) + + +def resolve_acs_pums_earnings_universe( + frame: Frame, + *, + columns: Sequence[str], + person_scope: Sequence[bool] | pd.Series | np.ndarray | None = None, + boundary: str, +) -> AcsPumsEarningsUniverse: + """Resolve and validate the exact age-based ACS earnings universe. + + ``person_scope`` identifies the rows whose source role is being consumed. + Primary-QRF preparation passes its clone-recipient rows; post-PUF QBI + reconciliation passes the live produced frame. The function never mutates + the person table. Raw PUMS values must be blank and mapped values must be + explicit zero below age 15. Eligible nulls are counted but deliberately + deferred to the caller's source-preflight boundary. + """ + + if frame.schema != US_SCHEMA: + raise ValueError(f"{boundary}: ACS earnings universes require the US schema.") + requested = tuple(dict.fromkeys(columns)) + unsupported = sorted(set(requested) - set(ACS_PUMS_EARNINGS_SOURCE_COLUMNS)) + if unsupported: + raise ValueError( + f"{boundary}: unsupported ACS earnings-universe column(s) {unsupported}." + ) + + person = frame.table("person") + required = { + "age", + "person_tax_unit_id", + support_channel_column("person"), + support_clone_index_column("person"), + *requested, + } + missing = sorted(required - set(person.columns)) + if missing: + raise ValueError( + f"{boundary}: ACS earnings universe cannot resolve required person " + f"column(s) {missing}." + ) + + scope = _aligned_scope(person, person_scope, boundary=boundary) + channel = person[support_channel_column("person")].astype(str) + acs_scope = scope & channel.eq(_ACS_SUPPORT_CHANNEL) + age = pd.to_numeric(person["age"], errors="coerce") + invalid_age = acs_scope & (~np.isfinite(age.to_numpy(dtype=np.float64))) + if invalid_age.any(): + raise ValueError( + f"{boundary}: ACS earnings universe has {int(invalid_age.sum())} " + "missing or nonfinite age value(s)." + ) + structurally_absent = acs_scope & age.lt(ACS_PUMS_EARNINGS_MINIMUM_AGE) + eligible_acs = acs_scope & ~structurally_absent + + masks: dict[str, pd.Series] = {} + rule_receipts: dict[str, dict[str, object]] = {} + failures: list[str] = [] + for column in requested: + values = person[column] + null = values.isna() + numeric = pd.to_numeric(values, errors="coerce") + in_universe_null = int((eligible_acs & null).sum()) + source_column = ACS_PUMS_EARNINGS_SOURCE_COLUMNS[column] + rule_id = f"acs_2024_pums_{source_column.lower()}_age_15_plus" + universe_zero_missing = int((structurally_absent & null).sum()) + out_of_universe_nonzero = int( + (structurally_absent & ~null & ~numeric.eq(0.0)).sum() + ) + mapped_universe_zero_rows = int( + (structurally_absent & ~null & numeric.eq(0.0)).sum() + ) + raw_source_present = source_column in person.columns + raw_in_universe_null_rows = 0 + raw_out_of_universe_nonblank_rows = 0 + if not raw_source_present and acs_scope.any(): + failures.append( + f"{rule_id}: {source_column} " + "raw_source_authority_missing_for_scoped_acs_rows" + ) + elif raw_source_present: + raw_null = person[source_column].isna() + raw_in_universe_null_rows = int((eligible_acs & raw_null).sum()) + raw_out_of_universe_nonblank_rows = int( + (structurally_absent & ~raw_null).sum() + ) + if ( + universe_zero_missing + or out_of_universe_nonzero + or raw_out_of_universe_nonblank_rows + ): + failures.append( + f"{rule_id}: universe_zero_missing_rows={universe_zero_missing}, " + "out_of_universe_mapped_nonzero_rows=" + f"{out_of_universe_nonzero}, raw_source_nonblank_rows=" + f"{raw_out_of_universe_nonblank_rows}" + ) + masks[column] = structurally_absent.copy() + rule_receipts[column] = { + "rule_id": rule_id, + "source_column": source_column, + "mapped_column": column, + "source_channel": _ACS_SUPPORT_CHANNEL, + "source_universe": _UNIVERSE_DESCRIPTION, + "produced_frame_semantics": _PRODUCED_FRAME_SEMANTICS, + "structurally_absent_person_rows": int(structurally_absent.sum()), + "eligible_acs_person_rows": int(eligible_acs.sum()), + "in_universe_null_rows": in_universe_null, + "raw_in_universe_null_rows": raw_in_universe_null_rows, + "mapped_universe_zero_rows": mapped_universe_zero_rows, + "universe_zero_missing_rows": universe_zero_missing, + "out_of_universe_mapped_nonzero_rows": out_of_universe_nonzero, + "raw_source_column_present": raw_source_present, + "raw_source_nonblank_rows": raw_out_of_universe_nonblank_rows, + "source_cells_sha256": ( + _source_cells_sha256( + person, + column=column, + raw_source_column=source_column, + scope=scope, + ) + if raw_source_present + else None + ), + } + if failures: + raise ValueError( + f"{boundary}: ACS earnings source-universe equation failed:\n " + + "\n ".join(failures) + ) + + affected_ids = person.loc[structurally_absent, "person_tax_unit_id"] + affected_unique = pd.Index(affected_ids.drop_duplicates()) + acs_scoped_people = person.loc[acs_scope, ["person_tax_unit_id"]].copy() + acs_scoped_people["eligible"] = eligible_acs.loc[acs_scope].to_numpy() + eligible_by_unit = acs_scoped_people.groupby("person_tax_unit_id", sort=False)[ + "eligible" + ].any() + empty_ids = pd.Index(eligible_by_unit.index[~eligible_by_unit.to_numpy()]) + mixed_ids = affected_unique.difference(empty_ids, sort=False) + lineage_column = support_source_id_column("person") + if lineage_column not in person: + lineage_column = frame.schema.entity_id_column("person") + structural_lineages = person.loc[ + structurally_absent, lineage_column + ].drop_duplicates() + clone_index = pd.to_numeric( + person[support_clone_index_column("person")], errors="raise" + ).astype("int64") + by_origin_role = { + f"{origin}/clone_{int(role)}": int(count) + for (origin, role), count in ( + pd.DataFrame( + { + "origin": channel.loc[structurally_absent], + "clone_index": clone_index.loc[structurally_absent], + } + ) + .groupby(["origin", "clone_index"], sort=True) + .size() + .items() + ) + } + receipt: dict[str, object] = { + "version": _RULE_VERSION, + "policy": "asec_consistent_receipted_universe_zero", + "source_dataset": "2024 ACS 1-year PUMS", + "source_document": "2024 ACS PUMS Data Dictionary (WAGP and SEMP)", + "source_url": ACS_PUMS_2024_DATA_DICTIONARY_URL, + "source_channel": _ACS_SUPPORT_CHANNEL, + "age_column": "age (mapped from AGEP)", + "minimum_age": ACS_PUMS_EARNINGS_MINIMUM_AGE, + "aggregation": _AGGREGATION_SEMANTICS, + "produced_frame_semantics": _PRODUCED_FRAME_SEMANTICS, + "raw_pums_source_cells_mutated": False, + "mapped_person_cells_materialized": bool( + requested and structurally_absent.any() + ), + "mapped_universe_zero_cells": int(len(requested) * structurally_absent.sum()), + "scoped_person_rows": int(scope.sum()), + "scoped_acs_person_rows": int(acs_scope.sum()), + "structurally_absent_person_rows": int(structurally_absent.sum()), + "affected_tax_unit_rows": int(len(affected_unique)), + "mixed_universe_tax_unit_rows": int(len(mixed_ids)), + "empty_universe_tax_unit_rows": int(len(empty_ids)), + "structurally_absent_person_lineages_sha256": _values_sha256( + structural_lineages + ), + "affected_tax_unit_ids_sha256": _values_sha256(affected_unique), + "empty_universe_tax_unit_ids_sha256": _values_sha256(empty_ids), + "by_origin_role": by_origin_role, + "rules": rule_receipts, + } + receipt["sha256"] = _mapping_sha256(receipt) + return AcsPumsEarningsUniverse( + structural_absence_masks=MappingProxyType(masks), + receipt=MappingProxyType(receipt), + ) + + +def _aligned_scope( + person: pd.DataFrame, + scope: Sequence[bool] | pd.Series | np.ndarray | None, + *, + boundary: str, +) -> pd.Series: + if scope is None: + return pd.Series(True, index=person.index, dtype=bool) + if isinstance(scope, pd.Series): + if not scope.index.equals(person.index): + raise ValueError(f"{boundary}: person scope index is not aligned.") + result = scope.astype(bool) + else: + values = np.asarray(scope) + if values.ndim != 1 or len(values) != len(person): + raise ValueError(f"{boundary}: person scope has the wrong shape.") + result = pd.Series(values, index=person.index, dtype=bool) + return result + + +def _source_cells_sha256( + person: pd.DataFrame, + *, + column: str, + raw_source_column: str, + scope: pd.Series, +) -> str: + lineage = support_source_id_column("person") + columns = ["person_tax_unit_id", column, raw_source_column] + if lineage in person: + columns.insert(0, lineage) + cells = person.loc[scope, columns] + header = { + "columns": columns, + "dtypes": [str(cells[name].dtype) for name in columns], + } + digest = hashlib.sha256(_canonical_json(header).encode()) + digest.update( + pd.util.hash_pandas_object(cells, index=False).to_numpy(dtype=" str: + normalized = sorted(str(value) for value in list(values)) + return hashlib.sha256(_canonical_json(normalized).encode()).hexdigest() + + +def _mapping_sha256(value: Mapping[str, object]) -> str: + return hashlib.sha256(_canonical_json(value).encode()).hexdigest() + + +def _canonical_json(value: object) -> str: + return json.dumps( + value, + sort_keys=True, + separators=(",", ":"), + allow_nan=False, + ) diff --git a/packages/microcosm-build/src/microcosm/build/us_runtime/multispine_pool.py b/packages/microcosm-build/src/microcosm/build/us_runtime/multispine_pool.py index 4653211e..4ffbdb0b 100644 --- a/packages/microcosm-build/src/microcosm/build/us_runtime/multispine_pool.py +++ b/packages/microcosm-build/src/microcosm/build/us_runtime/multispine_pool.py @@ -66,7 +66,11 @@ from microcosm.build.us_runtime.puf_qrf_chain import PRIMARY_QRF_TARGET_ORDER from microcosm.build.us_runtime.puf_support import clone_us_frame_for_puf_support from microcosm.build.us_runtime.qbi_inputs import ( - US_QBI_OUTPUT_COLUMNS, + bind_us_qbi_reconciliation_transition_authority, + us_qbi_post_reconciliation_person_columns, + us_qbi_reconciliation_change_receipt, + validate_us_qbi_reconciliation_live_output, + validate_us_qbi_reconciliation_transition, with_us_qbi_input_reconciliation, ) from microcosm.build.us_runtime.relationship_inputs import ( @@ -127,8 +131,12 @@ "derive_multispine_pool_inputs", "materialize_multispine_agreement_outputs", "materialize_pool_deferred_transfer_inputs", - "pool_transfer_target_families", "pool_input_surface", + "pool_post_puf_puf_producer_target_families", + "pool_post_puf_source_producer_target_families", + "pool_post_puf_transfer_target_families", + "pool_pre_clone_gap_fill_target_families", + "pool_transfer_target_families", "prepare_multispine_puf_predictors", "prepare_multispine_source_inputs_for_clone", "run_multispine_pool_path", @@ -237,6 +245,7 @@ class PoolStageOutput: frame: Frame receipt: Mapping[str, object] = field(default_factory=dict) + qbi_transition_authority_sha256: str | None = None def __post_init__(self) -> None: if not isinstance(self.frame, Frame): @@ -246,6 +255,14 @@ def __post_init__(self) -> None: ) if not isinstance(self.receipt, Mapping): raise TypeError("PoolStageOutput.receipt must be a mapping.") + if self.qbi_transition_authority_sha256 is not None and not isinstance( + self.qbi_transition_authority_sha256, + str, + ): + raise TypeError( + "PoolStageOutput.qbi_transition_authority_sha256 must be a " + "string when present." + ) @dataclass(frozen=True) @@ -262,6 +279,7 @@ class MultispinePoolCheckpoint: assembly_receipt: Mapping[str, object] stage_receipts: Mapping[str, Mapping[str, object]] simulation_frame: Frame | None = None + qbi_transition_authority_sha256: str | None = None def __post_init__(self) -> None: if self.stage not in POOL_CHECKPOINT_STAGE_ORDER: @@ -297,6 +315,14 @@ def __post_init__(self) -> None: "Only a simulated multispine pool checkpoint may carry a " "simulation_frame." ) + if self.qbi_transition_authority_sha256 is not None and not isinstance( + self.qbi_transition_authority_sha256, + str, + ): + raise TypeError( + "MultispinePoolCheckpoint.qbi_transition_authority_sha256 must " + "be a string when present." + ) @dataclass(frozen=True) @@ -308,6 +334,7 @@ class MultispinePoolResult: provenance_counts: Mapping[str, Mapping[str, object]] stage_receipts: Mapping[str, Mapping[str, object]] agreement_gate: GateResult + qbi_transition_authority_sha256: str | None = None @property def simulation_ready(self) -> bool: @@ -570,6 +597,113 @@ def pool_transfer_target_families() -> TargetFamilies: return plan +def _partition_pool_transfer_target_families( + target_families: TargetFamilies, +) -> tuple[TargetFamilies, TargetFamilies]: + """Split the legacy transfer surface at its declared producer boundary. + + Targets emitted by a pre-clone source operator are eligible for the early + cross-origin gap-fill. Every other target is produced only by the primary + PUF pass or post-clone source completion and therefore belongs to the + post-PUF transfer. The partition preserves the legacy entity, family, and + target order so fit identities stay deterministic. + """ + + pre_clone_outputs = { + (entity, column) + for operator_name in POOL_PRE_CLONE_SOURCE_OPERATOR_ORDER + for entity, columns in PRE_ASSEMBLY_OPERATOR_OUTPUT_FAMILIES[ + POOL_OPERATOR_CONTRACTS[operator_name].family + ].items() + for column in columns + } + early: dict[str, dict[str, tuple[str, ...]]] = {} + late: dict[str, dict[str, tuple[str, ...]]] = {} + for entity, families in target_families.items(): + for family, targets in families.items(): + early_targets = tuple( + target for target in targets if (entity, target) in pre_clone_outputs + ) + late_targets = tuple( + target + for target in targets + if (entity, target) not in pre_clone_outputs + ) + if early_targets: + early.setdefault(entity, {})[family] = early_targets + if late_targets: + late.setdefault(entity, {})[family] = late_targets + return early, late + + +def pool_pre_clone_gap_fill_target_families() -> TargetFamilies: + """Return targets available after native pre-clone source preparation.""" + + early, _late = _partition_pool_transfer_target_families( + pool_transfer_target_families() + ) + return early + + +def pool_post_puf_transfer_target_families() -> TargetFamilies: + """Return targets first available after PUF and source completion.""" + + _early, late = _partition_pool_transfer_target_families( + pool_transfer_target_families() + ) + return late + + +def _filter_target_families_by_outputs( + target_families: TargetFamilies, + outputs: set[tuple[str, str]], +) -> TargetFamilies: + """Preserve declaration order while retaining targets with named producers.""" + + filtered: dict[str, dict[str, tuple[str, ...]]] = {} + for entity, families in target_families.items(): + for family, targets in families.items(): + produced = tuple( + target for target in targets if (entity, target) in outputs + ) + if produced: + filtered.setdefault(entity, {})[family] = produced + return filtered + + +def pool_post_puf_puf_producer_target_families() -> TargetFamilies: + """Return late targets whose producer role is a live PUF clone.""" + + puf_outputs = { + (entity, target) + for entity, targets in PRE_ASSEMBLY_OPERATOR_OUTPUT_FAMILIES[ + "primary_puf_qrf" + ].items() + for target in targets + } + return _filter_target_families_by_outputs( + pool_post_puf_transfer_target_families(), + puf_outputs, + ) + + +def pool_post_puf_source_producer_target_families() -> TargetFamilies: + """Return late targets whose producer role is an ASEC source clone.""" + + source_outputs = { + (entity, target) + for operator_name in POOL_POST_CLONE_SOURCE_OPERATOR_ORDER + for entity, targets in PRE_ASSEMBLY_OPERATOR_OUTPUT_FAMILIES[ + POOL_OPERATOR_CONTRACTS[operator_name].family + ].items() + for target in targets + } + return _filter_target_families_by_outputs( + pool_post_puf_transfer_target_families(), + source_outputs, + ) + + def pool_input_surface() -> tuple[PoolInputSurfaceEntry, ...]: """Return the complete registry-derived pool input/imputation surface. @@ -1454,27 +1588,50 @@ def derive_multispine_pool_inputs(frame: Frame) -> PoolStageOutput: all-or-nothing identities on the imputed PUF-detail surface. """ + def reconcile_qbi_with_receipt(input_frame: Frame) -> PoolStageOutput: + reconciled = with_us_qbi_input_reconciliation(input_frame) + receipt = us_qbi_reconciliation_change_receipt(input_frame, reconciled) + validate_us_qbi_reconciliation_transition( + input_frame, + reconciled, + receipt, + boundary="multispine QBI reconciliation generation", + ) + return PoolStageOutput( + reconciled, + receipt, + ) + completed = _run_source_operator_chain( frame, phase=_POST_CLONE_PHASE, operator_names=POOL_DERIVE_OPERATOR_ORDER, operators={ "_complete_schedule_d_input": _complete_schedule_d_input, - "with_us_qbi_input_reconciliation": with_us_qbi_input_reconciliation, + "with_us_qbi_input_reconciliation": reconcile_qbi_with_receipt, }, ) schedule_d_receipt = completed.receipt["suboperators"][0]["kernel_receipt"] - return PoolStageOutput( + qbi_receipt = completed.receipt["suboperators"][1]["kernel_receipt"] + authorized = bind_us_qbi_reconciliation_transition_authority( completed.frame, + qbi_receipt, + ) + validate_us_qbi_reconciliation_live_output( + authorized, + qbi_receipt, + boundary="multispine pool derivation output", + expected_transition_authority_sha256=qbi_receipt["sha256"], + ) + return PoolStageOutput( + authorized, { "phase": _POST_CLONE_PHASE, "operator_order": list(POOL_DERIVE_OPERATOR_ORDER), "schedule_d_capital_gain_distributions": schedule_d_receipt, - "qbi_input_reconciliation": { - "columns": list(US_QBI_OUTPUT_COLUMNS), - "operation": "shared_all_or_nothing_identity_reconciliation", - }, + "qbi_input_reconciliation": dict(qbi_receipt), }, + qbi_transition_authority_sha256=qbi_receipt["sha256"], ) @@ -1930,7 +2087,57 @@ def _validated_resume_checkpoint( "Multispine pool simulated checkpoint evaluation provenance " "differs from its persistent frame." ) - return assembly_receipt, _checkpoint_stage_receipts(resume) + receipts = _checkpoint_stage_receipts(resume) + if resume.stage == "simulated": + _validate_qbi_stage_receipt( + resume.frame, + receipts, + boundary="multispine pool simulated checkpoint resume", + transition_authority_sha256=(resume.qbi_transition_authority_sha256), + ) + return assembly_receipt, receipts + + +def _qbi_receipt_from_stage_receipts( + stage_receipts: Mapping[str, Mapping[str, object]], + *, + boundary: str, +) -> Mapping[str, object]: + derive = stage_receipts.get("derive") + if not isinstance(derive, Mapping): + raise ValueError(f"{boundary}: stage receipts have no derive object.") + if "pool_derivation" in derive: + raise ValueError( + f"{boundary}: legacy checkpoint used the stacked derive receipt route." + ) + receipt = derive.get("qbi_input_reconciliation") + if not isinstance(receipt, Mapping): + raise ValueError( + f"{boundary}: derive receipt has no QBI reconciliation object." + ) + return receipt + + +def _validate_qbi_stage_receipt( + frame: Frame, + stage_receipts: Mapping[str, Mapping[str, object]], + *, + boundary: str, + transition_authority_sha256: str | None, +) -> None: + receipt = _qbi_receipt_from_stage_receipts( + stage_receipts, + boundary=boundary, + ) + validate_us_qbi_reconciliation_live_output( + frame, + receipt, + boundary=boundary, + expected_transition_authority_sha256=transition_authority_sha256, + allowed_post_reconciliation_person_columns=( + us_qbi_post_reconciliation_person_columns(stage_receipts.get("seed")) + ), + ) def _emit_pool_checkpoint( @@ -1941,9 +2148,17 @@ def _emit_pool_checkpoint( assembly_receipt: Mapping[str, object], stage_receipts: Mapping[str, Mapping[str, object]], simulation_frame: Frame | None = None, + qbi_transition_authority_sha256: str | None = None, ) -> None: if callback is None: return + if stage == "simulated": + _validate_qbi_stage_receipt( + frame, + stage_receipts, + boundary="multispine pool simulated checkpoint emission", + transition_authority_sha256=qbi_transition_authority_sha256, + ) callback( MultispinePoolCheckpoint( stage=stage, @@ -1953,6 +2168,7 @@ def _emit_pool_checkpoint( name: dict(receipt) for name, receipt in stage_receipts.items() }, simulation_frame=simulation_frame, + qbi_transition_authority_sha256=(qbi_transition_authority_sha256), ) ) @@ -2023,6 +2239,7 @@ def run_multispine_pool_path( boundary="multispine pool assembly", ) receipts: dict[str, Mapping[str, object]] = {} + qbi_transition_authority_sha256: str | None = None resume_stage = None _emit_pool_checkpoint( checkpoint, @@ -2037,6 +2254,7 @@ def run_multispine_pool_path( resume.frame, boundary=f"multispine pool {resume.stage} resume", ) + qbi_transition_authority_sha256 = resume.qbi_transition_authority_sha256 resume_stage = resume.stage if resume_stage in {None, "assembled"}: @@ -2120,6 +2338,10 @@ def run_multispine_pool_path( boundary=f"multispine pool {stage_name} output", ) receipts[stage_name] = dict(outcome.receipt) + if stage_name == "derive": + qbi_transition_authority_sha256 = ( + outcome.qbi_transition_authority_sha256 + ) simulated = operators["simulate"](current) if not isinstance(simulated, PoolStageOutput): @@ -2144,6 +2366,7 @@ def run_multispine_pool_path( assembly_receipt=assembly_receipt, stage_receipts=receipts, simulation_frame=simulation_frame, + qbi_transition_authority_sha256=(qbi_transition_authority_sha256), ) else: if resume.simulation_frame is None: # pragma: no cover - dataclass validates @@ -2177,4 +2400,5 @@ def run_multispine_pool_path( provenance_counts=counts, stage_receipts=receipts, agreement_gate=agreement, + qbi_transition_authority_sha256=qbi_transition_authority_sha256, ) diff --git a/packages/microcosm-build/src/microcosm/build/us_runtime/operator_boundary.py b/packages/microcosm-build/src/microcosm/build/us_runtime/operator_boundary.py index 1eb651f5..b5386fc0 100644 --- a/packages/microcosm-build/src/microcosm/build/us_runtime/operator_boundary.py +++ b/packages/microcosm-build/src/microcosm/build/us_runtime/operator_boundary.py @@ -73,7 +73,7 @@ PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS, PUF_TAX_DETAIL_DEFAULT_TAX_UNIT_OUTPUTS, ) -from microcosm.build.us_runtime.qbi_inputs import US_QBI_OUTPUT_COLUMNS +from microcosm.build.us_runtime.qbi_inputs import US_QBI_RECONCILED_PERSON_COLUMNS from microcosm.build.us_runtime.relationship_inputs import ( US_RELATIONSHIP_INPUTS_OUTPUT_COLUMNS, ) @@ -318,7 +318,7 @@ "person": frozenset(ACS_DERIVED_TRANSFER_INPUTS), }, "qbi_reconciliation": { - "person": frozenset(US_QBI_OUTPUT_COLUMNS), + "person": frozenset(US_QBI_RECONCILED_PERSON_COLUMNS), }, "housing_assistance": { "spm_unit": frozenset( diff --git a/packages/microcosm-build/src/microcosm/build/us_runtime/puf_qrf_chain.py b/packages/microcosm-build/src/microcosm/build/us_runtime/puf_qrf_chain.py index 318ca65c..4ea4ffd5 100644 --- a/packages/microcosm-build/src/microcosm/build/us_runtime/puf_qrf_chain.py +++ b/packages/microcosm-build/src/microcosm/build/us_runtime/puf_qrf_chain.py @@ -33,6 +33,7 @@ PufTaxDetailChainInputs, finalize_us_puf_tax_detail_predictions, prepare_us_puf_tax_detail_chain_inputs, + puf_recipient_predictor_universe_receipt, ) from microcosm.build.us_runtime.support_provenance import ( puf_tax_detail_clone_mask, @@ -70,7 +71,13 @@ # A non-legacy initialization writes both controls into each immutable bank, # the root manifest, and every target receipt so deletion or mutation cannot # silently downgrade a stacked chain to the legacy policy. -PRIMARY_QRF_CHECKPOINT_SCHEMA_VERSION = 5 +# Strict banks also bind the exact recipient source-universe/feature receipt; +# its optional v5 field invalidates pre-declaration strict banks without +# changing byte-identical legacy v5 manifests. +# v6 makes that recipient-universe authority a versioned chain semantic. Every +# v1--v5 root or target must rebuild rather than sharing a schema label with a +# chain whose root, banks, targets, and finalization bind the added receipt. +PRIMARY_QRF_CHECKPOINT_SCHEMA_VERSION = 6 PRIMARY_QRF_MANIFEST_FILENAME = "manifest.json" PRIMARY_QRF_DONOR_FILENAME = "donor.frame.h5" PRIMARY_QRF_RECIPIENT_FILENAME = "recipient.frame.h5" @@ -89,6 +96,7 @@ _RAW_DRAW_BITS_DATASET = "raw_draw_bits" _REQUIRE_COMPLETE_RECIPIENT_PREDICTORS = "require_complete_recipient_predictors" _ABSENT_CELLS = "absent_cells" +_RECIPIENT_PREDICTOR_UNIVERSE = "recipient_predictor_universe" _ABSENT_CELLS_POLICIES = ( PUF_ABSENT_CELLS_LEGACY_ZERO_FILL, PUF_ABSENT_CELLS_PRESERVE_NULLS, @@ -101,6 +109,7 @@ "finalize_primary_puf_qrf_chain", "initialize_primary_puf_qrf_chain", "load_primary_puf_qrf_predictions", + "primary_puf_qrf_recipient_predictor_universe_receipt", "run_primary_puf_qrf_chain", "run_primary_puf_qrf_target", ] @@ -123,9 +132,10 @@ def initialize_primary_puf_qrf_chain( The defaults preserve the historical two-spine chain exactly: recipient predictor absence is zero-filled, finalization globally zero-fills absent - outputs, and the two doctrine fields are omitted so legacy v5 bank bytes do - not change. Selecting either non-legacy control declares *both* settings - in every immutable bank, the root manifest, and per-target receipts. + outputs, and the three doctrine fields are omitted so historical legacy-bank + payload bytes do not change. Selecting either non-legacy control declares both settings + plus the recipient source-universe receipt in every immutable bank, the + root manifest, and per-target receipts. """ require_complete_recipient_predictors, absent_cells = ( @@ -135,17 +145,10 @@ def initialize_primary_puf_qrf_chain( source="Primary QRF initialization", ) ) - doctrine_receipt = _declared_chain_doctrine_receipt( - require_complete_recipient_predictors, - absent_cells, - ) - root = Path(checkpoint_dir) manifest_path = root / PRIMARY_QRF_MANIFEST_FILENAME if manifest_path.exists(): raise FileExistsError(f"Primary QRF manifest already exists: {manifest_path}") - root.mkdir(parents=True, exist_ok=True) - (root / PRIMARY_QRF_TARGETS_DIRNAME).mkdir(exist_ok=True) preparation_kwargs: dict[str, object] = {} if require_complete_recipient_predictors: @@ -158,6 +161,11 @@ def initialize_primary_puf_qrf_chain( tax_unit_outputs=tax_unit_outputs, **preparation_kwargs, ) + doctrine_receipt = _declared_chain_doctrine_receipt( + require_complete_recipient_predictors, + absent_cells, + inputs.recipient_predictor_universe, + ) target_order = inputs.target_order if ( tuple(person_outputs) == PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS @@ -166,6 +174,12 @@ def initialize_primary_puf_qrf_chain( ): raise AssertionError("The production primary QRF target order changed.") + # Preflight is intentionally complete before any checkpoint path exists: + # a failed source-universe/completeness contract cannot leave a poisoned, + # nonempty root that looks resumable to the stacked supervisor. + root.mkdir(parents=True, exist_ok=True) + (root / PRIMARY_QRF_TARGETS_DIRNAME).mkdir(exist_ok=True) + donor_frame = canonicalize_frame_string_dtypes( inputs.donor_frame, boundary="primary PUF QRF donor bank write", @@ -408,6 +422,35 @@ def load_primary_puf_qrf_predictions( return predictions +def primary_puf_qrf_recipient_predictor_universe_receipt( + checkpoint_dir: str | Path, +) -> dict[str, object]: + """Load the bank-authenticated strict recipient-universe receipt.""" + + root = Path(checkpoint_dir).resolve() + manifest = _load_manifest(root) + # Authenticate the recipient bank and require its doctrine fields to agree + # before exposing the manifest copy to an outer stacked receipt. + _load_bound_frame( + root, + manifest, + filename_key="recipient_filename", + digest_key="recipient_checkpoint_sha256", + role="recipient", + ) + require_complete, _absent_cells, universe, declared = ( + _resolve_chain_doctrine_controls( + manifest, + source="Primary QRF manifest", + ) + ) + if not declared or not require_complete or not universe: + raise ValueError( + "Primary QRF checkpoint is not a strict recipient-universe bank." + ) + return json.loads(json.dumps(universe)) + + def finalize_primary_puf_qrf_chain( frame: Frame, checkpoint_dir: str | Path, @@ -419,6 +462,7 @@ def finalize_primary_puf_qrf_chain( root = Path(checkpoint_dir).resolve() manifest = _load_manifest(root) _assert_live_recipient_identity(frame, root, manifest) + _assert_live_recipient_predictor_universe(frame, manifest) donor_frame = _load_bound_frame( root, manifest, @@ -706,7 +750,11 @@ def _load_manifest(root: Path) -> dict[str, object]: if manifest.get("artifact_kind") != _ARTIFACT_KIND: raise ValueError("Primary QRF manifest has the wrong artifact kind.") if manifest.get("schema_version") != PRIMARY_QRF_CHECKPOINT_SCHEMA_VERSION: - raise ValueError("Unsupported primary QRF checkpoint schema version.") + raise ValueError( + "Unsupported primary QRF checkpoint schema version: expected " + f"{PRIMARY_QRF_CHECKPOINT_SCHEMA_VERSION}, got " + f"{manifest.get('schema_version')!r}." + ) target_order = _manifest_strings(manifest, "target_order") if manifest.get("target_order_sha256") != _ordered_strings_sha256(target_order): raise ValueError("Primary QRF manifest target order digest is invalid.") @@ -783,6 +831,30 @@ def _assert_live_recipient_identity( ) +def _assert_live_recipient_predictor_universe( + frame: Frame, + manifest: Mapping[str, object], +) -> None: + require_complete, _absent_cells, expected, declared = ( + _resolve_chain_doctrine_controls( + manifest, + source="Primary QRF manifest", + ) + ) + if not declared or not require_complete: + return + live = puf_recipient_predictor_universe_receipt( + frame, + predictors=_manifest_strings(manifest, "predictors"), + person_outputs=_manifest_strings(manifest, "person_outputs"), + ) + if live != expected: + raise ValueError( + "Live finalization frame changed the PUF recipient predictor " + "source universe or feature values." + ) + + def _manifest_strings(manifest: Mapping[str, object], key: str) -> tuple[str, ...]: value = manifest.get(key) if not isinstance(value, list) or not all(isinstance(item, str) for item in value): @@ -818,6 +890,7 @@ def _validate_chain_doctrine_controls( def _declared_chain_doctrine_receipt( require_complete_recipient_predictors: bool, absent_cells: str, + recipient_predictor_universe: Mapping[str, object], ) -> dict[str, object]: if ( not require_complete_recipient_predictors @@ -827,6 +900,7 @@ def _declared_chain_doctrine_receipt( return { _REQUIRE_COMPLETE_RECIPIENT_PREDICTORS: (require_complete_recipient_predictors), _ABSENT_CELLS: absent_cells, + _RECIPIENT_PREDICTOR_UNIVERSE: dict(recipient_predictor_universe), } @@ -834,37 +908,56 @@ def _resolve_chain_doctrine_controls( payload: Mapping[str, object], *, source: str, -) -> tuple[bool, str, bool]: +) -> tuple[bool, str, dict[str, object], bool]: has_require_complete = _REQUIRE_COMPLETE_RECIPIENT_PREDICTORS in payload has_absent_cells = _ABSENT_CELLS in payload - if has_require_complete != has_absent_cells: + has_universe = _RECIPIENT_PREDICTOR_UNIVERSE in payload + if len({has_require_complete, has_absent_cells, has_universe}) != 1: raise ValueError( - f"{source} doctrine controls must declare both " + f"{source} doctrine controls must declare all of " f"{_REQUIRE_COMPLETE_RECIPIENT_PREDICTORS!r} and " - f"{_ABSENT_CELLS!r}, or neither for a legacy v5 bank." + f"{_ABSENT_CELLS!r} and {_RECIPIENT_PREDICTOR_UNIVERSE!r}, " + "or none for a legacy-mode bank." ) if not has_require_complete: - return False, PUF_ABSENT_CELLS_LEGACY_ZERO_FILL, False + return False, PUF_ABSENT_CELLS_LEGACY_ZERO_FILL, {}, False require_complete, absent_cells = _validate_chain_doctrine_controls( payload[_REQUIRE_COMPLETE_RECIPIENT_PREDICTORS], payload[_ABSENT_CELLS], source=source, ) - return require_complete, absent_cells, True + universe = payload[_RECIPIENT_PREDICTOR_UNIVERSE] + if not isinstance(universe, Mapping): + raise ValueError( + f"{source} {_RECIPIENT_PREDICTOR_UNIVERSE!r} must be an object." + ) + normalized = json.loads(json.dumps(universe)) + if require_complete: + digest = normalized.get("sha256") + body = dict(normalized) + body.pop("sha256", None) + if not isinstance(digest, str) or digest != _mapping_sha256(body): + raise ValueError( + f"{source} recipient predictor universe digest is invalid." + ) + return require_complete, absent_cells, normalized, True def _manifest_doctrine_receipt( manifest: Mapping[str, object], ) -> dict[str, object]: - require_complete, absent_cells, declared = _resolve_chain_doctrine_controls( - manifest, - source="Primary QRF manifest", + require_complete, absent_cells, universe, declared = ( + _resolve_chain_doctrine_controls( + manifest, + source="Primary QRF manifest", + ) ) if not declared: return {} return { _REQUIRE_COMPLETE_RECIPIENT_PREDICTORS: require_complete, _ABSENT_CELLS: absent_cells, + _RECIPIENT_PREDICTOR_UNIVERSE: universe, } @@ -874,22 +967,29 @@ def _assert_bound_doctrine_controls( *, role: str, ) -> None: - manifest_require_complete, manifest_absent_cells, manifest_declared = ( - _resolve_chain_doctrine_controls( - manifest, - source="Primary QRF manifest", - ) + ( + manifest_require_complete, + manifest_absent_cells, + manifest_universe, + manifest_declared, + ) = _resolve_chain_doctrine_controls( + manifest, + source="Primary QRF manifest", ) - receipt_require_complete, receipt_absent_cells, receipt_declared = ( - _resolve_chain_doctrine_controls( - receipt, - source=f"Primary QRF {role}", - ) + ( + receipt_require_complete, + receipt_absent_cells, + receipt_universe, + receipt_declared, + ) = _resolve_chain_doctrine_controls( + receipt, + source=f"Primary QRF {role}", ) if ( receipt_declared != manifest_declared or receipt_require_complete != manifest_require_complete or receipt_absent_cells != manifest_absent_cells + or receipt_universe != manifest_universe ): raise ValueError( f"Primary QRF {role} doctrine controls disagree with the manifest." @@ -899,9 +999,11 @@ def _assert_bound_doctrine_controls( def _finalization_doctrine_kwargs( manifest: Mapping[str, object], ) -> dict[str, object]: - _require_complete, absent_cells, declared = _resolve_chain_doctrine_controls( - manifest, - source="Primary QRF manifest", + _require_complete, absent_cells, _universe, declared = ( + _resolve_chain_doctrine_controls( + manifest, + source="Primary QRF manifest", + ) ) if not declared: return {} diff --git a/packages/microcosm-build/src/microcosm/build/us_runtime/puf_support.py b/packages/microcosm-build/src/microcosm/build/us_runtime/puf_support.py index db98fe6a..2a6f9bbe 100644 --- a/packages/microcosm-build/src/microcosm/build/us_runtime/puf_support.py +++ b/packages/microcosm-build/src/microcosm/build/us_runtime/puf_support.py @@ -19,6 +19,11 @@ import pandas as pd from microcosm.build.gates import FitWeightRecord +from microcosm.build.us_runtime.acs_income_universe import ( + ACS_PUMS_EARNINGS_MINIMUM_AGE, + ACS_PUMS_EARNINGS_SOURCE_COLUMNS, + resolve_acs_pums_earnings_universe, +) from microcosm.build.us_runtime.puf_e01000_reconciliation import ( PUF_SCHEDULE_D_JOINT_COLUMNS, puf_capital_gains_joint_metrics, @@ -71,6 +76,7 @@ "has_support_role_metadata", "impute_us_puf_tax_detail_support", "puf_tax_detail_clone_mask", + "puf_recipient_predictor_universe_receipt", "puf_tax_unit_donor_from_arrays", "prepare_us_puf_tax_detail_chain_inputs", "resolve_formula_owned_outputs", @@ -169,6 +175,7 @@ class PufTaxDetailChainInputs: predictors: tuple[str, ...] person_outputs: tuple[str, ...] tax_unit_outputs: tuple[str, ...] + recipient_predictor_universe: Mapping[str, object] @property def target_order(self) -> tuple[str, ...]: @@ -177,6 +184,15 @@ def target_order(self) -> tuple[str, ...]: return (*self.person_outputs, *self.tax_unit_outputs) +@dataclass(frozen=True) +class _PredictorSourcePlan: + """One fail-closed source choice shared by strict receipt and features.""" + + source_column: str + entity: str + columns: tuple[str, ...] + + PUF_TAX_DETAIL_DEFAULT_PREDICTORS = ( "puf_predictor_filing_status_code", "puf_predictor_tax_unit_person_count", @@ -466,6 +482,12 @@ def target_order(self) -> tuple[str, ...]: "estate_income", ), } +_PUF_EARNINGS_UNIVERSE_PERSON_OUTPUTS = frozenset( + { + *ACS_PUMS_EARNINGS_SOURCE_COLUMNS, + "sstb_self_employment_income_before_lsr", + } +) _PREDICTOR_LEAF_ALIASES: Mapping[str, tuple[str, ...]] = { "employment_income": ("employment_income_before_lsr",), "self_employment_income": ("self_employment_income_before_lsr",), @@ -1543,6 +1565,7 @@ def impute_us_puf_tax_detail_support( fit_records: list[FitWeightRecord] | None = None, raw_predictions_callback: Callable[[pd.DataFrame], None] | None = None, tail_bound_diagnostics: list[dict[str, object]] | None = None, + predictor_universe_receipts: list[dict[str, object]] | None = None, require_complete_recipient_predictors: bool = False, absent_cells: str = PUF_ABSENT_CELLS_LEGACY_ZERO_FILL, ) -> Frame: @@ -1573,6 +1596,9 @@ def impute_us_puf_tax_detail_support( tail_bound_diagnostics: Optional output sink for the per-target tail-bound records produced during finalization. Build callers publish these records with the QRF-finalization telemetry. + predictor_universe_receipts: Optional output sink for the exact + recipient predictor source-universe receipt. Stacked callers bind + this receipt into the PUF-pass provenance; legacy callers omit it. require_complete_recipient_predictors: Stacked-spine doctrine switch (microcosm#578): build recipient features null-preserving and fail closed by name when any recipient row is missing a predictor @@ -1603,8 +1629,18 @@ def impute_us_puf_tax_detail_support( ) if not puf_mask.any(): raise ValueError("PUF detail clone has no tax-unit rows.") + strict_features: pd.DataFrame | None = None if require_complete_recipient_predictors: - _require_complete_recipient_predictor_sources(frame, puf_mask, predictors) + strict_features, predictor_universe = _strict_recipient_predictor_surface( + frame, + puf_mask, + predictors, + person_outputs=person_outputs, + ) + if predictor_universe_receipts is not None: + predictor_universe_receipts.append( + json.loads(json.dumps(predictor_universe)) + ) donor_tax_units = donor_tax_units.copy() _add_predictor_aliases(donor_tax_units, predictors) missing_donor = [ @@ -1639,13 +1675,11 @@ def impute_us_puf_tax_detail_support( # fit did not silently resolve unweighted (microcosm #300). fit_records.append(FitWeightRecord(US_PUF_SUPPORT_FIT_NAME, fitted.weight_kind)) - features = _tax_unit_feature_frame( - frame, - predictors, - preserve_nulls=require_complete_recipient_predictors, + features = ( + strict_features + if strict_features is not None + else _tax_unit_feature_frame(frame, predictors, preserve_nulls=False) ) - if require_complete_recipient_predictors: - _require_complete_recipient_predictors(features, puf_mask, predictors) predictions = fitted.predict( features.loc[puf_mask, list(predictors)], release_models=True ) @@ -1702,8 +1736,15 @@ def prepare_us_puf_tax_detail_chain_inputs( ) if not puf_mask.any(): raise ValueError("PUF detail clone has no tax-unit rows.") + strict_features: pd.DataFrame | None = None + predictor_universe: Mapping[str, object] = {} if require_complete_recipient_predictors: - _require_complete_recipient_predictor_sources(frame, puf_mask, predictors) + strict_features, predictor_universe = _strict_recipient_predictor_surface( + frame, + puf_mask, + predictors, + person_outputs=person_outputs, + ) donor_tax_units = donor_tax_units.copy() _add_predictor_aliases(donor_tax_units, predictors) missing_donor = [ @@ -1720,13 +1761,11 @@ def prepare_us_puf_tax_detail_chain_inputs( for column in donor.columns: donor[column] = pd.to_numeric(donor[column], errors="coerce").fillna(0.0) donor_frame = _tax_unit_model_frame(donor) - features = _tax_unit_feature_frame( - frame, - predictors, - preserve_nulls=require_complete_recipient_predictors, + features = ( + strict_features + if strict_features is not None + else _tax_unit_feature_frame(frame, predictors, preserve_nulls=False) ) - if require_complete_recipient_predictors: - _require_complete_recipient_predictors(features, puf_mask, predictors) recipient_features = features.loc[puf_mask, list(predictors)].copy() recipient_tax_unit_ids = ( frame.table("tax_unit").loc[puf_mask, "tax_unit_id"].to_numpy(copy=True) @@ -1739,6 +1778,7 @@ def prepare_us_puf_tax_detail_chain_inputs( predictors=predictors, person_outputs=person_outputs, tax_unit_outputs=tax_unit_outputs, + recipient_predictor_universe=predictor_universe, ) @@ -1936,17 +1976,35 @@ def finalize_us_puf_tax_detail_predictions( tables["person"], entity="person", ) + earnings_eligible_mask: pd.Series | None = None + if set(person_outputs) & _PUF_EARNINGS_UNIVERSE_PERSON_OUTPUTS: + earnings_eligible_mask = _puf_earnings_allocation_mask( + tables["person"], + person_puf_mask=person_puf_mask, + ) for column in person_outputs: _ensure_float_output_column( tables["person"], column, preserve_nulls=preserve_nulls, ) + allocation_mask = person_puf_mask + if column in _PUF_EARNINGS_UNIVERSE_PERSON_OUTPUTS: + if earnings_eligible_mask is None: # pragma: no cover - loop invariant + raise AssertionError("PUF earnings allocation mask was not resolved.") + allocation_mask = earnings_eligible_mask + # This is the explicit produced-frame universe zero. It applies + # only outside the age-15-plus allocation universe; no fillna or + # first-person fallback can reach these rows. + tables["person"].loc[ + person_puf_mask & ~earnings_eligible_mask, + column, + ] = 0.0 totals = pd.Series(predictions[column].to_numpy(), index=tax_unit_ids) if column in _PUF_TAX_DETAIL_BOOLEAN_PERSON_OUTPUTS: _write_person_tax_unit_boolean_counts( tables["person"], - mask=person_puf_mask, + mask=allocation_mask, column=column, totals=totals, fallback_basis_columns=_PERSON_OUTPUT_DISTRIBUTION_BASIS.get( @@ -1956,7 +2014,7 @@ def finalize_us_puf_tax_detail_predictions( else: _write_person_tax_unit_totals( tables["person"], - mask=person_puf_mask, + mask=allocation_mask, column=column, totals=totals, nonnegative=column in _PUF_TAX_DETAIL_NONNEGATIVE_OUTPUTS, @@ -2586,21 +2644,272 @@ def _reject_formula_owned_outputs( ) +def puf_recipient_predictor_universe_receipt( + frame: Frame, + *, + predictors: Sequence[str] = PUF_TAX_DETAIL_DEFAULT_PREDICTORS, + person_outputs: Sequence[str] = PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS, +) -> dict[str, object]: + """Recompute the strict recipient feature/universe receipt for validation.""" + + tax_unit = frame.table("tax_unit") + puf_mask = puf_tax_detail_clone_mask(tax_unit, entity="tax_unit") + if not puf_mask.any(): + raise ValueError("PUF detail clone has no tax-unit rows.") + _features, receipt = _strict_recipient_predictor_surface( + frame, + puf_mask, + tuple(predictors), + person_outputs=tuple(person_outputs), + ) + return json.loads(json.dumps(receipt)) + + +def _strict_recipient_predictor_surface( + frame: Frame, + recipient_mask: np.ndarray, + predictors: Sequence[str], + *, + person_outputs: Sequence[str], +) -> tuple[pd.DataFrame, dict[str, object]]: + """Build strict features and the exact source-universe receipt together.""" + + tax_unit = frame.table("tax_unit") + person = frame.table("person") + recipient_ids = tax_unit.loc[recipient_mask, "tax_unit_id"] + person_scope = person["person_tax_unit_id"].isin(recipient_ids) + source_plans = { + str(predictor): _strict_predictor_source_plan( + str(predictor), + tax_unit=tax_unit, + person=person, + ) + for predictor in predictors + } + source_mapping = { + predictor: { + "entity": plan.entity, + "columns": list(plan.columns), + } + for predictor, plan in source_plans.items() + } + universe_columns = tuple( + dict.fromkeys( + source_column + for plan in source_plans.values() + if plan.entity == "person" + for source_column in plan.columns + if source_column in ACS_PUMS_EARNINGS_SOURCE_COLUMNS + ) + ) + structural_absence_masks: Mapping[str, pd.Series] = {} + source_universe_rules: Mapping[str, object] = {} + if universe_columns: + universe_predictors = [ + predictor + for predictor, plan in source_plans.items() + if plan.entity == "person" and set(plan.columns) & set(universe_columns) + ] + universe = resolve_acs_pums_earnings_universe( + frame, + columns=universe_columns, + person_scope=person_scope, + boundary=( + f"PUF recipient predictor source universe for {universe_predictors}" + ), + ) + structural_absence_masks = universe.structural_absence_masks + receipt = dict(universe.receipt) + source_universe_rules = receipt["rules"] + receipt["source_universe_sha256"] = receipt.pop("sha256") + else: + receipt = { + "version": 1, + "policy": "asec_consistent_receipted_universe_zero", + "rules": {}, + "structurally_absent_person_rows": 0, + "affected_tax_unit_rows": 0, + "empty_universe_tax_unit_rows": 0, + "raw_pums_source_cells_mutated": False, + "mapped_person_cells_materialized": False, + } + + _require_complete_recipient_predictor_sources( + frame, + recipient_mask, + predictors, + structural_absence_masks=structural_absence_masks, + source_universe_rules=source_universe_rules, + source_plans=source_plans, + ) + features = _tax_unit_feature_frame( + frame, + predictors, + preserve_nulls=True, + source_plans=source_plans, + ) + _require_complete_recipient_predictors(features, recipient_mask, predictors) + recipient_features = features.loc[recipient_mask, list(predictors)] + receipt.update( + { + "predictor_source_mapping": source_mapping, + "recipient_tax_unit_rows": int(recipient_mask.sum()), + "recipient_feature_values_sha256": _predictor_feature_values_sha256( + recipient_features, + tax_unit.loc[recipient_mask, "tax_unit_id"], + ), + "person_output_allocation": _earnings_allocation_receipt( + person, + person_scope=person_scope, + person_outputs=person_outputs, + ), + } + ) + receipt["sha256"] = _receipt_sha256(receipt) + return features, receipt + + +def _earnings_allocation_receipt( + person: pd.DataFrame, + *, + person_scope: pd.Series, + person_outputs: Sequence[str], +) -> dict[str, object]: + """Describe the exact cross-arm allocation universe bound into QRF banks.""" + + earnings_outputs = sorted( + set(person_outputs) & _PUF_EARNINGS_UNIVERSE_PERSON_OUTPUTS + ) + rule_ids = sorted( + f"acs_2024_pums_{source.lower()}_age_15_plus" + for source in ACS_PUMS_EARNINGS_SOURCE_COLUMNS.values() + ) + if not earnings_outputs: + return { + "minimum_age": ACS_PUMS_EARNINGS_MINIMUM_AGE, + "person_outputs": [], + "policy": "not_applicable_without_earnings_person_outputs", + "out_of_universe_person_rows": 0, + "empty_eligible_tax_unit_rows": 0, + "first_person_fallback_out_of_universe_rows": 0, + "rule_ids": rule_ids, + } + age = pd.to_numeric(person["age"], errors="coerce") + out_of_universe = person_scope & age.lt(ACS_PUMS_EARNINGS_MINIMUM_AGE) + scoped_people = person.loc[person_scope, ["person_tax_unit_id"]].copy() + scoped_people["eligible"] = ( + age.loc[person_scope].ge(ACS_PUMS_EARNINGS_MINIMUM_AGE).to_numpy() + ) + eligible_by_unit = scoped_people.groupby("person_tax_unit_id", sort=False)[ + "eligible" + ].any() + return { + "minimum_age": ACS_PUMS_EARNINGS_MINIMUM_AGE, + "person_outputs": earnings_outputs, + "policy": "allocate only to age-15-plus persons; explicit zero otherwise", + "out_of_universe_person_rows": int(out_of_universe.sum()), + "empty_eligible_tax_unit_rows": int((~eligible_by_unit).sum()), + "first_person_fallback_out_of_universe_rows": 0, + "rule_ids": rule_ids, + } + + +def _predictor_person_source_columns( + source: str, + person: pd.DataFrame, +) -> tuple[str, ...]: + if source == "filing_status_code" or source == "tax_unit_person_count": + return () + if source == "dividend_income" and source not in person.columns: + return ("non_qualified_dividend_income", "qualified_dividend_income") + if source not in person.columns and source in _PREDICTOR_LEAF_ALIASES: + return _PREDICTOR_LEAF_ALIASES[source] + return (source,) + + +def _strict_predictor_source_plan( + predictor: str, + *, + tax_unit: pd.DataFrame, + person: pd.DataFrame, +) -> _PredictorSourcePlan: + """Resolve one strict predictor source and reject ambiguous grain collisions.""" + + source = _predictor_source_column(predictor) + if source == "filing_status_code": + columns = ( + ("filing_status_input",) + if "filing_status_input" in tax_unit + else ("filing_status",) + ) + return _PredictorSourcePlan(source, "tax_unit", columns) + if source == "tax_unit_person_count": + return _PredictorSourcePlan(source, "derived", ("person_tax_unit_id",)) + + person_columns = _predictor_person_source_columns(source, person) + person_source_present = all(column in person for column in person_columns) + tax_unit_source_present = source in tax_unit + if tax_unit_source_present and person_source_present: + raise ValueError( + "PUF recipient predictor source is ambiguous across entity grains: " + f"{predictor!r} resolves to tax_unit.{source} and person column(s) " + f"{list(person_columns)}. Remove the colliding tax-unit column or " + "declare a single canonical source." + ) + if tax_unit_source_present: + return _PredictorSourcePlan(source, "tax_unit", (source,)) + return _PredictorSourcePlan(source, "person", person_columns) + + +def _predictor_feature_values_sha256( + features: pd.DataFrame, + tax_unit_ids: pd.Series, +) -> str: + values = features.copy() + values.insert(0, "tax_unit_id", tax_unit_ids.to_numpy(copy=True)) + header = { + "columns": list(values.columns), + "dtypes": [str(values[column].dtype) for column in values.columns], + } + digest = hashlib.sha256( + json.dumps(header, sort_keys=True, separators=(",", ":")).encode() + ) + digest.update( + pd.util.hash_pandas_object(values, index=False).to_numpy(dtype=" str: + return hashlib.sha256( + json.dumps( + receipt, + sort_keys=True, + separators=(",", ":"), + allow_nan=False, + ).encode() + ).hexdigest() + + def _tax_unit_feature_frame( frame: Frame, columns: Sequence[str], *, preserve_nulls: bool = False, + source_plans: Mapping[str, _PredictorSourcePlan] | None = None, ) -> pd.DataFrame: """Build the tax-unit predictor surface under the active absence policy. The legacy policy zero-fills every missing predictor cell — the exact boundary the microcosm#578 audit identified as collapsing recipient draws to zero-conditioned degenerates. Under ``preserve_nulls`` absence - propagates as null (a leaf-alias sum is null wherever any component is - null, and an entirely absent component column is null everywhere) so the - strict recipient check can fail closed by name instead of a silent fill. - Structural person counts are not absence and stay zero-filled. + propagates as null (a leaf-alias sum is null wherever any undeclared + component is null, and an entirely absent component column is null + everywhere) so the strict recipient check can fail closed by name instead + of a silent fill. Exact documented earnings-universe zeros are already + materialized on the person table by their named rule, so ordinary addition + handles child-only units without a special null coercion. Structural person + counts are not absence and stay zero-filled. """ tax_unit = frame.table("tax_unit") @@ -2608,7 +2917,10 @@ def _tax_unit_feature_frame( result = pd.DataFrame(index=tax_unit.index) for column in columns: source_column = _predictor_source_column(column) - if source_column == "filing_status_code": + plan = None if source_plans is None else source_plans[str(column)] + if (plan is not None and plan.source_column == "filing_status_code") or ( + plan is None and source_column == "filing_status_code" + ): source = tax_unit.get("filing_status_input") if source is None: source = tax_unit.get("filing_status") @@ -2618,7 +2930,9 @@ def _tax_unit_feature_frame( source, preserve_nulls=preserve_nulls, ) - elif source_column == "tax_unit_person_count": + elif (plan is not None and plan.entity == "derived") or ( + plan is None and source_column == "tax_unit_person_count" + ): result[column] = ( person.groupby("person_tax_unit_id", sort=False) .size() @@ -2626,9 +2940,12 @@ def _tax_unit_feature_frame( .fillna(0.0) .to_numpy(dtype=np.float64) ) - elif source_column in tax_unit.columns: + elif (plan is not None and plan.entity == "tax_unit") or ( + plan is None and source_column in tax_unit.columns + ): + tax_unit_source = plan.columns[0] if plan is not None else source_column numeric = pd.to_numeric( - tax_unit[source_column], + tax_unit[tax_unit_source], errors="raise" if preserve_nulls else "coerce", ) result[column] = numeric if preserve_nulls else numeric.fillna(0.0) @@ -2662,7 +2979,11 @@ def _person_tax_unit_sum( elif column not in person.columns and column in _PREDICTOR_LEAF_ALIASES: values = np.zeros(len(person), dtype=np.float64) for leaf in _PREDICTOR_LEAF_ALIASES[column]: - values += _optional_person(person, leaf, preserve_nulls=preserve_nulls) + values += _optional_person( + person, + leaf, + preserve_nulls=preserve_nulls, + ) else: if column not in person.columns: raise ValueError( @@ -2716,16 +3037,30 @@ def _require_complete_recipient_predictor_sources( frame: Frame, recipient_mask: np.ndarray, predictors: Sequence[str], + *, + structural_absence_masks: Mapping[str, pd.Series] | None = None, + source_universe_rules: Mapping[str, object] | None = None, + source_plans: Mapping[str, _PredictorSourcePlan] | None = None, ) -> None: """Reject raw recipient absence before feature conversion or coercion.""" tax_unit = frame.table("tax_unit") person = frame.table("person") recipient_ids = tax_unit.loc[recipient_mask, "tax_unit_id"] + absence_masks = structural_absence_masks or {} + universe_rules = source_universe_rules or {} offenders: dict[str, int] = {} + offender_rule_ids: dict[str, list[str]] = {} for predictor in predictors: - source = _predictor_source_column(predictor) - if source == "filing_status_code": + predictor_rule_ids: set[str] = set() + plan = ( + source_plans[str(predictor)] + if source_plans is not None + else _strict_predictor_source_plan( + str(predictor), tax_unit=tax_unit, person=person + ) + ) + if plan.source_column == "filing_status_code": values = tax_unit.get("filing_status_input") if values is None: values = tax_unit.get("filing_status") @@ -2734,20 +3069,12 @@ def _require_complete_recipient_predictor_sources( if values is None else values.loc[recipient_mask].isna().to_numpy() ) - elif source == "tax_unit_person_count": + elif plan.entity == "derived": missing = np.zeros(int(recipient_mask.sum()), dtype=bool) - elif source in tax_unit.columns: - missing = tax_unit.loc[recipient_mask, source].isna().to_numpy() + elif plan.entity == "tax_unit": + missing = tax_unit.loc[recipient_mask, plan.columns[0]].isna().to_numpy() else: - if source == "dividend_income" and source not in person.columns: - source_columns = ( - "non_qualified_dividend_income", - "qualified_dividend_income", - ) - elif source not in person.columns and source in _PREDICTOR_LEAF_ALIASES: - source_columns = _PREDICTOR_LEAF_ALIASES[source] - else: - source_columns = (source,) + source_columns = plan.columns absent_columns = [ column for column in source_columns if column not in person.columns ] @@ -2757,9 +3084,61 @@ def _require_complete_recipient_predictor_sources( link = person["person_tax_unit_id"] relevant = link.isin(recipient_ids) observed_ids = set(link.loc[relevant].tolist()) - null_source = ( - person.loc[relevant, list(source_columns)].isna().any(axis=1) - ) + unauthorized_nulls = pd.DataFrame(index=person.index[relevant]) + for source_column in source_columns: + source_null = person.loc[relevant, source_column].isna() + structural = absence_masks.get(source_column) + rule = universe_rules.get(source_column) + if structural is not None: + if not structural.index.equals(person.index): + raise ValueError( + "Recipient predictor structural-absence mask for " + f"{source_column!r} is not aligned." + ) + if not isinstance(rule, Mapping): + raise ValueError( + "Recipient predictor structural-absence mask for " + f"{source_column!r} has no named universe rule." + ) + rule_id = rule.get("rule_id") + raw_source_column = rule.get("source_column") + source_channel = rule.get("source_channel") + if not isinstance(rule_id, str) or not rule_id: + raise ValueError( + "Recipient predictor source-universe rule for " + f"{source_column!r} has no rule_id." + ) + predictor_rule_ids.add(rule_id) + if ( + not isinstance(raw_source_column, str) + or not raw_source_column + ): + raise ValueError( + f"{rule_id}: source-universe rule has no source_column." + ) + if not isinstance(source_channel, str) or not source_channel: + raise ValueError( + f"{rule_id}: source-universe rule has no source_channel." + ) + if raw_source_column not in person: + source_null |= True + else: + raw_null = person.loc[relevant, raw_source_column].isna() + raw_authority_scope = ( + person.loc[relevant, support_channel_column("person")] + .astype(str) + .eq(source_channel) + ) + # Raw PUMS blanks are authorized only below the named + # age floor. The mapped cell itself must already be the + # rule-applied numeric zero on those rows. + source_null |= ( + raw_null + & raw_authority_scope + & ~structural.loc[relevant] + ) + unauthorized_nulls[source_column] = source_null + null_source = unauthorized_nulls.any(axis=1) null_ids = set(link.loc[relevant].loc[null_source].tolist()) missing = np.asarray( [ @@ -2771,12 +3150,20 @@ def _require_complete_recipient_predictor_sources( count = int(missing.sum()) if count: offenders[str(predictor)] = count + if predictor_rule_ids: + offender_rule_ids[str(predictor)] = sorted(predictor_rule_ids) if offenders: + named_rules = ( + f" Named source-universe rules: {offender_rule_ids}." + if offender_rule_ids + else "" + ) raise ValueError( "PUF recipient predictor source(s) have missing values before " f"coercion: {offenders} (of {int(len(recipient_ids))} recipient rows). " - "This is a terminal stacked-spine failure; gap-fill the source " - "before the PUF pass." + "This is a terminal stacked-spine failure; complete the source or " + "declare its exact source universe before the PUF pass." + f"{named_rules}" ) @@ -2807,6 +3194,19 @@ def _require_complete_recipient_predictors( "The primary QRF must not zero-fill absence (microcosm#578); " "gap-fill the stacked spine before the PUF pass." ) + numeric = recipient.to_numpy(dtype=np.float64) + nonfinite_counts = (~np.isfinite(numeric)).sum(axis=0) + nonfinite_offenders = { + str(name): int(count) + for name, count in zip(recipient.columns, nonfinite_counts, strict=True) + if int(count) + } + if nonfinite_offenders: + raise ValueError( + "PUF recipient predictor(s) have nonfinite values on recipient " + f"rows: {nonfinite_offenders} (of {int(len(recipient))} recipient rows). " + "The primary QRF requires a finite feature surface before fitting." + ) def _ensure_float_output_column( @@ -2884,6 +3284,25 @@ def _snap_to_observed_values( return snapped +def _puf_earnings_allocation_mask( + person: pd.DataFrame, + *, + person_puf_mask: pd.Series, +) -> pd.Series: + """Return the age-15-plus mask for every PUF-allocated earnings output.""" + + if "age" not in person: + raise ValueError("PUF earnings allocation requires person.age.") + age = pd.to_numeric(person["age"], errors="coerce") + invalid = person_puf_mask & ~np.isfinite(age.to_numpy(dtype=np.float64)) + if invalid.any(): + raise ValueError( + "PUF earnings allocation has " + f"{int(invalid.sum())} missing or nonfinite age value(s)." + ) + return person_puf_mask & age.ge(ACS_PUMS_EARNINGS_MINIMUM_AGE) + + def _write_person_tax_unit_boolean_counts( person: pd.DataFrame, *, diff --git a/packages/microcosm-build/src/microcosm/build/us_runtime/qbi_inputs.py b/packages/microcosm-build/src/microcosm/build/us_runtime/qbi_inputs.py index 052e2e2a..13f5c35c 100644 --- a/packages/microcosm-build/src/microcosm/build/us_runtime/qbi_inputs.py +++ b/packages/microcosm-build/src/microcosm/build/us_runtime/qbi_inputs.py @@ -17,6 +17,10 @@ from __future__ import annotations +import hashlib +import json +import re +from collections.abc import Mapping, Sequence from importlib.resources import files import numpy as np @@ -24,11 +28,16 @@ from microcosm.build.gates import GateResult from microcosm.build.source_manifest import SourceStageSpec, load_source_manifest +from microcosm.build.us_runtime.acs_income_universe import ( + resolve_acs_pums_earnings_universe, +) from microcosm.build.us_runtime.support_provenance import ( BASE_ASEC_SUPPORT_CHANNEL, has_support_role_metadata, + support_clone_index_column, support_role_series, ) +from microcosm.build.us_runtime.take_up_contract import load_take_up_contract from microcosm.frame import US_SCHEMA, Frame __all__ = [ @@ -43,10 +52,19 @@ "US_QBI_NONCONSTANT_PERSON_COLUMNS", "US_QBI_NONNEGATIVE_OUTPUT_COLUMNS", "US_QBI_OUTPUT_COLUMNS", + "US_QBI_RECONCILED_PERSON_COLUMNS", "US_QBI_STAGE_NAME", + "bind_us_qbi_reconciliation_transition_authority", "us_qbi_inputs_signal_gate", "us_qbi_inputs_stage_spec", "us_qbi_inputs_summary", + "us_qbi_post_reconciliation_person_columns", + "us_qbi_reconciliation_change_receipt", + "us_qbi_reconciliation_contract_identity", + "us_qbi_reconciliation_universe_receipt", + "validate_us_qbi_reconciliation_live_output", + "validate_us_qbi_reconciliation_receipt", + "validate_us_qbi_reconciliation_transition", "with_us_qbi_input_reconciliation", ] @@ -107,6 +125,21 @@ _SELF_EMPLOYMENT_COLUMN = "self_employment_income_before_lsr" _SSTB_SELF_EMPLOYMENT_COLUMN = "sstb_self_employment_income_before_lsr" +US_QBI_RECONCILED_PERSON_COLUMNS: tuple[str, ...] = ( + *US_QBI_OUTPUT_COLUMNS, + _SELF_EMPLOYMENT_COLUMN, +) +_QBI_UNDECLARED_DRIVER_PERSON_COLUMNS: tuple[str, ...] = ( + "partnership_income", + "s_corp_income", + "estate_income", + "non_qualified_dividend_income", + "age", + "SEMP", + "person_tax_unit_id", + "person_support_clone_index", + "person_source_id", +) _BOOLEAN_SHARE_BANDS: dict[str, tuple[float, float]] = { "business_is_sstb": (0.001, 0.25), **{column: (0.01, 0.9999) for column in _GENERAL_QUALIFICATION_FLAGS}, @@ -122,6 +155,177 @@ "w2_wages_from_qualified_business": (0.001, 0.35), } _INVARIANT_ATOL = 1e-8 +_QBI_RECONCILIATION_RECEIPT_VERSION = 2 +_QBI_TRANSITION_AUTHORITY_METADATA_KEY = "us_qbi_reconciliation_transition_authority" +_QBI_RECONCILIATION_OPERATION = "shared_all_or_nothing_identity_reconciliation" +_LOWERCASE_SHA256 = re.compile(r"[0-9a-f]{64}") +_QBI_CHANGE_RECEIPT_KEYS = frozenset( + { + "version", + "operation", + "declared_person_columns", + "recipient_source_universe", + "input_person_columns", + "input_person_table_sha256", + "output_declared_person_values_sha256", + "live_driver_surface_sha256", + "changed_person_rows", + "base_self_employment_changed_rows", + "structurally_absent_base_source_changed_rows", + "undeclared_surface_preservation", + "sha256", + } +) +_QBI_TRANSITION_AUTHORITY_KEYS = frozenset( + { + "version", + "receipt_sha256", + "input_person_columns_sha256", + "input_person_table_sha256", + "output_declared_person_values_sha256", + "live_driver_surface_sha256", + "recipient_source_universe_sha256", + "sha256", + } +) +_QBI_PRESERVATION_RECEIPT_KEYS = frozenset( + { + "undeclared_person_columns_verified", + "non_person_entity_tables_verified", + "link_tables_verified", + "weight_vectors_verified", + "strata_verified", + "mass_log_verified", + "metadata_verified", + } +) +_QBI_STACKED_UNIVERSE_RECEIPT_KEYS = frozenset( + { + "version", + "policy", + "source_dataset", + "source_document", + "source_url", + "source_channel", + "age_column", + "minimum_age", + "aggregation", + "produced_frame_semantics", + "mapped_person_cells_materialized", + "mapped_universe_zero_cells", + "scoped_person_rows", + "scoped_acs_person_rows", + "structurally_absent_person_rows", + "affected_tax_unit_rows", + "mixed_universe_tax_unit_rows", + "empty_universe_tax_unit_rows", + "structurally_absent_person_lineages_sha256", + "affected_tax_unit_ids_sha256", + "empty_universe_tax_unit_ids_sha256", + "by_origin_role", + "rules", + "source_universe_sha256", + "source_universe_resolution_mutated_raw_pums_cells", + "operation", + "rows_excluded_from_base_self_employment_rewrite", + "rows_included_in_other_qbi_reconciliation", + "structurally_absent_base_source_cells_mutated", + "sha256", + } +) +_QBI_UNSTACKED_UNIVERSE_RECEIPT_KEYS = frozenset( + { + "version", + "policy", + "rows_excluded_from_base_self_employment_rewrite", + "rows_included_in_other_qbi_reconciliation", + "structurally_absent_base_source_cells_mutated", + "sha256", + } +) +_QBI_UNIVERSE_RULE_KEYS = frozenset( + { + "rule_id", + "source_column", + "mapped_column", + "source_channel", + "source_universe", + "produced_frame_semantics", + "structurally_absent_person_rows", + "eligible_acs_person_rows", + "in_universe_null_rows", + "raw_in_universe_null_rows", + "mapped_universe_zero_rows", + "universe_zero_missing_rows", + "out_of_universe_mapped_nonzero_rows", + "raw_source_column_present", + "raw_source_nonblank_rows", + "source_cells_sha256", + } +) + + +def us_qbi_reconciliation_contract_identity() -> dict[str, object]: + """Return the immutable QBI mutation/receipt semantics for base identity.""" + + return { + "version": _QBI_RECONCILIATION_RECEIPT_VERSION, + "operation": _QBI_RECONCILIATION_OPERATION, + "execution_scope": "whole_pool", + "declared_person_columns": list(US_QBI_RECONCILED_PERSON_COLUMNS), + "kernel_binding": "after_equals_deterministic_kernel_of_before", + "live_output_binding": ( + "declared_values_and_driver_surface_digests_plus_exact_live_" + "person_inventory_universe_kernel_fixed_point_and_independent_" + "frame_metadata_transition_authority" + ), + "receipt_schema": sorted(_QBI_CHANGE_RECEIPT_KEYS), + "transition_authority_metadata_key": (_QBI_TRANSITION_AUTHORITY_METADATA_KEY), + "transition_authority_schema": sorted(_QBI_TRANSITION_AUTHORITY_KEYS), + "acs_base_self_employment_policy": ( + "preserve_receipted_age_under_15_universe_zero_in_every_clone_role" + ), + } + + +def us_qbi_post_reconciliation_person_columns( + seed_receipt: Mapping[str, object] | None, +) -> tuple[str, ...]: + """Resolve canonical person outputs named by the later seed receipt.""" + + if seed_receipt is None: + return () + if not isinstance(seed_receipt, Mapping): + raise ValueError("QBI seed receipt must be an object.") + outputs: list[str] = [] + programs = seed_receipt.get("programs") + if programs is None: + return () + if not isinstance(programs, Mapping): + raise ValueError("QBI seed receipt programs must be an object.") + contract_entities = { + program.variable: program.entity for program in load_take_up_contract().programs + } + for variable, declaration in programs.items(): + if not isinstance(variable, str) or not variable: + raise ValueError("QBI seed receipt has a malformed program name.") + if variable not in contract_entities: + raise ValueError( + f"QBI seed receipt names non-contract program {variable!r}." + ) + if not isinstance(declaration, Mapping): + raise ValueError( + f"QBI seed receipt program {variable!r} must be an object." + ) + entity = declaration.get("entity") + if entity != contract_entities[variable]: + raise ValueError( + f"QBI seed receipt program {variable!r} has entity {entity!r}; " + f"expected {contract_entities[variable]!r}." + ) + if entity == "person": + outputs.append(variable) + return tuple(outputs) def us_qbi_inputs_stage_spec() -> SourceStageSpec: @@ -150,12 +354,877 @@ def _numeric(person: pd.DataFrame, column: str) -> np.ndarray: return values +def _numeric_in_scope( + person: pd.DataFrame, + column: str, + scope: np.ndarray, +) -> np.ndarray: + """Read finite values only inside a declared reconciliation scope.""" + + values = pd.to_numeric(person[column], errors="coerce").to_numpy(dtype=np.float64) + nonfinite = int(np.count_nonzero(scope & ~np.isfinite(values))) + if nonfinite: + raise ValueError( + f"US QBI input {column!r} contains {nonfinite} nonfinite in-scope value(s)." + ) + return values + + def _optional_numeric(person: pd.DataFrame, column: str) -> np.ndarray: if column not in person: return np.zeros(len(person), dtype=np.float64) return _numeric(person, column) +def _qbi_reconciliation_scope( + frame: Frame, +) -> tuple[np.ndarray, dict[str, object]]: + """Return active rows and an exact all-role ACS source-universe receipt.""" + + person = frame.table("person") + clone_column = support_clone_index_column("person") + if clone_column not in person: + receipt: dict[str, object] = { + "version": 1, + "policy": "all_rows_without_stacked_clone_provenance", + "rows_excluded_from_base_self_employment_rewrite": 0, + "rows_included_in_other_qbi_reconciliation": int(len(person)), + "structurally_absent_base_source_cells_mutated": False, + } + receipt["sha256"] = _qbi_receipt_sha256(receipt) + return np.ones(len(person), dtype=bool), receipt + + clone_index = pd.to_numeric(person[clone_column], errors="coerce") + if clone_index.isna().any(): + raise ValueError("US QBI reconciliation requires complete clone provenance.") + universe = resolve_acs_pums_earnings_universe( + frame, + columns=(_SELF_EMPLOYMENT_COLUMN,), + person_scope=np.ones(len(person), dtype=bool), + boundary="US QBI reconciliation source universe", + ) + universe_rule = universe.receipt["rules"][_SELF_EMPLOYMENT_COLUMN] + if ( + universe_rule["in_universe_null_rows"] + or universe_rule["raw_in_universe_null_rows"] + ): + raise ValueError( + "US QBI reconciliation has missing eligible source values under " + f"{universe_rule['rule_id']}: in_universe_null_rows=" + f"{universe_rule['in_universe_null_rows']}, " + "raw_in_universe_null_rows=" + f"{universe_rule['raw_in_universe_null_rows']}." + ) + structural = universe.structural_absence_masks[_SELF_EMPLOYMENT_COLUMN].to_numpy( + dtype=bool + ) + active = ~structural + receipt = dict(universe.receipt) + receipt["source_universe_sha256"] = receipt.pop("sha256") + receipt["source_universe_resolution_mutated_raw_pums_cells"] = receipt.pop( + "raw_pums_source_cells_mutated" + ) + receipt.update( + { + "operation": ( + "preserve every receipted ACS under-15 base self-employment zero; " + "reconcile every derived QBI identity on every person row" + ), + "rows_excluded_from_base_self_employment_rewrite": int(structural.sum()), + "rows_included_in_other_qbi_reconciliation": int(len(person)), + "structurally_absent_base_source_cells_mutated": False, + } + ) + receipt["sha256"] = _qbi_receipt_sha256(receipt) + return active, receipt + + +def us_qbi_reconciliation_universe_receipt(frame: Frame) -> dict[str, object]: + """Expose the exact QBI reconciliation scope receipt to the pool driver.""" + + _scope, receipt = _qbi_reconciliation_scope(frame) + return json.loads(json.dumps(receipt)) + + +def _qbi_receipt_sha256(receipt: Mapping[str, object]) -> str: + return hashlib.sha256( + json.dumps( + receipt, + sort_keys=True, + separators=(",", ":"), + allow_nan=False, + ).encode() + ).hexdigest() + + +def us_qbi_reconciliation_change_receipt( + before: Frame, + after: Frame, +) -> dict[str, object]: + """Generate a receipt only for the live deterministic QBI transition.""" + + expected = with_us_qbi_input_reconciliation(before) + _assert_qbi_kernel_output_matches( + expected, + after, + boundary="US QBI reconciliation receipt generation", + ) + receipt = _build_us_qbi_reconciliation_change_receipt(before, after) + validate_us_qbi_reconciliation_transition( + before, + after, + receipt, + boundary="US QBI reconciliation receipt generation", + ) + return receipt + + +def bind_us_qbi_reconciliation_transition_authority( + frame: Frame, + receipt: Mapping[str, object], +) -> Frame: + """Anchor one generated QBI transition independently in frame metadata.""" + + validated = validate_us_qbi_reconciliation_receipt( + receipt, + boundary="US QBI transition-authority binding", + ) + if _QBI_TRANSITION_AUTHORITY_METADATA_KEY in frame.metadata: + raise ValueError( + "US QBI transition authority is already bound; refusing to overwrite " + "the immutable generation anchor." + ) + authority = _qbi_transition_authority_receipt(validated) + tables = {entity: frame.table(entity).copy() for entity in frame.entities} + tables.update({link: frame.link(link).copy() for link in frame.links}) + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + metadata={ + **frame.metadata, + _QBI_TRANSITION_AUTHORITY_METADATA_KEY: authority, + }, + ) + + +def _build_us_qbi_reconciliation_change_receipt( + before: Frame, + after: Frame, +) -> dict[str, object]: + """Build the canonical envelope after the kernel transition is established.""" + + before_person = before.table("person") + after_person = after.table("person") + if not before_person.index.equals(after_person.index): + raise ValueError("US QBI reconciliation changed the person-row index.") + missing = sorted( + set(US_QBI_RECONCILED_PERSON_COLUMNS) + - set(before_person.columns).intersection(after_person.columns) + ) + if missing: + raise ValueError( + "US QBI reconciliation change receipt lacks declared person column(s): " + f"{missing}." + ) + preservation = _assert_qbi_undeclared_surface_unchanged(before, after) + base_scope, _input_universe_receipt = _qbi_reconciliation_scope(before) + _output_scope, universe_receipt = _qbi_reconciliation_scope(after) + structural = ~base_scope + before_values = before_person.loc[:, US_QBI_RECONCILED_PERSON_COLUMNS] + after_values = after_person.loc[:, US_QBI_RECONCILED_PERSON_COLUMNS] + equal = before_values.eq(after_values) | ( + before_values.isna() & after_values.isna() + ) + equal = equal.fillna(False) + changed_rows = ~equal.all(axis=1).to_numpy(dtype=bool) + before_base = before_person[_SELF_EMPLOYMENT_COLUMN] + after_base = after_person[_SELF_EMPLOYMENT_COLUMN] + base_equal = before_base.eq(after_base) | (before_base.isna() & after_base.isna()) + base_equal = base_equal.fillna(False) + structural_base_changes = int(np.count_nonzero(structural & ~base_equal.to_numpy())) + if structural_base_changes: + raise ValueError( + "US QBI reconciliation mutated " + f"{structural_base_changes} structurally absent ACS base source cell(s)." + ) + receipt: dict[str, object] = { + "version": _QBI_RECONCILIATION_RECEIPT_VERSION, + "operation": _QBI_RECONCILIATION_OPERATION, + "declared_person_columns": list(US_QBI_RECONCILED_PERSON_COLUMNS), + "recipient_source_universe": universe_receipt, + "input_person_columns": list(before_person.columns), + "input_person_table_sha256": _qbi_person_values_sha256( + before_person, + columns=tuple(before_person.columns), + ), + "output_declared_person_values_sha256": _qbi_person_values_sha256( + after_person, + columns=US_QBI_RECONCILED_PERSON_COLUMNS, + ), + "live_driver_surface_sha256": _qbi_live_driver_surface_sha256(after), + "changed_person_rows": int(changed_rows.sum()), + "base_self_employment_changed_rows": int((~base_equal).sum()), + "structurally_absent_base_source_changed_rows": structural_base_changes, + "undeclared_surface_preservation": preservation, + } + receipt["sha256"] = _qbi_receipt_sha256(receipt) + return receipt + + +def validate_us_qbi_reconciliation_receipt( + receipt: Mapping[str, object], + *, + boundary: str, +) -> dict[str, object]: + """Validate the exact authenticated QBI receipt envelope and nested scope.""" + + _require_exact_qbi_keys( + receipt, + _QBI_CHANGE_RECEIPT_KEYS, + boundary=boundary, + label="receipt", + ) + if receipt.get("version") != _QBI_RECONCILIATION_RECEIPT_VERSION: + raise ValueError( + f"{boundary}: QBI reconciliation receipt version must be " + f"{_QBI_RECONCILIATION_RECEIPT_VERSION}." + ) + if receipt.get("operation") != _QBI_RECONCILIATION_OPERATION: + raise ValueError(f"{boundary}: QBI reconciliation operation changed.") + if receipt.get("declared_person_columns") != list(US_QBI_RECONCILED_PERSON_COLUMNS): + raise ValueError(f"{boundary}: QBI declared person columns changed.") + input_columns = receipt.get("input_person_columns") + if ( + not isinstance(input_columns, list) + or not all(isinstance(column, str) and column for column in input_columns) + or len(input_columns) != len(set(input_columns)) + or not set(US_QBI_RECONCILED_PERSON_COLUMNS).issubset(input_columns) + ): + raise ValueError(f"{boundary}: QBI input person columns are malformed.") + for field in ( + "input_person_table_sha256", + "output_declared_person_values_sha256", + "live_driver_surface_sha256", + ): + _require_qbi_sha256(receipt.get(field), boundary=boundary, field=field) + for field in ( + "changed_person_rows", + "base_self_employment_changed_rows", + "structurally_absent_base_source_changed_rows", + ): + _require_qbi_nonnegative_integer( + receipt.get(field), boundary=boundary, field=field + ) + if receipt.get("structurally_absent_base_source_changed_rows") != 0: + raise ValueError( + f"{boundary}: QBI receipt reports a structurally absent base-source " + "mutation." + ) + preservation = receipt.get("undeclared_surface_preservation") + if not isinstance(preservation, Mapping): + raise ValueError(f"{boundary}: QBI preservation receipt must be an object.") + _require_exact_qbi_keys( + preservation, + _QBI_PRESERVATION_RECEIPT_KEYS, + boundary=boundary, + label="undeclared-surface preservation receipt", + ) + for field in ( + "undeclared_person_columns_verified", + "non_person_entity_tables_verified", + "link_tables_verified", + "weight_vectors_verified", + ): + _require_qbi_nonnegative_integer( + preservation.get(field), boundary=boundary, field=field + ) + for field in ("strata_verified", "mass_log_verified", "metadata_verified"): + if preservation.get(field) is not True: + raise ValueError(f"{boundary}: QBI preservation field {field!r} is false.") + universe = receipt.get("recipient_source_universe") + if not isinstance(universe, Mapping): + raise ValueError(f"{boundary}: QBI source-universe receipt must be an object.") + _validate_qbi_universe_receipt(universe, boundary=boundary) + changed_rows = receipt["changed_person_rows"] + base_changed_rows = receipt["base_self_employment_changed_rows"] + included_rows = universe["rows_included_in_other_qbi_reconciliation"] + if not (base_changed_rows <= changed_rows <= included_rows): + raise ValueError( + f"{boundary}: QBI changed-row counts violate " + "base <= any <= reconciliation-scope rows." + ) + observed_sha256 = _require_qbi_sha256( + receipt.get("sha256"), + boundary=boundary, + field="sha256", + ) + unsigned = dict(receipt) + unsigned.pop("sha256") + expected_sha256 = _qbi_receipt_sha256(unsigned) + if observed_sha256 != expected_sha256: + raise ValueError(f"{boundary}: QBI reconciliation receipt SHA-256 mismatch.") + return json.loads(json.dumps(receipt)) + + +def validate_us_qbi_reconciliation_transition( + before: Frame, + after: Frame, + receipt: Mapping[str, object], + *, + boundary: str, +) -> dict[str, object]: + """Authenticate receipt and output against the live deterministic transition.""" + + validated = validate_us_qbi_reconciliation_receipt(receipt, boundary=boundary) + expected_after = with_us_qbi_input_reconciliation(before) + _assert_qbi_kernel_output_matches( + expected_after, + after, + boundary=boundary, + ) + expected_receipt = _build_us_qbi_reconciliation_change_receipt( + before, + expected_after, + ) + if validated != expected_receipt: + raise ValueError( + f"{boundary}: QBI receipt is not bound to the live deterministic " + "input/output transition." + ) + return validated + + +def validate_us_qbi_reconciliation_live_output( + frame: Frame, + receipt: Mapping[str, object], + *, + boundary: str, + expected_transition_authority_sha256: str, + allowed_post_reconciliation_person_columns: Sequence[str] = (), +) -> dict[str, object]: + """Bind a persisted receipt to live output values and the kernel fixed point.""" + + validated = validate_us_qbi_reconciliation_receipt(receipt, boundary=boundary) + person = frame.table("person") + missing = sorted(set(US_QBI_RECONCILED_PERSON_COLUMNS) - set(person.columns)) + if missing: + raise ValueError( + f"{boundary}: live QBI output lacks declared person column(s) {missing}." + ) + allowed_columns = tuple(allowed_post_reconciliation_person_columns) + canonical_seed_person_columns = { + program.variable + for program in load_take_up_contract().programs + if program.entity == "person" + } + if ( + isinstance(allowed_post_reconciliation_person_columns, (str, bytes)) + or not all(isinstance(column, str) and column for column in allowed_columns) + or len(allowed_columns) != len(set(allowed_columns)) + or not set(allowed_columns).issubset(canonical_seed_person_columns) + ): + raise ValueError(f"{boundary}: allowed post-QBI person columns are malformed.") + input_columns = tuple(validated["input_person_columns"]) + input_set = set(input_columns) + live_columns = tuple(person.columns) + live_input_prefix = tuple(column for column in live_columns if column in input_set) + live_additions = set(live_columns) - input_set + expected_additions = set(allowed_columns) - input_set + if live_input_prefix != input_columns or live_additions != expected_additions: + raise ValueError( + f"{boundary}: QBI live person-column inventory changed outside " + "declared post-reconciliation outputs." + ) + live_digest = _qbi_person_values_sha256( + person, + columns=US_QBI_RECONCILED_PERSON_COLUMNS, + ) + if validated["output_declared_person_values_sha256"] != live_digest: + raise ValueError( + f"{boundary}: QBI receipt output digest does not match the live frame." + ) + live_driver_surface_sha256 = _qbi_live_driver_surface_sha256(frame) + if validated["live_driver_surface_sha256"] != live_driver_surface_sha256: + raise ValueError( + f"{boundary}: QBI receipt driver-surface digest does not match " + "the live frame." + ) + expected = with_us_qbi_input_reconciliation(frame) + _assert_qbi_kernel_output_matches(expected, frame, boundary=boundary) + + live_universe = us_qbi_reconciliation_universe_receipt(frame) + recorded_universe = validated["recipient_source_universe"] + if recorded_universe != live_universe: + raise ValueError( + f"{boundary}: QBI source-universe receipt does not exactly match " + "the live frame." + ) + expected_preservation = { + "undeclared_person_columns_verified": len( + set(input_columns) - set(US_QBI_RECONCILED_PERSON_COLUMNS) + ), + "non_person_entity_tables_verified": len(frame.entities) - 1, + "link_tables_verified": len(frame.links), + "weight_vectors_verified": len(frame.weighted_entities), + "strata_verified": True, + "mass_log_verified": True, + "metadata_verified": True, + } + if validated["undeclared_surface_preservation"] != expected_preservation: + raise ValueError( + f"{boundary}: QBI preservation receipt does not match the live frame " + "inventory and declared input surface." + ) + expected_authority_sha256 = _require_qbi_sha256( + expected_transition_authority_sha256, + boundary=boundary, + field="expected_transition_authority_sha256", + ) + if validated["sha256"] != expected_authority_sha256: + raise ValueError( + f"{boundary}: QBI receipt differs from the independently carried " + "transition authority." + ) + _validate_qbi_transition_authority( + frame, + validated, + boundary=boundary, + ) + return validated + + +def _qbi_transition_authority_receipt( + receipt: Mapping[str, object], +) -> dict[str, object]: + universe = receipt["recipient_source_universe"] + assert isinstance(universe, Mapping) + authority: dict[str, object] = { + "version": 1, + "receipt_sha256": receipt["sha256"], + "input_person_columns_sha256": _qbi_receipt_sha256( + {"input_person_columns": receipt["input_person_columns"]} + ), + "input_person_table_sha256": receipt["input_person_table_sha256"], + "output_declared_person_values_sha256": receipt[ + "output_declared_person_values_sha256" + ], + "live_driver_surface_sha256": receipt["live_driver_surface_sha256"], + "recipient_source_universe_sha256": universe["sha256"], + } + authority["sha256"] = _qbi_receipt_sha256(authority) + return authority + + +def _validate_qbi_transition_authority( + frame: Frame, + receipt: Mapping[str, object], + *, + boundary: str, +) -> None: + authority = frame.metadata.get(_QBI_TRANSITION_AUTHORITY_METADATA_KEY) + if not isinstance(authority, Mapping): + raise ValueError( + f"{boundary}: QBI transition authority is absent from live frame metadata." + ) + _require_exact_qbi_keys( + authority, + _QBI_TRANSITION_AUTHORITY_KEYS, + boundary=boundary, + label="transition-authority receipt", + ) + if authority.get("version") != 1: + raise ValueError(f"{boundary}: QBI transition-authority version changed.") + for field in _QBI_TRANSITION_AUTHORITY_KEYS - {"version"}: + _require_qbi_sha256(authority.get(field), boundary=boundary, field=field) + unsigned = dict(authority) + observed_sha256 = unsigned.pop("sha256") + if observed_sha256 != _qbi_receipt_sha256(unsigned): + raise ValueError(f"{boundary}: QBI transition-authority SHA-256 mismatch.") + expected = _qbi_transition_authority_receipt(receipt) + if dict(authority) != expected: + raise ValueError( + f"{boundary}: QBI receipt is not bound to the immutable live " + "transition authority." + ) + + +def _assert_qbi_kernel_output_matches( + expected: Frame, + actual: Frame, + *, + boundary: str, +) -> None: + _assert_qbi_undeclared_surface_unchanged(expected, actual) + expected_person = expected.table("person") + actual_person = actual.table("person") + if not expected_person.loc[:, US_QBI_RECONCILED_PERSON_COLUMNS].equals( + actual_person.loc[:, US_QBI_RECONCILED_PERSON_COLUMNS] + ): + changed = [ + column + for column in US_QBI_RECONCILED_PERSON_COLUMNS + if not expected_person[column].equals(actual_person[column]) + ] + raise ValueError( + f"{boundary}: QBI output differs from the deterministic kernel in " + f"declared person column(s) {changed}." + ) + + +def _validate_qbi_universe_receipt( + receipt: Mapping[str, object], + *, + boundary: str, +) -> None: + policy = receipt.get("policy") + keys = ( + _QBI_UNSTACKED_UNIVERSE_RECEIPT_KEYS + if policy == "all_rows_without_stacked_clone_provenance" + else _QBI_STACKED_UNIVERSE_RECEIPT_KEYS + ) + _require_exact_qbi_keys( + receipt, + keys, + boundary=boundary, + label="source-universe receipt", + ) + if receipt.get("version") != 1: + raise ValueError(f"{boundary}: QBI source-universe version changed.") + for field, value in receipt.items(): + if field.endswith(("_rows", "_cells")) and field != ( + "source_universe_resolution_mutated_raw_pums_cells" + ): + _require_qbi_nonnegative_integer(value, boundary=boundary, field=field) + if receipt.get("structurally_absent_base_source_cells_mutated") is not False: + raise ValueError( + f"{boundary}: QBI source-universe receipt reports a forbidden mutation." + ) + if policy != "all_rows_without_stacked_clone_provenance": + if policy != "asec_consistent_receipted_universe_zero": + raise ValueError(f"{boundary}: QBI source-universe policy changed.") + if receipt.get("minimum_age") != 15: + raise ValueError(f"{boundary}: QBI source-universe age floor changed.") + if receipt.get("operation") != ( + "preserve every receipted ACS under-15 base self-employment zero; " + "reconcile every derived QBI identity on every person row" + ): + raise ValueError(f"{boundary}: QBI source-universe operation changed.") + if ( + receipt.get("source_universe_resolution_mutated_raw_pums_cells") + is not False + ): + raise ValueError( + f"{boundary}: QBI source-universe resolution mutated raw PUMS cells." + ) + if not isinstance(receipt.get("mapped_person_cells_materialized"), bool): + raise ValueError( + f"{boundary}: QBI mapped-person materialization flag is malformed." + ) + rules = receipt.get("rules") + if not isinstance(rules, Mapping) or set(rules) != {_SELF_EMPLOYMENT_COLUMN}: + raise ValueError(f"{boundary}: QBI source-universe rules are not exact.") + rule = rules[_SELF_EMPLOYMENT_COLUMN] + if not isinstance(rule, Mapping): + raise ValueError(f"{boundary}: QBI source-universe rule must be an object.") + _require_exact_qbi_keys( + rule, + _QBI_UNIVERSE_RULE_KEYS, + boundary=boundary, + label="source-universe rule", + ) + if rule.get("rule_id") != "acs_2024_pums_semp_age_15_plus": + raise ValueError(f"{boundary}: QBI SEMP universe rule ID changed.") + if ( + rule.get("source_column") != "SEMP" + or rule.get("mapped_column") != _SELF_EMPLOYMENT_COLUMN + ): + raise ValueError(f"{boundary}: QBI SEMP universe mapping changed.") + if rule.get("raw_source_column_present") is not True: + raise ValueError(f"{boundary}: QBI SEMP raw authority is absent.") + for field, value in rule.items(): + if field.endswith("_rows"): + _require_qbi_nonnegative_integer( + value, + boundary=boundary, + field=f"source-universe rule {field}", + ) + for field in ( + "in_universe_null_rows", + "raw_in_universe_null_rows", + "universe_zero_missing_rows", + "out_of_universe_mapped_nonzero_rows", + "raw_source_nonblank_rows", + ): + if rule.get(field) != 0: + raise ValueError( + f"{boundary}: QBI source-universe rule field {field!r} " + "must be zero." + ) + if ( + rule.get("mapped_universe_zero_rows") + != rule.get("structurally_absent_person_rows") + or receipt.get("mapped_universe_zero_cells") + != receipt.get("structurally_absent_person_rows") + or receipt.get("rows_excluded_from_base_self_employment_rewrite") + != receipt.get("structurally_absent_person_rows") + ): + raise ValueError( + f"{boundary}: QBI source-universe zero/count equation failed." + ) + _require_qbi_sha256( + rule.get("source_cells_sha256"), + boundary=boundary, + field="source_cells_sha256", + ) + for field in ( + "structurally_absent_person_lineages_sha256", + "affected_tax_unit_ids_sha256", + "empty_universe_tax_unit_ids_sha256", + "source_universe_sha256", + ): + _require_qbi_sha256(receipt.get(field), boundary=boundary, field=field) + by_origin_role = receipt.get("by_origin_role") + if not isinstance(by_origin_role, Mapping) or not all( + isinstance(name, str) + and name + and isinstance(count, int) + and not isinstance(count, bool) + and count >= 0 + for name, count in by_origin_role.items() + ): + raise ValueError( + f"{boundary}: QBI universe by-origin counts are malformed." + ) + source_receipt = { + key: value + for key, value in receipt.items() + if key + not in { + "source_universe_sha256", + "source_universe_resolution_mutated_raw_pums_cells", + "operation", + "rows_excluded_from_base_self_employment_rewrite", + "rows_included_in_other_qbi_reconciliation", + "structurally_absent_base_source_cells_mutated", + "sha256", + } + } + source_receipt["raw_pums_source_cells_mutated"] = receipt[ + "source_universe_resolution_mutated_raw_pums_cells" + ] + if _qbi_receipt_sha256(source_receipt) != receipt["source_universe_sha256"]: + raise ValueError( + f"{boundary}: QBI nested ACS source-universe SHA-256 mismatch." + ) + observed_sha256 = _require_qbi_sha256( + receipt.get("sha256"), + boundary=boundary, + field="source-universe sha256", + ) + unsigned = dict(receipt) + unsigned.pop("sha256") + if observed_sha256 != _qbi_receipt_sha256(unsigned): + raise ValueError(f"{boundary}: QBI source-universe SHA-256 mismatch.") + + +def _require_exact_qbi_keys( + value: Mapping[str, object], + expected: frozenset[str], + *, + boundary: str, + label: str, +) -> None: + missing = sorted(expected - set(value)) + extra = sorted(set(value) - expected) + if missing or extra: + raise ValueError( + f"{boundary}: QBI {label} schema mismatch; missing={missing}, " + f"extra={extra}." + ) + + +def _require_qbi_sha256( + value: object, + *, + boundary: str, + field: str, +) -> str: + if not isinstance(value, str) or _LOWERCASE_SHA256.fullmatch(value) is None: + raise ValueError( + f"{boundary}: QBI receipt field {field!r} must be a lowercase SHA-256." + ) + return value + + +def _require_qbi_nonnegative_integer( + value: object, + *, + boundary: str, + field: str, +) -> int: + if not isinstance(value, int) or isinstance(value, bool) or value < 0: + raise ValueError( + f"{boundary}: QBI receipt field {field!r} must be a nonnegative integer." + ) + return value + + +def _assert_qbi_undeclared_surface_unchanged( + before: Frame, + after: Frame, +) -> dict[str, object]: + """Fail if the whole-pool kernel changed anything outside its declaration.""" + + if before.schema != after.schema: + raise ValueError("US QBI reconciliation changed the frame schema.") + if before.entities != after.entities: + raise ValueError("US QBI reconciliation changed the entity inventory.") + if before.links != after.links: + raise ValueError("US QBI reconciliation changed the link-table inventory.") + if before.weighted_entities != after.weighted_entities: + raise ValueError("US QBI reconciliation changed the weighted-entity inventory.") + + person_entity = before.schema.person_entity + before_person = before.table(person_entity) + after_person = after.table(person_entity) + if list(before_person.columns) != list(after_person.columns): + raise ValueError("US QBI reconciliation changed the person-column inventory.") + declared = set(US_QBI_RECONCILED_PERSON_COLUMNS) + undeclared_person_columns = [ + column for column in before_person.columns if column not in declared + ] + if not before_person.loc[:, undeclared_person_columns].equals( + after_person.loc[:, undeclared_person_columns] + ): + changed = [ + column + for column in undeclared_person_columns + if not before_person[column].equals(after_person[column]) + ] + raise ValueError( + f"US QBI reconciliation mutated undeclared person column(s): {changed}." + ) + + changed_entities = [ + entity + for entity in before.entities + if entity != person_entity + and not before.table(entity).equals(after.table(entity)) + ] + if changed_entities: + raise ValueError( + "US QBI reconciliation mutated undeclared entity table(s): " + f"{changed_entities}." + ) + changed_links = [ + link for link in before.links if not before.link(link).equals(after.link(link)) + ] + if changed_links: + raise ValueError( + f"US QBI reconciliation mutated undeclared link table(s): {changed_links}." + ) + changed_weights = [] + for entity in before.weighted_entities: + before_weights = before.weights_for(entity) + after_weights = after.weights_for(entity) + if ( + before_weights.kind != after_weights.kind + or before_weights.values.dtype != after_weights.values.dtype + or before_weights.values.shape != after_weights.values.shape + or before_weights.values.tobytes() != after_weights.values.tobytes() + ): + changed_weights.append(entity) + if changed_weights: + raise ValueError( + "US QBI reconciliation mutated undeclared weight vector(s): " + f"{changed_weights}." + ) + if not before.strata.equals(after.strata): + raise ValueError("US QBI reconciliation mutated undeclared strata.") + if before.mass_log != after.mass_log: + raise ValueError("US QBI reconciliation mutated the undeclared mass log.") + if before.metadata != after.metadata: + raise ValueError("US QBI reconciliation mutated undeclared frame metadata.") + return { + "undeclared_person_columns_verified": len(undeclared_person_columns), + "non_person_entity_tables_verified": len(before.entities) - 1, + "link_tables_verified": len(before.links), + "weight_vectors_verified": len(before.weighted_entities), + "strata_verified": True, + "mass_log_verified": True, + "metadata_verified": True, + } + + +def _qbi_person_values_sha256( + person: pd.DataFrame, + *, + columns: tuple[str, ...], +) -> str: + id_column = "person_id" if "person_id" in person else None + hashed_columns = list(columns) + if id_column is not None and id_column not in hashed_columns: + hashed_columns.insert(0, id_column) + values = person.loc[:, hashed_columns] + header = { + "columns": hashed_columns, + "dtypes": [str(values[column].dtype) for column in hashed_columns], + } + digest = hashlib.sha256( + json.dumps(header, sort_keys=True, separators=(",", ":")).encode() + ) + digest.update( + pd.util.hash_pandas_object(values, index=True).to_numpy(dtype=" str: + """Bind the non-output person fields that determine QBI reconciliation.""" + + person = frame.table(frame.schema.person_entity) + present_driver_columns = tuple( + column for column in _QBI_UNDECLARED_DRIVER_PERSON_COLUMNS if column in person + ) + support_role_sha256 = None + if has_support_role_metadata(person, entity=frame.schema.person_entity): + roles = support_role_series(person, entity=frame.schema.person_entity) + support_role_sha256 = _qbi_table_values_sha256( + roles.rename("support_role").to_frame() + ) + payload = { + "version": 1, + "driver_presence": { + column: column in person for column in _QBI_UNDECLARED_DRIVER_PERSON_COLUMNS + }, + "driver_values_sha256": _qbi_table_values_sha256( + person.loc[:, list(present_driver_columns)] + ), + "support_role_sha256": support_role_sha256, + } + return _qbi_receipt_sha256(payload) + + +def _qbi_table_values_sha256(table: pd.DataFrame) -> str: + header = { + "columns": [str(column) for column in table.columns], + "dtypes": [str(table[column].dtype) for column in table.columns], + "index_type": type(table.index).__name__, + "index_dtype": str(table.index.dtype), + "index_names": list(table.index.names), + } + digest = hashlib.sha256( + json.dumps(header, sort_keys=True, separators=(",", ":")).encode() + ) + digest.update( + pd.util.hash_pandas_object(table, index=True).to_numpy(dtype=" Frame: """Restore archived all-or-nothing SSTB split identities after PUF QRF.""" @@ -171,6 +1240,8 @@ def with_us_qbi_input_reconciliation(frame: Frame) -> Frame: ) result = person.copy(deep=True) + base_self_employment_scope, _universe_receipt = _qbi_reconciliation_scope(frame) + reconciliation_scope = np.ones(len(result), dtype=bool) asec_mask = np.zeros(len(result), dtype=bool) if has_support_role_metadata(result, entity="person"): asec_mask = ( @@ -188,18 +1259,31 @@ def with_us_qbi_input_reconciliation(frame: Frame) -> Frame: # This is a hermetic two-channel choice, not a claim about the retired # clone implementation (which imputed absent PUF variables onto both halves). for column in _GENERAL_QUALIFICATION_FLAGS: - flags[column][asec_mask] = True - flags["business_is_sstb"][asec_mask] = False - flags[_SSTB_QUALIFICATION_FLAG][asec_mask] = False + flags[column][asec_mask & reconciliation_scope] = True + flags["business_is_sstb"][asec_mask & reconciliation_scope] = False + flags[_SSTB_QUALIFICATION_FLAG][asec_mask & reconciliation_scope] = False - non_sstb_self_employment = _numeric(result, _SELF_EMPLOYMENT_COLUMN) - sstb_self_employment = _numeric(result, _SSTB_SELF_EMPLOYMENT_COLUMN) + non_sstb_self_employment = _numeric_in_scope( + result, + _SELF_EMPLOYMENT_COLUMN, + base_self_employment_scope, + ) + non_sstb_self_employment = np.where( + base_self_employment_scope, + non_sstb_self_employment, + 0.0, + ) + sstb_self_employment = np.where( + base_self_employment_scope, + _numeric(result, _SSTB_SELF_EMPLOYMENT_COLUMN), + 0.0, + ) total_self_employment = non_sstb_self_employment + sstb_self_employment base_self_employment_qualified = ( flags["self_employment_income_would_be_qualified"] | flags[_SSTB_QUALIFICATION_FLAG] - ) + ) & base_self_employment_scope partnership_s_corp_income = _optional_numeric( result, "partnership_income" ) + _optional_numeric(result, "s_corp_income") @@ -220,43 +1304,63 @@ def with_us_qbi_input_reconciliation(frame: Frame) -> Frame: ) flags[_SSTB_QUALIFICATION_FLAG] = business_is_sstb & base_self_employment_qualified - result[_SELF_EMPLOYMENT_COLUMN] = np.where( + reconciled_self_employment = np.where( business_is_sstb, 0.0, total_self_employment, ) - result[_SSTB_SELF_EMPLOYMENT_COLUMN] = np.where( + reconciled_sstb_self_employment = np.where( business_is_sstb, total_self_employment, 0.0, ) + result.loc[base_self_employment_scope, _SELF_EMPLOYMENT_COLUMN] = ( + reconciled_self_employment[base_self_employment_scope] + ) + result.loc[reconciliation_scope, _SSTB_SELF_EMPLOYMENT_COLUMN] = ( + reconciled_sstb_self_employment[reconciliation_scope] + ) w2_wages = _numeric(result, "w2_wages_from_qualified_business") ubia = _numeric(result, "unadjusted_basis_qualified_property") - result["sstb_w2_wages_from_qualified_business"] = np.where( + reconciled_sstb_w2 = np.where( business_is_sstb, w2_wages, 0.0, ) - result["sstb_unadjusted_basis_qualified_property"] = np.where( + reconciled_sstb_ubia = np.where( business_is_sstb, ubia, 0.0, ) + result.loc[reconciliation_scope, "sstb_w2_wages_from_qualified_business"] = ( + reconciled_sstb_w2[reconciliation_scope] + ) + result.loc[reconciliation_scope, "sstb_unadjusted_basis_qualified_property"] = ( + reconciled_sstb_ubia[reconciliation_scope] + ) non_qualified_dividends = np.maximum( _optional_numeric(result, "non_qualified_dividend_income"), 0.0, ) - result["qualified_bdc_income"] = np.minimum( + reconciled_bdc = np.minimum( _numeric(result, "qualified_bdc_income"), non_qualified_dividends, ) - result["qualified_reit_and_ptp_income"] = np.minimum( + reconciled_reit = np.minimum( _numeric(result, "qualified_reit_and_ptp_income"), non_qualified_dividends + np.maximum(partnership_s_corp_income, 0.0), ) + result.loc[reconciliation_scope, "qualified_bdc_income"] = reconciled_bdc[ + reconciliation_scope + ] + result.loc[reconciliation_scope, "qualified_reit_and_ptp_income"] = reconciled_reit[ + reconciliation_scope + ] for column, values in flags.items(): - result[column] = values.astype(bool) + updated = result[column].astype(bool).to_numpy(copy=True) + updated[reconciliation_scope] = values[reconciliation_scope] + result[column] = updated tables = {entity: frame.table(entity).copy() for entity in frame.entities} tables["person"] = result @@ -274,6 +1378,8 @@ def us_qbi_inputs_summary(frame: Frame) -> dict[str, object]: """Return weighted signal, validity, and split-identity diagnostics.""" person = frame.table("person") + base_self_employment_scope, universe_receipt = _qbi_reconciliation_scope(frame) + reconciliation_scope = np.ones(len(person), dtype=bool) weights = np.asarray(frame.resolve_weights("person").values, dtype=np.float64) total_weight = float(weights.sum()) columns: dict[str, dict[str, object]] = {} @@ -299,9 +1405,16 @@ def us_qbi_inputs_summary(frame: Frame) -> dict[str, object]: } business = person["business_is_sstb"].astype(bool).to_numpy() - self_employment = pd.to_numeric( - person[_SELF_EMPLOYMENT_COLUMN], errors="coerce" - ).to_numpy(dtype=np.float64) + self_employment = _numeric_in_scope( + person, + _SELF_EMPLOYMENT_COLUMN, + base_self_employment_scope, + ) + self_employment = np.where( + base_self_employment_scope, + self_employment, + 0.0, + ) sstb_self_employment = pd.to_numeric( person[_SSTB_SELF_EMPLOYMENT_COLUMN], errors="coerce" ).to_numpy(dtype=np.float64) @@ -330,51 +1443,53 @@ def us_qbi_inputs_summary(frame: Frame) -> dict[str, object]: ) + _optional_numeric(person, "s_corp_income") qualified_bdc_income = _numeric(person, "qualified_bdc_income") qualified_reit_and_ptp_income = _numeric(person, "qualified_reit_and_ptp_income") + + def scoped_count(condition: np.ndarray) -> int: + return int(np.count_nonzero(reconciliation_scope & condition)) + invariants = { - "sstb_rows_with_non_sstb_income": int( - np.count_nonzero(business & ~np.isclose(self_employment, 0.0)) + "sstb_rows_with_non_sstb_income": scoped_count( + business & ~np.isclose(self_employment, 0.0) ), - "non_sstb_rows_with_sstb_income": int( - np.count_nonzero(~business & ~np.isclose(sstb_self_employment, 0.0)) + "non_sstb_rows_with_sstb_income": scoped_count( + ~business & ~np.isclose(sstb_self_employment, 0.0) ), - "sstb_w2_split_mismatches": int( - np.count_nonzero( - ~np.isclose(sstb_w2, np.where(business, w2, 0.0), atol=_INVARIANT_ATOL) + "sstb_w2_split_mismatches": scoped_count( + ~np.isclose( + sstb_w2, + np.where(business, w2, 0.0), + atol=_INVARIANT_ATOL, ) ), - "sstb_ubia_split_mismatches": int( - np.count_nonzero( - ~np.isclose( - sstb_ubia, - np.where(business, ubia, 0.0), - atol=_INVARIANT_ATOL, - ) + "sstb_ubia_split_mismatches": scoped_count( + ~np.isclose( + sstb_ubia, + np.where(business, ubia, 0.0), + atol=_INVARIANT_ATOL, ) ), - "self_employment_qualification_overlap": int( - np.count_nonzero(self_qualified & sstb_qualified) + "self_employment_qualification_overlap": scoped_count( + self_qualified & sstb_qualified ), - "sstb_qualification_route_mismatches": int( - np.count_nonzero(sstb_qualified & ~business) + "sstb_qualification_route_mismatches": scoped_count(sstb_qualified & ~business), + "non_sstb_qualification_route_mismatches": scoped_count( + self_qualified & business ), - "non_sstb_qualification_route_mismatches": int( - np.count_nonzero(self_qualified & business) + "qualified_bdc_exposure_mismatches": scoped_count( + qualified_bdc_income > non_qualified_dividends + _INVARIANT_ATOL ), - "qualified_bdc_exposure_mismatches": int( - np.count_nonzero( - qualified_bdc_income > non_qualified_dividends + _INVARIANT_ATOL - ) - ), - "qualified_reit_ptp_exposure_mismatches": int( - np.count_nonzero( - qualified_reit_and_ptp_income - > non_qualified_dividends - + np.maximum(partnership_s_corp_income, 0.0) - + _INVARIANT_ATOL - ) + "qualified_reit_ptp_exposure_mismatches": scoped_count( + qualified_reit_and_ptp_income + > non_qualified_dividends + + np.maximum(partnership_s_corp_income, 0.0) + + _INVARIANT_ATOL ), } - return {"columns": columns, "invariants": invariants} + return { + "columns": columns, + "invariants": invariants, + "reconciliation_universe": universe_receipt, + } def us_qbi_inputs_signal_gate(frame: Frame) -> GateResult: diff --git a/packages/microcosm-build/src/microcosm/build/us_runtime/stacked_spine.py b/packages/microcosm-build/src/microcosm/build/us_runtime/stacked_spine.py index 843b212b..73496c26 100644 --- a/packages/microcosm-build/src/microcosm/build/us_runtime/stacked_spine.py +++ b/packages/microcosm-build/src/microcosm/build/us_runtime/stacked_spine.py @@ -47,8 +47,15 @@ import numpy as np import pandas as pd -from microcosm.build.gates import FitWeightRecord, GateResult +from microcosm.build.gates import ( + FitWeightRecord, + GateResult, + _sealed_stacked_gate_result, +) from microcosm.build.serialization_dtypes import canonicalize_table_string_dtypes +from microcosm.build.us_runtime.acs_income_universe import ( + apply_acs_pums_earnings_universe_zeros, +) from microcosm.build.us_runtime.acs_transfer import ( DEFAULT_ACS_TRANSFER_MAX_TARGETS_PER_FIT, AcsTransferResult, @@ -57,9 +64,18 @@ transfer_acs_inputs, ) from microcosm.build.us_runtime.multispine_pool import ( + POOL_OPERATOR_CONTRACTS, + POOL_PRE_CLONE_SOURCE_OPERATOR_ORDER, POOL_SPINE_AGREEMENT_REGISTRY, + pool_post_puf_puf_producer_target_families, + pool_post_puf_source_producer_target_families, + pool_post_puf_transfer_target_families, + pool_pre_clone_gap_fill_target_families, pool_transfer_target_families, ) +from microcosm.build.us_runtime.operator_boundary import ( + PRE_ASSEMBLY_OPERATOR_OUTPUT_FAMILIES, +) from microcosm.build.us_runtime.puf_capital_gains_tail import ( PUF_CAPITAL_GAINS_TAIL_APPLIED_COLUMN, PUF_CAPITAL_GAINS_TAIL_DONOR_AGI_BAND_COLUMN, @@ -77,10 +93,12 @@ PRIMARY_QRF_MANIFEST_FILENAME, finalize_primary_puf_qrf_chain, initialize_primary_puf_qrf_chain, + primary_puf_qrf_recipient_predictor_universe_receipt, run_primary_puf_qrf_chain, ) from microcosm.build.us_runtime.puf_support import ( PUF_ABSENT_CELLS_PRESERVE_NULLS, + PUF_CLONE_ATTACHMENT_MANIFEST_KEY, PUF_TAX_DETAIL_DEFAULT_PERSON_OUTPUTS, PUF_TAX_DETAIL_DEFAULT_TAX_UNIT_OUTPUTS, US_PUF_SUPPORT_FIT_NAME, @@ -92,6 +110,7 @@ from microcosm.build.us_runtime.spine_assembly import assemble_spines from microcosm.build.us_runtime.support_provenance import ( BASE_ASEC_SUPPORT_CHANNEL, + PUF_TAX_DETAIL_CLONE_INDEX, SPINE_ASSEMBLY_MANIFEST_KEY, spine_source_id_column, support_channel_column, @@ -109,16 +128,21 @@ "CANONICAL_STACKED_DECLARED_SURFACE", "CANONICAL_STACKED_GAP_FILL_SURFACE", "CANONICAL_STACKED_GAP_FILL_PLAN", + "CANONICAL_STACKED_POST_PUF_PUF_PRODUCER_SURFACE", + "CANONICAL_STACKED_POST_PUF_SOURCE_PRODUCER_SURFACE", + "CANONICAL_STACKED_POST_PUF_TRANSFER_SURFACE", "DEFAULT_STACKED_HOUSEHOLD_MASS_SHARES", "ORIGIN_BATTERY_METRIC_KINDS", "STACKED_PILOT_ACS_SAMPLE_FRACTION", "STACKED_PILOT_ACS_SAMPLE_SEED", "STACKED_SPINE_MANIFEST_KEY", "AbsenceProof", + "GapFillAbsenceRule", "GapFillDirection", "GapFillResult", "OriginBatterySpec", "StackedPufPassResult", + "StackedPostPufTransferResult", "StackedSpineResult", "assemble_stacked_spine", "assert_stacked_tail_cells_preserved", @@ -129,7 +153,10 @@ "sample_acs_households", "stacked_completeness_gate", "stacked_gap_fill_plan", + "stacked_gap_fill_producer_schedule_receipt", "stacked_spine_authority_receipt", + "transfer_stacked_post_puf_inputs", + "validate_stacked_post_puf_transfer_receipt", "validate_stacked_spine_frame", ] @@ -146,12 +173,14 @@ STACKED_SPINE_MANIFEST_KEY = "us_stacked_spine_manifest" _LEGACY_STACKED_SPINE_MANIFEST_VERSION = 1 -_STACKED_SPINE_MANIFEST_VERSION = 2 +_STACKED_SPINE_MANIFEST_VERSION = 4 _SUPPORTED_STACKED_SPINE_MANIFEST_VERSIONS = { _LEGACY_STACKED_SPINE_MANIFEST_VERSION, _STACKED_SPINE_MANIFEST_VERSION, } _EXACT_COUNT_RULE = "floor(fraction * eligible)" +_ACS_NATIVE_GQ_LINEAGE_VERSION = 1 +_ACS_NATIVE_GQ_SELECTION = "TYPEHUGQ in {2,3} on sampled native ACS rows" _MASS_RTOL = 1e-9 #: The ratified pilot stack configuration (#578 revision): a seeded 10% ACS @@ -284,6 +313,167 @@ def sample_acs_households( ) +def _acs_native_group_quarters_receipt( + sampled_acs: Frame, + assembled: Frame, +) -> dict[str, object]: + """Bind source GQ evidence to immutable assembly-time ACS lineages.""" + + household = sampled_acs.table("household") + person = sampled_acs.table("person") + missing_household = sorted({"household_id", "TYPEHUGQ"} - set(household)) + missing_person = sorted({"person_id", "person_household_id"} - set(person)) + if missing_household or missing_person: + raise ValueError( + "Stacked ACS assembly cannot bind native group-quarters lineage; " + f"missing_household={missing_household}, missing_person={missing_person}." + ) + kind = pd.to_numeric(household["TYPEHUGQ"], errors="coerce") + invalid = ~kind.isin((1, 2, 3)) + if invalid.any(): + raise ValueError( + "Stacked ACS assembly requires every sampled household to carry a " + f"TYPEHUGQ 1/2/3 source classification; found {int(invalid.sum())} " + "invalid row(s)." + ) + gq_households = kind.isin((2, 3)) + household_ids = np.sort( + household.loc[gq_households, "household_id"].to_numpy(dtype=np.int64) + ) + gq_people = person["person_household_id"].isin(household_ids) + person_lineages = person.loc[ + gq_people, ["person_id", "person_household_id"] + ].to_numpy(dtype=np.int64) + if len(person_lineages): + person_lineages = person_lineages[ + np.lexsort((person_lineages[:, 1], person_lineages[:, 0])) + ] + else: + person_lineages = person_lineages.reshape(0, 2) + linked_counts = person.loc[gq_people, "person_household_id"].value_counts() + if len(linked_counts) != len(household_ids) or not linked_counts.eq(1).all(): + raise ValueError( + "Stacked ACS assembly requires exactly one native person per " + "TYPEHUGQ 2/3 group-quarters household." + ) + + assembled_household = assembled.table("household") + assembled_person = assembled.table("person") + household_channel = assembled_household[support_channel_column("household")].astype( + str + ) + household_clone = pd.to_numeric( + assembled_household[support_clone_index_column("household")], + errors="raise", + ).astype("int64") + native_acs_household = household_channel.eq(ACS_STACKED_SUPPORT_CHANNEL) & ( + household_clone.eq(0) + ) + assembled_kind = pd.to_numeric(assembled_household["TYPEHUGQ"], errors="coerce") + native_household_lineages = np.column_stack( + ( + assembled_household.loc[native_acs_household, "household_id"].to_numpy( + dtype=np.int64 + ), + assembled_household.loc[ + native_acs_household, + support_source_id_column("household"), + ].to_numpy(dtype=np.int64), + assembled_household.loc[ + native_acs_household, + spine_source_id_column("household"), + ].to_numpy(dtype=np.int64), + assembled_kind.loc[native_acs_household].to_numpy(dtype=np.int64), + ) + ) + sorted_household_lineages = native_household_lineages[ + np.lexsort( + tuple( + native_household_lineages[:, column] + for column in reversed(range(native_household_lineages.shape[1])) + ) + ) + ] + + native_household_support_by_id = pd.Series( + assembled_household.loc[ + native_acs_household, + support_source_id_column("household"), + ].to_numpy(dtype=np.int64), + index=assembled_household.loc[native_acs_household, "household_id"].to_numpy( + dtype=np.int64 + ), + ) + native_household_kind_by_id = pd.Series( + assembled_kind.loc[native_acs_household].to_numpy(dtype=np.int64), + index=assembled_household.loc[native_acs_household, "household_id"].to_numpy( + dtype=np.int64 + ), + ) + person_channel = assembled_person[support_channel_column("person")].astype(str) + person_clone = pd.to_numeric( + assembled_person[support_clone_index_column("person")], errors="raise" + ).astype("int64") + native_acs_person = person_channel.eq( + ACS_STACKED_SUPPORT_CHANNEL + ) & person_clone.eq(0) + parent_support = assembled_person.loc[native_acs_person, "person_household_id"].map( + native_household_support_by_id + ) + parent_kind = assembled_person.loc[native_acs_person, "person_household_id"].map( + native_household_kind_by_id + ) + if parent_support.isna().any() or parent_kind.isna().any(): + raise ValueError( + "Stacked ACS assembly cannot bind native person-to-household lineages." + ) + native_person_lineages = np.column_stack( + ( + assembled_person.loc[native_acs_person, "person_id"].to_numpy( + dtype=np.int64 + ), + assembled_person.loc[ + native_acs_person, + support_source_id_column("person"), + ].to_numpy(dtype=np.int64), + assembled_person.loc[ + native_acs_person, + spine_source_id_column("person"), + ].to_numpy(dtype=np.int64), + parent_support.to_numpy(dtype=np.int64), + parent_kind.to_numpy(dtype=np.int64), + ) + ) + sorted_person_lineages = native_person_lineages[ + np.lexsort( + tuple( + native_person_lineages[:, column] + for column in reversed(range(native_person_lineages.shape[1])) + ) + ) + ] + return { + "version": _ACS_NATIVE_GQ_LINEAGE_VERSION, + "source_channel": ACS_STACKED_SUPPORT_CHANNEL, + "selection": _ACS_NATIVE_GQ_SELECTION, + "one_person_per_household": True, + "household_count": int(len(household_ids)), + "person_count": int(len(person_lineages)), + "household_spine_source_ids_sha256": _ids_sha256(household_ids), + "person_spine_lineages_sha256": _integer_rows_sha256(person_lineages), + "native_household_count": int(len(native_household_lineages)), + "native_person_count": int(len(native_person_lineages)), + "native_household_mapping_sha256": _integer_rows_sha256( + sorted_household_lineages + ), + "native_household_order_sha256": _integer_rows_sha256( + native_household_lineages + ), + "native_person_mapping_sha256": _integer_rows_sha256(sorted_person_lineages), + "native_person_order_sha256": _integer_rows_sha256(native_person_lineages), + } + + def _normalize_sampled_household_mass( sampled: Frame, *, @@ -462,6 +652,10 @@ def assemble_stacked_spine( household_mass_shares=shares, mass_anchor_channel=mass_anchor_channel, ) + acs_native_group_quarters = _acs_native_group_quarters_receipt( + sampled_acs, + assembled, + ) harmonization = _harmonization_receipt( assembled, @@ -477,6 +671,7 @@ def assemble_stacked_spine( }, "mass_anchor_channel": mass_anchor_channel, "weight_harmonization": harmonization, + "acs_native_group_quarters": acs_native_group_quarters, } # The assembly metadata is preserved in full and augmented with the stack # manifest; the mass history is carried unchanged from the same source. @@ -604,6 +799,587 @@ def _validate_survey_sample_receipt( ) +def _validated_acs_native_group_quarters_masks( + frame: Frame, + manifest: Mapping[str, object], + *, + boundary: str, +) -> tuple[pd.Series, pd.Series]: + """Prove live ACS GQ classifications against assembly-bound lineages.""" + + receipt = manifest.get("acs_native_group_quarters") + required_receipt_keys = { + "version", + "source_channel", + "selection", + "one_person_per_household", + "household_count", + "person_count", + "household_spine_source_ids_sha256", + "person_spine_lineages_sha256", + "native_household_count", + "native_person_count", + "native_household_mapping_sha256", + "native_household_order_sha256", + "native_person_mapping_sha256", + "native_person_order_sha256", + } + if not isinstance(receipt, Mapping) or set(receipt) != required_receipt_keys: + raise ValueError( + f"{boundary}: stacked spine native ACS group-quarters lineage " + "receipt is absent or malformed." + ) + if ( + receipt.get("version") != _ACS_NATIVE_GQ_LINEAGE_VERSION + or receipt.get("source_channel") != ACS_STACKED_SUPPORT_CHANNEL + or receipt.get("selection") != _ACS_NATIVE_GQ_SELECTION + or receipt.get("one_person_per_household") is not True + ): + raise ValueError( + f"{boundary}: stacked spine native ACS group-quarters lineage " + "receipt declares unsupported authority." + ) + for receipt_field in ( + "household_count", + "person_count", + "native_household_count", + "native_person_count", + ): + value = receipt.get(receipt_field) + if isinstance(value, bool) or not isinstance(value, int) or value < 0: + raise ValueError( + f"{boundary}: stacked spine native ACS group-quarters " + f"{receipt_field} " + f"must be a non-negative integer, got {value!r}." + ) + for receipt_field in ( + "household_spine_source_ids_sha256", + "person_spine_lineages_sha256", + "native_household_mapping_sha256", + "native_household_order_sha256", + "native_person_mapping_sha256", + "native_person_order_sha256", + ): + value = receipt.get(receipt_field) + if ( + not isinstance(value, str) + or len(value) != 64 + or any(character not in "0123456789abcdef" for character in value) + ): + raise ValueError( + f"{boundary}: stacked spine native ACS group-quarters " + f"{receipt_field} " + "must be a lowercase SHA-256 digest." + ) + + household = frame.table("household") + person = frame.table("person") + required_household = { + "household_id", + "TYPEHUGQ", + spine_source_id_column("household"), + support_source_id_column("household"), + support_channel_column("household"), + support_clone_index_column("household"), + } + required_person = { + "person_id", + "person_household_id", + spine_source_id_column("person"), + support_source_id_column("person"), + support_channel_column("person"), + support_clone_index_column("person"), + } + missing_household = sorted(required_household - set(household)) + missing_person = sorted(required_person - set(person)) + if missing_household or missing_person: + raise ValueError( + f"{boundary}: live native ACS group-quarters lineage cannot be " + f"validated; missing_household={missing_household}, " + f"missing_person={missing_person}." + ) + + household_channel = household[support_channel_column("household")].astype(str) + household_clone = pd.to_numeric( + household[support_clone_index_column("household")], errors="raise" + ).astype("int64") + acs_household = household_channel.eq(ACS_STACKED_SUPPORT_CHANNEL) + native_acs_household = acs_household & household_clone.eq(0) + kind = pd.to_numeric(household["TYPEHUGQ"], errors="coerce") + invalid_kind = acs_household & ~kind.isin((1, 2, 3)) + if invalid_kind.any(): + raise ValueError( + f"{boundary}: live ACS group-quarters lineage has " + f"{int(invalid_kind.sum())} row(s) without TYPEHUGQ 1/2/3." + ) + gq_household = acs_household & kind.isin((2, 3)) + native_gq_household = native_acs_household & gq_household + native_household_ids = np.sort( + household.loc[ + native_gq_household, + spine_source_id_column("household"), + ].to_numpy(dtype=np.int64) + ) + if ( + len(native_household_ids) != receipt["household_count"] + or _ids_sha256(native_household_ids) + != receipt["household_spine_source_ids_sha256"] + ): + raise ValueError( + f"{boundary}: live native ACS group-quarters household lineage " + "differs from its assembly-bound count or digest." + ) + + native_support_ids = household.loc[ + native_acs_household, + support_source_id_column("household"), + ] + if native_support_ids.duplicated().any(): + raise ValueError( + f"{boundary}: live native ACS household support lineages are not unique." + ) + native_household_lineages = np.column_stack( + ( + household.loc[native_acs_household, "household_id"].to_numpy( + dtype=np.int64 + ), + native_support_ids.to_numpy(dtype=np.int64), + household.loc[ + native_acs_household, + spine_source_id_column("household"), + ].to_numpy(dtype=np.int64), + kind.loc[native_acs_household].to_numpy(dtype=np.int64), + ) + ) + sorted_household_lineages = native_household_lineages[ + np.lexsort( + tuple( + native_household_lineages[:, column] + for column in reversed(range(native_household_lineages.shape[1])) + ) + ) + ] + if ( + len(native_household_lineages) != receipt["native_household_count"] + or _integer_rows_sha256(sorted_household_lineages) + != receipt["native_household_mapping_sha256"] + or _integer_rows_sha256(native_household_lineages) + != receipt["native_household_order_sha256"] + ): + raise ValueError( + f"{boundary}: live native ACS household support/raw/classification " + "mapping differs from its assembly-bound digest." + ) + native_classification = pd.Series( + gq_household.loc[native_acs_household].to_numpy(dtype=bool), + index=native_support_ids.to_numpy(dtype=np.int64), + ) + native_spine_source_by_support = pd.Series( + household.loc[ + native_acs_household, + spine_source_id_column("household"), + ].to_numpy(dtype=np.int64), + index=native_support_ids.to_numpy(dtype=np.int64), + ) + live_acs_support_ids = household.loc[ + acs_household, + support_source_id_column("household"), + ] + expected_classification = live_acs_support_ids.map(native_classification) + expected_spine_source = live_acs_support_ids.map(native_spine_source_by_support) + if expected_classification.isna().any() or not np.array_equal( + expected_classification.to_numpy(dtype=bool), + gq_household.loc[acs_household].to_numpy(dtype=bool), + ): + raise ValueError( + f"{boundary}: live ACS group-quarters classification differs across " + "clone roles from its assembly-bound native lineage." + ) + if expected_spine_source.isna().any() or not np.array_equal( + expected_spine_source.to_numpy(dtype=np.int64), + household.loc[ + acs_household, + spine_source_id_column("household"), + ].to_numpy(dtype=np.int64), + ): + raise ValueError( + f"{boundary}: live ACS support/raw household lineage pairs differ " + "from their assembly-bound native lineage." + ) + native_pairs = set( + zip( + native_support_ids.to_numpy(dtype=np.int64), + household.loc[ + native_acs_household, + spine_source_id_column("household"), + ].to_numpy(dtype=np.int64), + strict=True, + ) + ) + for clone_role in sorted(int(value) for value in household_clone.unique()): + role = acs_household & household_clone.eq(clone_role) + role_pairs = list( + zip( + household.loc[ + role, + support_source_id_column("household"), + ].to_numpy(dtype=np.int64), + household.loc[ + role, + spine_source_id_column("household"), + ].to_numpy(dtype=np.int64), + strict=True, + ) + ) + if len(role_pairs) != len(set(role_pairs)): + raise ValueError( + f"{boundary}: live ACS household lineages are not unique in " + f"clone role {clone_role}." + ) + if ( + clone_role == PUF_TAX_DETAIL_CLONE_INDEX + and PUF_CLONE_ATTACHMENT_MANIFEST_KEY not in frame.metadata + and set(role_pairs) != native_pairs + ): + raise ValueError( + f"{boundary}: unreceipted ACS clone role {clone_role} does not " + "exactly preserve every native support/raw lineage pair." + ) + attachment = frame.metadata.get(PUF_CLONE_ATTACHMENT_MANIFEST_KEY) + if attachment is not None: + if not isinstance(attachment, Mapping): + raise ValueError(f"{boundary}: clone attachment receipt is malformed.") + detail = household_clone.eq(PUF_TAX_DETAIL_CLONE_INDEX) + selected_support_ids = np.sort( + household.loc[ + detail, + support_source_id_column("household"), + ].to_numpy(dtype=np.int64) + ) + if attachment.get("realized_household_count") != int( + len(selected_support_ids) + ) or attachment.get("selected_household_source_ids_sha256") != _ids_sha256( + selected_support_ids + ): + raise ValueError( + f"{boundary}: live clone-1 household lineages differ from the " + "attachment-bound selection count or digest." + ) + + person_channel = person[support_channel_column("person")].astype(str) + person_clone = pd.to_numeric( + person[support_clone_index_column("person")], errors="raise" + ).astype("int64") + native_acs_person = person_channel.eq( + ACS_STACKED_SUPPORT_CHANNEL + ) & person_clone.eq(0) + native_household_support_by_live_id = pd.Series( + household.loc[ + native_acs_household, + support_source_id_column("household"), + ].to_numpy(dtype=np.int64), + index=household.loc[native_acs_household, "household_id"].to_numpy( + dtype=np.int64 + ), + ) + native_household_kind_by_live_id = pd.Series( + kind.loc[native_acs_household].to_numpy(dtype=np.int64), + index=household.loc[native_acs_household, "household_id"].to_numpy( + dtype=np.int64 + ), + ) + native_parent_support = person.loc[native_acs_person, "person_household_id"].map( + native_household_support_by_live_id + ) + native_parent_kind = person.loc[native_acs_person, "person_household_id"].map( + native_household_kind_by_live_id + ) + if native_parent_support.isna().any() or native_parent_kind.isna().any(): + raise ValueError( + f"{boundary}: live native ACS person-to-household lineage cannot be " + "resolved." + ) + native_person_mapping = np.column_stack( + ( + person.loc[native_acs_person, "person_id"].to_numpy(dtype=np.int64), + person.loc[ + native_acs_person, + support_source_id_column("person"), + ].to_numpy(dtype=np.int64), + person.loc[ + native_acs_person, + spine_source_id_column("person"), + ].to_numpy(dtype=np.int64), + native_parent_support.to_numpy(dtype=np.int64), + native_parent_kind.to_numpy(dtype=np.int64), + ) + ) + sorted_person_mapping = native_person_mapping[ + np.lexsort( + tuple( + native_person_mapping[:, column] + for column in reversed(range(native_person_mapping.shape[1])) + ) + ) + ] + if ( + len(native_person_mapping) != receipt["native_person_count"] + or _integer_rows_sha256(sorted_person_mapping) + != receipt["native_person_mapping_sha256"] + or _integer_rows_sha256(native_person_mapping) + != receipt["native_person_order_sha256"] + ): + raise ValueError( + f"{boundary}: live native ACS person support/raw/parent mapping " + "differs from its assembly-bound digest." + ) + + native_person_support = person.loc[ + native_acs_person, + support_source_id_column("person"), + ] + if native_person_support.duplicated().any(): + raise ValueError( + f"{boundary}: live native ACS person support lineages are not unique." + ) + native_person_raw_by_support = pd.Series( + person.loc[ + native_acs_person, + spine_source_id_column("person"), + ].to_numpy(dtype=np.int64), + index=native_person_support.to_numpy(dtype=np.int64), + ) + native_person_parent_by_support = pd.Series( + native_parent_support.to_numpy(dtype=np.int64), + index=native_person_support.to_numpy(dtype=np.int64), + ) + native_person_parent_kind_by_support = pd.Series( + native_parent_kind.to_numpy(dtype=np.int64), + index=native_person_support.to_numpy(dtype=np.int64), + ) + + acs_person = person_channel.eq(ACS_STACKED_SUPPORT_CHANNEL) + acs_household_live_ids = household.loc[acs_household, "household_id"] + if acs_household_live_ids.duplicated().any(): + raise ValueError(f"{boundary}: live ACS household IDs are not unique.") + household_support_by_live_id = pd.Series( + household.loc[ + acs_household, + support_source_id_column("household"), + ].to_numpy(dtype=np.int64), + index=acs_household_live_ids.to_numpy(dtype=np.int64), + ) + household_kind_by_live_id = pd.Series( + kind.loc[acs_household].to_numpy(dtype=np.int64), + index=acs_household_live_ids.to_numpy(dtype=np.int64), + ) + household_clone_by_live_id = pd.Series( + household_clone.loc[acs_household].to_numpy(dtype=np.int64), + index=acs_household_live_ids.to_numpy(dtype=np.int64), + ) + live_person_support = person.loc[ + acs_person, + support_source_id_column("person"), + ] + expected_person_raw = live_person_support.map(native_person_raw_by_support) + expected_parent_support = live_person_support.map(native_person_parent_by_support) + expected_parent_kind = live_person_support.map(native_person_parent_kind_by_support) + live_parent_support = person.loc[acs_person, "person_household_id"].map( + household_support_by_live_id + ) + live_parent_kind = person.loc[acs_person, "person_household_id"].map( + household_kind_by_live_id + ) + live_parent_clone = person.loc[acs_person, "person_household_id"].map( + household_clone_by_live_id + ) + unresolved_person_lineage = any( + values.isna().any() + for values in ( + expected_person_raw, + expected_parent_support, + expected_parent_kind, + live_parent_support, + live_parent_kind, + live_parent_clone, + ) + ) + if unresolved_person_lineage or not ( + np.array_equal( + expected_person_raw.to_numpy(dtype=np.int64), + person.loc[ + acs_person, + spine_source_id_column("person"), + ].to_numpy(dtype=np.int64), + ) + and np.array_equal( + expected_parent_support.to_numpy(dtype=np.int64), + live_parent_support.to_numpy(dtype=np.int64), + ) + and np.array_equal( + expected_parent_kind.to_numpy(dtype=np.int64), + live_parent_kind.to_numpy(dtype=np.int64), + ) + and np.array_equal( + person_clone.loc[acs_person].to_numpy(dtype=np.int64), + live_parent_clone.to_numpy(dtype=np.int64), + ) + ): + raise ValueError( + f"{boundary}: live ACS person support/raw/parent/classification " + "lineages differ from their assembly-bound native mappings." + ) + + for clone_role in sorted(int(value) for value in household_clone.unique()): + role_households = acs_household & household_clone.eq(clone_role) + role_parent_support = set( + household.loc[ + role_households, + support_source_id_column("household"), + ].to_numpy(dtype=np.int64) + ) + expected_role_people = set( + native_person_parent_by_support.index[ + native_person_parent_by_support.isin(role_parent_support) + ].to_numpy(dtype=np.int64) + ) + role_people = acs_person & person_clone.eq(clone_role) + role_person_support = person.loc[ + role_people, + support_source_id_column("person"), + ].to_numpy(dtype=np.int64) + if ( + len(role_person_support) != len(set(role_person_support)) + or set(role_person_support) != expected_role_people + ): + raise ValueError( + f"{boundary}: live ACS clone role {clone_role} person lineages " + "do not exactly cover its assembly-bound household selection." + ) + native_gq_household_live_ids = household.loc[native_gq_household, "household_id"] + native_gq_person = native_acs_person & person["person_household_id"].isin( + native_gq_household_live_ids + ) + household_spine_source_by_live_id = pd.Series( + household[spine_source_id_column("household")].to_numpy(dtype=np.int64), + index=household["household_id"].to_numpy(dtype=np.int64), + ) + native_person_parent_sources = person.loc[ + native_gq_person, "person_household_id" + ].map(household_spine_source_by_live_id) + if native_person_parent_sources.isna().any(): + raise ValueError( + f"{boundary}: live native ACS group-quarters person parent lineage " + "cannot be resolved." + ) + native_person_lineages = np.column_stack( + ( + person.loc[ + native_gq_person, + spine_source_id_column("person"), + ].to_numpy(dtype=np.int64), + native_person_parent_sources.to_numpy(dtype=np.int64), + ) + ) + if len(native_person_lineages): + native_person_lineages = native_person_lineages[ + np.lexsort((native_person_lineages[:, 1], native_person_lineages[:, 0])) + ] + else: + native_person_lineages = native_person_lineages.reshape(0, 2) + if ( + len(native_person_lineages) != receipt["person_count"] + or _integer_rows_sha256(native_person_lineages) + != receipt["person_spine_lineages_sha256"] + ): + raise ValueError( + f"{boundary}: live native ACS group-quarters person lineage differs " + "from its assembly-bound count or digest." + ) + + gq_household_live_ids = household.loc[gq_household, "household_id"] + gq_person = person_channel.eq(ACS_STACKED_SUPPORT_CHANNEL) & person[ + "person_household_id" + ].isin(gq_household_live_ids) + linked_counts = person.loc[gq_person, "person_household_id"].value_counts() + if len(linked_counts) != int(gq_household.sum()) or not linked_counts.eq(1).all(): + raise ValueError( + f"{boundary}: live ACS group-quarters lineage requires exactly one " + "linked person per household in every clone role." + ) + return gq_household, gq_person + + +def _validate_stacked_clone_role_lifecycle( + frame: Frame, + *, + boundary: str, +) -> None: + """Require one exact clone-role set authorized by the live lifecycle.""" + + role_sets: dict[str, set[int]] = {} + for entity in frame.entities: + table = frame.table(entity) + clone_column = support_clone_index_column(entity) + if clone_column not in table: + raise ValueError( + f"{boundary}: live stacked {entity} rows lack {clone_column!r}." + ) + numeric = pd.to_numeric(table[clone_column], errors="raise") + if numeric.isna().any() or not np.equal(numeric, np.floor(numeric)).all(): + raise ValueError( + f"{boundary}: live stacked {entity} clone roles must be integers." + ) + role_sets[entity] = set(numeric.to_numpy(dtype=np.int64).tolist()) + + household_roles = role_sets["household"] + attachment = frame.metadata.get(PUF_CLONE_ATTACHMENT_MANIFEST_KEY) + if attachment is None: + if household_roles == {0}: + expected_roles = {0} + elif household_roles == {0, PUF_TAX_DETAIL_CLONE_INDEX}: + validate_puf_clone_attachment( + frame, + boundary=f"{boundary} full clone identity", + expected_fraction=1.0, + # Full-clone frames intentionally carry no attachment + # metadata, so the seed is not part of frame identity. The + # validator uses this value only in its returned receipt. + expected_seed=0, + ) + expected_roles = {0, PUF_TAX_DETAIL_CLONE_INDEX} + else: + raise ValueError( + f"{boundary}: unreceipted stacked clone roles " + f"{sorted(household_roles)} are unauthorized." + ) + else: + validated_attachment = validate_puf_clone_attachment( + frame, + boundary=f"{boundary} clone attachment", + ) + version = validated_attachment.get("version") + if version == 1: + expected_roles = {0, PUF_TAX_DETAIL_CLONE_INDEX} + elif version == 2: + expected_roles = {0, PUF_TAX_DETAIL_CLONE_INDEX, 2} + else: # pragma: no cover - the attachment validator rejects this first. + raise ValueError( + f"{boundary}: clone attachment authorizes no known role lifecycle." + ) + + inconsistent = { + entity: sorted(roles) + for entity, roles in role_sets.items() + if roles != expected_roles + } + if inconsistent: + raise ValueError( + f"{boundary}: stacked clone roles must exactly equal " + f"{sorted(expected_roles)} in every entity; got {inconsistent}." + ) + + def validate_stacked_spine_frame( frame: Frame, *, @@ -643,6 +1419,7 @@ def validate_stacked_spine_frame( f"{boundary}: stacked spine requires exactly the channels " f"{sorted(expected_channels)}; assembly declares {sorted(channels)}." ) + _validate_stacked_clone_role_lifecycle(frame, boundary=boundary) if version == _LEGACY_STACKED_SPINE_MANIFEST_VERSION: fraction = manifest.get("acs_sample_fraction") @@ -702,6 +1479,12 @@ def validate_stacked_spine_frame( require_normalization=version == _STACKED_SPINE_MANIFEST_VERSION, ) + _validated_acs_native_group_quarters_masks( + frame, + manifest, + boundary=boundary, + ) + household = frame.table("household") channel_values = household[support_channel_column("household")].astype(str) @@ -830,6 +1613,17 @@ def _ids_sha256(ids: np.ndarray) -> str: return hashlib.sha256(payload.encode()).hexdigest() +def _integer_rows_sha256(rows: np.ndarray) -> str: + values = np.asarray(rows, dtype=np.int64) + if values.ndim != 2: + raise ValueError("Lineage digest rows must be a two-dimensional array.") + payload = json.dumps( + [[int(value) for value in row] for row in values.tolist()], + separators=(",", ":"), + ) + return hashlib.sha256(payload.encode()).hexdigest() + + def _validate_fraction(fraction: float, *, boundary: str | None = None) -> None: prefix = f"{boundary}: " if boundary else "" if ( @@ -871,12 +1665,22 @@ def thaw(item: object) -> object: # --------------------------------------------------------------------------- _GAP_FILL_ASEC_TO_ACS = "asec_survey_to_acs" -_GAP_FILL_ACS_TO_ASEC = "acs_housing_to_asec" +_GAP_FILL_ASEC_HOUSING_TO_ACS = "asec_housing_to_acs" _GAP_FILL_HOUSING_FAMILY = "housing" _STACKED_AUTHORITY_ID = "us_stacked_spine_authority" -_STACKED_AUTHORITY_VERSION = 2 +# v6 binds the ASEC-consistent ACS earnings-universe application and the +# authenticated whole-pool QBI mutation semantics into the outer identity. +_STACKED_AUTHORITY_VERSION = 6 _CANONICAL_AUTHORITY_FORM = "CANONICAL" _NONCANONICAL_AUTHORITY_FORM = "NON-CANONICAL" +_PRE_CLONE_PREPARATION_STAGE = "prepare_multispine_source_inputs_for_clone" +_POST_GAP_FILL_STAGE = "after_gap_fill_stacked_spine" +_ACS_GQ_RENT_ABSENCE_RULE_ID = "acs_native_group_quarters_without_housing_unit" +_ACS_GQ_RENT_ABSENCE_SELECTION = "acs_typehugq_2_or_3_person" +_ACS_GQ_RENT_ABSENCE_REASON = ( + "ACS TYPEHUGQ 2/3 rows have no observed housing unit; rent must remain " + "structurally absent rather than be synthesized as zero or donor housing." +) def _freeze_target_families(target_families: TargetFamilies) -> TargetFamilies: @@ -908,22 +1712,47 @@ def _freeze_target_families(target_families: TargetFamilies) -> TargetFamilies: @dataclass(frozen=True) -class GapFillDirection: - """One declared cross-origin fill: recipient origin <- donor origin. +class GapFillAbsenceRule: + """One digest-bound exact recipient-universe rule for structural nulls.""" - Activation authority is declared here, not inferred from nullness: the - named recipient channel's rows are the only rows the direction may fill, - and the named donor channel's native rows are the only donor evidence. - The transfer machinery itself stays spine-blind; this owner-level - declaration is what makes the run-7 silent-skip class impossible — a - direction either fills its declared families on its declared rows or - fails by name. - """ + rule_id: str + entity: str + column: str + selection: str + reason: str - name: str - recipient_channel: str - donor_channel: str + def __post_init__(self) -> None: + for label, value in ( + ("rule_id", self.rule_id), + ("entity", self.entity), + ("column", self.column), + ("selection", self.selection), + ("reason", self.reason), + ): + if not isinstance(value, str) or not value.strip(): + raise ValueError( + f"GapFillAbsenceRule.{label} must be a non-empty string." + ) + + +@dataclass(frozen=True) +class GapFillDirection: + """One declared cross-origin fill: recipient origin <- donor origin. + + Activation authority is declared here, not inferred from nullness: the + named recipient channel's rows are the only rows the direction may fill, + and the named donor channel's native rows are the only donor evidence. + The transfer machinery itself stays spine-blind; this owner-level + declaration is what makes the run-7 silent-skip class impossible — a + direction either fills its declared families on its declared rows or + fails by name. + """ + + name: str + recipient_channel: str + donor_channel: str target_families: TargetFamilies + recipient_absence_rules: tuple[GapFillAbsenceRule, ...] = () def __post_init__(self) -> None: for label, value in ( @@ -951,6 +1780,43 @@ def __post_init__(self) -> None: "target_families", _freeze_target_families(self.target_families), ) + rules = tuple(self.recipient_absence_rules) + if any(not isinstance(rule, GapFillAbsenceRule) for rule in rules): + raise TypeError( + "GapFillDirection recipient_absence_rules require " + "GapFillAbsenceRule values." + ) + target_keys = { + (entity, target) + for entity, families in self.target_families.items() + for targets in families.values() + for target in targets + } + rule_keys = [(rule.entity, rule.column) for rule in rules] + outside = sorted(set(rule_keys) - target_keys) + duplicates = sorted( + key for key, count in Counter(rule_keys).items() if count > 1 + ) + if outside or duplicates: + raise ValueError( + f"GapFillDirection {self.name!r} has invalid recipient absence " + f"rules; outside_targets={outside}, duplicate_targets={duplicates}." + ) + object.__setattr__(self, "recipient_absence_rules", rules) + + +@dataclass(frozen=True) +class _GapFillProducerRecord: + """One channel-aware proof that a declared target exists before its check.""" + + entity: str + family: str + target: str + operator: str + operator_order_index: int + execution_scope: str + produced_channel: str + producer_stage: str @dataclass(frozen=True) @@ -990,10 +1856,19 @@ def _build_stacked_gap_fill_plan( if housing_families: directions.append( GapFillDirection( - name=_GAP_FILL_ACS_TO_ASEC, - recipient_channel=BASE_ASEC_SUPPORT_CHANNEL, - donor_channel=ACS_STACKED_SUPPORT_CHANNEL, + name=_GAP_FILL_ASEC_HOUSING_TO_ACS, + recipient_channel=ACS_STACKED_SUPPORT_CHANNEL, + donor_channel=BASE_ASEC_SUPPORT_CHANNEL, target_families=housing_families, + recipient_absence_rules=( + GapFillAbsenceRule( + rule_id=_ACS_GQ_RENT_ABSENCE_RULE_ID, + entity="person", + column="pre_subsidy_rent", + selection=_ACS_GQ_RENT_ABSENCE_SELECTION, + reason=_ACS_GQ_RENT_ABSENCE_REASON, + ), + ), ) ) return tuple(directions) @@ -1021,6 +1896,9 @@ class _StackedAuthority: authority_id: str version: int gap_fill_plan: tuple[GapFillDirection, ...] + post_puf_transfer_surface: TargetFamilies + post_puf_puf_producer_surface: TargetFamilies + post_puf_source_producer_surface: TargetFamilies declared_surface: TargetFamilies metric_registry: Mapping[tuple[str, str, str, int], str] joint_metric_registry: Mapping[tuple[str, str, tuple[str, ...], int], str] @@ -1038,6 +1916,21 @@ def __post_init__(self) -> None: if any(not isinstance(direction, GapFillDirection) for direction in plan): raise TypeError("Stacked authority plans require GapFillDirection values.") object.__setattr__(self, "gap_fill_plan", plan) + object.__setattr__( + self, + "post_puf_transfer_surface", + _freeze_target_families(self.post_puf_transfer_surface), + ) + object.__setattr__( + self, + "post_puf_puf_producer_surface", + _freeze_target_families(self.post_puf_puf_producer_surface), + ) + object.__setattr__( + self, + "post_puf_source_producer_surface", + _freeze_target_families(self.post_puf_source_producer_surface), + ) object.__setattr__( self, "declared_surface", @@ -1060,6 +1953,7 @@ def __post_init__(self) -> None: component_digests = dict(self.declared_component_sha256) if set(component_digests) != { "gap_fill_plan", + "post_puf_transfer_surface", "declared_surface", "metric_registry", "joint_metric_registry", @@ -1187,11 +2081,128 @@ def _plan_payload(plan: Sequence[GapFillDirection]) -> list[dict[str, object]]: "recipient_channel": direction.recipient_channel, "donor_channel": direction.donor_channel, "target_families": _surface_payload(direction.target_families), + "recipient_absence_rules": [ + { + "rule_id": rule.rule_id, + "entity": rule.entity, + "column": rule.column, + "selection": rule.selection, + "reason": rule.reason, + } + for rule in direction.recipient_absence_rules + ], } for direction in plan ] +def _build_gap_fill_producer_schedule( + surface: TargetFamilies, +) -> tuple[_GapFillProducerRecord, ...]: + """Resolve every early target to its actual pre-clone producer contract.""" + + records: list[_GapFillProducerRecord] = [] + for entity, families in surface.items(): + for family, targets in families.items(): + for target in targets: + for operator_order_index, operator in enumerate( + POOL_PRE_CLONE_SOURCE_OPERATOR_ORDER + ): + contract = POOL_OPERATOR_CONTRACTS[operator] + outputs = PRE_ASSEMBLY_OPERATOR_OUTPUT_FAMILIES[contract.family] + if target not in outputs.get(entity, ()): + continue + if contract.execution_scope == "cps_source": + produced_channel = BASE_ASEC_SUPPORT_CHANNEL + elif contract.execution_scope == "whole_pool": + produced_channel = "*" + else: + produced_channel = f"" + records.append( + _GapFillProducerRecord( + entity=entity, + family=family, + target=target, + operator=operator, + operator_order_index=operator_order_index, + execution_scope=contract.execution_scope, + produced_channel=produced_channel, + producer_stage=_PRE_CLONE_PREPARATION_STAGE, + ) + ) + return tuple(records) + + +def _gap_fill_activation_stage(direction_name: str) -> str: + return f"gap_fill_stacked_spine.activation[{direction_name}]" + + +def _gap_fill_producer_precedence_failures( + plan: Sequence[GapFillDirection], + schedule: Sequence[_GapFillProducerRecord], +) -> list[str]: + """Fail unless every direction reads a donor its producer already populated.""" + + stage_order = { + _PRE_CLONE_PREPARATION_STAGE: 0, + **{ + _gap_fill_activation_stage(direction.name): index + 1 + for index, direction in enumerate(plan) + }, + _POST_GAP_FILL_STAGE: len(plan) + 1, + } + producer_index: dict[tuple[str, str, str], list[_GapFillProducerRecord]] = {} + for record in schedule: + producer_index.setdefault( + (record.entity, record.family, record.target), [] + ).append(record) + + failures: list[str] = [] + for direction in plan: + check_stage = _gap_fill_activation_stage(direction.name) + check_order = stage_order[check_stage] + for entity, families in direction.target_families.items(): + for family, targets in families.items(): + for target in targets: + label = f"{direction.name}/{entity}/{family}/{target}" + records = producer_index.get((entity, family, target), []) + if len(records) != 1: + failures.append( + f"{label}: expected exactly one declared pre-clone " + f"producer, found {len(records)}." + ) + continue + record = records[0] + if record.execution_scope not in {"cps_source", "whole_pool"}: + failures.append( + f"{label}: producer {record.operator!r} declares " + f"unknown execution scope {record.execution_scope!r}." + ) + continue + producer_order = stage_order.get(record.producer_stage) + if producer_order is None: + failures.append( + f"{label}: producer {record.operator!r} declares " + f"unknown stage {record.producer_stage!r}." + ) + elif producer_order >= check_order: + failures.append( + f"{label}: producer {record.operator!r} runs at " + f"{record.producer_stage!r}, which does not precede " + f"activation stage {check_stage!r}." + ) + if record.produced_channel not in { + direction.donor_channel, + "*", + }: + failures.append( + f"{label}: producer {record.operator!r} populates " + f"channel {record.produced_channel!r}, but activation " + f"declares donor {direction.donor_channel!r}." + ) + return failures + + def _metric_registry_payload( registry: Mapping[tuple[str, str, str, int], str], ) -> list[dict[str, object]]: @@ -1233,6 +2244,9 @@ def _support_profile_payload(profile: _BatterySupportProfile) -> dict[str, objec def _authority_component_payloads( *, gap_fill_plan: Sequence[GapFillDirection], + post_puf_transfer_surface: TargetFamilies, + post_puf_puf_producer_surface: TargetFamilies, + post_puf_source_producer_surface: TargetFamilies, declared_surface: TargetFamilies, metric_registry: Mapping[tuple[str, str, str, int], str], joint_metric_registry: Mapping[tuple[str, str, tuple[str, ...], int], str], @@ -1240,6 +2254,18 @@ def _authority_component_payloads( ) -> dict[str, object]: return { "gap_fill_plan": _plan_payload(gap_fill_plan), + "post_puf_transfer_surface": { + "donor_channel": BASE_ASEC_SUPPORT_CHANNEL, + "donor_clone_index": PUF_TAX_DETAIL_CLONE_INDEX, + "recipient_selection": ( + "target_specific_complement_of_declared_producer_rows" + ), + "producer_surfaces": { + "puf_clone": _surface_payload(post_puf_puf_producer_surface), + "post_clone_source": _surface_payload(post_puf_source_producer_surface), + }, + "target_families": _surface_payload(post_puf_transfer_surface), + }, "declared_surface": _surface_payload(declared_surface), "metric_registry": _metric_registry_payload(metric_registry), "joint_metric_registry": _joint_metric_registry_payload(joint_metric_registry), @@ -1263,6 +2289,9 @@ def _authority_live_digests( ) -> tuple[dict[str, str], str]: payloads = _authority_component_payloads( gap_fill_plan=authority.gap_fill_plan, + post_puf_transfer_surface=authority.post_puf_transfer_surface, + post_puf_puf_producer_surface=(authority.post_puf_puf_producer_surface), + post_puf_source_producer_surface=(authority.post_puf_source_producer_surface), declared_surface=authority.declared_surface, metric_registry=authority.metric_registry, joint_metric_registry=authority.joint_metric_registry, @@ -1286,6 +2315,9 @@ def _make_stacked_authority( authority_id: str, version: int, gap_fill_plan: Sequence[GapFillDirection], + post_puf_transfer_surface: TargetFamilies, + post_puf_puf_producer_surface: TargetFamilies, + post_puf_source_producer_surface: TargetFamilies, declared_surface: TargetFamilies, metric_registry: Mapping[tuple[str, str, str, int], str], support_profile: _BatterySupportProfile, @@ -1296,6 +2328,13 @@ def _make_stacked_authority( declared_sha256: str | None = None, ) -> _StackedAuthority: frozen_plan = tuple(gap_fill_plan) + frozen_post_puf_surface = _freeze_target_families(post_puf_transfer_surface) + frozen_post_puf_puf_producer_surface = _freeze_target_families( + post_puf_puf_producer_surface + ) + frozen_post_puf_source_producer_surface = _freeze_target_families( + post_puf_source_producer_surface + ) frozen_surface = _freeze_target_families(declared_surface) frozen_registry = _freeze_metric_registry(metric_registry) frozen_joint_registry = _freeze_joint_metric_registry( @@ -1303,6 +2342,9 @@ def _make_stacked_authority( ) component_payloads = _authority_component_payloads( gap_fill_plan=frozen_plan, + post_puf_transfer_surface=frozen_post_puf_surface, + post_puf_puf_producer_surface=frozen_post_puf_puf_producer_surface, + post_puf_source_producer_surface=frozen_post_puf_source_producer_surface, declared_surface=frozen_surface, metric_registry=frozen_registry, joint_metric_registry=frozen_joint_registry, @@ -1322,6 +2364,9 @@ def _make_stacked_authority( authority_id=authority_id, version=version, gap_fill_plan=frozen_plan, + post_puf_transfer_surface=frozen_post_puf_surface, + post_puf_puf_producer_surface=frozen_post_puf_puf_producer_surface, + post_puf_source_producer_surface=frozen_post_puf_source_producer_surface, declared_surface=frozen_surface, metric_registry=frozen_registry, joint_metric_registry=frozen_joint_registry, @@ -1839,12 +2884,33 @@ def _explicit_origin_battery_metric_registry( CANONICAL_STACKED_GAP_FILL_SURFACE = _freeze_target_families( - pool_transfer_target_families() + pool_pre_clone_gap_fill_target_families() +) +CANONICAL_STACKED_POST_PUF_TRANSFER_SURFACE = _freeze_target_families( + pool_post_puf_transfer_target_families() +) +CANONICAL_STACKED_POST_PUF_PUF_PRODUCER_SURFACE = _freeze_target_families( + pool_post_puf_puf_producer_target_families() +) +CANONICAL_STACKED_POST_PUF_SOURCE_PRODUCER_SURFACE = _freeze_target_families( + pool_post_puf_source_producer_target_families() ) CANONICAL_STACKED_DECLARED_SURFACE = _terminal_surface_from_pool_registry() CANONICAL_STACKED_GAP_FILL_PLAN = _build_stacked_gap_fill_plan( CANONICAL_STACKED_GAP_FILL_SURFACE ) +_CANONICAL_STACKED_GAP_FILL_PRODUCER_SCHEDULE = _build_gap_fill_producer_schedule( + CANONICAL_STACKED_GAP_FILL_SURFACE +) +_canonical_producer_precedence_failures = _gap_fill_producer_precedence_failures( + CANONICAL_STACKED_GAP_FILL_PLAN, + _CANONICAL_STACKED_GAP_FILL_PRODUCER_SCHEDULE, +) +if _canonical_producer_precedence_failures: + raise RuntimeError( + "Canonical stacked gap-fill producer precedence is invalid:\n " + + "\n ".join(_canonical_producer_precedence_failures) + ) CANONICAL_ORIGIN_BATTERY_METRIC_REGISTRY = _explicit_origin_battery_metric_registry( CANONICAL_STACKED_DECLARED_SURFACE ) @@ -1869,6 +2935,15 @@ def _explicit_origin_battery_metric_registry( _CANONICAL_STACKED_DECLARED_SURFACE_ANCHOR = CANONICAL_STACKED_DECLARED_SURFACE _CANONICAL_STACKED_GAP_FILL_SURFACE_ANCHOR = CANONICAL_STACKED_GAP_FILL_SURFACE _CANONICAL_STACKED_GAP_FILL_PLAN_ANCHOR = CANONICAL_STACKED_GAP_FILL_PLAN +_CANONICAL_STACKED_POST_PUF_TRANSFER_SURFACE_ANCHOR = ( + CANONICAL_STACKED_POST_PUF_TRANSFER_SURFACE +) +_CANONICAL_STACKED_POST_PUF_PUF_PRODUCER_SURFACE_ANCHOR = ( + CANONICAL_STACKED_POST_PUF_PUF_PRODUCER_SURFACE +) +_CANONICAL_STACKED_POST_PUF_SOURCE_PRODUCER_SURFACE_ANCHOR = ( + CANONICAL_STACKED_POST_PUF_SOURCE_PRODUCER_SURFACE +) _CANONICAL_ORIGIN_BATTERY_METRIC_REGISTRY_ANCHOR = ( CANONICAL_ORIGIN_BATTERY_METRIC_REGISTRY ) @@ -1885,6 +2960,11 @@ def _explicit_origin_battery_metric_registry( _STACKED_DECLARED_SURFACE = CANONICAL_STACKED_DECLARED_SURFACE _STACKED_GAP_FILL_SURFACE = CANONICAL_STACKED_GAP_FILL_SURFACE _STACKED_GAP_FILL_PLAN = CANONICAL_STACKED_GAP_FILL_PLAN +_STACKED_POST_PUF_TRANSFER_SURFACE = CANONICAL_STACKED_POST_PUF_TRANSFER_SURFACE +_STACKED_POST_PUF_PUF_PRODUCER_SURFACE = CANONICAL_STACKED_POST_PUF_PUF_PRODUCER_SURFACE +_STACKED_POST_PUF_SOURCE_PRODUCER_SURFACE = ( + CANONICAL_STACKED_POST_PUF_SOURCE_PRODUCER_SURFACE +) _BATTERY_METRIC_REGISTRY = CANONICAL_ORIGIN_BATTERY_METRIC_REGISTRY _BATTERY_JOINT_METRIC_REGISTRY = CANONICAL_ORIGIN_BATTERY_JOINT_METRIC_REGISTRY _BATTERY_SUPPORT_PROFILE = CANONICAL_ORIGIN_BATTERY_SUPPORT_PROFILE @@ -1893,6 +2973,13 @@ def _explicit_origin_battery_metric_registry( authority_id=_STACKED_AUTHORITY_ID, version=_STACKED_AUTHORITY_VERSION, gap_fill_plan=_CANONICAL_STACKED_GAP_FILL_PLAN_ANCHOR, + post_puf_transfer_surface=(_CANONICAL_STACKED_POST_PUF_TRANSFER_SURFACE_ANCHOR), + post_puf_puf_producer_surface=( + _CANONICAL_STACKED_POST_PUF_PUF_PRODUCER_SURFACE_ANCHOR + ), + post_puf_source_producer_surface=( + _CANONICAL_STACKED_POST_PUF_SOURCE_PRODUCER_SURFACE_ANCHOR + ), declared_surface=_CANONICAL_STACKED_DECLARED_SURFACE_ANCHOR, metric_registry=_CANONICAL_ORIGIN_BATTERY_METRIC_REGISTRY_ANCHOR, joint_metric_registry=_CANONICAL_ORIGIN_BATTERY_JOINT_METRIC_REGISTRY_ANCHOR, @@ -1907,6 +2994,15 @@ def _production_stacked_authority( _canonical_authority: _StackedAuthority = _CANONICAL_STACKED_AUTHORITY, _canonical_plan: tuple[GapFillDirection, ...] = CANONICAL_STACKED_GAP_FILL_PLAN, _canonical_gap_surface: TargetFamilies = CANONICAL_STACKED_GAP_FILL_SURFACE, + _canonical_post_puf_surface: TargetFamilies = ( + CANONICAL_STACKED_POST_PUF_TRANSFER_SURFACE + ), + _canonical_post_puf_puf_producer_surface: TargetFamilies = ( + CANONICAL_STACKED_POST_PUF_PUF_PRODUCER_SURFACE + ), + _canonical_post_puf_source_producer_surface: TargetFamilies = ( + CANONICAL_STACKED_POST_PUF_SOURCE_PRODUCER_SURFACE + ), _canonical_surface: TargetFamilies = CANONICAL_STACKED_DECLARED_SURFACE, _canonical_registry: Mapping[ tuple[str, str, str, int], str @@ -1921,6 +3017,11 @@ def _production_stacked_authority( identity = ( _STACKED_GAP_FILL_PLAN is _canonical_plan and _STACKED_GAP_FILL_SURFACE is _canonical_gap_surface + and _STACKED_POST_PUF_TRANSFER_SURFACE is _canonical_post_puf_surface + and _STACKED_POST_PUF_PUF_PRODUCER_SURFACE + is _canonical_post_puf_puf_producer_surface + and _STACKED_POST_PUF_SOURCE_PRODUCER_SURFACE + is _canonical_post_puf_source_producer_surface and _STACKED_DECLARED_SURFACE is _canonical_surface and _BATTERY_METRIC_REGISTRY is _canonical_registry and _BATTERY_JOINT_METRIC_REGISTRY is _canonical_joint_registry @@ -1932,6 +3033,9 @@ def _production_stacked_authority( authority_id=_STACKED_AUTHORITY_ID, version=_STACKED_AUTHORITY_VERSION, gap_fill_plan=_STACKED_GAP_FILL_PLAN, + post_puf_transfer_surface=_STACKED_POST_PUF_TRANSFER_SURFACE, + post_puf_puf_producer_surface=_STACKED_POST_PUF_PUF_PRODUCER_SURFACE, + post_puf_source_producer_surface=(_STACKED_POST_PUF_SOURCE_PRODUCER_SURFACE), declared_surface=_STACKED_DECLARED_SURFACE, metric_registry=_BATTERY_METRIC_REGISTRY, joint_metric_registry=_BATTERY_JOINT_METRIC_REGISTRY, @@ -1957,10 +3061,31 @@ def _metric_registry_for_surface( ) +def _restrict_surface_to_declared_targets( + surface: TargetFamilies, + declared: TargetFamilies, +) -> TargetFamilies: + declared_keys = set(_surface_target_keys(declared)) + restricted: dict[str, dict[str, tuple[str, ...]]] = {} + for entity, families in surface.items(): + for family, targets in families.items(): + retained = tuple( + target + for target in targets + if (entity, family, target, 0) in declared_keys + ) + if retained: + restricted.setdefault(entity, {})[family] = retained + return restricted + + def _make_test_stacked_authority( *, declared_surface: TargetFamilies | None = None, gap_fill_plan: Sequence[GapFillDirection] | None = None, + post_puf_transfer_surface: TargetFamilies | None = None, + post_puf_puf_producer_surface: TargetFamilies | None = None, + post_puf_source_producer_surface: TargetFamilies | None = None, metric_registry: Mapping[tuple[str, str, str, int], str] | None = None, joint_metric_registry: Mapping[tuple[str, str, tuple[str, ...], int], str] | None = None, @@ -1974,6 +3099,25 @@ def _make_test_stacked_authority( else declared_surface ) plan = CANONICAL_STACKED_GAP_FILL_PLAN if gap_fill_plan is None else gap_fill_plan + post_puf_surface = ( + {} if post_puf_transfer_surface is None else post_puf_transfer_surface + ) + puf_producer_surface = ( + _restrict_surface_to_declared_targets( + CANONICAL_STACKED_POST_PUF_PUF_PRODUCER_SURFACE, + post_puf_surface, + ) + if post_puf_puf_producer_surface is None + else post_puf_puf_producer_surface + ) + source_producer_surface = ( + _restrict_surface_to_declared_targets( + CANONICAL_STACKED_POST_PUF_SOURCE_PRODUCER_SURFACE, + post_puf_surface, + ) + if post_puf_source_producer_surface is None + else post_puf_source_producer_surface + ) registry = ( _metric_registry_for_surface(surface) if metric_registry is None @@ -1995,6 +3139,9 @@ def _make_test_stacked_authority( authority_id=f"{_STACKED_AUTHORITY_ID}.test", version=_STACKED_AUTHORITY_VERSION, gap_fill_plan=plan, + post_puf_transfer_surface=post_puf_surface, + post_puf_puf_producer_surface=puf_producer_surface, + post_puf_source_producer_surface=source_producer_surface, declared_surface=surface, metric_registry=registry, joint_metric_registry=joints, @@ -2013,6 +3160,59 @@ def stacked_gap_fill_plan() -> tuple[GapFillDirection, ...]: return _STACKED_GAP_FILL_PLAN +def stacked_gap_fill_producer_schedule_receipt() -> Mapping[str, object]: + """Return the live channel-aware proof that every producer precedes its check.""" + + plan = stacked_gap_fill_plan() + schedule = _build_gap_fill_producer_schedule( + pool_pre_clone_gap_fill_target_families() + ) + failures = _gap_fill_producer_precedence_failures(plan, schedule) + if failures: + raise ValueError( + "Stacked gap-fill producer precedence failed:\n " + "\n ".join(failures) + ) + schedule_by_target = { + (record.entity, record.family, record.target): record for record in schedule + } + directions: list[dict[str, object]] = [] + for index, direction in enumerate(plan): + targets: list[dict[str, object]] = [] + for entity, families in direction.target_families.items(): + for family, columns in families.items(): + for column in columns: + record = schedule_by_target[(entity, family, column)] + targets.append( + { + "entity": entity, + "family": family, + "column": column, + "producer": record.operator, + "producer_order_index": record.operator_order_index, + "execution_scope": record.execution_scope, + "produced_channel": record.produced_channel, + "producer_stage": record.producer_stage, + } + ) + directions.append( + { + "name": direction.name, + "order_index": index, + "donor_channel": direction.donor_channel, + "activation_stage": _gap_fill_activation_stage(direction.name), + "target_count": len(targets), + "targets": targets, + } + ) + payload: dict[str, object] = { + "status": "all_producers_precede_activation", + "direction_count": len(directions), + "target_count": len(schedule), + "directions": directions, + } + return {**payload, "sha256": _canonical_sha256(payload)} + + def stacked_spine_authority_receipt() -> Mapping[str, object]: """Return the live-digested canonical authority for build identity binding.""" @@ -2082,6 +3282,27 @@ def _authority_receipt( "direction_count": len(authority.gap_fill_plan), "digest_matches_declared": component_integrity["gap_fill_plan"], }, + "post_puf_transfer_surface": { + "sha256": live_components["post_puf_transfer_surface"], + "declared_sha256": authority.declared_component_sha256[ + "post_puf_transfer_surface" + ], + "target_count": len( + _surface_target_keys(authority.post_puf_transfer_surface) + ), + "puf_producer_target_count": len( + _surface_target_keys(authority.post_puf_puf_producer_surface) + ), + "source_producer_target_count": len( + _surface_target_keys(authority.post_puf_source_producer_surface) + ), + "donor_channel": BASE_ASEC_SUPPORT_CHANNEL, + "donor_clone_index": PUF_TAX_DETAIL_CLONE_INDEX, + "recipient_selection": ( + "target_specific_complement_of_declared_producer_rows" + ), + "digest_matches_declared": component_integrity["post_puf_transfer_surface"], + }, "declared_surface": { "sha256": live_components["declared_surface"], "declared_sha256": authority.declared_component_sha256["declared_surface"], @@ -2148,12 +3369,42 @@ def _authority_validation_failures( failures: list[str] = [] surface_targets = _surface_target_keys(authority.declared_surface) plan_targets = _plan_target_keys(authority.gap_fill_plan) + post_puf_targets = _surface_target_keys(authority.post_puf_transfer_surface) + post_puf_puf_producer_targets = _surface_target_keys( + authority.post_puf_puf_producer_surface + ) + post_puf_source_producer_targets = _surface_target_keys( + authority.post_puf_source_producer_surface + ) duplicate_surface_targets = sorted( target for target, count in Counter(surface_targets).items() if count > 1 ) duplicate_plan_targets = sorted( target for target, count in Counter(plan_targets).items() if count > 1 ) + duplicate_post_puf_targets = sorted( + target for target, count in Counter(post_puf_targets).items() if count > 1 + ) + duplicate_post_puf_producer_targets = sorted( + { + target + for targets in ( + post_puf_puf_producer_targets, + post_puf_source_producer_targets, + ) + for target, count in Counter(targets).items() + if count > 1 + } + ) + if production: + failures.extend( + _gap_fill_producer_precedence_failures( + authority.gap_fill_plan, + _build_gap_fill_producer_schedule( + pool_pre_clone_gap_fill_target_families() + ), + ) + ) if duplicate_surface_targets: failures.append( "declared surface repeats target(s): " @@ -2170,8 +3421,58 @@ def _authority_validation_failures( ) + "." ) + if duplicate_post_puf_targets: + failures.append( + "post-PUF transfer surface repeats target(s): " + + ", ".join( + _battery_target_label(target) for target in duplicate_post_puf_targets + ) + + "." + ) + if duplicate_post_puf_producer_targets: + failures.append( + "post-PUF producer surfaces repeat target(s) within a role: " + + ", ".join( + _battery_target_label(target) + for target in duplicate_post_puf_producer_targets + ) + + "." + ) + overlap = sorted(set(plan_targets) & set(post_puf_targets)) + if overlap: + failures.append( + "early gap-fill and post-PUF transfer surfaces overlap: " + + ", ".join(_battery_target_label(target) for target in overlap) + + "." + ) + outside_declared = sorted(set(post_puf_targets) - set(surface_targets)) + if outside_declared: + failures.append( + "post-PUF transfer targets are absent from the declared terminal " + "surface: " + + ", ".join(_battery_target_label(target) for target in outside_declared) + + "." + ) + producer_targets = set(post_puf_puf_producer_targets) | set( + post_puf_source_producer_targets + ) + unowned_post_puf = sorted(set(post_puf_targets) - producer_targets) + outside_post_puf = sorted(producer_targets - set(post_puf_targets)) + if unowned_post_puf: + failures.append( + "post-PUF transfer targets have no declared producer role: " + + ", ".join(_battery_target_label(target) for target in unowned_post_puf) + + "." + ) + if outside_post_puf: + failures.append( + "post-PUF producer roles name targets outside the transfer surface: " + + ", ".join(_battery_target_label(target) for target in outside_post_puf) + + "." + ) for name, label in ( ("gap_fill_plan", "gap-fill plan"), + ("post_puf_transfer_surface", "post-PUF transfer surface"), ("declared_surface", "declared surface"), ("metric_registry", "metric registry"), ("joint_metric_registry", "joint metric registry"), @@ -2269,6 +3570,24 @@ def _validate_production_authority_receipt( ) +def validate_stacked_post_puf_transfer_receipt( + receipt: Mapping[str, object], + *, + boundary: str, +) -> None: + """Reject a late-transfer receipt unless it carries canonical authority.""" + + if not isinstance(receipt, Mapping): + raise ValueError(f"{boundary}: stacked post-PUF transfer receipt is absent.") + authority = receipt.get("authority") + if not isinstance(authority, Mapping): + raise ValueError( + f"{boundary}: stacked post-PUF transfer receipt has no authority; " + "production manifest emission is forbidden." + ) + _validate_production_authority_receipt(authority, boundary=boundary) + + def _validate_test_authority(authority: _StackedAuthority, *, boundary: str) -> None: """Keep the explicit fixture seam visibly and terminally non-production.""" @@ -2285,6 +3604,7 @@ def _validate_stacked_gate_manifest_details( gate_name: str, details: Mapping[str, object], *, + passed: bool, _canonical_surface: TargetFamilies = CANONICAL_STACKED_DECLARED_SURFACE, _canonical_plan: tuple[GapFillDirection, ...] = CANONICAL_STACKED_GAP_FILL_PLAN, _canonical_registry: Mapping[ @@ -2306,6 +3626,9 @@ def _validate_stacked_gate_manifest_details( _validate_production_authority_receipt(authority, boundary=boundary) authority_sha256 = authority["sha256"] plan_sha256 = authority["components"]["gap_fill_plan"]["sha256"] + post_puf_surface_sha256 = authority["components"]["post_puf_transfer_surface"][ + "sha256" + ] surface_sha256 = authority["components"]["declared_surface"]["sha256"] expected_keys = _surface_target_keys(_canonical_surface) @@ -2314,6 +3637,94 @@ def reject(reason: str) -> None: f"{boundary}: {reason}; production manifest emission is forbidden." ) + def nonnegative_int( + receipt: Mapping[str, object], + field_name: str, + *, + label: str, + ) -> int: + value = receipt.get(field_name) + if isinstance(value, bool) or not isinstance(value, int) or value < 0: + reject(f"{label} {field_name} must be a non-negative integer") + return value + + def validate_structural_absence_receipt( + raw_receipt: object, + *, + label: str, + battery: bool, + ) -> tuple[int, dict[str, int]]: + if not isinstance(raw_receipt, Mapping): + reject(f"{label} must carry canonical recipient-absence authority") + expected_fields = { + "rule_id", + "selection", + "reason", + "status", + "rows", + "by_origin_role", + "unexpected_null_rows", + "structural_rows_filled", + } + if battery: + expected_fields.update( + {"comparison_clone_index", "rows_excluded_from_scope"} + ) + if set(raw_receipt) != expected_fields: + reject(f"{label} structural-absence receipt schema mismatch") + if ( + raw_receipt.get("rule_id") != _ACS_GQ_RENT_ABSENCE_RULE_ID + or raw_receipt.get("selection") != _ACS_GQ_RENT_ABSENCE_SELECTION + or raw_receipt.get("reason") != _ACS_GQ_RENT_ABSENCE_REASON + or raw_receipt.get("status") != "exact_structural_absence" + ): + reject(f"{label} structural-absence doctrine mismatch") + rows = nonnegative_int(raw_receipt, "rows", label=label) + unexpected = nonnegative_int( + raw_receipt, + "unexpected_null_rows", + label=label, + ) + filled = nonnegative_int( + raw_receipt, + "structural_rows_filled", + label=label, + ) + raw_by_role = raw_receipt.get("by_origin_role") + if not isinstance(raw_by_role, Mapping): + reject(f"{label} structural by-origin-role counts are not a mapping") + by_role: dict[str, int] = {} + for cell, count in raw_by_role.items(): + cell_channel, separator, clone_role = str(cell).partition("/clone_") + if ( + not separator + or cell_channel != ACS_STACKED_SUPPORT_CHANNEL + or not clone_role.isdigit() + or isinstance(count, bool) + or not isinstance(count, int) + or count <= 0 + ): + reject(f"{label} structural by-origin-role receipt is malformed") + by_role[str(cell)] = count + if rows != sum(by_role.values()): + reject(f"{label} structural row count does not equal its role counts") + if passed and (unexpected != 0 or filled != 0): + reject(f"{label} passing structural-absence equation is not exact") + if battery: + if raw_receipt.get("comparison_clone_index") != 0: + reject(f"{label} structural battery clone scope must be clone 0") + excluded = nonnegative_int( + raw_receipt, + "rows_excluded_from_scope", + label=label, + ) + if excluded != by_role.get(f"{ACS_STACKED_SUPPORT_CHANNEL}/clone_0", 0): + reject( + f"{label} structural battery exclusion count differs from " + "its clone-0 authority" + ) + return rows, by_role + if gate_name == _COMPLETENESS_GATE_NAME: expected_labels = { f"{entity}/{family}/{column}" @@ -2349,6 +3760,8 @@ def reject(reason: str) -> None: if ( target_receipt.get("authority_sha256") != authority_sha256 or target_receipt.get("plan_sha256") != plan_sha256 + or target_receipt.get("post_puf_surface_sha256") + != post_puf_surface_sha256 or target_receipt.get("surface_sha256") != surface_sha256 ): reject(f"{label} target receipt is not bound to canonical authority") @@ -2358,6 +3771,10 @@ def reject(reason: str) -> None: proven = target_receipt.get("proven", {}) if not isinstance(proven, Mapping): reject(f"{label} proven-absence receipts are not a mapping") + unproven = target_receipt.get("unproven", {}) + if not isinstance(unproven, Mapping): + reject(f"{label} unproven-absence counts are not a mapping") + proven_rows = 0 for cell, proof_receipt in proven.items(): if not isinstance(proof_receipt, Mapping): reject(f"{label} {cell} proof receipt is not a mapping") @@ -2370,10 +3787,13 @@ def reject(reason: str) -> None: cell_channel, separator, _clone_role = str(cell).partition("/clone_") if ( not separator + or not _clone_role.isdigit() or cell_channel != direction.recipient_channel or proof_receipt.get("authority_form") != "origin_exact_recipient" or proof_receipt.get("authority_sha256") != authority_sha256 or proof_receipt.get("plan_sha256") != plan_sha256 + or proof_receipt.get("post_puf_surface_sha256") + != post_puf_surface_sha256 or proof_receipt.get("surface_sha256") != surface_sha256 or proof_receipt.get("declared_direction") != direction.name or proof_receipt.get("declared_donor_channel") @@ -2384,6 +3804,114 @@ def reject(reason: str) -> None: reject( f"{label} {cell} proof is not recipient-exact canonical authority" ) + proven_rows += nonnegative_int( + proof_receipt, + "null_rows", + label=f"{label} {cell} proof", + ) + unproven_rows = 0 + for cell, count in unproven.items(): + cell_channel, separator, clone_role = str(cell).partition("/clone_") + if ( + not separator + or not clone_role.isdigit() + or not isinstance(cell_channel, str) + or not cell_channel + or isinstance(count, bool) + or not isinstance(count, int) + or count <= 0 + ): + reject(f"{label} unproven-absence receipt is malformed") + unproven_rows += count + status = target_receipt.get("status") + null_rows = target_receipt.get("null_rows") + if passed: + if status not in {"complete", "proven_absent"}: + reject(f"{label} passing completeness status is {status!r}") + if ( + isinstance(null_rows, bool) + or not isinstance(null_rows, int) + or null_rows < 0 + ): + reject(f"{label} passing null_rows is not a non-negative integer") + if unproven_rows: + reject(f"{label} passing receipt carries unproven nulls") + invalid_rows = nonnegative_int( + target_receipt, + "invalid_rows", + label=label, + ) + if invalid_rows: + reject(f"{label} passing receipt carries invalid values") + if status == "complete" and ( + null_rows != 0 + or proven_rows != 0 + or authority_form != "observed_complete" + ): + reject(f"{label} complete receipt has contradictory null authority") + if status == "proven_absent" and ( + null_rows <= 0 + or proven_rows != null_rows + or authority_form != "origin_exact_recipient" + ): + reject(f"{label} proven-absence count arithmetic is inconsistent") + + rent_label = "person/housing/pre_subsidy_rent" + rent_target = targets[rent_label] + validate_rent_structure = passed or rent_target.get("status") not in { + "missing", + "missing_entity", + } + if validate_rent_structure: + rent_rows, rent_by_role = validate_structural_absence_receipt( + rent_target.get("recipient_absence_authority"), + label=rent_label, + battery=False, + ) + else: + rent_rows, rent_by_role = 0, {} + if passed: + rent_null_rows = nonnegative_int( + rent_target, + "null_rows", + label=rent_label, + ) + if rent_null_rows != rent_rows: + reject(f"{rent_label} null count differs from structural authority") + rent_proven = rent_target.get("proven", {}) + if not isinstance(rent_proven, Mapping): + reject(f"{rent_label} proven-absence receipts are not a mapping") + if set(rent_proven) != set(rent_by_role): + reject(f"{rent_label} proven roles differ from structural authority") + direction = direction_by_label[rent_label] + expected_proof_fields = { + "null_rows", + "reason", + "authority_form", + "authority_sha256", + "plan_sha256", + "post_puf_surface_sha256", + "surface_sha256", + "declared_direction", + "declared_donor_channel", + "declared_recipient_channel", + "structural_absence_rule_id", + "structural_absence_selection", + } + for cell, count in rent_by_role.items(): + proof = rent_proven[cell] + if ( + not isinstance(proof, Mapping) + or set(proof) != expected_proof_fields + or proof.get("null_rows") != count + or proof.get("reason") != _ACS_GQ_RENT_ABSENCE_REASON + or proof.get("structural_absence_rule_id") + != _ACS_GQ_RENT_ABSENCE_RULE_ID + or proof.get("structural_absence_selection") + != _ACS_GQ_RENT_ABSENCE_SELECTION + or proof.get("declared_direction") != direction.name + ): + reject(f"{rent_label} {cell} structural proof is not canonical") return if gate_name == _BATTERY_GATE_NAME: @@ -2434,6 +3962,46 @@ def reject(reason: str) -> None: or comparison.get("metric") != metric ): reject(f"{label} comparison must use canonical metric {metric!r}") + if passed: + allowed_statuses = {"tested", "insufficient_support"} + statuses = { + label: comparison.get("status") + for label, comparison in comparisons.items() + if isinstance(comparison, Mapping) + } + invalid_statuses = { + label: status + for label, status in statuses.items() + if status not in allowed_statuses + } + if invalid_statuses: + reject( + "passing battery carries failing comparison statuses " + f"{invalid_statuses}" + ) + tested_labels = { + label for label, status in statuses.items() if status == "tested" + } + untestable_labels = sorted( + label + for label, status in statuses.items() + if status == "insufficient_support" + ) + if details.get("tested_comparisons") != len(tested_labels): + reject("passing battery tested-comparison count is inconsistent") + if details.get("untestable_comparisons") != untestable_labels: + reject("passing battery untestable-comparison list is inconsistent") + + rent_label = "person/housing/pre_subsidy_rent[clone_0]" + rent_comparison = comparisons[rent_label] + if not isinstance(rent_comparison, Mapping): + reject(f"{rent_label} comparison is not a mapping") + if passed or rent_comparison.get("status") != "missing_column": + validate_structural_absence_receipt( + rent_comparison.get("recipient_absence_authority"), + label=rent_label, + battery=True, + ) return reject("unknown stacked authority gate") @@ -2442,15 +4010,39 @@ def reject(reason: str) -> None: _canonical_surface_keys = _surface_target_keys( _CANONICAL_STACKED_DECLARED_SURFACE_ANCHOR ) +_canonical_early_transfer_keys = set( + _surface_target_keys(_CANONICAL_STACKED_GAP_FILL_SURFACE_ANCHOR) +) +_canonical_late_transfer_keys = set( + _surface_target_keys(_CANONICAL_STACKED_POST_PUF_TRANSFER_SURFACE_ANCHOR) +) +_canonical_late_puf_producer_keys = set( + _surface_target_keys(_CANONICAL_STACKED_POST_PUF_PUF_PRODUCER_SURFACE_ANCHOR) +) +_canonical_late_source_producer_keys = set( + _surface_target_keys(_CANONICAL_STACKED_POST_PUF_SOURCE_PRODUCER_SURFACE_ANCHOR) +) +_canonical_full_transfer_keys = set( + _surface_target_keys(_freeze_target_families(pool_transfer_target_families())) +) if ( len(_canonical_surface_keys) != 131 or len(set(_canonical_surface_keys)) != 131 - or len(_surface_target_keys(_CANONICAL_STACKED_GAP_FILL_SURFACE_ANCHOR)) != 118 + or len(_canonical_early_transfer_keys) != 48 + or len(_canonical_late_transfer_keys) != 70 + or len(_canonical_late_puf_producer_keys) != 43 + or len(_canonical_late_source_producer_keys) != 30 + or len(_canonical_late_puf_producer_keys & _canonical_late_source_producer_keys) + != 3 + or _canonical_late_puf_producer_keys | _canonical_late_source_producer_keys + != _canonical_late_transfer_keys + or _canonical_early_transfer_keys & _canonical_late_transfer_keys + or _canonical_early_transfer_keys | _canonical_late_transfer_keys + != _canonical_full_transfer_keys + or len(_canonical_full_transfer_keys) != 118 or set(_plan_target_keys(_CANONICAL_STACKED_GAP_FILL_PLAN_ANCHOR)) - != set(_surface_target_keys(_CANONICAL_STACKED_GAP_FILL_SURFACE_ANCHOR)) - or not set(_plan_target_keys(_CANONICAL_STACKED_GAP_FILL_PLAN_ANCHOR)).issubset( - _canonical_surface_keys - ) + != _canonical_early_transfer_keys + or not _canonical_full_transfer_keys.issubset(_canonical_surface_keys) or set(_CANONICAL_ORIGIN_BATTERY_METRIC_REGISTRY_ANCHOR) != set(_canonical_surface_keys) or len(_CANONICAL_ORIGIN_BATTERY_JOINT_METRIC_REGISTRY_ANCHOR) != 1 @@ -2467,8 +4059,11 @@ def reject(reason: str) -> None: ) ): raise RuntimeError( - "Canonical stacked authority must bind an exact 118-target gap-fill " - "plan inside an exact 131-target terminal surface and metric registry." + "Canonical stacked authority must partition the exact 118-target " + "transfer surface into 48 early gap-fill and 70 post-PUF targets " + "inside an exact 131-target terminal surface and metric registry; " + "the late surface must be exactly covered by 43 PUF-clone and 30 " + "ASEC-source producer targets with their declared three-target overlap." ) @@ -2648,26 +4243,121 @@ def _direction_entity_targets( return result -def _origin_projection(frame: Frame, *, channel: str) -> Frame: - """Project one origin's native lineages as a standalone donor frame.""" - - person = frame.table("person") - mask = ( - person[support_channel_column("person")].astype(str).eq(channel) - & person[support_clone_index_column("person")].eq(0) - ).to_numpy() - if not mask.any(): - raise ValueError( - f"Stacked spine has no native person rows for origin {channel!r}." - ) - return frame.select(mask) +def _direction_absence_rule_index( + direction: GapFillDirection, +) -> dict[tuple[str, str], GapFillAbsenceRule]: + return { + (rule.entity, rule.column): rule for rule in direction.recipient_absence_rules + } -def _direction_targets_snapshot( +def _gap_fill_absence_rule_mask( frame: Frame, *, - entity: str, - targets: Sequence[str], + direction: GapFillDirection, + rule: GapFillAbsenceRule, +) -> tuple[pd.Series, dict[str, object]]: + """Resolve one structural-null rule to an exact, live row mask and receipt.""" + + if ( + rule.rule_id != _ACS_GQ_RENT_ABSENCE_RULE_ID + or rule.selection != _ACS_GQ_RENT_ABSENCE_SELECTION + or rule.entity != "person" + or rule.column != "pre_subsidy_rent" + or direction.recipient_channel != ACS_STACKED_SUPPORT_CHANNEL + ): + raise ValueError( + f"Gap-fill direction {direction.name!r} declares unsupported " + f"recipient absence rule {rule.rule_id!r}." + ) + + manifest = frame.metadata.get(STACKED_SPINE_MANIFEST_KEY) + if not isinstance(manifest, Mapping): + raise ValueError( + f"Gap-fill absence rule {rule.rule_id!r} requires the stacked " + "assembly manifest." + ) + gq_households, mask = _validated_acs_native_group_quarters_masks( + frame, + manifest, + boundary=f"gap-fill absence rule {rule.rule_id!r}", + ) + household = frame.table("household") + person = frame.table("person") + required_household = { + "tenure_type", + support_channel_column("household"), + } + required_person = { + "person_household_id", + support_channel_column("person"), + support_clone_index_column("person"), + } + missing_household = sorted(required_household - set(household.columns)) + missing_person = sorted(required_person - set(person.columns)) + if missing_household or missing_person: + raise ValueError( + f"Gap-fill absence rule {rule.rule_id!r} cannot resolve its exact " + "universe; " + f"missing_household={missing_household}, missing_person={missing_person}." + ) + + nonnull_gq_tenure = gq_households & household["tenure_type"].notna() + if nonnull_gq_tenure.any(): + raise ValueError( + f"Gap-fill absence rule {rule.rule_id!r} found " + f"{int(nonnull_gq_tenure.sum())} ACS group-quarters household row(s) " + "with synthesized tenure." + ) + person_channel = person[support_channel_column("person")].astype(str) + + clone_index = pd.to_numeric( + person[support_clone_index_column("person")], errors="raise" + ).astype("int64") + by_origin_role = { + f"{channel}/clone_{int(clone)}": int(count) + for (channel, clone), count in ( + pd.DataFrame( + { + "channel": person_channel.loc[mask], + "clone_index": clone_index.loc[mask], + } + ) + .groupby(["channel", "clone_index"], sort=True) + .size() + .items() + ) + } + return mask, { + "rule_id": rule.rule_id, + "selection": rule.selection, + "reason": rule.reason, + "status": "exact_structural_absence", + "rows": int(mask.sum()), + "by_origin_role": by_origin_role, + } + + +def _origin_projection(frame: Frame, *, channel: str) -> Frame: + """Project one origin's native lineages as a standalone donor frame.""" + + person = frame.table("person") + mask = ( + person[support_channel_column("person")].astype(str).eq(channel) + & person[support_clone_index_column("person")].eq(0) + ).to_numpy() + if not mask.any(): + raise ValueError( + f"Stacked spine has no native person rows for origin {channel!r}." + ) + return frame.select(mask) + + +def _direction_targets_snapshot( + frame: Frame, + *, + entity: str, + targets: Sequence[str], channel: str, ) -> pd.DataFrame: table = frame.table(entity) @@ -2760,219 +4450,672 @@ def _semantic_scalar_payload( ), ) ) - if not supported: - raise TypeError( - f"{boundary}: donor byte identity found unsupported semantic " - f"scalar type {_qualified_type_name(value)!r}." + if not supported: + raise TypeError( + f"{boundary}: donor byte identity found unsupported semantic " + f"scalar type {_qualified_type_name(value)!r}." + ) + return ( + _qualified_type_name(value), + pickle.dumps(value, protocol=5), + ) + + +def _semantic_scalar_sequence_payload( + values: Sequence[object], + *, + boundary: str, +) -> tuple[tuple[str, bytes], ...]: + payloads: list[tuple[str, bytes]] = [] + for position, value in enumerate(values): + payload = _semantic_scalar_payload( + value, + boundary=f"{boundary} position {position}", + ) + payloads.append(payload) + return tuple(payloads) + + +def _index_identity_payload( + index: pd.Index, + *, + boundary: str, +) -> tuple[object, ...]: + """Return exact index authority without list-wide object serialization.""" + + names = tuple( + _semantic_scalar_payload( + name, + boundary=f"{boundary} name {position}", + ) + for position, name in enumerate(index.names) + ) + index_type = _qualified_type_name(index) + if isinstance(index, pd.MultiIndex): + levels = tuple( + _index_identity_payload( + level, + boundary=f"{boundary} level {position}", + ) + for position, level in enumerate(index.levels) + ) + codes = tuple( + ( + code.dtype.str, + code.shape, + np.ascontiguousarray(code).tobytes(order="C"), + ) + for code in index.codes + ) + return (index_type, names, "multiindex_levels_codes", levels, codes) + + dtype = index.dtype + dtype_authority = ( + _qualified_type_name(dtype), + pickle.dumps(dtype, protocol=5), + ) + values = index.to_numpy(copy=False) + if not pd.api.types.is_extension_array_dtype(dtype) and not values.dtype.hasobject: + encoding = "raw_numpy_c_order" + value_payload: object = ( + values.shape, + np.ascontiguousarray(values).tobytes(order="C"), + ) + else: + encoding = "independent_scalar_pickle_protocol_5" + semantic_values = index.to_numpy(dtype=object, copy=True).tolist() + value_payload = _semantic_scalar_sequence_payload( + semantic_values, + boundary=f"{boundary} values", + ) + return (index_type, names, dtype_authority, encoding, value_payload) + + +def _verify_gap_fill_activation_authority( + frame: Frame, + *, + direction: GapFillDirection, +) -> dict[tuple[str, str], dict[str, int]]: + """Verify declared activation authority before any modeling runs.""" + + failures: list[str] = [] + counts: dict[tuple[str, str], dict[str, int]] = {} + for entity, families in direction.target_families.items(): + table = frame.table(entity) + channel = table[support_channel_column(entity)].astype(str) + recipient_rows = channel.eq(direction.recipient_channel) + donor_rows = channel.eq(direction.donor_channel) + recipient_count = int(recipient_rows.sum()) + donor_count = int(donor_rows.sum()) + if recipient_count == 0: + failures.append( + f"{direction.name}/{entity}: declared recipient channel " + f"{direction.recipient_channel!r} has no live rows." + ) + if donor_count == 0: + failures.append( + f"{direction.name}/{entity}: declared donor channel " + f"{direction.donor_channel!r} has no live rows." + ) + for family, targets in families.items(): + for target in targets: + label = f"{direction.name}/{entity}/{family}/{target}" + if target not in table.columns: + failures.append( + f"{label}: declared gap-fill target column is absent " + "from the stacked spine." + ) + continue + null_mask = table[target].isna() + donor_nulls = int((null_mask & donor_rows).sum()) + unauthorized_nulls = int( + (null_mask & ~recipient_rows & ~donor_rows).sum() + ) + if donor_nulls: + failures.append( + f"{label}: donor origin {direction.donor_channel!r} " + f"has {donor_nulls} null cell(s); donors must observe " + "every declared target." + ) + if unauthorized_nulls: + failures.append( + f"{label}: {unauthorized_nulls} null cell(s) lie " + "outside the declared recipient origin " + f"{direction.recipient_channel!r}." + ) + counts[(entity, target)] = { + "authorized_null_rows": int((null_mask & recipient_rows).sum()), + "recipient_rows": int(recipient_rows.sum()), + "donor_rows": int(donor_rows.sum()), + } + if failures: + raise ValueError( + "Stacked gap-fill activation authority failed:\n " + "\n ".join(failures) + ) + return counts + + +def _verify_gap_fill_outcome( + frame: Frame, + *, + direction: GapFillDirection, + pre_counts: Mapping[tuple[str, str], Mapping[str, int]], + donor_snapshot: Mapping[str, pd.DataFrame], + result: AcsTransferResult, +) -> dict[str, object]: + """Verify donor invariance and residual nulls; build the direction receipt.""" + + failures: list[str] = [] + imputed_by_target = { + (record.entity, record.column): record for record in result.imputed_inputs + } + target_receipts: dict[str, dict[str, object]] = {} + absence_rules = _direction_absence_rule_index(direction) + for entity, families in direction.target_families.items(): + table = frame.table(entity) + channel = table[support_channel_column(entity)].astype(str) + recipient_rows = channel.eq(direction.recipient_channel) + donor_after = { + entity_name: _direction_targets_snapshot( + frame, + entity=entity_name, + targets=targets, + channel=direction.donor_channel, + ) + for entity_name, targets in _direction_entity_targets(direction).items() + }[entity] + for family, targets in families.items(): + for target in targets: + label = f"{direction.name}/{entity}/{family}/{target}" + before = donor_snapshot[entity].get(target) + after = donor_after.get(target) + if ( + before is None + or after is None + or _canonical_donor_series_payload( + before, + boundary=f"{label} donor identity before transfer", + ) + != _canonical_donor_series_payload( + after, + boundary=f"{label} donor identity after transfer", + ) + ): + failures.append( + f"{label}: donor byte identity failed for origin " + f"{direction.donor_channel!r}; canonical donor payload " + "changed during gap-fill transfer." + ) + null_mask = table[target].isna() + residual_mask = null_mask & recipient_rows + residual_nulls = int(residual_mask.sum()) + outside_nulls = int((null_mask & ~recipient_rows).sum()) + if outside_nulls: + failures.append( + f"{label}: {outside_nulls} null cell(s) appeared " + "outside the declared recipient origin during the " + "transfer." + ) + record = imputed_by_target.get((entity, target)) + imputed = record.imputed_recipient_rows if record else 0 + unmodeled = record.unmodeled_recipient_rows if record else 0 + pre = pre_counts[(entity, target)] + authorized = pre["authorized_null_rows"] + if residual_nulls != unmodeled: + failures.append( + f"{label}: residual-null equation failed: " + f"residual_null_rows={residual_nulls} != " + f"unmodeled_rows={unmodeled}." + ) + if authorized != imputed + unmodeled: + failures.append( + f"{label}: activation accounting equation failed: " + f"authorized_null_rows={authorized} != " + f"imputed_rows={imputed} + unmodeled_rows={unmodeled}." + ) + target_receipt: dict[str, object] = { + "authorized_null_rows": authorized, + "imputed_rows": imputed, + "unmodeled_rows": unmodeled, + "residual_null_rows": residual_nulls, + } + rule = absence_rules.get((entity, target)) + if rule is not None: + expected_absence, absence_receipt = _gap_fill_absence_rule_mask( + frame, + direction=direction, + rule=rule, + ) + unexpected = int((residual_mask & ~expected_absence).sum()) + synthesized = int((expected_absence & ~residual_mask).sum()) + if unexpected or synthesized: + failures.append( + f"{label}: exact structural-absence equation failed; " + f"unexpected_null_rows={unexpected}, " + f"structural_rows_filled={synthesized}." + ) + target_receipt["recipient_absence_authority"] = { + **absence_receipt, + "unexpected_null_rows": unexpected, + "structural_rows_filled": synthesized, + } + elif unmodeled or residual_nulls: + failures.append( + f"{label}: undeclared gap-fill residual is forbidden; " + f"unmodeled_rows={unmodeled}, " + f"residual_null_rows={residual_nulls}. Every downstream " + "consumer requires this early target complete." + ) + target_receipts[f"{entity}/{family}/{target}"] = target_receipt + if failures: + raise ValueError( + "Stacked gap-fill outcome verification failed:\n " + "\n ".join(failures) + ) + return { + "recipient_channel": direction.recipient_channel, + "donor_channel": direction.donor_channel, + "donor_selection": "owner_projection_of_native_donor_rows", + "resolved_donor_channel": result.resolved_donor_channel, + "targets": target_receipts, + "deferred_inputs": list(result.deferred_inputs), + "fit_records": [ + {"fit_name": record.fit_name, "weight_kind": record.weight_kind} + for record in result.fit_records + ], + } + + +# --------------------------------------------------------------------------- +# Post-PUF transfer of outputs that do not exist at the early gap-fill stage +# --------------------------------------------------------------------------- + + +@dataclass(frozen=True) +class StackedPostPufTransferResult: + """The completed stacked frame plus late-transfer provenance.""" + + frame: Frame + receipt: Mapping[str, object] + transfer_result: AcsTransferResult + + +def transfer_stacked_post_puf_inputs( + frame: Frame, + *, + seed: int = 0, + n_estimators: int = 100, + max_targets_per_fit: int = DEFAULT_ACS_TRANSFER_MAX_TARGETS_PER_FIT, + target_bank: AcsTransferTargetBank | None = None, +) -> StackedPostPufTransferResult: + """Transfer canonical late-produced inputs after source completion.""" + + return _transfer_stacked_post_puf_inputs_evaluate( + frame, + authority=_production_stacked_authority(), + production=True, + seed=seed, + n_estimators=n_estimators, + max_targets_per_fit=max_targets_per_fit, + target_bank=target_bank, + ) + + +def _transfer_stacked_post_puf_inputs_with_test_authority( + frame: Frame, + *, + authority: _StackedAuthority, + seed: int = 0, + n_estimators: int = 100, + max_targets_per_fit: int = DEFAULT_ACS_TRANSFER_MAX_TARGETS_PER_FIT, + target_bank: AcsTransferTargetBank | None = None, +) -> StackedPostPufTransferResult: + """Explicit non-production seam for fixture-sized late surfaces.""" + + _validate_test_authority(authority, boundary="post-PUF transfer test seam") + return _transfer_stacked_post_puf_inputs_evaluate( + frame, + authority=authority, + production=False, + seed=seed, + n_estimators=n_estimators, + max_targets_per_fit=max_targets_per_fit, + target_bank=target_bank, + ) + + +def _transfer_stacked_post_puf_inputs_evaluate( + frame: Frame, + *, + authority: _StackedAuthority, + production: bool, + seed: int, + n_estimators: int, + max_targets_per_fit: int, + target_bank: AcsTransferTargetBank | None, +) -> StackedPostPufTransferResult: + """Run the late transfer from the one role carrying every declared target.""" + + authority_receipt = _authority_receipt(authority) + authority_failures = _authority_validation_failures( + authority, + production=production, + ) + if authority_failures: + raise ValueError( + "Stacked post-PUF transfer authority validation failed:\n " + + "\n ".join(authority_failures) + ) + if production: + _validate_production_authority_receipt( + authority_receipt, + boundary="stacked post-PUF transfer entry", + ) + validate_stacked_spine_frame(frame, boundary="stacked post-PUF transfer entry") + validate_puf_clone_attachment( + frame, + boundary="stacked post-PUF transfer attachment", + ) + surface = authority.post_puf_transfer_surface + if not _surface_target_keys(surface): + raise ValueError("Stacked post-PUF transfer requires at least one target.") + + pre_counts = _verify_post_puf_transfer_activation_authority( + frame, + target_families=surface, + puf_producer_families=authority.post_puf_puf_producer_surface, + source_producer_families=authority.post_puf_source_producer_surface, + ) + donor = _post_puf_donor_projection(frame) + puf_producer_keys = set( + _surface_target_keys(authority.post_puf_puf_producer_surface) + ) + source_producer_keys = set( + _surface_target_keys(authority.post_puf_source_producer_surface) + ) + producer_snapshot = { + (entity, target): _post_puf_producer_snapshot( + frame, + entity=entity, + target=target, + puf_produced=(entity, family, target, 0) in puf_producer_keys, + source_produced=(entity, family, target, 0) in source_producer_keys, ) - return ( - _qualified_type_name(value), - pickle.dumps(value, protocol=5), + for entity, families in surface.items() + for family, targets in families.items() + for target in targets + } + transfer = transfer_acs_inputs( + frame, + donor, + target_families=surface, + donor_channel=None, + seed=seed, + n_estimators=n_estimators, + max_targets_per_fit=max_targets_per_fit, + target_bank=target_bank, + ) + target_receipts = _verify_post_puf_transfer_outcome( + transfer.frame, + target_families=surface, + puf_producer_families=authority.post_puf_puf_producer_surface, + source_producer_families=authority.post_puf_source_producer_surface, + pre_counts=pre_counts, + producer_snapshot=producer_snapshot, + result=transfer, + ) + validate_stacked_spine_frame( + transfer.frame, + boundary="stacked post-PUF transfer output", + ) + return StackedPostPufTransferResult( + frame=transfer.frame, + receipt={ + "authority": authority_receipt, + "donor_selection": "owner_projection_of_asec_origin_clone_1", + "donor_channel": BASE_ASEC_SUPPORT_CHANNEL, + "donor_clone_index": PUF_TAX_DETAIL_CLONE_INDEX, + "recipient_selection": ( + "target_specific_complement_of_declared_producer_rows" + ), + "resolved_donor_channel": transfer.resolved_donor_channel, + "targets": target_receipts, + "fit_records": [ + {"fit_name": record.fit_name, "weight_kind": record.weight_kind} + for record in transfer.fit_records + ], + }, + transfer_result=transfer, ) -def _semantic_scalar_sequence_payload( - values: Sequence[object], - *, - boundary: str, -) -> tuple[tuple[str, bytes], ...]: - payloads: list[tuple[str, bytes]] = [] - for position, value in enumerate(values): - payload = _semantic_scalar_payload( - value, - boundary=f"{boundary} position {position}", +def _post_puf_role_mask(frame: Frame, *, entity: str) -> pd.Series: + table = frame.table(entity) + return table[support_channel_column(entity)].astype(str).eq( + BASE_ASEC_SUPPORT_CHANNEL + ) & pd.to_numeric( + table[support_clone_index_column(entity)], + errors="raise", + ).eq(PUF_TAX_DETAIL_CLONE_INDEX) + + +def _post_puf_donor_projection(frame: Frame) -> Frame: + person_mask = _post_puf_role_mask(frame, entity=frame.schema.person_entity) + if not person_mask.any(): + raise ValueError( + "Stacked post-PUF transfer has no ASEC-origin clone-1 donor rows." ) - payloads.append(payload) - return tuple(payloads) + return frame.select(person_mask.to_numpy(dtype=bool)) -def _index_identity_payload( - index: pd.Index, - *, - boundary: str, -) -> tuple[object, ...]: - """Return exact index authority without list-wide object serialization.""" +def _verify_post_puf_cross_grain_clone_provenance(frame: Frame) -> None: + """Require each person to share the clone role of every parent entity.""" - names = tuple( - _semantic_scalar_payload( - name, - boundary=f"{boundary} name {position}", - ) - for position, name in enumerate(index.names) + person_entity = frame.schema.person_entity + person = frame.table(person_entity) + person_clone = pd.to_numeric( + person[support_clone_index_column(person_entity)], + errors="raise", ) - index_type = _qualified_type_name(index) - if isinstance(index, pd.MultiIndex): - levels = tuple( - _index_identity_payload( - level, - boundary=f"{boundary} level {position}", - ) - for position, level in enumerate(index.levels) + failures: list[str] = [] + for group in frame.schema.group_entities: + membership_column = frame.schema.membership_column(group) + group_id_column = frame.schema.entity_id_column(group) + group_table = frame.table(group) + group_clone = pd.to_numeric( + group_table.set_index(group_id_column)[support_clone_index_column(group)], + errors="raise", ) - codes = tuple( - ( - code.dtype.str, - code.shape, - np.ascontiguousarray(code).tobytes(order="C"), + expected = person[membership_column].map(group_clone) + mismatch = expected.isna() | expected.ne(person_clone) + if mismatch.any(): + failures.append( + f"{int(mismatch.sum())} person/{group} link(s) disagree on " + "support clone index" ) - for code in index.codes + if failures: + raise ValueError( + "Stacked post-PUF transfer cross-grain clone provenance failed:\n " + + "\n ".join(failures) ) - return (index_type, names, "multiindex_levels_codes", levels, codes) - dtype = index.dtype - dtype_authority = ( - _qualified_type_name(dtype), - pickle.dumps(dtype, protocol=5), - ) - values = index.to_numpy(copy=False) - if not pd.api.types.is_extension_array_dtype(dtype) and not values.dtype.hasobject: - encoding = "raw_numpy_c_order" - value_payload: object = ( - values.shape, - np.ascontiguousarray(values).tobytes(order="C"), - ) - else: - encoding = "independent_scalar_pickle_protocol_5" - semantic_values = index.to_numpy(dtype=object, copy=True).tolist() - value_payload = _semantic_scalar_sequence_payload( - semantic_values, - boundary=f"{boundary} values", + +def _post_puf_producer_mask( + frame: Frame, + *, + entity: str, + puf_produced: bool, + source_produced: bool, +) -> pd.Series: + table = frame.table(entity) + producer_rows = pd.Series(False, index=table.index, dtype=bool) + if puf_produced: + producer_rows |= pd.to_numeric( + table[support_clone_index_column(entity)], + errors="raise", + ).gt(0) + if source_produced: + producer_rows |= ( + table[support_channel_column(entity)] + .astype(str) + .eq(BASE_ASEC_SUPPORT_CHANNEL) ) - return (index_type, names, dtype_authority, encoding, value_payload) + return producer_rows -def _verify_gap_fill_activation_authority( +def _post_puf_producer_snapshot( frame: Frame, *, - direction: GapFillDirection, + entity: str, + target: str, + puf_produced: bool, + source_produced: bool, +) -> pd.Series: + table = frame.table(entity) + producer_rows = _post_puf_producer_mask( + frame, + entity=entity, + puf_produced=puf_produced, + source_produced=source_produced, + ) + return table.loc[producer_rows, target].copy(deep=True) + + +def _verify_post_puf_transfer_activation_authority( + frame: Frame, + *, + target_families: TargetFamilies, + puf_producer_families: TargetFamilies, + source_producer_families: TargetFamilies, ) -> dict[tuple[str, str], dict[str, int]]: - """Verify declared activation authority before any modeling runs.""" + """Require complete declared producers before authorizing recipient nulls.""" failures: list[str] = [] counts: dict[tuple[str, str], dict[str, int]] = {} - for entity, families in direction.target_families.items(): + _verify_post_puf_cross_grain_clone_provenance(frame) + puf_producer_keys = set(_surface_target_keys(puf_producer_families)) + source_producer_keys = set(_surface_target_keys(source_producer_families)) + for entity, families in target_families.items(): table = frame.table(entity) - channel = table[support_channel_column(entity)].astype(str) - recipient_rows = channel.eq(direction.recipient_channel) - donor_rows = channel.eq(direction.donor_channel) - recipient_count = int(recipient_rows.sum()) + donor_rows = _post_puf_role_mask(frame, entity=entity) donor_count = int(donor_rows.sum()) - if recipient_count == 0: - failures.append( - f"{direction.name}/{entity}: declared recipient channel " - f"{direction.recipient_channel!r} has no live rows." - ) if donor_count == 0: failures.append( - f"{direction.name}/{entity}: declared donor channel " - f"{direction.donor_channel!r} has no live rows." + f"post_puf_transfer/{entity}: declared ASEC clone-1 donor role " + "has no live rows." ) for family, targets in families.items(): for target in targets: - label = f"{direction.name}/{entity}/{family}/{target}" + label = f"post_puf_transfer/{entity}/{family}/{target}" + key = (entity, family, target, 0) + puf_produced = key in puf_producer_keys + source_produced = key in source_producer_keys + producer_rows = _post_puf_producer_mask( + frame, + entity=entity, + puf_produced=puf_produced, + source_produced=source_produced, + ) + recipient_rows = ~producer_rows if target not in table.columns: failures.append( - f"{label}: declared gap-fill target column is absent " - "from the stacked spine." + f"{label}: declared post-PUF transfer target column is " + "absent from the stacked spine." ) continue null_mask = table[target].isna() - donor_nulls = int((null_mask & donor_rows).sum()) - unauthorized_nulls = int( - (null_mask & ~recipient_rows & ~donor_rows).sum() - ) - if donor_nulls: - failures.append( - f"{label}: donor origin {direction.donor_channel!r} " - f"has {donor_nulls} null cell(s); donors must observe " - "every declared target." + producer_nulls = int((null_mask & producer_rows).sum()) + if producer_nulls: + roles = ", ".join( + role + for role, active in ( + ("PUF clone", puf_produced), + ("ASEC source", source_produced), + ) + if active ) - if unauthorized_nulls: failures.append( - f"{label}: {unauthorized_nulls} null cell(s) lie " - "outside the declared recipient origin " - f"{direction.recipient_channel!r}." + f"{label}: declared {roles} producer role(s) have " + f"{producer_nulls} null cell(s); upstream producers must " + "observe every producer-owned target." ) counts[(entity, target)] = { "authorized_null_rows": int((null_mask & recipient_rows).sum()), "recipient_rows": int(recipient_rows.sum()), - "donor_rows": int(donor_rows.sum()), + "producer_rows": int(producer_rows.sum()), + "donor_rows": donor_count, } if failures: raise ValueError( - "Stacked gap-fill activation authority failed:\n " + "\n ".join(failures) + "Stacked post-PUF transfer activation authority failed:\n " + + "\n ".join(failures) ) return counts -def _verify_gap_fill_outcome( +def _verify_post_puf_transfer_outcome( frame: Frame, *, - direction: GapFillDirection, + target_families: TargetFamilies, + puf_producer_families: TargetFamilies, + source_producer_families: TargetFamilies, pre_counts: Mapping[tuple[str, str], Mapping[str, int]], - donor_snapshot: Mapping[str, pd.DataFrame], + producer_snapshot: Mapping[tuple[str, str], pd.Series], result: AcsTransferResult, -) -> dict[str, object]: - """Verify donor invariance and residual nulls; build the direction receipt.""" +) -> dict[str, dict[str, object]]: + """Prove producers were preserved and every authorized null was filled.""" failures: list[str] = [] imputed_by_target = { (record.entity, record.column): record for record in result.imputed_inputs } target_receipts: dict[str, dict[str, object]] = {} - for entity, families in direction.target_families.items(): + puf_producer_keys = set(_surface_target_keys(puf_producer_families)) + source_producer_keys = set(_surface_target_keys(source_producer_families)) + for entity, families in target_families.items(): table = frame.table(entity) - channel = table[support_channel_column(entity)].astype(str) - recipient_rows = channel.eq(direction.recipient_channel) - donor_after = { - entity_name: _direction_targets_snapshot( - frame, - entity=entity_name, - targets=targets, - channel=direction.donor_channel, - ) - for entity_name, targets in _direction_entity_targets(direction).items() - }[entity] - for family, targets in families.items(): - for target in targets: - label = f"{direction.name}/{entity}/{family}/{target}" - before = donor_snapshot[entity].get(target) - after = donor_after.get(target) - if ( - before is None - or after is None - or _canonical_donor_series_payload( - before, - boundary=f"{label} donor identity before transfer", - ) - != _canonical_donor_series_payload( - after, - boundary=f"{label} donor identity after transfer", - ) + for family, family_targets in families.items(): + for target in family_targets: + label = f"post_puf_transfer/{entity}/{family}/{target}" + key = (entity, family, target, 0) + puf_produced = key in puf_producer_keys + source_produced = key in source_producer_keys + producer_rows = _post_puf_producer_mask( + frame, + entity=entity, + puf_produced=puf_produced, + source_produced=source_produced, + ) + recipient_rows = ~producer_rows + before = producer_snapshot.get((entity, target)) + after = table.loc[producer_rows, target].copy(deep=True) + if before is None or _canonical_donor_series_payload( + before, + boundary=f"{label} producer identity before transfer", + ) != _canonical_donor_series_payload( + after, + boundary=f"{label} producer identity after transfer", ): failures.append( - f"{label}: donor byte identity failed for origin " - f"{direction.donor_channel!r}; canonical donor payload " - "changed during gap-fill transfer." + f"{label}: producer byte identity failed; a declared " + "producer payload changed during transfer." ) null_mask = table[target].isna() - residual_nulls = int((null_mask & recipient_rows).sum()) - outside_nulls = int((null_mask & ~recipient_rows).sum()) - if outside_nulls: - failures.append( - f"{label}: {outside_nulls} null cell(s) appeared " - "outside the declared recipient origin during the " - "transfer." - ) + producer_nulls = int((null_mask & producer_rows).sum()) + recipient_nulls = int((null_mask & recipient_rows).sum()) + residual_nulls = int(null_mask.sum()) record = imputed_by_target.get((entity, target)) imputed = record.imputed_recipient_rows if record else 0 unmodeled = record.unmodeled_recipient_rows if record else 0 - pre = pre_counts[(entity, target)] - authorized = pre["authorized_null_rows"] - if residual_nulls != unmodeled: + authorized = pre_counts[(entity, target)]["authorized_null_rows"] + if producer_nulls: + failures.append( + f"{label}: {producer_nulls} producer null cell(s) appeared " + "during transfer." + ) + if recipient_nulls != unmodeled: failures.append( f"{label}: residual-null equation failed: " - f"residual_null_rows={residual_nulls} != " + f"recipient_null_rows={recipient_nulls} != " f"unmodeled_rows={unmodeled}." ) if authorized != imputed + unmodeled: @@ -2981,7 +5124,22 @@ def _verify_gap_fill_outcome( f"authorized_null_rows={authorized} != " f"imputed_rows={imputed} + unmodeled_rows={unmodeled}." ) + if unmodeled or residual_nulls: + failures.append( + f"{label}: declared post-PUF transfer left " + f"unmodeled_rows={unmodeled}, " + f"residual_null_rows={residual_nulls}; zero are allowed." + ) target_receipts[f"{entity}/{family}/{target}"] = { + "producer_roles": [ + role + for role, active in ( + ("puf_clone", puf_produced), + ("asec_source", source_produced), + ) + if active + ], + "producer_rows": pre_counts[(entity, target)]["producer_rows"], "authorized_null_rows": authorized, "imputed_rows": imputed, "unmodeled_rows": unmodeled, @@ -2989,20 +5147,10 @@ def _verify_gap_fill_outcome( } if failures: raise ValueError( - "Stacked gap-fill outcome verification failed:\n " + "\n ".join(failures) + "Stacked post-PUF transfer outcome verification failed:\n " + + "\n ".join(failures) ) - return { - "recipient_channel": direction.recipient_channel, - "donor_channel": direction.donor_channel, - "donor_selection": "owner_projection_of_native_donor_rows", - "resolved_donor_channel": result.resolved_donor_channel, - "targets": target_receipts, - "deferred_inputs": list(result.deferred_inputs), - "fit_records": [ - {"fit_name": record.fit_name, "weight_kind": record.weight_kind} - for record in result.fit_records - ], - } + return target_receipts # --------------------------------------------------------------------------- @@ -3101,8 +5249,12 @@ def _run_stacked_puf_pass_evaluate( "The stacked PUF pass owns clone attachment; found nonzero person " "support clone indices on its input." ) - cloned = clone_us_frame_for_puf_support( + universe_application = apply_acs_pums_earnings_universe_zeros( frame, + boundary="stacked PUF pass ACS earnings universe", + ) + cloned = clone_us_frame_for_puf_support( + universe_application.frame, clone_attachment_fraction=clone_attachment_fraction, clone_attachment_seed=clone_attachment_seed, ) @@ -3121,6 +5273,7 @@ def _run_stacked_puf_pass_evaluate( if tax_unit_outputs is not None: kwargs["tax_unit_outputs"] = tuple(tax_unit_outputs) if primary_qrf_checkpoint_dir is None: + predictor_universe_receipts: list[dict[str, object]] = [] imputed = impute_us_puf_tax_detail_support( cloned, donor_tax_units, @@ -3128,13 +5281,20 @@ def _run_stacked_puf_pass_evaluate( n_estimators=n_estimators, fit_records=fit_records, tail_bound_diagnostics=tail_bound_diagnostics, + predictor_universe_receipts=predictor_universe_receipts, require_complete_recipient_predictors=True, absent_cells=PUF_ABSENT_CELLS_PRESERVE_NULLS, **kwargs, ) + if len(predictor_universe_receipts) != 1: + raise AssertionError( + "Stacked monolithic PUF pass emitted the wrong number of " + "recipient predictor universe receipts." + ) primary_qrf_receipt: dict[str, object] = { "mode": "monolithic", "resume_status": "not_applicable", + "recipient_predictor_universe": predictor_universe_receipts[0], } else: checkpoint_dir = Path(primary_qrf_checkpoint_dir) @@ -3158,6 +5318,9 @@ def _run_stacked_puf_pass_evaluate( **kwargs, ) resume_status = "initialized" + predictor_universe_receipt = ( + primary_puf_qrf_recipient_predictor_universe_receipt(checkpoint_dir) + ) run_primary_puf_qrf_chain(checkpoint_dir) imputed, weight_kind = finalize_primary_puf_qrf_chain( cloned, @@ -3170,6 +5333,7 @@ def _run_stacked_puf_pass_evaluate( "mode": "checkpoint_chain", "resume_status": resume_status, "checkpoint_manifest": str(manifest_path.resolve()), + "recipient_predictor_universe": predictor_universe_receipt, } validate_stacked_spine_frame(imputed, boundary="stacked primary PUF output") @@ -3187,7 +5351,18 @@ def _run_stacked_puf_pass_evaluate( seed=seed, ) validate_puf_capital_gains_tail_manifest(tail_receipt) - tail_receipt = _bind_stacked_tail_origin_receipt(output, tail_receipt) + # The tail producer creates clone role 2 before its final origin + # receipt can be added to the tail manifest. Bind the producer's + # original manifest into a provisional attachment first, so no + # unreceipted clone role crosses the stacked validation boundary used + # to derive that origin receipt. Rebind once more to the final tail + # digest below. + provisional = bind_puf_clone_attachment_tail_descendant( + output, + attachment_receipt=attachment, + tail_manifest=tail_receipt, + ) + tail_receipt = _bind_stacked_tail_origin_receipt(provisional, tail_receipt) output = bind_puf_clone_attachment_tail_descendant( output, attachment_receipt=attachment, @@ -3224,6 +5399,9 @@ def _run_stacked_puf_pass_evaluate( return StackedPufPassResult( frame=output, receipt={ + "acs_earnings_universe_application": _json_ready( + universe_application.receipt + ), "clone_attachment": _json_ready(attachment), "doctrines": { "require_complete_recipient_predictors": True, @@ -3887,6 +6065,9 @@ def _stacked_completeness_gate_evaluate( _canonical_gap_fill_plan: tuple[ GapFillDirection, ... ] = CANONICAL_STACKED_GAP_FILL_PLAN, + _canonical_post_puf_surface: TargetFamilies = ( + CANONICAL_STACKED_POST_PUF_TRANSFER_SURFACE + ), ) -> GateResult: """Prove every declared target is filled or carries absence authority. @@ -3907,7 +6088,7 @@ def _stacked_completeness_gate_evaluate( if declared_count == 0: failures.append("declared stacked surface contains zero targets.") if failures: - return GateResult( + return _sealed_stacked_gate_result( name=_COMPLETENESS_GATE_NAME, passed=False, failures=tuple(failures), @@ -3919,6 +6100,9 @@ def _stacked_completeness_gate_evaluate( ) declared_directions: dict[tuple[str, str], GapFillDirection] = {} + declared_absence_rules: dict[ + tuple[str, str], tuple[GapFillDirection, GapFillAbsenceRule] + ] = {} for direction in authority.gap_fill_plan: for entity, targets in _direction_entity_targets(direction).items(): for target in targets: @@ -3931,6 +6115,8 @@ def _stacked_completeness_gate_evaluate( f"{direction.name!r}." ) declared_directions[key] = direction + for rule in direction.recipient_absence_rules: + declared_absence_rules[(rule.entity, rule.column)] = (direction, rule) canonical_direction_keys: set[tuple[str, str]] = set() for direction in _canonical_gap_fill_plan: for entity, targets in _direction_entity_targets(direction).items(): @@ -3938,6 +6124,21 @@ def _stacked_completeness_gate_evaluate( key = (entity, target) canonical_direction_keys.add(key) declared_directions[key] = direction + for rule in direction.recipient_absence_rules: + declared_absence_rules[(rule.entity, rule.column)] = (direction, rule) + declared_post_puf_keys = { + (entity, target) + for entity, families in authority.post_puf_transfer_surface.items() + for targets in families.values() + for target in targets + } + canonical_post_puf_keys = { + (entity, target) + for entity, families in _canonical_post_puf_surface.items() + for targets in families.values() + for target in targets + } + post_puf_keys = declared_post_puf_keys | canonical_post_puf_keys proof_index: dict[tuple[str, str], dict[tuple[str, int], str]] = {} for proof in absence_proofs: @@ -3959,6 +6160,9 @@ def _stacked_completeness_gate_evaluate( } authority_sha256 = authority_receipt["sha256"] plan_sha256 = authority_receipt["components"]["gap_fill_plan"]["sha256"] + post_puf_surface_sha256 = authority_receipt["components"][ + "post_puf_transfer_surface" + ]["sha256"] surface_sha256 = authority_receipt["components"]["declared_surface"]["sha256"] def authority_binding(authority_form: str) -> dict[str, object]: @@ -3966,6 +6170,7 @@ def authority_binding(authority_form: str) -> dict[str, object]: "authority_form": authority_form, "authority_sha256": authority_sha256, "plan_sha256": plan_sha256, + "post_puf_surface_sha256": post_puf_surface_sha256, "surface_sha256": surface_sha256, } @@ -4007,6 +6212,32 @@ def authority_binding(authority_form: str) -> dict[str, object]: } continue null_mask = table[target].isna() + structural_absence_receipt: dict[str, object] | None = None + structural_absence_mismatch = False + structural_rule = declared_absence_rules.get((entity, target)) + if structural_rule is not None: + rule_direction, rule = structural_rule + structural_mask, structural_absence_receipt = ( + _gap_fill_absence_rule_mask( + frame, + direction=rule_direction, + rule=rule, + ) + ) + unexpected = int((null_mask & ~structural_mask).sum()) + synthesized = int((structural_mask & ~null_mask).sum()) + structural_absence_receipt = { + **structural_absence_receipt, + "unexpected_null_rows": unexpected, + "structural_rows_filled": synthesized, + } + structural_absence_mismatch = bool(unexpected or synthesized) + if structural_absence_mismatch: + failures.append( + f"{label}: exact structural-absence equation failed; " + f"unexpected_null_rows={unexpected}, " + f"structural_rows_filled={synthesized}." + ) metric = metric_by_target[(entity, family, target)] invalid_mask, invalidity = _declared_metric_invalidity( table[target], @@ -4038,7 +6269,13 @@ def authority_binding(authority_form: str) -> dict[str, object]: ) if not null_mask.any(): target_receipts[label] = { - "status": ("invalid_values" if invalid_rows else "complete"), + "status": ( + "structural_absence_mismatch" + if structural_absence_mismatch + else "invalid_values" + if invalid_rows + else "complete" + ), "null_rows": 0, "metric": metric, "invalid_rows": invalid_rows, @@ -4049,10 +6286,39 @@ def authority_binding(authority_form: str) -> dict[str, object]: if invalid_rows else "observed_complete" ), + **( + { + "recipient_absence_authority": ( + structural_absence_receipt + ) + } + if structural_absence_receipt is not None + else {} + ), } continue - proofs = proof_index.get((entity, target), {}) + proofs = dict(proof_index.get((entity, target), {})) + if structural_rule is not None: + proofs = {} + if not structural_absence_mismatch: + _rule_direction, rule = structural_rule + structural_mask, _receipt = _gap_fill_absence_rule_mask( + frame, + direction=_rule_direction, + rule=rule, + ) + structural_channels = channel.loc[structural_mask] + structural_clones = clone_index.loc[structural_mask] + for cell_channel, cell_clone in set( + zip( + structural_channels.astype(str), + structural_clones.astype(int), + strict=True, + ) + ): + proofs[(str(cell_channel), int(cell_clone))] = rule.reason declared_direction = declared_directions.get((entity, target)) + post_puf_target = (entity, target) in post_puf_keys null_channels = channel.loc[null_mask] null_clones = clone_index.loc[null_mask] unproven: dict[str, int] = {} @@ -4068,6 +6334,15 @@ def authority_binding(authority_form: str) -> dict[str, object]: cell_clone = int(cell_clone) reason = proofs.get((cell_channel, cell_clone)) authority_form = "origin_exact" + if post_puf_target: + if reason is not None or (_ANY_CHANNEL, cell_clone) in proofs: + failures.append( + f"{label}: absence authority is forbidden because " + "the declared post-PUF transfer requires zero " + f"residual nulls; found {int(count)} null cell(s) " + f"on {cell_channel}/clone_{cell_clone}." + ) + reason = None if declared_direction is not None and reason is not None: if cell_channel != declared_direction.recipient_channel: failures.append( @@ -4080,7 +6355,11 @@ def authority_binding(authority_form: str) -> dict[str, object]: reason = None else: authority_form = "origin_exact_recipient" - if reason is None and declared_direction is None: + if ( + reason is None + and declared_direction is None + and not post_puf_target + ): reason = proofs.get((_ANY_CHANNEL, cell_clone)) authority_form = "wildcard_no_declared_donor_plan" key = f"{cell_channel}/clone_{cell_clone}" @@ -4097,6 +6376,12 @@ def authority_binding(authority_form: str) -> dict[str, object]: f"{declared_direction.recipient_channel!r}; wildcard " "authority is forbidden." ) + elif post_puf_target: + failures.append( + f"{label}: {int(count)} null cell(s) on {key} " + "violate the zero-residual post-PUF transfer " + "contract; absence proofs are forbidden." + ) else: proof_receipt: dict[str, object] = { "null_rows": int(count), @@ -4115,6 +6400,14 @@ def authority_binding(authority_form: str) -> dict[str, object]: ), } ) + if structural_rule is not None: + _rule_direction, rule = structural_rule + proof_receipt.update( + { + "structural_absence_rule_id": rule.rule_id, + "structural_absence_selection": rule.selection, + } + ) proven[key] = proof_receipt target_authority_forms.add(authority_form) if unproven: @@ -4149,8 +6442,13 @@ def authority_binding(authority_form: str) -> dict[str, object]: if invalid_rows and not unproven else target_authority_form ), + **( + {"recipient_absence_authority": (structural_absence_receipt)} + if structural_absence_receipt is not None + else {} + ), } - return GateResult( + return _sealed_stacked_gate_result( name=_COMPLETENESS_GATE_NAME, passed=not failures, failures=tuple(failures), @@ -4276,6 +6574,9 @@ def _by_origin_battery_evaluate( *, authority: _StackedAuthority, production: bool, + _canonical_gap_fill_plan: tuple[ + GapFillDirection, ... + ] = CANONICAL_STACKED_GAP_FILL_PLAN, ) -> GateResult: """Compare declared statistics between origins within the one spine. @@ -4339,7 +6640,7 @@ def _by_origin_battery_evaluate( "support_profile": support_profile_receipt, } if registration_failures: - return GateResult( + return _sealed_stacked_gate_result( name=_BATTERY_GATE_NAME, passed=False, failures=tuple(registration_failures), @@ -4355,8 +6656,15 @@ def _by_origin_battery_evaluate( failures: list[str] = [] comparisons: dict[str, object] = {} + structural_absence_receipts: dict[str, dict[str, object]] = {} untestable: list[str] = [] tested = 0 + declared_absence_rules: dict[ + tuple[str, str], tuple[GapFillDirection, GapFillAbsenceRule] + ] = {} + for direction in (*authority.gap_fill_plan, *_canonical_gap_fill_plan): + for rule in direction.recipient_absence_rules: + declared_absence_rules[(rule.entity, rule.column)] = (direction, rule) for spec in specs: table = frame.table(spec.entity) channel = table[support_channel_column(spec.entity)].astype(str) @@ -4369,8 +6677,6 @@ def _by_origin_battery_evaluate( dtype=np.float64, ) scope = (clone_index.eq(spec.clone_index)).to_numpy() & (weights > 0.0) - left_rows = scope & channel.eq(BASE_ASEC_SUPPORT_CHANNEL).to_numpy() - right_rows = scope & channel.eq(ACS_STACKED_SUPPORT_CHANNEL).to_numpy() for column, metric in spec.column_metrics.items(): label = f"{spec.entity}/{spec.family}/{column}[clone_{spec.clone_index}]" if column not in table.columns: @@ -4381,7 +6687,44 @@ def _by_origin_battery_evaluate( } continue series = table[column] - scoped_nulls = int(series.isna().to_numpy(dtype=bool)[scope].sum()) + target_scope = scope.copy() + structural_rule = declared_absence_rules.get((spec.entity, column)) + if structural_rule is not None: + rule_direction, rule = structural_rule + structural_mask, structural_receipt = _gap_fill_absence_rule_mask( + frame, + direction=rule_direction, + rule=rule, + ) + scoped_structural = structural_mask.to_numpy(dtype=bool) & scope + scoped_null_mask = series.isna().to_numpy(dtype=bool) & scope + unexpected = int((scoped_null_mask & ~scoped_structural).sum()) + synthesized = int((scoped_structural & ~scoped_null_mask).sum()) + structural_receipt = { + **structural_receipt, + "comparison_clone_index": spec.clone_index, + "rows_excluded_from_scope": int(scoped_structural.sum()), + "unexpected_null_rows": unexpected, + "structural_rows_filled": synthesized, + } + structural_absence_receipts[label] = structural_receipt + if unexpected or synthesized: + failures.append( + f"{label}: exact structural-absence equation failed; " + f"unexpected_null_rows={unexpected}, " + f"structural_rows_filled={synthesized}." + ) + comparisons[label] = { + "status": "structural_absence_mismatch", + "metric": metric, + } + continue + target_scope &= ~scoped_structural + left_rows = target_scope & channel.eq(BASE_ASEC_SUPPORT_CHANNEL).to_numpy() + right_rows = ( + target_scope & channel.eq(ACS_STACKED_SUPPORT_CHANNEL).to_numpy() + ) + scoped_nulls = int(series.isna().to_numpy(dtype=bool)[target_scope].sum()) if scoped_nulls: failures.append( f"{label}: {scoped_nulls} null value(s) inside the " @@ -4397,7 +6740,7 @@ def _by_origin_battery_evaluate( invalid_mask, invalidity = _declared_metric_invalidity( series, metric=metric, - scope=scope, + scope=target_scope, ) invalid_rows = int(invalid_mask.sum()) if invalid_rows: @@ -4420,7 +6763,7 @@ def _by_origin_battery_evaluate( label, series, metric=metric, - scope=scope, + scope=target_scope, failures=failures, ) if values is None: @@ -4540,7 +6883,11 @@ def _by_origin_battery_evaluate( failures=failures, comparisons=comparisons, ) - return GateResult( + for label, receipt in structural_absence_receipts.items(): + comparison = comparisons.get(label) + if isinstance(comparison, dict): + comparison["recipient_absence_authority"] = receipt + return _sealed_stacked_gate_result( name=_BATTERY_GATE_NAME, passed=not failures, failures=tuple(failures), diff --git a/packages/microcosm-build/tests/test_puf_qrf_chain.py b/packages/microcosm-build/tests/test_puf_qrf_chain.py index 93633272..08ff48ec 100644 --- a/packages/microcosm-build/tests/test_puf_qrf_chain.py +++ b/packages/microcosm-build/tests/test_puf_qrf_chain.py @@ -466,6 +466,77 @@ def test_target_subprocess_chain_matches_monolith_raw_bits_and_final_frame( run_primary_puf_qrf_chain(checkpoint_dir) +def test_resumed_earnings_chain_keeps_children_out_of_allocation( + tmp_path: Path, + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setenv("POPULACE_FIT_N_JOBS", "1") + monkeypatch.setenv("POPULACE_FIT_PREDICT_WORKERS", "1") + frame = _expanded_frame() + person = frame.table("person") + source_age = {1: 12.0, 2: 40.0, 3: 10.0} + person["age"] = person["person_source_id"].map(source_age).astype("float64") + donor = _donor() + donor["employment_income_before_lsr"] = np.linspace( + 1_000.0, + 40_000.0, + len(donor), + ) + donor["self_employment_income_before_lsr"] = np.linspace( + 100.0, + 4_000.0, + len(donor), + ) + checkpoint_dir = tmp_path / "primary_qrf_earnings_resume" + person_outputs = ( + "employment_income_before_lsr", + "self_employment_income_before_lsr", + ) + manifest = initialize_primary_puf_qrf_chain( + frame, + donor, + checkpoint_dir, + predictors=_PREDICTORS, + person_outputs=person_outputs, + tax_unit_outputs=(), + n_estimators=2, + seed=19, + require_complete_recipient_predictors=True, + absent_cells=PUF_ABSENT_CELLS_PRESERVE_NULLS, + ) + allocation = manifest["recipient_predictor_universe"]["person_output_allocation"] + assert allocation["person_outputs"] == sorted(person_outputs) + assert allocation["out_of_universe_person_rows"] == 2 + assert allocation["empty_eligible_tax_unit_rows"] == 1 + assert allocation["first_person_fallback_out_of_universe_rows"] == 0 + + completed_target = run_primary_puf_qrf_target(checkpoint_dir, 0) + completed_bytes = completed_target.read_bytes() + pending_target = puf_qrf_chain_module._target_path(checkpoint_dir, manifest, 1) + assert not pending_target.exists() + + # The supervisor preserves the completed prefix and resumes the missing target. + run_primary_puf_qrf_chain(checkpoint_dir) + assert completed_target.read_bytes() == completed_bytes + assert pending_target.is_file() + resumed, weight_kind = finalize_primary_puf_qrf_chain(frame, checkpoint_dir) + + assert weight_kind == "design" + resumed_person = resumed.table("person") + detail = resumed_person["person_support_channel"].eq("puf_tax_detail") + children = detail & resumed_person["age"].lt(15) + adults = detail & resumed_person["age"].ge(15) + assert int(children.sum()) == 2 + assert resumed_person.loc[children, list(person_outputs)].eq(0.0).all().all() + assert resumed_person.loc[adults, list(person_outputs)].gt(0.0).all().all() + all_child_totals = ( + resumed_person.loc[detail] + .groupby("person_tax_unit_id", sort=False)[list(person_outputs)] + .sum() + ) + assert all_child_totals.loc[100020].eq(0.0).all() + + def test_primary_qrf_chain_manifest_binds_stacked_doctrines( tmp_path: Path, monkeypatch: pytest.MonkeyPatch, @@ -489,7 +560,9 @@ def test_primary_qrf_chain_manifest_binds_stacked_doctrines( expected_controls = { "require_complete_recipient_predictors": True, "absent_cells": PUF_ABSENT_CELLS_PRESERVE_NULLS, + "recipient_predictor_universe": manifest["recipient_predictor_universe"], } + assert len(expected_controls["recipient_predictor_universe"]["sha256"]) == 64 assert {name: manifest[name] for name in expected_controls} == expected_controls for filename in ( puf_qrf_chain_module.PRIMARY_QRF_DONOR_FILENAME, @@ -514,6 +587,14 @@ def test_primary_qrf_chain_manifest_binds_stacked_doctrines( with pytest.raises(ValueError, match="doctrine controls"): run_primary_puf_qrf_target(checkpoint_dir, 1) + tampered_manifest = dict(manifest) + tampered_universe = dict(manifest["recipient_predictor_universe"]) + tampered_universe["recipient_tax_unit_rows"] += 1 + tampered_manifest["recipient_predictor_universe"] = tampered_universe + manifest_path.write_text(json.dumps(tampered_manifest), encoding="utf-8") + with pytest.raises(ValueError, match="universe digest is invalid"): + run_primary_puf_qrf_target(checkpoint_dir, 1) + def test_primary_qrf_chain_stacked_finalization_preserves_unowned_nulls( tmp_path: Path, @@ -553,7 +634,7 @@ def test_primary_qrf_chain_stacked_finalization_preserves_unowned_nulls( assert tax_unit.loc[tax_unit_clone.eq(1), "domestic_production_ald"].notna().all() -def test_primary_qrf_chain_legacy_v5_manifest_defaults_remain_loadable( +def test_primary_qrf_chain_legacy_doctrine_defaults_remain_loadable( tmp_path: Path, monkeypatch: pytest.MonkeyPatch, ) -> None: @@ -572,8 +653,10 @@ def test_primary_qrf_chain_legacy_v5_manifest_defaults_remain_loadable( n_estimators=2, seed=3, ) + assert manifest["schema_version"] == 6 assert "require_complete_recipient_predictors" not in manifest assert "absent_cells" not in manifest + assert "recipient_predictor_universe" not in manifest for filename in ( puf_qrf_chain_module.PRIMARY_QRF_DONOR_FILENAME, puf_qrf_chain_module.PRIMARY_QRF_RECIPIENT_FILENAME, @@ -581,6 +664,7 @@ def test_primary_qrf_chain_legacy_v5_manifest_defaults_remain_loadable( metadata = load_frame_checkpoint(checkpoint_dir / filename).metadata assert "require_complete_recipient_predictors" not in metadata assert "absent_cells" not in metadata + assert "recipient_predictor_universe" not in metadata for target_index in range(len((*_PERSON_OUTPUTS, *_TAX_UNIT_OUTPUTS))): run_primary_puf_qrf_target(checkpoint_dir, target_index) @@ -615,15 +699,53 @@ def test_primary_qrf_chain_rejects_incomplete_stacked_predictors( require_complete_recipient_predictors=True, absent_cells=PUF_ABSENT_CELLS_PRESERVE_NULLS, ) + assert not (tmp_path / "primary_qrf").exists() + + +def test_primary_qrf_finalization_rejects_changed_recipient_features( + tmp_path: Path, + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setenv("POPULACE_FIT_N_JOBS", "1") + monkeypatch.setenv("POPULACE_FIT_PREDICT_WORKERS", "1") + frame = _expanded_frame() + checkpoint_dir = tmp_path / "primary_qrf" + initialize_primary_puf_qrf_chain( + frame, + _donor(), + checkpoint_dir, + predictors=_PREDICTORS, + person_outputs=_PERSON_OUTPUTS, + tax_unit_outputs=_TAX_UNIT_OUTPUTS, + n_estimators=2, + seed=3, + require_complete_recipient_predictors=True, + absent_cells=PUF_ABSENT_CELLS_PRESERVE_NULLS, + ) + for target_index in range(len((*_PERSON_OUTPUTS, *_TAX_UNIT_OUTPUTS))): + run_primary_puf_qrf_target(checkpoint_dir, target_index) + tax_unit = frame.table("tax_unit") + detail = tax_unit["tax_unit_support_clone_index"].eq(1) + row = tax_unit.index[detail][0] + tax_unit.loc[row, "filing_status_input"] = "SINGLE" + + with pytest.raises( + ValueError, + match="changed the PUF recipient predictor source universe or feature values", + ): + finalize_primary_puf_qrf_chain(frame, checkpoint_dir) -def test_primary_qrf_rejects_stale_donor_construction_schema_versions( +@pytest.mark.parametrize("stale_version", (1, 2, 3, 4, 5)) +def test_primary_qrf_rejects_every_stale_schema_version( tmp_path: Path, monkeypatch: pytest.MonkeyPatch, + stale_version: int, ) -> None: - # Loading validates no donor-construction identity, so stale roots and - # target checkpoints must reject every pre-field-local schema: pre-carve - # v1, pre-screen v2, national-carve v3, and whole-row-quarantine v4. + # Roots and target checkpoints must reject every schema predating the v6 + # recipient-universe authority. In particular, the v5 case models the old + # strict two-control payload that omitted recipient_predictor_universe; its + # stale schema is rejected before that payload can be interpreted. monkeypatch.setenv("POPULACE_FIT_N_JOBS", "1") monkeypatch.setenv("POPULACE_FIT_PREDICT_WORKERS", "1") checkpoint_dir = tmp_path / "primary_qrf" @@ -648,34 +770,45 @@ def test_primary_qrf_rejects_stale_donor_construction_schema_versions( manifest_path = checkpoint_dir / "manifest.json" original_manifest = json.loads(manifest_path.read_text()) assert original_manifest["schema_version"] == PRIMARY_QRF_CHECKPOINT_SCHEMA_VERSION - # This protects custom target-order standalone checkpoints whose manifests - # do not hash donor construction. - for stale_version in (1, 2, 3, 4): - stale_manifest = dict(original_manifest) - stale_manifest["schema_version"] = stale_version - manifest_path.write_text(json.dumps(stale_manifest)) - with pytest.raises(ValueError, match="schema version"): - load_primary_puf_qrf_predictions(checkpoint_dir) - with pytest.raises(ValueError, match="schema version"): - run_primary_puf_qrf_chain(checkpoint_dir) + stale_manifest = dict(original_manifest) + stale_manifest["schema_version"] = stale_version + if stale_version == 5: + stale_manifest["require_complete_recipient_predictors"] = True + stale_manifest["absent_cells"] = PUF_ABSENT_CELLS_PRESERVE_NULLS + manifest_path.write_text(json.dumps(stale_manifest)) + with pytest.raises( + ValueError, + match=rf"schema version: expected 6, got {stale_version}", + ): + load_primary_puf_qrf_predictions(checkpoint_dir) + with pytest.raises( + ValueError, + match=rf"schema version: expected 6, got {stale_version}", + ): + run_primary_puf_qrf_chain(checkpoint_dir) manifest_path.write_text(json.dumps(original_manifest)) target_path = sorted((checkpoint_dir / "targets").glob("*.h5"))[0] with h5py.File(target_path, mode="r") as h5: pristine_metadata = json.loads(bytes(h5["metadata_json"][...]).decode()) assert pristine_metadata["schema_version"] == PRIMARY_QRF_CHECKPOINT_SCHEMA_VERSION - for stale_version in (1, 2, 3, 4): - metadata = dict(pristine_metadata) - metadata["schema_version"] = stale_version - with h5py.File(target_path, mode="r+") as h5: - del h5["metadata_json"] - h5.create_dataset( - "metadata_json", - data=np.frombuffer(json.dumps(metadata).encode(), dtype=np.uint8), - track_times=False, - ) - with pytest.raises(ValueError, match="invalid schema_version"): - load_primary_puf_qrf_predictions(checkpoint_dir) + metadata = dict(pristine_metadata) + metadata["schema_version"] = stale_version + if stale_version == 5: + metadata["require_complete_recipient_predictors"] = True + metadata["absent_cells"] = PUF_ABSENT_CELLS_PRESERVE_NULLS + with h5py.File(target_path, mode="r+") as h5: + del h5["metadata_json"] + h5.create_dataset( + "metadata_json", + data=np.frombuffer(json.dumps(metadata).encode(), dtype=np.uint8), + track_times=False, + ) + with pytest.raises( + ValueError, + match=rf"invalid schema_version: expected 6, got {stale_version}", + ): + load_primary_puf_qrf_predictions(checkpoint_dir) def test_primary_qrf_resume_rejects_a_gap( diff --git a/packages/microcosm-build/tests/test_us_multispine_pool.py b/packages/microcosm-build/tests/test_us_multispine_pool.py index 59cc7103..3b831342 100644 --- a/packages/microcosm-build/tests/test_us_multispine_pool.py +++ b/packages/microcosm-build/tests/test_us_multispine_pool.py @@ -1,6 +1,7 @@ from __future__ import annotations import ast +import copy import hashlib import importlib import inspect @@ -42,6 +43,10 @@ materialize_multispine_agreement_outputs, materialize_pool_deferred_transfer_inputs, pool_input_surface, + pool_post_puf_puf_producer_target_families, + pool_post_puf_source_producer_target_families, + pool_post_puf_transfer_target_families, + pool_pre_clone_gap_fill_target_families, pool_transfer_target_families, prepare_multispine_puf_predictors, prepare_multispine_source_inputs_for_clone, @@ -59,6 +64,12 @@ PUF_SUPPORT_MAX_CLONE_SAFE_SOURCE_ID, clone_us_frame_for_puf_support, ) +from microcosm.build.us_runtime.qbi_inputs import ( + US_QBI_OUTPUT_COLUMNS, + bind_us_qbi_reconciliation_transition_authority, + us_qbi_reconciliation_change_receipt, + with_us_qbi_input_reconciliation, +) from microcosm.build.us_runtime.spine_agreement import ( SpineAgreementSpec, default_spine_agreement_registry, @@ -76,6 +87,8 @@ PolicyEngineUSVariableMetadataIndex, ) +_FIXTURE_SEED_PERSON_COLUMN = "takes_up_medicaid_if_eligible" + _EXPECTED_POOL_SOURCE_OPERATOR_ORDER = ( "derive_us_cps_carried_inputs", "with_us_hours_worked_inputs", @@ -138,6 +151,7 @@ def _installed_variable_metadata_index() -> PolicyEngineUSVariableMetadataIndex: _EXPECTED_SOURCE_OPERATOR_WRAPPERS = { "with_us_hours_worked_inputs": "_with_gated_us_hours_worked_inputs", + "with_us_qbi_input_reconciliation": "reconcile_qbi_with_receipt", } @@ -402,22 +416,49 @@ def apply(frame: Frame) -> PoolStageOutput: def _fixture_pool_operators( order: list[str], ) -> dict[str, Callable[[Frame], PoolStageOutput]]: + def derive(frame: Frame) -> PoolStageOutput: + order.append("derive") + person = frame.table("person").copy() + person["derived"] = person["transferred"] * 2 + person["SEMP"] = 0.0 + person["self_employment_income_before_lsr"] = 0.0 + for column in US_QBI_OUTPUT_COLUMNS: + person[column] = 0.0 + before = _replace_person(frame, person) + after = with_us_qbi_input_reconciliation(before) + receipt = us_qbi_reconciliation_change_receipt(before, after) + after = bind_us_qbi_reconciliation_transition_authority(after, receipt) + return PoolStageOutput( + after, + { + "operator": "derive", + "qbi_input_reconciliation": receipt, + }, + qbi_transition_authority_sha256=receipt["sha256"], + ) + + def seed(frame: Frame) -> PoolStageOutput: + order.append("seed") + person = frame.table("person").copy() + person[_FIXTURE_SEED_PERSON_COLUMN] = person["age"] >= 40 + return PoolStageOutput( + _replace_person(frame, person), + { + "operator": "seed", + "programs": { + _FIXTURE_SEED_PERSON_COLUMN: {"entity": "person"}, + }, + }, + ) + return { "impute": _operator( "impute", order, lambda person: person.__setitem__("transferred", person["age"]), ), - "derive": _operator( - "derive", - order, - lambda person: person.__setitem__("derived", person["transferred"] * 2), - ), - "seed": _operator( - "seed", - order, - lambda person: person.__setitem__("seeded", person["age"] >= 40), - ), + "derive": derive, + "seed": seed, "simulate": _operator( "simulate", order, @@ -591,6 +632,131 @@ def test_pool_checkpoint_callbacks_capture_each_fixed_boundary() -> None: assert checkpoint.assembly_receipt == result.assembly_receipt +def test_pool_checkpoint_emission_requires_qbi_receipt() -> None: + order: list[str] = [] + operators = _fixture_pool_operators(order) + operators["derive"] = _operator( + "derive", + order, + lambda person: person.__setitem__( + "derived", + person["transferred"] * 2, + ), + ) + + with pytest.raises( + ValueError, + match=( + "multispine pool simulated checkpoint emission: derive receipt " + "has no QBI reconciliation object" + ), + ): + run_multispine_pool_path( + _source_frame(), + _source_frame(), + **operators, + agreement_gate=lambda _frame: GateResult("fixture", True), + checkpoint=lambda _checkpoint: None, + ) + + +def test_legacy_runtime_qbi_route_rejects_present_stacked_none() -> None: + with pytest.raises( + ValueError, + match="legacy checkpoint used the stacked derive receipt route", + ): + multispine_pool_module._qbi_receipt_from_stage_receipts( + { + "derive": { + "qbi_input_reconciliation": {"fixture": "legacy"}, + "pool_derivation": None, + } + }, + boundary="ambiguous legacy QBI route test", + ) + + +def test_pool_simulated_resume_rejects_forged_qbi_receipt() -> None: + checkpoints: list[MultispinePoolCheckpoint] = [] + run_multispine_pool_path( + _source_frame(), + _source_frame(), + **_fixture_pool_operators([]), + agreement_gate=lambda _frame: GateResult("fixture", True), + checkpoint=checkpoints.append, + ) + simulated = checkpoints[-1] + stage_receipts = copy.deepcopy(simulated.stage_receipts) + stage_receipts["derive"]["qbi_input_reconciliation"]["sha256"] = "0" * 64 + forged = MultispinePoolCheckpoint( + stage="simulated", + frame=simulated.frame, + assembly_receipt=simulated.assembly_receipt, + stage_receipts=stage_receipts, + simulation_frame=simulated.simulation_frame, + qbi_transition_authority_sha256=(simulated.qbi_transition_authority_sha256), + ) + + with pytest.raises( + ValueError, + match=( + "multispine pool simulated checkpoint resume: QBI reconciliation " + "receipt SHA-256" + ), + ): + run_multispine_pool_path( + None, + None, + **_fixture_pool_operators([]), + agreement_gate=lambda _frame: GateResult("fixture", True), + resume=forged, + ) + + +def test_pool_simulated_resume_rejects_reissued_qbi_receipt() -> None: + checkpoints: list[MultispinePoolCheckpoint] = [] + run_multispine_pool_path( + _source_frame(), + _source_frame(), + **_fixture_pool_operators([]), + agreement_gate=lambda _frame: GateResult("fixture", True), + checkpoint=checkpoints.append, + ) + simulated = checkpoints[-1] + persistent_person = simulated.frame.table("person").copy() + persistent_person["non_qualified_dividend_income"] = 100.0 + persistent_person["qualified_bdc_income"] = 50.0 + persistent = _replace_person(simulated.frame, persistent_person) + simulation_person = simulated.simulation_frame.table("person").copy() + simulation_person["non_qualified_dividend_income"] = 100.0 + simulation_person["qualified_bdc_income"] = 50.0 + simulation = _replace_person(simulated.simulation_frame, simulation_person) + receipts = copy.deepcopy(simulated.stage_receipts) + receipts["derive"]["qbi_input_reconciliation"] = ( + us_qbi_reconciliation_change_receipt(persistent, persistent) + ) + forged = MultispinePoolCheckpoint( + stage="simulated", + frame=persistent, + assembly_receipt=simulated.assembly_receipt, + stage_receipts=receipts, + simulation_frame=simulation, + qbi_transition_authority_sha256=(simulated.qbi_transition_authority_sha256), + ) + + with pytest.raises( + ValueError, + match="independently carried transition authority", + ): + run_multispine_pool_path( + None, + None, + **_fixture_pool_operators([]), + agreement_gate=lambda _frame: GateResult("fixture", True), + resume=forged, + ) + + def test_fresh_pool_path_rejects_missing_source_frames() -> None: with pytest.raises( TypeError, @@ -829,6 +995,48 @@ def test_pool_transfer_plan_extends_legacy_except_receipted_asset_deferrals() -> ) +def test_pool_transfer_plan_partitions_at_the_declared_producer_boundary() -> None: + def keys(target_families): + return { + (entity, family, target) + for entity, families in target_families.items() + for family, targets in families.items() + for target in targets + } + + full = keys(pool_transfer_target_families()) + early = keys(pool_pre_clone_gap_fill_target_families()) + late = keys(pool_post_puf_transfer_target_families()) + puf_producers = keys(pool_post_puf_puf_producer_target_families()) + source_producers = keys(pool_post_puf_source_producer_target_families()) + + assert len(early) == 48 + assert len(late) == 70 + assert early.isdisjoint(late) + assert early | late == full + assert len(puf_producers) == 43 + assert len(source_producers) == 30 + assert len(puf_producers & source_producers) == 3 + assert puf_producers | source_producers == late + assert ("person", "source_operator_cps_carried", "strike_benefits") in early + assert ("person", "model_required_boolean", "is_pregnant") in late + assert ( + "tax_unit", + "puf_tax_itemization", + "health_savings_account_ald", + ) in late + assert ( + "tax_unit", + "puf_tax_itemization", + "health_savings_account_ald", + ) in puf_producers + assert ( + "person", + "model_required_boolean", + "is_pregnant", + ) in source_producers + + def test_pool_input_surface_normalizes_all_four_source_registries() -> None: surface = pool_input_surface() by_name = {entry.variable: entry for entry in surface} @@ -1751,7 +1959,17 @@ def test_production_operator_invocations_are_total_and_guarded( ( multispine_pool_module.derive_multispine_pool_inputs, POOL_DERIVE_OPERATOR_ORDER, - {"PoolStageOutput", "_run_source_operator_chain", "list"}, + { + "PoolStageOutput", + "_run_source_operator_chain", + "bind_us_qbi_reconciliation_transition_authority", + "dict", + "list", + "us_qbi_reconciliation_change_receipt", + "validate_us_qbi_reconciliation_live_output", + "validate_us_qbi_reconciliation_transition", + "with_us_qbi_input_reconciliation", + }, ), ) for ( @@ -1791,7 +2009,15 @@ def observe_guarded_chain( "phase": phase, "operator_order": list(operator_names), "suboperators": [ - {"operator": name, "kernel_receipt": {}} for name in operator_names + { + "operator": name, + "kernel_receipt": ( + {"sha256": "0" * 64} + if name == "with_us_qbi_input_reconciliation" + else {} + ), + } + for name in operator_names ], }, ) @@ -1801,6 +2027,16 @@ def observe_guarded_chain( "_run_source_operator_chain", observe_guarded_chain, ) + monkeypatch.setattr( + multispine_pool_module, + "validate_us_qbi_reconciliation_live_output", + lambda _frame, _receipt, *, boundary, expected_transition_authority_sha256: {}, + ) + monkeypatch.setattr( + multispine_pool_module, + "bind_us_qbi_reconciliation_transition_authority", + lambda current, _receipt: current, + ) frame = _source_frame() multispine_pool_module.prepare_multispine_source_inputs_for_clone( frame, @@ -1869,9 +2105,10 @@ def test_derive_stage_keeps_whole_pool_qbi_reconciliation() -> None: person = frame.table("person").copy() person["long_term_capital_gains_before_response"] = 100.0 person["non_sch_d_capital_gains"] = 0.0 - for column in multispine_pool_module.US_QBI_OUTPUT_COLUMNS: + for column in US_QBI_OUTPUT_COLUMNS: person[column] = 0.0 person["self_employment_income_before_lsr"] = 10.0 + person["SEMP"] = 10.0 person["sstb_self_employment_income_before_lsr"] = 5.0 frame = _replace_person(frame, person) @@ -1879,11 +2116,81 @@ def test_derive_stage_keeps_whole_pool_qbi_reconciliation() -> None: derived = result.frame.table("person") assert result.receipt["operator_order"] == list(POOL_DERIVE_OPERATOR_ORDER) + assert ( + result.receipt["qbi_input_reconciliation"]["recipient_source_universe"][ + "rows_excluded_from_base_self_employment_rewrite" + ] + == 0 + ) + assert result.receipt["qbi_input_reconciliation"][ + "base_self_employment_changed_rows" + ] == len(derived) + assert ( + result.receipt["qbi_input_reconciliation"][ + "structurally_absent_base_source_changed_rows" + ] + == 0 + ) assert derived["schedule_d_capital_gain_distributions"].notna().all() assert derived["self_employment_income_before_lsr"].eq(15.0).all() assert derived["sstb_self_employment_income_before_lsr"].eq(0.0).all() +def _qbi_ready_derive_frame() -> Frame: + assembled = assemble_spines( + {"asec": _source_frame(), "acs": _source_frame()}, + household_mass_shares={"asec": 0.5, "acs": 0.5}, + ) + frame = clone_us_frame_for_puf_support(assembled) + person = frame.table("person").copy() + person["long_term_capital_gains_before_response"] = 100.0 + person["non_sch_d_capital_gains"] = 0.0 + for column in US_QBI_OUTPUT_COLUMNS: + person[column] = 0.0 + person["self_employment_income_before_lsr"] = 10.0 + person["SEMP"] = 10.0 + person["sstb_self_employment_income_before_lsr"] = 5.0 + return _replace_person(frame, person) + + +def test_derive_stage_rejects_forged_qbi_kernel_receipt( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setattr( + multispine_pool_module, + "us_qbi_reconciliation_change_receipt", + lambda _before, _after: { + "version": 2, + "sha256": "0" * 64, + "changed_person_rows": -1, + "tampered": True, + }, + ) + + with pytest.raises(ValueError, match="QBI.*schema mismatch"): + multispine_pool_module.derive_multispine_pool_inputs(_qbi_ready_derive_frame()) + + +def test_derive_stage_rejects_mutated_qbi_output_with_fresh_receipt( + monkeypatch: pytest.MonkeyPatch, +) -> None: + real_kernel = multispine_pool_module.with_us_qbi_input_reconciliation + + def mutate_kernel(frame: Frame) -> Frame: + result = real_kernel(frame) + result.table("person").loc[0, "qualified_bdc_income"] = 0.25 + return result + + monkeypatch.setattr( + multispine_pool_module, + "with_us_qbi_input_reconciliation", + mutate_kernel, + ) + + with pytest.raises(ValueError, match="deterministic kernel"): + multispine_pool_module.derive_multispine_pool_inputs(_qbi_ready_derive_frame()) + + def test_prior_year_contract_fails_on_raw_clone_and_succeeds_in_clone_stage( monkeypatch: pytest.MonkeyPatch, ) -> None: diff --git a/packages/microcosm-build/tests/test_us_multispine_pool_tool.py b/packages/microcosm-build/tests/test_us_multispine_pool_tool.py index 6b6a7d13..986b01e5 100644 --- a/packages/microcosm-build/tests/test_us_multispine_pool_tool.py +++ b/packages/microcosm-build/tests/test_us_multispine_pool_tool.py @@ -22,6 +22,8 @@ import pytest import microcosm.build.us_runtime.acs_transfer as acs_transfer_module +import microcosm.build.us_runtime.multispine_pool as multispine_pool_module +import microcosm.build.us_runtime.stacked_spine as stacked_spine_module from microcosm.build.gates import GateReport, GateResult from microcosm.build.logbook import LOGBOOK_ROW_FIELDS, load_logbook_row from microcosm.build.serialization_dtypes import CANONICAL_STRING_DTYPE @@ -29,13 +31,23 @@ from microcosm.build.us_runtime.acs_transfer_bank import ( ACS_TRANSFER_TARGET_BANK_MATERIALIZER_VERSION, ) -from microcosm.build.us_runtime.multispine_pool import PoolStageOutput +from microcosm.build.us_runtime.multispine_pool import ( + MultispinePoolCheckpoint, + PoolStageOutput, +) from microcosm.build.us_runtime.operator_boundary import ( PRE_ASSEMBLY_OPERATOR_OUTPUT_FAMILIES, ) from microcosm.build.us_runtime.puf_support import ( PUF_SUPPORT_MAX_CLONE_SAFE_SOURCE_ID, ) +from microcosm.build.us_runtime.qbi_inputs import ( + US_QBI_OUTPUT_COLUMNS, + bind_us_qbi_reconciliation_transition_authority, + us_qbi_reconciliation_change_receipt, + with_us_qbi_input_reconciliation, +) +from microcosm.build.us_runtime.stacked_spine import GapFillDirection from microcosm.build.us_runtime.support_provenance import ( support_channel_column, support_clone_index_column, @@ -46,6 +58,8 @@ ) from microcosm.frame import US_SCHEMA, Frame, WeightKind, Weights +_FIXTURE_SEED_PERSON_COLUMN = "takes_up_medicaid_if_eligible" + @pytest.fixture(scope="module") def pool_tool() -> ModuleType: @@ -130,6 +144,8 @@ def _many_household_source_frame( for entity in US_SCHEMA.group_entities }, } + if measured_offset: + tables["household"]["TYPEHUGQ"] = 1 return Frame( tables, US_SCHEMA, @@ -156,6 +172,49 @@ def _replace_person( ) +def _fixture_qbi_stage_output( + frame: Frame, + receipt: Mapping[str, object], +) -> PoolStageOutput: + """Attach a real, live-frame-bound QBI receipt to a tiny derive fixture.""" + + person = frame.table("person").copy() + if "age" not in person: + person["age"] = pd.to_numeric(person["A_AGE"], errors="raise") + if "SEMP" not in person: + person["SEMP"] = 0.0 + if "self_employment_income_before_lsr" not in person: + person["self_employment_income_before_lsr"] = 0.0 + if "non_qualified_dividend_income" not in person: + person["non_qualified_dividend_income"] = 0.0 + for column in US_QBI_OUTPUT_COLUMNS: + if column not in person: + person[column] = 0.0 + before = _replace_person(frame, person) + after = with_us_qbi_input_reconciliation(before) + qbi_receipt = us_qbi_reconciliation_change_receipt(before, after) + after = bind_us_qbi_reconciliation_transition_authority(after, qbi_receipt) + return PoolStageOutput( + after, + { + **dict(receipt), + "qbi_input_reconciliation": qbi_receipt, + }, + qbi_transition_authority_sha256=qbi_receipt["sha256"], + ) + + +def _with_fixture_pre_clone_strike_benefits(frame: Frame) -> Frame: + """Fixture producer for one real pre-clone operator-owned target.""" + + person = frame.table("person").copy() + assert "strike_benefits" not in person + channel = person[support_channel_column("person")].astype(str) + person["strike_benefits"] = np.nan + person.loc[channel.eq("asec"), "strike_benefits"] = 125.0 + return _replace_person(frame, person) + + def _semantic_string_columns(table: pd.DataFrame) -> tuple[str, ...]: return tuple( column @@ -237,7 +296,12 @@ def _transfer_source_frame(targets: list[float]) -> Frame: ) -def _red_pool_result(pool_tool: ModuleType, tmp_path: Path): +def _red_pool_result( + pool_tool: ModuleType, + tmp_path: Path, + *, + authenticated_qbi: bool = False, +): order: list[str] = [] def stage( @@ -256,9 +320,19 @@ def apply(frame: Frame) -> PoolStageOutput: "acs", } transform(person) + if name == "derive" and authenticated_qbi: + return _fixture_qbi_stage_output( + _replace_person(frame, person), + {"fixture_stage": name}, + ) + receipt: dict[str, object] = {"fixture_stage": name} + if name == "seed" and authenticated_qbi: + receipt["programs"] = { + _FIXTURE_SEED_PERSON_COLUMN: {"entity": "person"}, + } return PoolStageOutput( _replace_person(frame, person), - {"fixture_stage": name}, + receipt, ) return apply @@ -270,7 +344,8 @@ def derive(person: pd.DataFrame) -> None: person["fixture_derived"] = person["fixture_transfer"] + 1.0 def seed(person: pd.DataFrame) -> None: - person["fixture_seed"] = person["fixture_derived"] > 0.0 + column = _FIXTURE_SEED_PERSON_COLUMN if authenticated_qbi else "fixture_seed" + person[column] = person["fixture_derived"] > 0.0 def simulate(person: pd.DataFrame) -> None: channels = person[support_channel_column("person")] @@ -303,8 +378,14 @@ def simulate(person: pd.DataFrame) -> None: def _output_context( pool_tool: ModuleType, tmp_path: Path, + *, + authenticated_qbi: bool = True, ): - result = _red_pool_result(pool_tool, tmp_path) + result = _red_pool_result( + pool_tool, + tmp_path, + authenticated_qbi=authenticated_qbi, + ) outputs = pool_tool._output_paths(tmp_path / "pool.h5") source_manifest = pool_tool.load_acs_source_manifest() verified_inputs = {} @@ -430,6 +511,7 @@ def _run_checkpoint_fixture( resume=None, target_bank_receipt: Mapping[str, object] | None = None, primary_qrf_manifest_path: Path | None = None, + authenticated_qbi: bool = True, ): order: list[str] = [] @@ -468,6 +550,15 @@ def apply(frame: Frame) -> PoolStageOutput: }, "weights_audit": {"passed": True}, } + if name == "derive" and authenticated_qbi: + return _fixture_qbi_stage_output( + _replace_person(frame, person), + receipt, + ) + if name == "seed" and authenticated_qbi: + receipt["programs"] = { + _FIXTURE_SEED_PERSON_COLUMN: {"entity": "person"}, + } return PoolStageOutput(_replace_person(frame, person), receipt) return apply @@ -498,7 +589,7 @@ def apply(frame: Frame) -> PoolStageOutput: seed=stage( "seed", lambda person: person.__setitem__( - "fixture_seed", + (_FIXTURE_SEED_PERSON_COLUMN if authenticated_qbi else "fixture_seed"), person["fixture_derived"] > 0.0, ), ), @@ -531,13 +622,26 @@ def _run_production_impute_checkpoint_fixture( def stage( transform: Callable[[pd.DataFrame], None], + *, + reconcile_qbi: bool = False, + seed_person_output: str | None = None, ) -> Callable[[Frame], PoolStageOutput]: def apply(frame: Frame) -> PoolStageOutput: person = frame.table("person").copy() transform(person) + if reconcile_qbi: + return _fixture_qbi_stage_output( + _replace_person(frame, person), + {"fixture_stage": True}, + ) + receipt: dict[str, object] = {"fixture_stage": True} + if seed_person_output is not None: + receipt["programs"] = { + seed_person_output: {"entity": "person"}, + } return PoolStageOutput( _replace_person(frame, person), - {"fixture_stage": True}, + receipt, ) return apply @@ -554,9 +658,11 @@ def apply(frame: Frame) -> PoolStageOutput: prepare_clone=stage(lambda _person: None), derive=stage( lambda person: person.__setitem__("fixture_derived", 1.0), + reconcile_qbi=True, ), seed=stage( - lambda person: person.__setitem__("fixture_seed", True), + lambda person: person.__setitem__(_FIXTURE_SEED_PERSON_COLUMN, True), + seed_person_output=_FIXTURE_SEED_PERSON_COLUMN, ), simulate=stage( lambda person: person.__setitem__( @@ -780,12 +886,23 @@ def _stacked_main_argv( return arguments +def _noncanonical_post_puf_authority_receipt() -> dict[str, object]: + surface = {"person": {"model_required_boolean": ("is_pregnant",)}} + test_authority = stacked_spine_module._make_test_stacked_authority( + declared_surface=surface, + gap_fill_plan=(), + post_puf_transfer_surface=surface, + ) + return stacked_spine_module._authority_receipt(test_authority) + + def _install_stacked_entrypoint_stubs( pool_tool: ModuleType, monkeypatch: pytest.MonkeyPatch, tmp_path: Path, *, terminal: str, + post_puf_authority: Mapping[str, object] | None = None, ) -> tuple[list[str], int]: order: list[str] = [] verified = _verified_inputs_fixture(pool_tool, tmp_path / "pins") @@ -839,27 +956,55 @@ def build_stacked_pool(*args, **kwargs): monkeypatch.setattr(pool_tool, "build_stacked_pool", build_stacked_pool) - def prepare(frame: Frame, *, acs_rent_donor: pd.DataFrame): + def fixture_pre_clone_source_chain( + frame: Frame, + *, + phase: str, + operator_names: tuple[str, ...], + operators: Mapping[str, object], + **_kwargs: object, + ) -> PoolStageOutput: order.append("prepare") - assert len(acs_rent_donor) == 1 - return PoolStageOutput(frame, {"fixture": "prepare"}) + assert phase == "pre_clone" + assert operator_names + assert set(operator_names) == set(operators) + return PoolStageOutput( + _with_fixture_pre_clone_strike_benefits(frame), + {"fixture": "pre_clone_source_chain"}, + ) monkeypatch.setattr( - pool_tool, - "prepare_multispine_source_inputs_for_clone", - prepare, + multispine_pool_module, + "_run_source_operator_chain", + fixture_pre_clone_source_chain, ) directions = ( - SimpleNamespace(name="asec_survey_to_acs"), - SimpleNamespace(name="acs_housing_to_asec"), + GapFillDirection( + name="asec_survey_to_acs", + recipient_channel="acs", + donor_channel="asec", + target_families={ + "person": { + "source_operator_cps_carried": ("strike_benefits",), + } + }, + ), ) monkeypatch.setattr(pool_tool, "stacked_gap_fill_plan", lambda: directions) def gap_fill(frame: Frame, **kwargs): order.append("gap") - assert set(kwargs["target_banks"]) == { - "asec_survey_to_acs", - "acs_housing_to_asec", + assert set(kwargs["target_banks"]) == {"asec_survey_to_acs"} + counts = stacked_spine_module._verify_gap_fill_activation_authority( + frame, + direction=directions[0], + ) + assert counts == { + ("person", "strike_benefits"): { + "authorized_null_rows": 1, + "recipient_rows": 1, + "donor_rows": 1, + } } return SimpleNamespace( frame=frame, @@ -911,6 +1056,28 @@ def complete(frame: Frame): return PoolStageOutput(frame, {"fixture": "complete"}) monkeypatch.setattr(pool_tool, "complete_multispine_source_inputs", complete) + + def post_puf_transfer(frame: Frame, **kwargs: object): + order.append("post_puf_transfer") + assert kwargs["target_bank"] is not None + return SimpleNamespace( + frame=frame, + receipt={ + "fixture": "post_puf_transfer", + "authority": dict( + pool_tool.stacked_spine_authority_receipt() + if post_puf_authority is None + else post_puf_authority + ), + }, + transfer_result=SimpleNamespace(fit_records=()), + ) + + monkeypatch.setattr( + pool_tool, + "transfer_stacked_post_puf_inputs", + post_puf_transfer, + ) monkeypatch.setattr( pool_tool, "assert_stacked_tail_cells_preserved", @@ -926,6 +1093,8 @@ def tail_prepare(frame: Frame): def identity_stage(name: str): def stage(frame: Frame): order.append(name) + if name == "derive": + return _fixture_qbi_stage_output(frame, {"fixture": name}) return PoolStageOutput(frame, {"fixture": name}) return stage @@ -979,6 +1148,51 @@ def publish(*args, **kwargs): return order, len(puf_donor) +def test_stacked_operator_target_requires_preparation_before_activation_authority( + pool_tool: ModuleType, +) -> None: + stacked = pool_tool.assemble_stacked_spine( + _many_household_source_frame(), + _many_household_source_frame(measured_offset=1_000.0), + sample_fraction=0.01, + sample_seed=578, + ).frame + direction = GapFillDirection( + name="asec_survey_to_acs", + recipient_channel="acs", + donor_channel="asec", + target_families={ + "person": { + "source_operator_cps_carried": ("strike_benefits",), + } + }, + ) + + with pytest.raises( + ValueError, + match=( + r"asec_survey_to_acs/person/source_operator_cps_carried/" + r"strike_benefits: declared gap-fill target column is absent" + ), + ): + stacked_spine_module._verify_gap_fill_activation_authority( + stacked, + direction=direction, + ) + + prepared = _with_fixture_pre_clone_strike_benefits(stacked) + assert stacked_spine_module._verify_gap_fill_activation_authority( + prepared, + direction=direction, + ) == { + ("person", "strike_benefits"): { + "authorized_null_rows": 1, + "recipient_rows": 1, + "donor_rows": 1, + } + } + + @pytest.mark.parametrize( ("terminal", "expected_code", "disposition"), [ @@ -1032,6 +1246,7 @@ def test_stacked_tool_entrypoint_fixture_e2e_emits_one_logbook_row_at_every_term "gap", "puf", "complete", + "post_puf_transfer", "tail_prepare", "derive", "seed", @@ -1057,6 +1272,20 @@ def test_stacked_tool_entrypoint_fixture_e2e_emits_one_logbook_row_at_every_term (tmp_path / "stacked-pool.manifest.json").read_text(encoding="utf-8") ) assert manifest["pipeline"] == "us-stacked-pool" + assert manifest["operator_order"] == [ + "assemble_stacked_spine", + "prepare_multispine_source_inputs_for_clone", + "gap_fill_stacked_spine", + "run_stacked_puf_pass", + "complete_multispine_source_inputs", + "transfer_stacked_post_puf_inputs", + "prepare_stacked_tail_derivation", + "derive_multispine_pool_inputs", + "seed_multispine_pool_inputs", + "materialize_multispine_agreement_outputs", + "stacked_completeness_gate", + "by_origin_battery", + ] assert manifest["sampling"] == { **manifest["sampling"], "sample_fraction": 0.01, @@ -1066,6 +1295,83 @@ def test_stacked_tool_entrypoint_fixture_e2e_emits_one_logbook_row_at_every_term } +def test_stacked_entrypoint_rejects_noncanonical_post_puf_transfer_receipt( + pool_tool: ModuleType, + monkeypatch: pytest.MonkeyPatch, + tmp_path: Path, +) -> None: + pytest.importorskip("tables") + noncanonical = _noncanonical_post_puf_authority_receipt() + assert noncanonical["authority_form"] == "NON-CANONICAL" + assert noncanonical["production_manifest_permitted"] is False + order, _full_puf_rows = _install_stacked_entrypoint_stubs( + pool_tool, + monkeypatch, + tmp_path, + terminal="success", + post_puf_authority=noncanonical, + ) + + with pytest.raises( + ValueError, + match=( + "stacked cold-build post-PUF transfer: non-canonical stacked " + "authority is forbidden" + ), + ): + pool_tool.main(_stacked_main_argv(tmp_path)) + + assert order == [ + "stack", + "build_stacked_pool", + "prepare", + "gap", + "puf", + "complete", + "post_puf_transfer", + ] + assert not (tmp_path / "stacked-pool.h5").exists() + assert not (tmp_path / "stacked-pool.manifest.json").exists() + + +def test_stacked_publication_rejects_noncanonical_receipt_before_any_write( + pool_tool: ModuleType, + tmp_path: Path, +) -> None: + noncanonical = _noncanonical_post_puf_authority_receipt() + outputs = pool_tool._stacked_output_paths(tmp_path / "stacked-pool.h5") + result = SimpleNamespace( + stage_receipts={ + "impute": { + "stacked_post_puf_transfer": {"authority": noncanonical}, + } + } + ) + + with pytest.raises( + ValueError, + match=( + "stacked publication entry: non-canonical stacked authority is forbidden" + ), + ): + pool_tool._write_stacked_outputs( + result, + outputs=outputs, + verified_inputs={}, + acs_source_manifest=pool_tool.load_acs_source_manifest(), + input_receipts={}, + checkpoint_provenance={}, + sample_fraction=0.01, + sample_seed=578, + clone_attachment_fraction=1.0, + clone_attachment_seed=579, + ) + + assert not outputs.pool_h5.exists() + assert not outputs.manifest.exists() + assert not outputs.agreement_diagnostics.exists() + + @pytest.mark.parametrize( "failure", ("negative_seed", "oversized_seed", "invalid_output", "code_pin"), @@ -1290,6 +1596,10 @@ def identity( "clone_seed": identity(clone_seed=579), "stack_manifest": identity(stack=mutated_stack), } + producer_schedule = identities["base"]["pool_code"]["gap_fill_producer_schedule"] + assert producer_schedule["status"] == "all_producers_precede_activation" + assert producer_schedule["direction_count"] == 2 + assert producer_schedule["target_count"] == 48 digests = { name: pool_tool._pool_checkpoint_identity_sha256(value) for name, value in identities.items() @@ -1371,11 +1681,200 @@ def identity( assert changed_store.load_deepest() is None -def test_stacked_materializer_v1_checkpoint_is_not_discovered( +def test_stacked_checkpoint_identity_binds_v6_semantic_contracts( + pool_tool: ModuleType, + monkeypatch: pytest.MonkeyPatch, + tmp_path: Path, + capsys: pytest.CaptureFixture[str], +) -> None: + verified = _verified_inputs_fixture(pool_tool, tmp_path / "pins") + stack = pool_tool.assemble_stacked_spine( + _many_household_source_frame(), + _many_household_source_frame(measured_offset=1_000.0), + sample_fraction=0.10, + sample_seed=578, + ) + + def identity() -> dict[str, object]: + return pool_tool._stacked_checkpoint_base_identity( + verified, + stack_receipt=stack.receipt, + sample_fraction=0.10, + sample_seed=578, + clone_attachment_fraction=1.0, + clone_attachment_seed=578, + policyengine_us_version="fixture-engine", + ) + + current = identity() + pool_code = current["pool_code"] + assert current["materializer_version"] == 6 + assert current["stacked_authority"]["version"] == 6 + assert pool_code["primary_qrf_checkpoint_schema_version"] == 6 + assert pool_code["acs_pums_earnings_universe_contract"] == ( + pool_tool.acs_pums_earnings_universe_contract_identity() + ) + assert pool_code["us_qbi_reconciliation_contract"] == ( + pool_tool.us_qbi_reconciliation_contract_identity() + ) + + with monkeypatch.context() as changed: + changed.setattr(pool_tool, "PRIMARY_QRF_CHECKPOINT_SCHEMA_VERSION", 5) + stale_qrf = identity() + with monkeypatch.context() as changed: + acs_contract = copy.deepcopy( + pool_tool.acs_pums_earnings_universe_contract_identity() + ) + acs_contract["minimum_age"] = 14 + acs_body = dict(acs_contract) + acs_body.pop("sha256") + acs_contract["sha256"] = hashlib.sha256( + pool_tool._canonical_json_bytes(acs_body) + ).hexdigest() + changed.setattr( + pool_tool, + "acs_pums_earnings_universe_contract_identity", + lambda: acs_contract, + ) + stale_acs = identity() + with monkeypatch.context() as changed: + qbi_contract = copy.deepcopy( + pool_tool.us_qbi_reconciliation_contract_identity() + ) + qbi_contract["execution_scope"] = "recipient_subset" + changed.setattr( + pool_tool, + "us_qbi_reconciliation_contract_identity", + lambda: qbi_contract, + ) + stale_qbi = identity() + + digests = { + pool_tool._pool_checkpoint_identity_sha256(candidate) + for candidate in (current, stale_qrf, stale_acs, stale_qbi) + } + assert len(digests) == 4 + + # A checkpoint produced by the current materializer with the prior QRF + # schema is not merely identity-distinct: discovery must refuse it as stale. + checkpoint_root = tmp_path / "mixed-qrf-version-checkpoints" + stale_qrf_store = pool_tool._PoolStageCheckpointStore( + checkpoint_root, + base_identity=stale_qrf, + ) + stale_qrf_store.bind_input_receipts(_checkpoint_fixture_input_receipts()) + stale_qrf_store.write( + pool_tool.MultispinePoolCheckpoint( + stage="assembled", + frame=stack.frame, + assembly_receipt=stack.frame.metadata[ + pool_tool.SPINE_ASSEMBLY_MANIFEST_KEY + ], + stage_receipts={}, + ) + ) + + assert current["materializer_version"] == stale_qrf["materializer_version"] == 6 + assert stale_qrf["pool_code"]["primary_qrf_checkpoint_schema_version"] == 5 + assert ( + pool_tool._discover_stacked_checkpoint_identity( + checkpoint_root, + verified_inputs=verified, + sample_fraction=0.10, + sample_seed=578, + clone_attachment_fraction=1.0, + clone_attachment_seed=578, + ) + is None + ) + assert "checkpoint base identity is stale" in capsys.readouterr().out + + +@pytest.mark.parametrize( + ("route", "stage_receipts"), + ( + ( + "legacy", + {"derive": {"qbi_input_reconciliation": {"fixture": "receipt"}}}, + ), + ( + "stacked", + { + "derive": { + "pool_derivation": { + "qbi_input_reconciliation": {"fixture": "receipt"} + } + } + }, + ), + ), +) +def test_qbi_receipt_route_resolution_is_exact( + pool_tool: ModuleType, + route: str, + stage_receipts: Mapping[str, Mapping[str, object]], +) -> None: + assert pool_tool._qbi_receipt_from_stage_receipts( + stage_receipts, + route=route, + boundary="fixture QBI route", + ) == {"fixture": "receipt"} + + +@pytest.mark.parametrize( + ("route", "stage_receipts", "message"), + ( + ( + "legacy", + { + "derive": { + "pool_derivation": { + "qbi_input_reconciliation": {"fixture": "stacked"} + } + } + }, + "legacy QBI receipt used the stacked derive route", + ), + ( + "stacked", + {"derive": {"qbi_input_reconciliation": {"fixture": "legacy"}}}, + "stacked derive receipts have no pool_derivation object", + ), + ( + "stacked", + { + "derive": { + "qbi_input_reconciliation": {"fixture": "legacy"}, + "pool_derivation": { + "qbi_input_reconciliation": {"fixture": "stacked"} + }, + } + }, + "stacked QBI receipt also appears at the legacy route", + ), + ), +) +def test_qbi_receipt_route_resolution_rejects_wrong_or_ambiguous_paths( + pool_tool: ModuleType, + route: str, + stage_receipts: Mapping[str, Mapping[str, object]], + message: str, +) -> None: + with pytest.raises(ValueError, match=message): + pool_tool._qbi_receipt_from_stage_receipts( + stage_receipts, + route=route, + boundary="fixture QBI route", + ) + + +@pytest.mark.parametrize("legacy_version", (1, 2, 3, 4, 5)) +def test_legacy_stacked_materializer_checkpoint_is_not_discovered( pool_tool: ModuleType, monkeypatch: pytest.MonkeyPatch, tmp_path: Path, capsys: pytest.CaptureFixture[str], + legacy_version: int, ) -> None: verified = _verified_inputs_fixture(pool_tool, tmp_path / "pins") stack = pool_tool.assemble_stacked_spine( @@ -1387,7 +1886,11 @@ def test_stacked_materializer_v1_checkpoint_is_not_discovered( checkpoint_root = tmp_path / "stacked-materializer-checkpoints" with monkeypatch.context() as legacy: - legacy.setattr(pool_tool, "_STACKED_CHECKPOINT_MATERIALIZER_VERSION", 1) + legacy.setattr( + pool_tool, + "_STACKED_CHECKPOINT_MATERIALIZER_VERSION", + legacy_version, + ) legacy_identity = pool_tool._stacked_checkpoint_base_identity( verified, stack_receipt=stack.receipt, @@ -1413,7 +1916,7 @@ def test_stacked_materializer_v1_checkpoint_is_not_discovered( ) ) - assert pool_tool._STACKED_CHECKPOINT_MATERIALIZER_VERSION == 2 + assert pool_tool._STACKED_CHECKPOINT_MATERIALIZER_VERSION == 6 assert ( pool_tool._discover_stacked_checkpoint_identity( checkpoint_root, @@ -1428,6 +1931,53 @@ def test_stacked_materializer_v1_checkpoint_is_not_discovered( assert "checkpoint base identity is stale" in capsys.readouterr().out +def test_stacked_resume_rejects_noncanonical_post_puf_transfer_receipt( + pool_tool: ModuleType, + tmp_path: Path, +) -> None: + stack = pool_tool.assemble_stacked_spine( + _many_household_source_frame(), + _many_household_source_frame(measured_offset=1_000.0), + sample_fraction=0.10, + sample_seed=578, + ) + noncanonical = _noncanonical_post_puf_authority_receipt() + resume = pool_tool.MultispinePoolCheckpoint( + stage="transferred", + frame=stack.frame, + assembly_receipt=stack.frame.metadata[pool_tool.SPINE_ASSEMBLY_MANIFEST_KEY], + stage_receipts={ + "impute": { + "stacked_post_puf_transfer": {"authority": noncanonical}, + } + }, + ) + + with pytest.raises( + ValueError, + match=( + "stacked transferred checkpoint resume: non-canonical stacked " + "authority is forbidden" + ), + ): + pool_tool.build_stacked_pool( + stack.frame, + expected_stack_receipt=stack.receipt, + release_id=( + "populace-us-2024-stacked-f010-s578-asec4-acs1-" + "20260807T000000Z-deadbeef" + ), + puf_donor=None, + acs_rent_donor=None, + primary_qrf_checkpoint_dir=tmp_path / "primary-qrf", + acs_transfer_checkpoint_dir=tmp_path / "acs-transfer", + checkpoint_identity={}, + clone_attachment_fraction=1.0, + clone_attachment_seed=578, + resume=resume, + ) + + def test_stacked_entrypoint_resumes_each_checkpoint_boundary( pool_tool: ModuleType, monkeypatch: pytest.MonkeyPatch, @@ -1492,6 +2042,7 @@ def test_stacked_entrypoint_resumes_each_checkpoint_boundary( "gap", "puf", "complete", + "post_puf_transfer", "tail_prepare", "derive", "seed", @@ -1522,6 +2073,7 @@ def test_stacked_entrypoint_resumes_each_checkpoint_boundary( "gap", "puf", "complete", + "post_puf_transfer", "tail_prepare", "derive", "seed", @@ -1608,14 +2160,31 @@ def test_stacked_release_id_carries_rung_seed_and_realized_counts( def test_legacy_two_spine_fixture_is_origin_main_byte_exact( pool_tool: ModuleType, + monkeypatch: pytest.MonkeyPatch, tmp_path: Path, ) -> None: + # Keep this byte-level legacy-output golden on its pre-authentication + # synthetic fixture. QBI authentication has independent generation, + # checkpoint, resume, manifest, and publication tamper tests. + monkeypatch.setattr( + multispine_pool_module, + "_validate_qbi_stage_receipt", + lambda _frame, _stage_receipts, *, boundary, transition_authority_sha256: None, + ) + monkeypatch.setattr( + pool_tool, + "_validate_qbi_stage_receipt", + lambda _frame, _stage_receipts, *, route, boundary, transition_authority_sha256: ( + None + ), + ) store = _checkpoint_fixture_store(pool_tool, tmp_path / "checkpoints") store.bind_input_receipts(_checkpoint_fixture_input_receipts()) result, order = _run_checkpoint_fixture( pool_tool, tmp_path, store=store, + authenticated_qbi=False, ) artifact = tmp_path / "legacy-origin-main.canonical.h5" pool_tool.write_frame_checkpoint(artifact, result.frame) @@ -1639,6 +2208,7 @@ def test_legacy_entrypoint_publication_matches_origin_main_golden( result, _unused_outputs, verified, source_manifest, loaded = _output_context( pool_tool, tmp_path, + authenticated_qbi=False, ) ready = replace( result, @@ -1692,6 +2262,16 @@ def deterministic_fixture_h5( "write_nullable_us_h5", deterministic_fixture_h5, ) + # This golden deliberately preserves the pre-authentication synthetic + # output byte-for-byte. Dedicated QBI boundary tests below exercise the + # authenticated production contract against real frame-bound receipts. + monkeypatch.setattr( + pool_tool, + "_validate_qbi_stage_receipt", + lambda _frame, _stage_receipts, *, route, boundary, transition_authority_sha256: ( + None + ), + ) output = tmp_path / "legacy-pool.h5" checkpoint_root = tmp_path / "checkpoints" @@ -3445,6 +4025,110 @@ def test_invalid_deep_checkpoint_receipts_do_not_poison_valid_fallback( assert "invalid SSI output binding" in output +def test_durable_checkpoint_write_rejects_forged_qbi_receipt( + pool_tool: ModuleType, + tmp_path: Path, +) -> None: + checkpoint_root = tmp_path / "checkpoints" + store = _checkpoint_fixture_store(pool_tool, checkpoint_root) + store.bind_input_receipts(_checkpoint_fixture_input_receipts()) + + def forge_after_emission(checkpoint: MultispinePoolCheckpoint) -> None: + if checkpoint.stage != "simulated": + store.write(checkpoint) + return + receipts = copy.deepcopy(checkpoint.stage_receipts) + receipts["derive"]["qbi_input_reconciliation"]["sha256"] = "0" * 64 + store.write(replace(checkpoint, stage_receipts=receipts)) + + with pytest.raises( + ValueError, + match=( + "pool simulated durable checkpoint write: QBI reconciliation " + "receipt SHA-256" + ), + ): + _run_checkpoint_fixture( + pool_tool, + tmp_path, + store=SimpleNamespace(write=forge_after_emission), + ) + + assert not store.checkpoint_path("simulated").exists() + + +def test_durable_checkpoint_write_rejects_forged_qbi_transition_authority( + pool_tool: ModuleType, + tmp_path: Path, +) -> None: + checkpoint_root = tmp_path / "checkpoints" + store = _checkpoint_fixture_store(pool_tool, checkpoint_root) + store.bind_input_receipts(_checkpoint_fixture_input_receipts()) + + def forge_after_emission(checkpoint: MultispinePoolCheckpoint) -> None: + if checkpoint.stage != "simulated": + store.write(checkpoint) + return + store.write( + replace( + checkpoint, + qbi_transition_authority_sha256="0" * 64, + ) + ) + + with pytest.raises( + ValueError, + match="independently carried transition authority", + ): + _run_checkpoint_fixture( + pool_tool, + tmp_path, + store=SimpleNamespace(write=forge_after_emission), + ) + + assert not store.checkpoint_path("simulated").exists() + + +def test_durable_checkpoint_load_rejects_forged_qbi_receipt_and_falls_back( + pool_tool: ModuleType, + tmp_path: Path, + capsys: pytest.CaptureFixture[str], +) -> None: + checkpoint_root = tmp_path / "checkpoints" + cold_store = _checkpoint_fixture_store(pool_tool, checkpoint_root) + cold_store.bind_input_receipts(_checkpoint_fixture_input_receipts()) + _run_checkpoint_fixture(pool_tool, tmp_path, store=cold_store) + capsys.readouterr() + + simulated_path = cold_store.checkpoint_path("simulated") + simulated = pool_tool.load_frame_checkpoint(simulated_path) + poisoned_metadata = copy.deepcopy(simulated.metadata) + poisoned_metadata["stage_receipts"]["derive"]["qbi_input_reconciliation"][ + "sha256" + ] = "0" * 64 + pool_tool.write_frame_checkpoint( + simulated_path, + simulated.frame, + metadata=poisoned_metadata, + ) + simulated_manifest_path = cold_store.checkpoint_manifest_path("simulated") + simulated_manifest = pool_tool._read_json_object(simulated_manifest_path) + simulated_manifest["checkpoint"]["sha256"] = pool_tool._file_sha256(simulated_path) + simulated_manifest["checkpoint"]["size_bytes"] = simulated_path.stat().st_size + pool_tool._atomic_write_json(simulated_manifest_path, simulated_manifest) + + warm_store = _checkpoint_fixture_store(pool_tool, checkpoint_root) + resume = warm_store.load_deepest() + + assert resume is not None + assert resume.stage == "transferred" + output = capsys.readouterr().out + assert "Ignored corrupt pool checkpoint 'simulated'" in output + assert ( + "pool simulated durable checkpoint load: QBI reconciliation receipt SHA-256" + ) in output + + def test_synthetic_two_spine_path_reaches_fixed_red_terminal_gate( pool_tool: ModuleType, tmp_path: Path, @@ -3540,6 +4224,172 @@ def simulate(frame: Frame) -> PoolStageOutput: assert "ssi" not in person +def test_manifest_and_publication_reject_forged_qbi_receipt( + pool_tool: ModuleType, + tmp_path: Path, +) -> None: + result, outputs, verified_inputs, source_manifest, loaded = _output_context( + pool_tool, + tmp_path, + ) + receipts = copy.deepcopy(result.stage_receipts) + receipts["derive"]["qbi_input_reconciliation"]["sha256"] = "0" * 64 + forged = replace(result, stage_receipts=receipts) + + with pytest.raises( + ValueError, + match=("legacy production manifest: QBI reconciliation receipt SHA-256"), + ): + pool_tool._manifest_payload( + result=forged, + outputs=outputs, + verified_inputs=verified_inputs, + acs_source_manifest=source_manifest, + input_receipts={}, + checkpoint_provenance={}, + publication_run_id="forged-manifest", + ) + + with pytest.raises( + ValueError, + match="legacy publication entry: QBI reconciliation receipt SHA-256", + ): + pool_tool._write_outputs( + forged, + outputs=outputs, + verified_inputs=verified_inputs, + acs_source_manifest=source_manifest, + loaded=loaded, + ) + + assert not outputs.pool_h5.exists() + assert not outputs.manifest.exists() + assert not outputs.agreement_diagnostics.exists() + + +def test_stacked_manifest_and_publication_reject_forged_qbi_receipt( + pool_tool: ModuleType, + tmp_path: Path, +) -> None: + legacy, _outputs, verified_inputs, source_manifest, _loaded = _output_context( + pool_tool, + tmp_path, + ) + receipt = copy.deepcopy(legacy.stage_receipts["derive"]["qbi_input_reconciliation"]) + receipt["sha256"] = "0" * 64 + stacked = SimpleNamespace( + frame=legacy.frame, + qbi_transition_authority_sha256=(legacy.qbi_transition_authority_sha256), + stage_receipts={ + "impute": { + "stacked_post_puf_transfer": { + "authority": dict(pool_tool.stacked_spine_authority_receipt()) + } + }, + "derive": {"pool_derivation": {"qbi_input_reconciliation": receipt}}, + }, + ) + outputs = pool_tool._stacked_output_paths(tmp_path / "stacked-pool.h5") + + with pytest.raises( + ValueError, + match=("stacked production manifest: QBI reconciliation receipt SHA-256"), + ): + pool_tool._stacked_manifest_payload( + result=stacked, + outputs=outputs, + verified_inputs=verified_inputs, + acs_source_manifest=source_manifest, + input_receipts={}, + checkpoint_provenance={}, + publication_run_id="forged-stacked-manifest", + sample_fraction=0.01, + sample_seed=578, + clone_attachment_fraction=1.0, + clone_attachment_seed=579, + ) + + with pytest.raises( + ValueError, + match="stacked publication entry: QBI reconciliation receipt SHA-256", + ): + pool_tool._write_stacked_outputs( + stacked, + outputs=outputs, + verified_inputs=verified_inputs, + acs_source_manifest=source_manifest, + input_receipts={}, + checkpoint_provenance={}, + sample_fraction=0.01, + sample_seed=578, + clone_attachment_fraction=1.0, + clone_attachment_seed=579, + ) + + assert not outputs.pool_h5.exists() + assert not outputs.manifest.exists() + assert not outputs.agreement_diagnostics.exists() + + +def test_publication_rejects_mutated_qbi_output_with_regenerated_receipt( + pool_tool: ModuleType, + tmp_path: Path, +) -> None: + result, outputs, verified_inputs, source_manifest, loaded = _output_context( + pool_tool, + tmp_path, + ) + person = result.frame.table("person").copy() + person.loc[person.index[0], "non_qualified_dividend_income"] = 100.0 + person.loc[person.index[0], "qualified_bdc_income"] = 50.0 + mutated_frame = _replace_person(result.frame, person) + receipts = copy.deepcopy(result.stage_receipts) + receipts["derive"]["qbi_input_reconciliation"] = ( + us_qbi_reconciliation_change_receipt(mutated_frame, mutated_frame) + ) + forged = replace( + result, + frame=mutated_frame, + stage_receipts=receipts, + ) + + with pytest.raises( + ValueError, + match=( + "legacy production manifest: QBI receipt differs from the " + "independently carried transition authority" + ), + ): + pool_tool._manifest_payload( + result=forged, + outputs=outputs, + verified_inputs=verified_inputs, + acs_source_manifest=source_manifest, + input_receipts={}, + checkpoint_provenance={}, + publication_run_id="reissued-qbi-manifest", + ) + + with pytest.raises( + ValueError, + match=( + "legacy publication entry: QBI receipt differs from the " + "independently carried transition authority" + ), + ): + pool_tool._write_outputs( + forged, + outputs=outputs, + verified_inputs=verified_inputs, + acs_source_manifest=source_manifest, + loaded=loaded, + ) + + assert not outputs.pool_h5.exists() + assert not outputs.manifest.exists() + assert not outputs.agreement_diagnostics.exists() + + def test_red_outputs_preserve_receipts_and_exclude_simulation_output( pool_tool: ModuleType, tmp_path: Path, diff --git a/packages/microcosm-build/tests/test_us_puf_support.py b/packages/microcosm-build/tests/test_us_puf_support.py index c14e3acc..f4c8375b 100644 --- a/packages/microcosm-build/tests/test_us_puf_support.py +++ b/packages/microcosm-build/tests/test_us_puf_support.py @@ -115,6 +115,7 @@ def _minimal_us_frame() -> Frame: "person_spm_unit_id": np.asarray([100, 100, 200], dtype="int64"), "person_family_id": np.asarray([1000, 1000, 2000], dtype="int64"), "person_marital_unit_id": np.asarray([10000, 10000, 20000], dtype="int64"), + "age": np.asarray([42, 40, 51], dtype="int64"), "employment_income_before_lsr": np.asarray( [50_000, 20_000, 125_000], dtype="int64", diff --git a/packages/microcosm-build/tests/test_us_puf_support_base_builder.py b/packages/microcosm-build/tests/test_us_puf_support_base_builder.py index 7e5f74cd..47c67c04 100644 --- a/packages/microcosm-build/tests/test_us_puf_support_base_builder.py +++ b/packages/microcosm-build/tests/test_us_puf_support_base_builder.py @@ -250,6 +250,14 @@ def _support_donor() -> pd.DataFrame: ) +def _minimal_us_puf_support_frame() -> Frame: + return _with_person_column( + _minimal_us_frame(), + "age", + np.asarray([42, 40, 51], dtype="int64"), + ) + + def test_equivalence_h5_metadata_mode_disables_all_leaf_timestamps( tmp_path: Path, ) -> None: @@ -300,7 +308,7 @@ def test_base_build_records_design_weight_kind_in_the_summary(self) -> None: builder = _load_support_builder_module() _imputed, weights_audit = builder.impute_and_audit_us_puf_support( - clone_us_frame_for_puf_support(_minimal_us_frame()), + clone_us_frame_for_puf_support(_minimal_us_puf_support_frame()), _support_donor(), **_SUPPORT_FIT_KWARGS, ) @@ -316,7 +324,7 @@ def test_base_build_summary_json_carries_the_audit(self) -> None: builder = _load_support_builder_module() _imputed, weights_audit = builder.impute_and_audit_us_puf_support( - clone_us_frame_for_puf_support(_minimal_us_frame()), + clone_us_frame_for_puf_support(_minimal_us_puf_support_frame()), _support_donor(), **_SUPPORT_FIT_KWARGS, ) diff --git a/packages/microcosm-build/tests/test_us_qbi_inputs.py b/packages/microcosm-build/tests/test_us_qbi_inputs.py index a606016c..06b0ddd8 100644 --- a/packages/microcosm-build/tests/test_us_qbi_inputs.py +++ b/packages/microcosm-build/tests/test_us_qbi_inputs.py @@ -2,6 +2,7 @@ from __future__ import annotations +import copy import importlib.util import numpy as np @@ -9,6 +10,7 @@ import pytest from microcosm.build.us_runtime import puf_support as puf_support_module +from microcosm.build.us_runtime import qbi_inputs as qbi_inputs_module from microcosm.build.us_runtime.qbi_inputs import ( QBI_ARCHIVED_ASSUMPTIONS_URL, QBI_ARCHIVED_CLONE_URL, @@ -23,6 +25,10 @@ us_qbi_inputs_signal_gate, us_qbi_inputs_stage_spec, us_qbi_inputs_summary, + us_qbi_reconciliation_change_receipt, + us_qbi_reconciliation_universe_receipt, + validate_us_qbi_reconciliation_live_output, + validate_us_qbi_reconciliation_receipt, with_us_qbi_input_reconciliation, ) from microcosm.frame import US_SCHEMA, Frame, WeightKind, Weights @@ -223,6 +229,369 @@ def test_reconciliation_restores_sstb_routes_total_pools_and_exposure_caps() -> assert reconciled.loc[100, "qualified_reit_and_ptp_income"] == 1_000.0 +def _stacked_qbi_universe_frame(*, child_age: float = 12.0) -> Frame: + person = _qbi_person(20) + person["person_support_channel"] = np.where( + np.arange(len(person)) < 10, + "asec", + "acs", + ) + person["person_support_clone_index"] = 0 + person["person_spine_source_id"] = np.arange(len(person), dtype=np.int64) + person["age"] = 40.0 + person["SEMP"] = person["self_employment_income_before_lsr"] + child = 10 + person.loc[child, "age"] = child_age + person.loc[child, "self_employment_income_before_lsr"] = 0.0 + person.loc[child, "SEMP"] = np.nan + person.loc[child, "business_is_sstb"] = True + person.loc[child, "qualified_bdc_income"] = 321.0 + return _frame(person) + + +def test_reconciliation_preserves_child_universe_zero_and_other_qbi_cells() -> None: + frame = _stacked_qbi_universe_frame() + + result = with_us_qbi_input_reconciliation(frame) + after = result.table("person").loc[10] + receipt = us_qbi_reconciliation_universe_receipt(frame) + + assert after["self_employment_income_before_lsr"] == 0.0 + assert after["business_is_sstb"] == 0 + assert after["sstb_self_employment_income_before_lsr"] == 0.0 + assert after["sstb_w2_wages_from_qualified_business"] == 0.0 + assert receipt["rows_excluded_from_base_self_employment_rewrite"] == 1 + assert receipt["rows_included_in_other_qbi_reconciliation"] == 20 + assert ( + receipt["rules"]["self_employment_income_before_lsr"]["source_column"] == "SEMP" + ) + assert receipt["structurally_absent_base_source_cells_mutated"] is False + assert len(receipt["sha256"]) == 64 + + +def test_reconciliation_rejects_native_acs_in_universe_null() -> None: + frame = _stacked_qbi_universe_frame(child_age=15.0) + + with pytest.raises( + ValueError, + match="acs_2024_pums_semp_age_15_plus.*raw_in_universe_null_rows=1", + ): + with_us_qbi_input_reconciliation(frame) + + +def test_change_receipt_binds_read_only_qbi_input_columns() -> None: + baseline = _stacked_qbi_universe_frame() + changed = _stacked_qbi_universe_frame() + changed.table("person").loc[0, "partnership_income"] = -500.0 + + baseline_output = with_us_qbi_input_reconciliation(baseline) + changed_output = with_us_qbi_input_reconciliation(changed) + baseline_receipt = us_qbi_reconciliation_change_receipt(baseline, baseline_output) + changed_receipt = us_qbi_reconciliation_change_receipt(changed, changed_output) + + pd.testing.assert_frame_equal( + baseline_output.table("person").loc[:, US_QBI_OUTPUT_COLUMNS], + changed_output.table("person").loc[:, US_QBI_OUTPUT_COLUMNS], + ) + assert ( + baseline_receipt["output_declared_person_values_sha256"] + == changed_receipt["output_declared_person_values_sha256"] + ) + assert ( + baseline_receipt["input_person_table_sha256"] + != changed_receipt["input_person_table_sha256"] + ) + assert baseline_receipt["sha256"] != changed_receipt["sha256"] + + +def test_change_receipt_rejects_undeclared_output_mutation() -> None: + frame = _stacked_qbi_universe_frame() + reconciled = with_us_qbi_input_reconciliation(frame) + reconciled.table("person").loc[0, "partnership_income"] = 999.0 + + with pytest.raises(ValueError, match="undeclared person column"): + us_qbi_reconciliation_change_receipt(frame, reconciled) + + +def test_change_receipt_rejects_structural_universe_zero_mutation() -> None: + frame = _stacked_qbi_universe_frame() + person = frame.table("person") + person["self_employment_income_before_lsr"] = person[ + "self_employment_income_before_lsr" + ].astype("Float64") + reconciled = with_us_qbi_input_reconciliation(frame) + reconciled.table("person").loc[10, "self_employment_income_before_lsr"] = 1.0 + + with pytest.raises(ValueError, match="deterministic kernel"): + us_qbi_reconciliation_change_receipt(frame, reconciled) + + +@pytest.mark.parametrize("tamper", ["sha256", "negative_count", "extra_field"]) +def test_change_receipt_rejects_forged_envelopes(tamper: str) -> None: + frame = _stacked_qbi_universe_frame() + reconciled = with_us_qbi_input_reconciliation(frame) + receipt = us_qbi_reconciliation_change_receipt(frame, reconciled) + forged = copy.deepcopy(receipt) + if tamper == "sha256": + forged["sha256"] = "0" * 64 + elif tamper == "negative_count": + forged["changed_person_rows"] = -1 + else: + forged["tampered"] = True + + with pytest.raises(ValueError, match="QBI"): + validate_us_qbi_reconciliation_receipt( + forged, + boundary="forged QBI receipt test", + ) + + +def test_change_receipt_rejects_mutated_output_even_when_regenerated() -> None: + frame = _stacked_qbi_universe_frame() + reconciled = with_us_qbi_input_reconciliation(frame) + reconciled.table("person").loc[0, "qualified_bdc_income"] = 0.25 + + with pytest.raises(ValueError, match="deterministic kernel"): + us_qbi_reconciliation_change_receipt(frame, reconciled) + + +def _authorized_qbi_output(frame: Frame) -> tuple[Frame, dict[str, object]]: + reconciled = with_us_qbi_input_reconciliation(frame) + receipt = us_qbi_reconciliation_change_receipt(frame, reconciled) + authorized = qbi_inputs_module.bind_us_qbi_reconciliation_transition_authority( + reconciled, + receipt, + ) + return authorized, receipt + + +def test_live_output_validation_rejects_post_receipt_mutation() -> None: + frame = _stacked_qbi_universe_frame() + reconciled, receipt = _authorized_qbi_output(frame) + reconciled.table("person").loc[0, "qualified_bdc_income"] = 0.25 + + with pytest.raises(ValueError, match="output digest"): + validate_us_qbi_reconciliation_live_output( + reconciled, + receipt, + boundary="mutated live QBI output test", + expected_transition_authority_sha256=receipt["sha256"], + ) + + +def _rehash_forged_stacked_universe_receipt( + receipt: dict[str, object], +) -> None: + universe = receipt["recipient_source_universe"] + assert isinstance(universe, dict) + source_receipt = { + key: value + for key, value in universe.items() + if key + not in { + "source_universe_sha256", + "source_universe_resolution_mutated_raw_pums_cells", + "operation", + "rows_excluded_from_base_self_employment_rewrite", + "rows_included_in_other_qbi_reconciliation", + "structurally_absent_base_source_cells_mutated", + "sha256", + } + } + source_receipt["raw_pums_source_cells_mutated"] = universe[ + "source_universe_resolution_mutated_raw_pums_cells" + ] + universe["source_universe_sha256"] = qbi_inputs_module._qbi_receipt_sha256( + source_receipt + ) + unsigned_universe = dict(universe) + unsigned_universe.pop("sha256") + universe["sha256"] = qbi_inputs_module._qbi_receipt_sha256(unsigned_universe) + unsigned_receipt = dict(receipt) + unsigned_receipt.pop("sha256") + receipt["sha256"] = qbi_inputs_module._qbi_receipt_sha256(unsigned_receipt) + + +def test_live_output_validation_rejects_rehashed_universe_semantic_forgery() -> None: + frame = _stacked_qbi_universe_frame() + reconciled, receipt = _authorized_qbi_output(frame) + forged = copy.deepcopy(receipt) + universe = forged["recipient_source_universe"] + assert isinstance(universe, dict) + universe["scoped_person_rows"] += 1 + _rehash_forged_stacked_universe_receipt(forged) + + with pytest.raises(ValueError, match="does not exactly match the live frame"): + validate_us_qbi_reconciliation_live_output( + reconciled, + forged, + boundary="rehashed QBI universe forgery test", + expected_transition_authority_sha256=receipt["sha256"], + ) + + +def test_live_output_validation_rejects_rebound_driver_and_fixed_point_output() -> None: + frame = _stacked_qbi_universe_frame() + reconciled, receipt = _authorized_qbi_output(frame) + person = reconciled.table("person") + person.loc[0, "non_qualified_dividend_income"] = 100.0 + person.loc[0, "qualified_bdc_income"] = 50.0 + forged = copy.deepcopy(receipt) + forged["output_declared_person_values_sha256"] = ( + qbi_inputs_module._qbi_person_values_sha256( + person, + columns=qbi_inputs_module.US_QBI_RECONCILED_PERSON_COLUMNS, + ) + ) + unsigned = dict(forged) + unsigned.pop("sha256") + forged["sha256"] = qbi_inputs_module._qbi_receipt_sha256(unsigned) + + with pytest.raises(ValueError, match="driver-surface digest"): + validate_us_qbi_reconciliation_live_output( + reconciled, + forged, + boundary="rebound QBI driver/output test", + expected_transition_authority_sha256=receipt["sha256"], + ) + + +def test_live_output_validation_rejects_fresh_receipt_for_alternate_fixed_point() -> ( + None +): + frame = _stacked_qbi_universe_frame() + reconciled, receipt = _authorized_qbi_output(frame) + person = reconciled.table("person") + person.loc[0, "non_qualified_dividend_income"] = 100.0 + person.loc[0, "qualified_bdc_income"] = 50.0 + fresh = us_qbi_reconciliation_change_receipt(reconciled, reconciled) + + with pytest.raises(ValueError, match="independently carried transition authority"): + validate_us_qbi_reconciliation_live_output( + reconciled, + fresh, + boundary="alternate QBI fixed-point receipt test", + expected_transition_authority_sha256=receipt["sha256"], + ) + + +def test_qbi_universe_receipt_rejects_boolean_count() -> None: + frame = _stacked_qbi_universe_frame() + reconciled = with_us_qbi_input_reconciliation(frame) + receipt = us_qbi_reconciliation_change_receipt(frame, reconciled) + forged = copy.deepcopy(receipt) + universe = forged["recipient_source_universe"] + assert isinstance(universe, dict) + universe["scoped_person_rows"] = True + + with pytest.raises(ValueError, match="scoped_person_rows.*nonnegative integer"): + validate_us_qbi_reconciliation_receipt( + forged, + boundary="Boolean QBI count test", + ) + + +def test_live_output_validation_rejects_undeclared_added_person_column() -> None: + frame = _stacked_qbi_universe_frame() + reconciled, receipt = _authorized_qbi_output(frame) + reconciled.table("person")["post_receipt_undeclared_tamper"] = 1 + + with pytest.raises(ValueError, match="person-column inventory changed"): + validate_us_qbi_reconciliation_live_output( + reconciled, + receipt, + boundary="added QBI person-column test", + expected_transition_authority_sha256=receipt["sha256"], + ) + + +def test_live_output_validation_rejects_receipt_laundered_added_column() -> None: + frame = _stacked_qbi_universe_frame() + reconciled, receipt = _authorized_qbi_output(frame) + added = "post_receipt_undeclared_tamper" + person = reconciled.table("person") + person[added] = 1 + forged = copy.deepcopy(receipt) + forged["input_person_columns"].append(added) + forged["input_person_table_sha256"] = qbi_inputs_module._qbi_person_values_sha256( + person, + columns=tuple(forged["input_person_columns"]), + ) + preservation = forged["undeclared_surface_preservation"] + assert isinstance(preservation, dict) + preservation["undeclared_person_columns_verified"] += 1 + unsigned = dict(forged) + unsigned.pop("sha256") + forged["sha256"] = qbi_inputs_module._qbi_receipt_sha256(unsigned) + + with pytest.raises(ValueError, match="independently carried transition authority"): + validate_us_qbi_reconciliation_live_output( + reconciled, + forged, + boundary="laundered QBI person-column test", + expected_transition_authority_sha256=receipt["sha256"], + ) + + +def test_live_output_validation_rejects_rehashed_preservation_count() -> None: + frame = _stacked_qbi_universe_frame() + reconciled, receipt = _authorized_qbi_output(frame) + forged = copy.deepcopy(receipt) + preservation = forged["undeclared_surface_preservation"] + assert isinstance(preservation, dict) + preservation["undeclared_person_columns_verified"] += 1 + unsigned = dict(forged) + unsigned.pop("sha256") + forged["sha256"] = qbi_inputs_module._qbi_receipt_sha256(unsigned) + + with pytest.raises(ValueError, match="preservation receipt does not match"): + validate_us_qbi_reconciliation_live_output( + reconciled, + forged, + boundary="rehashed QBI preservation-count test", + expected_transition_authority_sha256=receipt["sha256"], + ) + + +def test_live_output_validation_allows_only_canonical_seed_person_output() -> None: + frame = _stacked_qbi_universe_frame() + reconciled, receipt = _authorized_qbi_output(frame) + seed_column = "takes_up_medicaid_if_eligible" + reconciled.table("person")[seed_column] = False + + assert ( + qbi_inputs_module.validate_us_qbi_reconciliation_live_output( + reconciled, + receipt, + boundary="canonical post-QBI seed output test", + expected_transition_authority_sha256=receipt["sha256"], + allowed_post_reconciliation_person_columns=(seed_column,), + ) + == receipt + ) + with pytest.raises(ValueError, match="allowed post-QBI person columns"): + qbi_inputs_module.validate_us_qbi_reconciliation_live_output( + reconciled, + receipt, + boundary="non-contract post-QBI seed output test", + expected_transition_authority_sha256=receipt["sha256"], + allowed_post_reconciliation_person_columns=( + "post_receipt_undeclared_tamper", + ), + ) + + +def test_qbi_seed_receipt_rejects_non_contract_person_output() -> None: + with pytest.raises(ValueError, match="non-contract program"): + qbi_inputs_module.us_qbi_post_reconciliation_person_columns( + { + "programs": { + "post_receipt_undeclared_tamper": {"entity": "person"}, + } + } + ) + + def test_sstb_requires_a_positive_qualified_mapped_source() -> None: person = _qbi_person(200) row = 100 @@ -312,7 +681,11 @@ def test_summary_does_not_mutate_source_frame() -> None: summary = us_qbi_inputs_summary(reconciled) - assert set(summary) == {"columns", "invariants"} + assert set(summary) == { + "columns", + "invariants", + "reconciliation_universe", + } pd.testing.assert_frame_equal(before, reconciled.table("person")) diff --git a/packages/microcosm-build/tests/test_us_spine_blindness.py b/packages/microcosm-build/tests/test_us_spine_blindness.py index beee3865..f4237030 100644 --- a/packages/microcosm-build/tests/test_us_spine_blindness.py +++ b/packages/microcosm-build/tests/test_us_spine_blindness.py @@ -89,6 +89,8 @@ # reviewed contract change. _SOURCE_SPINE_PROVENANCE_OWNERS = frozenset( { + # Declares and receipts exact ACS source universes; never mutates rows. + "acs_income_universe.py", "base_pool.py", # Legacy late-spine assembly. # Enumerates provenance columns only to reject preassembled source frames. "operator_boundary.py", @@ -172,6 +174,8 @@ _OTHER_US_RUNTIME_MODULES = frozenset( { "__init__.py", + # Exact source-universe validator/receipt owner; no population treatment. + "acs_income_universe.py", "acs_inputs.py", "acs_multispine.py", "acs_pums.py", @@ -3261,8 +3265,8 @@ def test_pool_build_tool_import_graph_is_source_spine_blind() -> None: for tool in _SPINE_BLIND_BUILD_TOOLS: runtime_graph, missing_modules = _us_runtime_import_graph(tool) - assert len(runtime_graph) == 60, ( - f"{tool.name} must reach the pinned 60-module runtime graph; " + assert len(runtime_graph) == 61, ( + f"{tool.name} must reach the pinned 61-module runtime graph; " f"reached {len(runtime_graph)}" ) assert not missing_modules, ( diff --git a/packages/microcosm-build/tests/test_us_stacked_spine.py b/packages/microcosm-build/tests/test_us_stacked_spine.py index 0a81ba03..3dd15290 100644 --- a/packages/microcosm-build/tests/test_us_stacked_spine.py +++ b/packages/microcosm-build/tests/test_us_stacked_spine.py @@ -16,18 +16,28 @@ from collections import Counter from copy import deepcopy from dataclasses import FrozenInstanceError, replace +from pathlib import Path import numpy as np import pandas as pd import pytest from pandas.testing import assert_frame_equal +import microcosm.build.us_runtime.acs_income_universe as universe_module import microcosm.build.us_runtime.puf_support as puf_support_module import microcosm.build.us_runtime.stacked_spine as stacked_spine_module +from microcosm.build.frame_checkpoint import ( + load_frame_checkpoint, + write_frame_checkpoint, +) from microcosm.build.gates import GateReport, GateResult from microcosm.build.serialization_dtypes import CANONICAL_STRING_DTYPE +from microcosm.build.us_runtime.acs_income_universe import ( + apply_acs_pums_earnings_universe_zeros, +) from microcosm.build.us_runtime.acs_transfer_bank import AcsTransferTargetBankStore from microcosm.build.us_runtime.multispine_pool import ( + derive_multispine_pool_inputs, pool_transfer_target_families, ) from microcosm.build.us_runtime.puf_capital_gains_tail import ( @@ -43,6 +53,10 @@ prepare_us_puf_tax_detail_chain_inputs, validate_puf_clone_attachment, ) +from microcosm.build.us_runtime.qbi_inputs import ( + US_QBI_BOOLEAN_OUTPUT_COLUMNS, + US_QBI_OUTPUT_COLUMNS, +) from microcosm.build.us_runtime.spine_assembly import assemble_spines from microcosm.build.us_runtime.stacked_spine import ( DEFAULT_STACKED_HOUSEHOLD_MASS_SHARES, @@ -50,6 +64,7 @@ STACKED_PILOT_ACS_SAMPLE_SEED, STACKED_SPINE_MANIFEST_KEY, AbsenceProof, + GapFillAbsenceRule, GapFillDirection, OriginBatterySpec, assemble_stacked_spine, @@ -59,6 +74,8 @@ sample_acs_households, stacked_completeness_gate, stacked_gap_fill_plan, + stacked_gap_fill_producer_schedule_receipt, + transfer_stacked_post_puf_inputs, validate_stacked_spine_frame, ) from microcosm.build.us_runtime.support_provenance import ( @@ -148,8 +165,12 @@ def _acs_source() -> Frame: household_ids=list(range(101, 111)), persons_per_household={103: 3, 107: 2}, weights=[float(10 * position) for position in range(1, 11)], - extra_person_columns={"acs_native_aggregate": 15.0}, - extra_household_columns={"puma": "0600101"}, + extra_person_columns={"acs_native_aggregate": 15.0, "WAGP": 15.0}, + extra_household_columns={ + "puma": "0600101", + "TYPEHUGQ": 1, + "tenure_type": "RENTED", + }, stratum="acs_2024_1yr", ) @@ -232,7 +253,7 @@ def test_production_sampling_is_uniform_and_composition_preserving() -> None: sample_seed=578, ) - assert result.receipt["version"] == 2 + assert result.receipt["version"] == 4 assert result.receipt["sample_fraction"] == 0.5 assert result.receipt["sample_seed"] == 578 samples = result.receipt["survey_samples"] @@ -779,6 +800,445 @@ def test_strict_recipient_predictors_fail_closed_on_absence() -> None: assert zero_filled.eq(0.0).all() +def _cloned_acs_earnings_universe_fixture() -> Frame: + """Apply receipted zeros to multi-child mixed and all-child ACS units.""" + + cloned = _cloned_stacked_fixture() + person = cloned.table("person") + channel = person[support_channel_column("person")].astype(str) + clone_index = person[support_clone_index_column("person")] + acs_detail = channel.eq("acs") & clone_index.eq(1) + detail_groups = list( + person.loc[acs_detail].groupby("person_tax_unit_id", sort=False).groups.values() + ) + mixed_detail = list(next(group for group in detail_groups if len(group) >= 3)) + all_child_detail = list( + next( + group + for group in detail_groups + if len(group) > 1 and list(group) != mixed_detail + ) + ) + source_id = support_source_id_column("person") + all_child_lineages = set(person.loc[all_child_detail, source_id]) + mixed_child_lineages = set(person.loc[mixed_detail[:-1], source_id]) + + person["self_employment_income_before_lsr"] = 0.0 + person["WAGP"] = person["employment_income_before_lsr"] + person["SEMP"] = person["self_employment_income_before_lsr"] + acs = channel.eq("acs") + person.loc[acs, "age"] = 40.0 + person.loc[acs, "employment_income_before_lsr"] = 100.0 + person.loc[acs, "self_employment_income_before_lsr"] = 10.0 + person.loc[acs, "WAGP"] = 100.0 + person.loc[acs, "SEMP"] = 10.0 + structural = acs & ( + person[source_id].isin(all_child_lineages) + | person[source_id].isin(mixed_child_lineages) + ) + person.loc[structural, "age"] = 12.0 + for column in ( + "employment_income_before_lsr", + "self_employment_income_before_lsr", + "WAGP", + "SEMP", + ): + person.loc[structural, column] = np.nan + return apply_acs_pums_earnings_universe_zeros( + cloned, + boundary="stacked ACS earnings-universe fixture", + ).frame + + +def test_strict_recipient_predictors_apply_exact_acs_age_universe() -> None: + cloned = _cloned_acs_earnings_universe_fixture() + before = cloned.table("person")[ + ["employment_income_before_lsr", "self_employment_income_before_lsr"] + ].copy(deep=True) + donor = pd.DataFrame( + { + "employment_income": [45_000.0, 8_000.0], + "self_employment_income": [1_000.0, 0.0], + "taxable_interest_income": [120.0, 30.0], + "weight": [1.0, 1.0], + } + ) + + prepared = prepare_us_puf_tax_detail_chain_inputs( + cloned, + donor, + predictors=( + "puf_predictor_employment_income", + "puf_predictor_self_employment_income", + ), + person_outputs=("taxable_interest_income",), + tax_unit_outputs=(), + require_complete_recipient_predictors=True, + ) + + receipt = prepared.recipient_predictor_universe + assert receipt["structurally_absent_person_rows"] == 4 + assert receipt["affected_tax_unit_rows"] == 2 + assert receipt["mixed_universe_tax_unit_rows"] == 1 + assert receipt["empty_universe_tax_unit_rows"] == 1 + assert receipt["raw_pums_source_cells_mutated"] is False + assert receipt["mapped_person_cells_materialized"] is True + assert len(receipt["sha256"]) == 64 + employment = prepared.recipient_features["puf_predictor_employment_income"] + self_employment = prepared.recipient_features[ + "puf_predictor_self_employment_income" + ] + recipient_tax_units = cloned.table("tax_unit").loc[ + prepared.recipient_features.index + ] + acs_recipient = recipient_tax_units[support_channel_column("tax_unit")].eq("acs") + assert employment.loc[acs_recipient].eq(0.0).sum() == 1 + assert self_employment.loc[acs_recipient].eq(0.0).sum() == 1 + assert employment.notna().all() + assert self_employment.notna().all() + assert receipt["predictor_source_mapping"] == { + "puf_predictor_employment_income": { + "entity": "person", + "columns": ["employment_income_before_lsr"], + }, + "puf_predictor_self_employment_income": { + "entity": "person", + "columns": ["self_employment_income_before_lsr"], + }, + } + pd.testing.assert_frame_equal( + cloned.table("person")[before.columns], + before, + ) + + +def test_acs_universe_raw_blanks_survive_checkpoint_round_trip( + tmp_path: Path, +) -> None: + frame = _cloned_acs_earnings_universe_fixture() + person = frame.table("person") + raw_columns = ["WAGP", "SEMP"] + mapped_columns = [ + "employment_income_before_lsr", + "self_employment_income_before_lsr", + ] + raw_before = person[raw_columns].copy(deep=True) + mapped_before = person[mapped_columns].copy(deep=True) + assert raw_before.isna().sum().eq(8).all() + assert mapped_before.loc[raw_before.isna().any(axis=1)].eq(0.0).all().all() + + checkpoint_path = tmp_path / "acs-earnings-universe.frame.h5" + write_frame_checkpoint(checkpoint_path, frame) + restored = load_frame_checkpoint(checkpoint_path).frame.table("person") + + pd.testing.assert_frame_equal( + restored[raw_columns], + raw_before, + check_dtype=True, + check_exact=True, + ) + pd.testing.assert_frame_equal( + restored[mapped_columns], + mapped_before, + check_dtype=True, + check_exact=True, + ) + + +def test_acs_universe_application_rejects_unreceipted_preexisting_zero() -> None: + applied = _cloned_acs_earnings_universe_fixture() + + with pytest.raises( + ValueError, + match=("acs_2024_pums_wagp_age_15_plus.*unreceipted_preexisting_mapped_rows"), + ): + apply_acs_pums_earnings_universe_zeros( + applied, + boundary="reapplied ACS universe fixture", + ) + + +def test_strict_recipient_predictors_reject_cross_grain_source_collision() -> None: + cloned = _cloned_acs_earnings_universe_fixture() + cloned.table("tax_unit")["employment_income"] = 777.0 + donor = pd.DataFrame( + { + "employment_income": [45_000.0, 8_000.0], + "taxable_interest_income": [120.0, 30.0], + "weight": [1.0, 1.0], + } + ) + + with pytest.raises(ValueError, match="ambiguous across entity grains"): + prepare_us_puf_tax_detail_chain_inputs( + cloned, + donor, + predictors=("puf_predictor_employment_income",), + person_outputs=("taxable_interest_income",), + tax_unit_outputs=(), + require_complete_recipient_predictors=True, + ) + + +def test_strict_recipient_predictors_reject_infinite_feature() -> None: + cloned = _cloned_acs_earnings_universe_fixture() + person = cloned.table("person") + candidate = ( + person[support_channel_column("person")].eq("acs") + & person[support_clone_index_column("person")].eq(1) + & person["age"].ge(15.0) + ) + row = person.index[candidate][0] + person.loc[row, "employment_income_before_lsr"] = np.inf + person.loc[row, "WAGP"] = np.inf + donor = pd.DataFrame( + { + "employment_income": [45_000.0, 8_000.0], + "taxable_interest_income": [120.0, 30.0], + "weight": [1.0, 1.0], + } + ) + + with pytest.raises(ValueError, match="nonfinite values.*employment"): + prepare_us_puf_tax_detail_chain_inputs( + cloned, + donor, + predictors=("puf_predictor_employment_income",), + person_outputs=("taxable_interest_income",), + tax_unit_outputs=(), + require_complete_recipient_predictors=True, + ) + + +@pytest.mark.parametrize( + ("raw_source", "predictor"), + [ + ("WAGP", "puf_predictor_employment_income"), + ("SEMP", "puf_predictor_self_employment_income"), + ], +) +def test_strict_recipient_predictors_require_raw_acs_universe_authority( + raw_source: str, + predictor: str, +) -> None: + cloned = _cloned_acs_earnings_universe_fixture() + cloned.table("person").drop(columns=[raw_source], inplace=True) + donor_source = predictor.removeprefix("puf_predictor_") + donor = pd.DataFrame( + { + donor_source: [45_000.0, 8_000.0], + "taxable_interest_income": [120.0, 30.0], + "weight": [1.0, 1.0], + } + ) + + with pytest.raises(ValueError, match="raw_source_authority_missing"): + prepare_us_puf_tax_detail_chain_inputs( + cloned, + donor, + predictors=(predictor,), + person_outputs=("taxable_interest_income",), + tax_unit_outputs=(), + require_complete_recipient_predictors=True, + ) + + +def test_puf_finalize_masks_earnings_allocation_to_age_15_plus() -> None: + frame = _cloned_acs_earnings_universe_fixture() + person = frame.table("person") + for column in US_QBI_OUTPUT_COLUMNS: + person[column] = False if column in US_QBI_BOOLEAN_OUTPUT_COLUMNS else 0.0 + person["long_term_capital_gains_before_response"] = 0.0 + person["non_sch_d_capital_gains"] = 0.0 + tax_unit = frame.table("tax_unit") + detail_tax_units = tax_unit[support_clone_index_column("tax_unit")].eq(1) + predictions = pd.DataFrame( + { + "employment_income_before_lsr": 1_000.0, + "self_employment_income_before_lsr": 100.0, + }, + index=tax_unit.index[detail_tax_units], + ) + donor = pd.DataFrame( + { + "employment_income_before_lsr": [1_000.0, 2_000.0], + "self_employment_income_before_lsr": [100.0, 200.0], + "weight": [1.0, 1.0], + } + ) + + finalized = finalize_us_puf_tax_detail_predictions( + frame, + donor, + predictions, + person_outputs=tuple(predictions.columns), + tax_unit_outputs=(), + absent_cells=PUF_ABSENT_CELLS_PRESERVE_NULLS, + ) + finalized_person = finalized.table("person") + acs = finalized_person[support_channel_column("person")].eq("acs") + child = finalized_person["age"].lt(15) + native_child = ( + acs & child & finalized_person[support_clone_index_column("person")].eq(0) + ) + detail_child = ( + acs & child & finalized_person[support_clone_index_column("person")].eq(1) + ) + assert int(native_child.sum()) == int(detail_child.sum()) == 4 + for column in predictions: + assert finalized_person.loc[native_child, column].eq(0.0).all() + assert finalized_person.loc[detail_child, column].eq(0.0).all() + + detail = finalized_person[support_clone_index_column("person")].eq(1) + detail_units = finalized_person.loc[ + detail, ["person_tax_unit_id", "age", *predictions.columns] + ] + mixed = detail_units.groupby("person_tax_unit_id", sort=False).filter( + lambda group: group["age"].lt(15).any() and group["age"].ge(15).any() + ) + assert not mixed.empty + mixed_child_counts = mixed.groupby("person_tax_unit_id", sort=False)["age"].apply( + lambda age: int(age.lt(15).sum()) + ) + assert mixed_child_counts.eq(2).all() + assert mixed.loc[mixed["age"].lt(15), list(predictions)].eq(0.0).all().all() + mixed_adult_totals = ( + mixed.loc[mixed["age"].ge(15)] + .groupby("person_tax_unit_id", sort=False)[list(predictions)] + .sum() + ) + assert mixed_adult_totals["employment_income_before_lsr"].eq(1_000.0).all() + assert mixed_adult_totals["self_employment_income_before_lsr"].eq(100.0).all() + all_child = detail_units.groupby("person_tax_unit_id", sort=False).filter( + lambda group: group["age"].lt(15).all() + ) + assert len(all_child) == 2 + assert all_child[list(predictions)].eq(0.0).all().all() + + derived = derive_multispine_pool_inputs(finalized) + derived_person = derived.frame.table("person") + qbi_receipt = derived.receipt["qbi_input_reconciliation"] + + assert ( + derived_person.loc[native_child, "self_employment_income_before_lsr"] + .eq(0.0) + .all() + ) + assert ( + derived_person.loc[detail_child, "self_employment_income_before_lsr"] + .eq(0.0) + .all() + ) + assert ( + qbi_receipt["recipient_source_universe"][ + "rows_excluded_from_base_self_employment_rewrite" + ] + == 8 + ) + assert qbi_receipt["structurally_absent_base_source_changed_rows"] == 0 + assert len(qbi_receipt["input_person_table_sha256"]) == 64 + assert len(qbi_receipt["output_declared_person_values_sha256"]) == 64 + + +@pytest.mark.parametrize( + ("age", "value", "message"), + [ + ( + 15.0, + np.nan, + "missing values before coercion.*acs_2024_pums_wagp_age_15_plus", + ), + ( + 12.0, + 1.0, + "acs_2024_pums_wagp_age_15_plus.*out_of_universe_mapped_nonzero_rows=1", + ), + ], +) +def test_strict_recipient_predictors_reject_acs_universe_mismatch( + age: float, + value: float, + message: str, +) -> None: + cloned = _cloned_acs_earnings_universe_fixture() + person = cloned.table("person") + candidate = ( + person[support_channel_column("person")].eq("acs") + & person[support_clone_index_column("person")].eq(1) + & person["age"].eq(12.0) + ) + row = person.index[candidate][0] + person.loc[row, "age"] = age + person.loc[row, "employment_income_before_lsr"] = value + person.loc[row, "WAGP"] = value + donor = pd.DataFrame( + { + "employment_income": [45_000.0, 8_000.0], + "taxable_interest_income": [120.0, 30.0], + "weight": [1.0, 1.0], + } + ) + + with pytest.raises(ValueError, match=message): + prepare_us_puf_tax_detail_chain_inputs( + cloned, + donor, + predictors=("puf_predictor_employment_income",), + person_outputs=("taxable_interest_income",), + tax_unit_outputs=(), + require_complete_recipient_predictors=True, + ) + + +@pytest.mark.parametrize( + ("minimum_age", "row_age", "value"), + [ + (14, 14.0, np.nan), + (16, 15.0, 100.0), + ], +) +def test_acs_universe_age_rule_mutations_fail_closed_by_rule_id( + monkeypatch: pytest.MonkeyPatch, + minimum_age: int, + row_age: float, + value: float, +) -> None: + frame = _cloned_acs_earnings_universe_fixture() + person = frame.table("person") + candidate = person[support_channel_column("person")].eq("acs") & person[ + support_clone_index_column("person") + ].eq(1) + row = person.index[candidate][0] + person.loc[row, "age"] = row_age + person.loc[row, "employment_income_before_lsr"] = value + person.loc[row, "WAGP"] = value + monkeypatch.setattr( + universe_module, + "ACS_PUMS_EARNINGS_MINIMUM_AGE", + minimum_age, + ) + donor = pd.DataFrame( + { + "employment_income": [45_000.0, 8_000.0], + "taxable_interest_income": [120.0, 30.0], + "weight": [1.0, 1.0], + } + ) + + with pytest.raises( + ValueError, + match="acs_2024_pums_wagp_age_15_plus", + ): + prepare_us_puf_tax_detail_chain_inputs( + frame, + donor, + predictors=("puf_predictor_employment_income",), + person_outputs=("taxable_interest_income",), + tax_unit_outputs=(), + require_complete_recipient_predictors=True, + ) + + def test_strict_recipient_predictors_reject_null_filing_status_before_coercion( monkeypatch: pytest.MonkeyPatch, ) -> None: @@ -833,6 +1293,10 @@ def _asec_gap_source() -> Frame: person["is_female"] = np.asarray([False, True, True, False, True, False]) person["is_household_head"] = np.asarray([True, False, True, True, False, True]) person["employment_income_before_lsr"] = np.linspace(10_000.0, 60_000.0, count) + person["self_employment_income_before_lsr"] = 0.0 + person["pre_subsidy_rent"] = np.asarray( + [12_000.0, 0.0, 0.0, 9_600.0, 0.0, 14_400.0] + ) person["unemployment_compensation"] = np.asarray( [0.0, 1_200.0, 0.0, 3_600.0, 0.0, 2_400.0] ) @@ -876,25 +1340,14 @@ def _acs_gap_source() -> Frame: person["is_female"] = np.asarray([position % 2 == 0 for position in range(count)]) person["is_household_head"] = ~person["person_household_id"].duplicated() person["employment_income_before_lsr"] = np.linspace(8_000.0, 90_000.0, count) + person["WAGP"] = person["employment_income_before_lsr"] + person["self_employment_income_before_lsr"] = 0.0 + person["SEMP"] = 0.0 person["acs_interest_dividend_rental_income"] = np.asarray( [0.0, 400.0, 0.0, 150.0, 900.0, 0.0, 250.0, 0.0, 3_000.0, 120.0, 60.0] ) - person["pre_subsidy_rent"] = np.asarray( - [ - 0.0, - 14_400.0, - 0.0, - 0.0, - 9_600.0, - 12_000.0, - 0.0, - 18_000.0, - 0.0, - 7_200.0, - 15_600.0, - ] - ) household = frame.table("household").copy() + household["TYPEHUGQ"] = np.asarray([1] * 9 + [2], dtype=np.int64) household["tenure_type"] = pd.Series( [ "OWNED_OUTRIGHT", @@ -906,7 +1359,7 @@ def _acs_gap_source() -> Frame: "OWNED_WITH_MORTGAGE", "RENTED", "OWNED_OUTRIGHT", - "RENTED", + np.nan, ], dtype=object, ) @@ -921,6 +1374,14 @@ def _acs_gap_source() -> Frame: ) +_CANONICAL_RENT_ABSENCE_RULE = GapFillAbsenceRule( + **{ + field: getattr(stacked_gap_fill_plan()[1].recipient_absence_rules[0], field) + for field in ("rule_id", "entity", "column", "selection", "reason") + } +) + + _GAP_FILL_TEST_PLAN = ( GapFillDirection( name="asec_survey_to_acs", @@ -934,10 +1395,11 @@ def _acs_gap_source() -> Frame: }, ), GapFillDirection( - name="acs_housing_to_asec", - recipient_channel="asec", - donor_channel="acs", + name="asec_housing_to_acs", + recipient_channel="acs", + donor_channel="asec", target_families={"person": {"housing": ("pre_subsidy_rent",)}}, + recipient_absence_rules=(_CANONICAL_RENT_ABSENCE_RULE,), ), ) @@ -945,6 +1407,7 @@ def _acs_gap_source() -> Frame: def test_production_entrypoints_take_no_authority_parameters() -> None: production_entrypoints = ( gap_fill_stacked_spine, + transfer_stacked_post_puf_inputs, stacked_completeness_gate, by_origin_battery, ) @@ -975,6 +1438,13 @@ def test_production_entrypoints_take_no_authority_parameters() -> None: def test_canonical_authority_objects_are_deeply_immutable() -> None: plan = stacked_spine_module.CANONICAL_STACKED_GAP_FILL_PLAN + post_puf_surface = stacked_spine_module.CANONICAL_STACKED_POST_PUF_TRANSFER_SURFACE + puf_producer_surface = ( + stacked_spine_module.CANONICAL_STACKED_POST_PUF_PUF_PRODUCER_SURFACE + ) + source_producer_surface = ( + stacked_spine_module.CANONICAL_STACKED_POST_PUF_SOURCE_PRODUCER_SURFACE + ) surface = stacked_spine_module.CANONICAL_STACKED_DECLARED_SURFACE registry = stacked_spine_module.CANONICAL_ORIGIN_BATTERY_METRIC_REGISTRY joint_registry = stacked_spine_module.CANONICAL_ORIGIN_BATTERY_JOINT_METRIC_REGISTRY @@ -983,6 +1453,12 @@ def test_canonical_authority_objects_are_deeply_immutable() -> None: assert isinstance(plan, tuple) with pytest.raises(TypeError): plan[0].target_families["person"]["model_required_numeric"] = () + with pytest.raises(TypeError): + post_puf_surface["person"]["puf_tax_itemization"] = () + with pytest.raises(TypeError): + puf_producer_surface["person"]["puf_tax_itemization"] = () + with pytest.raises(TypeError): + source_producer_surface["person"]["model_required_boolean"] = () with pytest.raises(TypeError): surface["person"]["model_required_numeric"] = () with pytest.raises(TypeError): @@ -1046,8 +1522,39 @@ def test_canonical_metric_registry_covers_the_declared_131_target_split() -> Non for family, targets in families.items() for target in targets } - assert len(gap_targets) == 118 - assert gap_targets < surface_targets + post_puf_targets = { + (entity, family, target, 0) + for entity, families in ( + stacked_spine_module.CANONICAL_STACKED_POST_PUF_TRANSFER_SURFACE.items() + ) + for family, targets in families.items() + for target in targets + } + puf_producer_targets = { + (entity, family, target, 0) + for entity, families in ( + stacked_spine_module.CANONICAL_STACKED_POST_PUF_PUF_PRODUCER_SURFACE.items() + ) + for family, targets in families.items() + for target in targets + } + source_producer_targets = { + (entity, family, target, 0) + for entity, families in ( + stacked_spine_module.CANONICAL_STACKED_POST_PUF_SOURCE_PRODUCER_SURFACE.items() + ) + for family, targets in families.items() + for target in targets + } + assert len(gap_targets) == 48 + assert len(post_puf_targets) == 70 + assert len(puf_producer_targets) == 43 + assert len(source_producer_targets) == 30 + assert len(puf_producer_targets & source_producer_targets) == 3 + assert puf_producer_targets | source_producer_targets == post_puf_targets + assert gap_targets.isdisjoint(post_puf_targets) + assert len(gap_targets | post_puf_targets) == 118 + assert gap_targets | post_puf_targets < surface_targets assert not { "bank_account_assets", "bond_assets", @@ -1160,33 +1667,406 @@ def _stacked_gap_fixture() -> Frame: ).frame -def test_gap_fill_plan_covers_declared_families_exactly() -> None: - plan = stacked_gap_fill_plan() - assert [direction.name for direction in plan] == [ - "asec_survey_to_acs", - "acs_housing_to_asec", - ] - survey, housing = plan - assert survey.recipient_channel == "acs" - assert survey.donor_channel == "asec" - assert housing.recipient_channel == "asec" - assert housing.donor_channel == "acs" - assert set(housing.target_families) == {"person"} - assert housing.target_families["person"] == {"housing": ("pre_subsidy_rent",)} +def test_stacked_assembly_binds_native_acs_group_quarters_lineage() -> None: + stacked = _stacked_gap_fixture() + receipt = stacked.metadata[STACKED_SPINE_MANIFEST_KEY]["acs_native_group_quarters"] - declared = pool_transfer_target_families() - recombined: dict[str, dict[str, tuple[str, ...]]] = {} - for direction in plan: - for entity, families in direction.target_families.items(): - for family, targets in families.items(): - recombined.setdefault(entity, {})[family] = tuple(targets) - assert recombined == { - entity: {family: tuple(targets) for family, targets in families.items()} - for entity, families in declared.items() - } + assert receipt["version"] == 1 + assert receipt["household_count"] == 1 + assert receipt["person_count"] == 1 + assert len(receipt["household_spine_source_ids_sha256"]) == 64 + assert len(receipt["person_spine_lineages_sha256"]) == 64 + validate_stacked_spine_frame(stacked, boundary="native ACS GQ receipt fixture") -def test_gap_fill_fills_both_directions_with_authority_receipts() -> None: +@pytest.mark.parametrize( + ("from_kind", "to_kind"), + ((1, 2), (2, 1)), +) +def test_native_acs_group_quarters_reclassification_fails_closed( + from_kind: int, + to_kind: int, +) -> None: + stacked = _stacked_gap_fixture() + household = stacked.table("household").copy() + channel = household[support_channel_column("household")].astype(str) + kind = pd.to_numeric(household["TYPEHUGQ"], errors="coerce") + row = household.index[channel.eq("acs") & kind.eq(from_kind)][0] + household.loc[row, "TYPEHUGQ"] = to_kind + household.loc[row, "tenure_type"] = np.nan if to_kind in (2, 3) else "RENTED" + tables = {entity: stacked.table(entity) for entity in stacked.entities} + tables["household"] = household + corrupted = Frame( + tables, + stacked.schema, + {entity: stacked.weights_for(entity) for entity in stacked.weighted_entities}, + stacked.strata, + mass_log=stacked.mass_log, + metadata=stacked.metadata, + ) + + with pytest.raises( + ValueError, + match="native ACS group-quarters household lineage differs", + ): + validate_stacked_spine_frame( + corrupted, + boundary="reclassified ACS GQ fixture", + ) + + +def test_acs_group_quarters_reclassification_across_clone_roles_fails_closed() -> None: + attached = clone_us_frame_for_puf_support( + _stacked_gap_fixture(), + clone_attachment_fraction=1.0, + clone_attachment_seed=578, + ) + household = attached.table("household").copy() + channel = household[support_channel_column("household")].astype(str) + clone = household[support_clone_index_column("household")] + kind = pd.to_numeric(household["TYPEHUGQ"], errors="coerce") + row = household.index[channel.eq("acs") & clone.eq(1) & kind.isin((2, 3))][0] + household.loc[row, "TYPEHUGQ"] = 1 + tables = {entity: attached.table(entity) for entity in attached.entities} + tables["household"] = household + corrupted = Frame( + tables, + attached.schema, + {entity: attached.weights_for(entity) for entity in attached.weighted_entities}, + attached.strata, + mass_log=attached.mass_log, + metadata=attached.metadata, + ) + + with pytest.raises( + ValueError, + match="classification differs across clone roles", + ): + validate_stacked_spine_frame( + corrupted, + boundary="reclassified ACS GQ clone fixture", + ) + + +def test_acs_clone_cannot_substitute_group_quarters_support_lineage() -> None: + attached = clone_us_frame_for_puf_support( + _stacked_gap_fixture(), + clone_attachment_fraction=1.0, + clone_attachment_seed=578, + ) + household = attached.table("household").copy() + channel = household[support_channel_column("household")].astype(str) + clone = household[support_clone_index_column("household")] + kind = pd.to_numeric(household["TYPEHUGQ"], errors="coerce") + gq_row = household.index[channel.eq("acs") & clone.eq(1) & kind.isin((2, 3))][0] + housing_unit_row = household.index[channel.eq("acs") & clone.eq(1) & kind.eq(1)][0] + household.loc[ + housing_unit_row, + support_source_id_column("household"), + ] = household.loc[gq_row, support_source_id_column("household")] + household.loc[housing_unit_row, "TYPEHUGQ"] = 2 + household.loc[housing_unit_row, "tenure_type"] = np.nan + tables = {entity: attached.table(entity) for entity in attached.entities} + tables["household"] = household + corrupted = Frame( + tables, + attached.schema, + {entity: attached.weights_for(entity) for entity in attached.weighted_entities}, + attached.strata, + mass_log=attached.mass_log, + metadata=attached.metadata, + ) + + with pytest.raises( + ValueError, + match=("full-clone identity failed|support/raw household lineage pairs differ"), + ): + validate_stacked_spine_frame( + corrupted, + boundary="substituted ACS GQ clone fixture", + ) + + +def test_acs_housing_unit_cannot_substitute_group_quarters_raw_identity() -> None: + filled = _gap_fill_with_test_authority( + _stacked_gap_fixture(), + plan=_GAP_FILL_TEST_PLAN, + seed=578, + n_estimators=10, + ).frame + household = filled.table("household").copy() + person = filled.table("person").copy() + channel = household[support_channel_column("household")].astype(str) + clone = household[support_clone_index_column("household")] + kind = pd.to_numeric(household["TYPEHUGQ"], errors="coerce") + acs_native = channel.eq("acs") & clone.eq(0) + person_counts = person["person_household_id"].value_counts() + one_person_households = household["household_id"].map(person_counts).eq(1) + housing_unit_row = household.index[acs_native & kind.eq(1) & one_person_households][ + 0 + ] + gq_row = household.index[acs_native & kind.isin((2, 3))][0] + housing_unit_id = household.loc[housing_unit_row, "household_id"] + gq_id = household.loc[gq_row, "household_id"] + housing_unit_person = person.index[ + person["person_household_id"].eq(housing_unit_id) + ][0] + gq_person = person.index[person["person_household_id"].eq(gq_id)][0] + + for column in ("TYPEHUGQ", "tenure_type", spine_source_id_column("household")): + housing_unit_value = household.loc[housing_unit_row, column] + household.loc[housing_unit_row, column] = household.loc[gq_row, column] + household.loc[gq_row, column] = housing_unit_value + housing_unit_person_source = person.loc[ + housing_unit_person, spine_source_id_column("person") + ] + person.loc[housing_unit_person, spine_source_id_column("person")] = person.loc[ + gq_person, spine_source_id_column("person") + ] + person.loc[gq_person, spine_source_id_column("person")] = housing_unit_person_source + housing_unit_rent = person.loc[housing_unit_person, "pre_subsidy_rent"] + person.loc[housing_unit_person, "pre_subsidy_rent"] = np.nan + person.loc[gq_person, "pre_subsidy_rent"] = housing_unit_rent + + tables = {entity: filled.table(entity) for entity in filled.entities} + tables["household"] = household + tables["person"] = person + corrupted = Frame( + tables, + filled.schema, + {entity: filled.weights_for(entity) for entity in filled.weighted_entities}, + filled.strata, + mass_log=filled.mass_log, + metadata=filled.metadata, + ) + + with pytest.raises( + ValueError, + match="support/raw/classification mapping differs", + ): + validate_stacked_spine_frame( + corrupted, + boundary="substituted native ACS GQ identity fixture", + ) + + +def test_acs_clone_person_cannot_substitute_group_quarters_parent() -> None: + filled = _gap_fill_with_test_authority( + _stacked_gap_fixture(), + plan=_GAP_FILL_TEST_PLAN, + seed=578, + n_estimators=10, + ).frame + attached = clone_us_frame_for_puf_support( + filled, + clone_attachment_fraction=1.0, + clone_attachment_seed=578, + ) + household = attached.table("household") + person = attached.table("person").copy() + household_channel = household[support_channel_column("household")].astype(str) + household_clone = household[support_clone_index_column("household")] + kind = pd.to_numeric(household["TYPEHUGQ"], errors="coerce") + clone_one = household_channel.eq("acs") & household_clone.eq(1) + person_counts = person["person_household_id"].value_counts() + one_person_households = household["household_id"].map(person_counts).eq(1) + housing_unit_id = household.loc[ + clone_one & kind.eq(1) & one_person_households, + "household_id", + ].iloc[0] + gq_id = household.loc[clone_one & kind.isin((2, 3)), "household_id"].iloc[0] + housing_unit_person = person.index[ + person["person_household_id"].eq(housing_unit_id) + ][0] + gq_person = person.index[person["person_household_id"].eq(gq_id)][0] + + person.loc[housing_unit_person, "person_household_id"] = gq_id + person.loc[gq_person, "person_household_id"] = housing_unit_id + housing_unit_rent = person.loc[housing_unit_person, "pre_subsidy_rent"] + person.loc[housing_unit_person, "pre_subsidy_rent"] = np.nan + person.loc[gq_person, "pre_subsidy_rent"] = housing_unit_rent + tables = {entity: attached.table(entity) for entity in attached.entities} + tables["person"] = person + corrupted = Frame( + tables, + attached.schema, + {entity: attached.weights_for(entity) for entity in attached.weighted_entities}, + attached.strata, + mass_log=attached.mass_log, + metadata=attached.metadata, + ) + + with pytest.raises( + ValueError, + match="person support/raw/parent/classification lineages differ", + ): + validate_stacked_spine_frame( + corrupted, + boundary="substituted clone ACS GQ parent fixture", + ) + + +def test_gap_fill_plan_covers_declared_families_exactly() -> None: + plan = stacked_gap_fill_plan() + assert [direction.name for direction in plan] == [ + "asec_survey_to_acs", + "asec_housing_to_acs", + ] + survey, housing = plan + assert survey.recipient_channel == "acs" + assert survey.donor_channel == "asec" + assert housing.recipient_channel == "acs" + assert housing.donor_channel == "asec" + assert set(housing.target_families) == {"person"} + assert housing.target_families["person"] == {"housing": ("pre_subsidy_rent",)} + + declared = stacked_spine_module.CANONICAL_STACKED_GAP_FILL_SURFACE + recombined: dict[str, dict[str, tuple[str, ...]]] = {} + for direction in plan: + for entity, families in direction.target_families.items(): + for family, targets in families.items(): + recombined.setdefault(entity, {})[family] = tuple(targets) + assert recombined == { + entity: {family: tuple(targets) for family, targets in families.items()} + for entity, families in declared.items() + } + early_keys = { + (entity, family, target) + for entity, families in recombined.items() + for family, targets in families.items() + for target in targets + } + late_keys = { + (entity, family, target) + for entity, families in ( + stacked_spine_module.CANONICAL_STACKED_POST_PUF_TRANSFER_SURFACE.items() + ) + for family, targets in families.items() + for target in targets + } + full_keys = { + (entity, family, target) + for entity, families in pool_transfer_target_families().items() + for family, targets in families.items() + for target in targets + } + assert early_keys.isdisjoint(late_keys) + assert early_keys | late_keys == full_keys + + +def test_every_declared_direction_producer_precedes_its_activation_check() -> None: + receipt = stacked_gap_fill_producer_schedule_receipt() + assert receipt["status"] == "all_producers_precede_activation" + assert receipt["direction_count"] == 2 + assert receipt["target_count"] == 48 + assert [direction["target_count"] for direction in receipt["directions"]] == [ + 47, + 1, + ] + + schedule = stacked_spine_module._build_gap_fill_producer_schedule( + stacked_spine_module.CANONICAL_STACKED_GAP_FILL_SURFACE + ) + rent_record = next( + record for record in schedule if record.target == "pre_subsidy_rent" + ) + late_schedule = tuple( + replace( + record, + producer_stage=stacked_spine_module._POST_GAP_FILL_STAGE, + ) + if record == rent_record + else record + for record in schedule + ) + failures = stacked_spine_module._gap_fill_producer_precedence_failures( + stacked_gap_fill_plan(), + late_schedule, + ) + assert len(failures) == 1 + assert "asec_housing_to_acs/person/housing/pre_subsidy_rent" in failures[0] + assert "does not precede activation stage" in failures[0] + + +def test_gap_fill_producer_guard_rejects_unknown_execution_scope( + monkeypatch: pytest.MonkeyPatch, +) -> None: + schedule = stacked_spine_module._build_gap_fill_producer_schedule( + stacked_spine_module.CANONICAL_STACKED_GAP_FILL_SURFACE + ) + rent_record = next( + record for record in schedule if record.target == "pre_subsidy_rent" + ) + contract = stacked_spine_module.POOL_OPERATOR_CONTRACTS[rent_record.operator] + monkeypatch.setattr( + stacked_spine_module, + "POOL_OPERATOR_CONTRACTS", + { + **stacked_spine_module.POOL_OPERATOR_CONTRACTS, + rent_record.operator: replace( + contract, + execution_scope="bogus_source", + ), + }, + ) + + mutated = stacked_spine_module._build_gap_fill_producer_schedule( + stacked_spine_module.CANONICAL_STACKED_GAP_FILL_SURFACE + ) + failures = stacked_spine_module._gap_fill_producer_precedence_failures( + stacked_gap_fill_plan(), + mutated, + ) + + assert len(failures) == 3 + assert all( + "unknown execution scope 'bogus_source'" in failure for failure in failures + ) + assert any( + "asec_housing_to_acs/person/housing/pre_subsidy_rent" in failure + for failure in failures + ) + + +def test_housing_activation_requires_asec_rent_producer_to_run_first() -> None: + produced = _stacked_gap_fixture() + tables = {entity: produced.table(entity) for entity in produced.entities} + tables["person"] = tables["person"].drop(columns=["pre_subsidy_rent"]) + unproduced = Frame( + tables, + produced.schema, + {entity: produced.weights_for(entity) for entity in produced.weighted_entities}, + produced.strata, + mass_log=produced.mass_log, + metadata=produced.metadata, + ) + housing = _GAP_FILL_TEST_PLAN[1] + + with pytest.raises( + ValueError, + match=( + r"asec_housing_to_acs/person/housing/pre_subsidy_rent: " + r"declared gap-fill target column is absent" + ), + ): + stacked_spine_module._verify_gap_fill_activation_authority( + unproduced, + direction=housing, + ) + + assert stacked_spine_module._verify_gap_fill_activation_authority( + produced, + direction=housing, + ) == { + ("person", "pre_subsidy_rent"): { + "authorized_null_rows": 11, + "recipient_rows": 11, + "donor_rows": 6, + } + } + + +def test_gap_fill_fills_both_directions_with_authority_receipts() -> None: stacked = _stacked_gap_fixture() result = _gap_fill_with_test_authority( stacked, @@ -1199,9 +2079,17 @@ def test_gap_fill_fills_both_directions_with_authority_receipts() -> None: channel = person[support_channel_column("person")] acs_rows = channel.eq("acs") asec_rows = channel.eq("asec") + gq_household_ids = result.frame.table("household").loc[ + pd.to_numeric( + result.frame.table("household")["TYPEHUGQ"], errors="coerce" + ).isin((2, 3)), + "household_id", + ] + acs_gq_rows = acs_rows & person["person_household_id"].isin(gq_household_ids) for column in ("unemployment_compensation", "is_disabled"): assert person.loc[acs_rows, column].notna().all() - assert person.loc[asec_rows, "pre_subsidy_rent"].notna().all() + assert person.loc[acs_rows & ~acs_gq_rows, "pre_subsidy_rent"].notna().all() + assert person.loc[acs_gq_rows, "pre_subsidy_rent"].isna().all() before_person = stacked.table("person") for column in ("unemployment_compensation", "is_disabled"): @@ -1211,8 +2099,8 @@ def test_gap_fill_fills_both_directions_with_authority_receipts() -> None: check_names=False, ) pd.testing.assert_series_equal( - person.loc[acs_rows, "pre_subsidy_rent"], - before_person.loc[acs_rows.to_numpy(), "pre_subsidy_rent"], + person.loc[asec_rows, "pre_subsidy_rent"], + before_person.loc[asec_rows.to_numpy(), "pre_subsidy_rent"], check_names=False, ) @@ -1226,10 +2114,17 @@ def test_gap_fill_fills_both_directions_with_authority_receipts() -> None: assert unemployment["authorized_null_rows"] == int(acs_rows.sum()) assert unemployment["imputed_rows"] == int(acs_rows.sum()) assert unemployment["residual_null_rows"] == 0 - housing = directions["acs_housing_to_asec"] + housing = directions["asec_housing_to_acs"] rent = housing["targets"]["person/housing/pre_subsidy_rent"] - assert rent["imputed_rows"] == int(asec_rows.sum()) - assert rent["residual_null_rows"] == 0 + assert rent["imputed_rows"] == int((acs_rows & ~acs_gq_rows).sum()) + assert rent["residual_null_rows"] == int(acs_gq_rows.sum()) == 1 + assert rent["unmodeled_rows"] == 1 + absence = rent["recipient_absence_authority"] + assert absence["rule_id"] == "acs_native_group_quarters_without_housing_unit" + assert absence["status"] == "exact_structural_absence" + assert absence["rows"] == 1 + assert absence["unexpected_null_rows"] == 0 + assert absence["structural_rows_filled"] == 0 survey_transfer = result.transfer_results["asec_survey_to_acs"] native_predictor_used = any( @@ -1283,6 +2178,54 @@ def forge_unmodeled_rows(*args: object, **kwargs: object) -> object: ) +def test_gap_fill_rejects_honestly_accounted_undeclared_residual( + monkeypatch: pytest.MonkeyPatch, +) -> None: + """Accounting alone cannot authorize a residual that starves later gates.""" + + transfer = stacked_spine_module.transfer_acs_inputs + + def leave_one_unmodeled(*args: object, **kwargs: object) -> object: + result = transfer(*args, **kwargs) + person = result.frame.table("person") + recipient = person[support_channel_column("person")].eq("acs") + row = person.index[recipient][0] + person.loc[row, "unemployment_compensation"] = np.nan + return replace( + result, + imputed_inputs=tuple( + replace( + record, + imputed_recipient_rows=record.imputed_recipient_rows - 1, + unmodeled_recipient_rows=1, + ) + if record.column == "unemployment_compensation" + else record + for record in result.imputed_inputs + ), + ) + + monkeypatch.setattr( + stacked_spine_module, + "transfer_acs_inputs", + leave_one_unmodeled, + ) + + with pytest.raises( + ValueError, + match=( + "undeclared gap-fill residual is forbidden; " + "unmodeled_rows=1, residual_null_rows=1" + ), + ): + _gap_fill_with_test_authority( + _stacked_gap_fixture(), + plan=_GAP_FILL_TEST_PLAN, + seed=578, + n_estimators=10, + ) + + def test_gap_fill_rejects_signed_zero_donor_byte_change( monkeypatch: pytest.MonkeyPatch, ) -> None: @@ -1471,31 +2414,237 @@ def test_gap_fill_activation_authority_rejects_nonlive_declared_channel( with pytest.raises( ValueError, - match=f"declared {role} channel {missing_channel!r} has no live rows", + match=f"declared {role} channel {missing_channel!r} has no live rows", + ): + _gap_fill_with_test_authority(_stacked_gap_fixture(), plan=plan, seed=578) + + +def test_gap_fill_fails_closed_on_missing_target_column() -> None: + stacked = _stacked_gap_fixture() + plan = ( + GapFillDirection( + name="asec_survey_to_acs", + recipient_channel="acs", + donor_channel="asec", + target_families={ + "person": {"model_required_numeric": ("veterans_benefits",)} + }, + ), + ) + with pytest.raises(ValueError, match="veterans_benefits.*absent"): + _gap_fill_with_test_authority(stacked, plan=plan, seed=578) + + +def test_gap_fill_rejects_cloned_frames() -> None: + cloned = clone_us_frame_for_puf_support(_stacked_gap_fixture()) + with pytest.raises(ValueError, match="before clone operators"): + _gap_fill_with_test_authority(cloned, plan=_GAP_FILL_TEST_PLAN, seed=578) + + +def _post_puf_transfer_fixture() -> Frame: + attached = clone_us_frame_for_puf_support( + _stacked_gap_fixture(), + clone_attachment_fraction=0.75, + clone_attachment_seed=578, + ) + person = attached.table("person").copy() + source_producer_rows = ( + person[support_channel_column("person")].astype(str).eq("asec") + ) + person["is_pregnant"] = pd.Series( + pd.NA, + index=person.index, + dtype="boolean", + ) + person.loc[source_producer_rows, "is_pregnant"] = np.resize( + np.asarray([True, False]), + int(source_producer_rows.sum()), + ) + tables = {entity: attached.table(entity) for entity in attached.entities} + tables["person"] = person + return Frame( + tables, + attached.schema, + {entity: attached.weights_for(entity) for entity in attached.weighted_entities}, + attached.strata, + mass_log=attached.mass_log, + metadata=attached.metadata, + ) + + +def test_post_puf_transfer_preserves_complete_asec_source_producers() -> None: + frame = _post_puf_transfer_fixture() + surface = {"person": {"model_required_boolean": ("is_pregnant",)}} + authority = stacked_spine_module._make_test_stacked_authority( + declared_surface=surface, + gap_fill_plan=(), + post_puf_transfer_surface=surface, + ) + before = frame.table("person")["is_pregnant"].copy(deep=True) + result = stacked_spine_module._transfer_stacked_post_puf_inputs_with_test_authority( + frame, + authority=authority, + seed=578, + n_estimators=10, + ) + + person = result.frame.table("person") + producer_rows = person[support_channel_column("person")].astype(str).eq("asec") + assert person["is_pregnant"].notna().all() + pd.testing.assert_series_equal( + person.loc[producer_rows, "is_pregnant"], + before.loc[producer_rows], + ) + receipt = result.receipt["targets"]["person/model_required_boolean/is_pregnant"] + assert receipt["producer_roles"] == ["asec_source"] + assert receipt["producer_rows"] == int(producer_rows.sum()) + assert receipt["authorized_null_rows"] == int((~producer_rows).sum()) + assert receipt["imputed_rows"] == int((~producer_rows).sum()) + assert receipt["unmodeled_rows"] == 0 + assert receipt["residual_null_rows"] == 0 + + +@pytest.mark.parametrize( + "clone_index", + ( + pytest.param(0, id="source_nondonor"), + pytest.param(1, id="model_donor"), + ), +) +def test_post_puf_transfer_rejects_incomplete_source_producer( + clone_index: int, +) -> None: + frame = _post_puf_transfer_fixture() + person = frame.table("person").copy() + source_producer_rows = person[support_channel_column("person")].astype(str).eq( + "asec" + ) & person[support_clone_index_column("person")].eq(clone_index) + person.loc[person.index[source_producer_rows][0], "is_pregnant"] = pd.NA + tables = {entity: frame.table(entity) for entity in frame.entities} + tables["person"] = person + incomplete = Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + metadata=frame.metadata, + ) + surface = {"person": {"model_required_boolean": ("is_pregnant",)}} + authority = stacked_spine_module._make_test_stacked_authority( + declared_surface=surface, + gap_fill_plan=(), + post_puf_transfer_surface=surface, + ) + + with pytest.raises( + ValueError, + match=r"post_puf_transfer/person/model_required_boolean/is_pregnant:.*" + r"upstream producers must observe every producer-owned target", + ): + stacked_spine_module._transfer_stacked_post_puf_inputs_with_test_authority( + incomplete, + authority=authority, + seed=578, + n_estimators=10, + ) + + +def _post_puf_puf_producer_fixture() -> Frame: + frame = _post_puf_transfer_fixture() + person = frame.table("person").copy() + clone_index = pd.to_numeric( + person[support_clone_index_column("person")], + errors="raise", + ) + producer_rows = clone_index.gt(0) + person["educator_expense"] = np.nan + person.loc[producer_rows, "educator_expense"] = 10.0 * np.arange( + 1, + int(producer_rows.sum()) + 1, + ) + tables = {entity: frame.table(entity) for entity in frame.entities} + tables["person"] = person + return Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + metadata=frame.metadata, + ) + + +def test_post_puf_transfer_preserves_every_live_puf_clone_producer() -> None: + frame = _post_puf_puf_producer_fixture() + surface = {"person": {"puf_tax_itemization": ("educator_expense",)}} + authority = stacked_spine_module._make_test_stacked_authority( + declared_surface=surface, + gap_fill_plan=(), + post_puf_transfer_surface=surface, + ) + before = frame.table("person")["educator_expense"].copy(deep=True) + + result = stacked_spine_module._transfer_stacked_post_puf_inputs_with_test_authority( + frame, + authority=authority, + seed=578, + n_estimators=10, + ) + + person = result.frame.table("person") + producer_rows = pd.to_numeric( + person[support_clone_index_column("person")], errors="raise" + ).gt(0) + assert person["educator_expense"].notna().all() + pd.testing.assert_series_equal( + person.loc[producer_rows, "educator_expense"], + before.loc[producer_rows], + ) + receipt = result.receipt["targets"]["person/puf_tax_itemization/educator_expense"] + assert receipt["producer_roles"] == ["puf_clone"] + assert receipt["producer_rows"] == int(producer_rows.sum()) + assert receipt["authorized_null_rows"] == int((~producer_rows).sum()) + assert receipt["imputed_rows"] == int((~producer_rows).sum()) + assert receipt["unmodeled_rows"] == 0 + assert receipt["residual_null_rows"] == 0 + + +def test_post_puf_transfer_rejects_incomplete_acs_puf_clone_producer() -> None: + frame = _post_puf_puf_producer_fixture() + person = frame.table("person").copy() + corrupted_rows = person[support_channel_column("person")].astype(str).eq( + "acs" + ) & person[support_clone_index_column("person")].eq(1) + person.loc[person.index[corrupted_rows][0], "educator_expense"] = np.nan + tables = {entity: frame.table(entity) for entity in frame.entities} + tables["person"] = person + incomplete = Frame( + tables, + frame.schema, + {entity: frame.weights_for(entity) for entity in frame.weighted_entities}, + frame.strata, + mass_log=frame.mass_log, + metadata=frame.metadata, + ) + surface = {"person": {"puf_tax_itemization": ("educator_expense",)}} + authority = stacked_spine_module._make_test_stacked_authority( + declared_surface=surface, + gap_fill_plan=(), + post_puf_transfer_surface=surface, + ) + + with pytest.raises( + ValueError, + match=r"post_puf_transfer/person/puf_tax_itemization/educator_expense:.*" + r"upstream producers must observe every producer-owned target", ): - _gap_fill_with_test_authority(_stacked_gap_fixture(), plan=plan, seed=578) - - -def test_gap_fill_fails_closed_on_missing_target_column() -> None: - stacked = _stacked_gap_fixture() - plan = ( - GapFillDirection( - name="asec_survey_to_acs", - recipient_channel="acs", - donor_channel="asec", - target_families={ - "person": {"model_required_numeric": ("veterans_benefits",)} - }, - ), - ) - with pytest.raises(ValueError, match="veterans_benefits.*absent"): - _gap_fill_with_test_authority(stacked, plan=plan, seed=578) - - -def test_gap_fill_rejects_cloned_frames() -> None: - cloned = clone_us_frame_for_puf_support(_stacked_gap_fixture()) - with pytest.raises(ValueError, match="before clone operators"): - _gap_fill_with_test_authority(cloned, plan=_GAP_FILL_TEST_PLAN, seed=578) + stacked_spine_module._transfer_stacked_post_puf_inputs_with_test_authority( + incomplete, + authority=authority, + seed=578, + n_estimators=10, + ) def test_gap_fill_banks_per_target_via_608_store(tmp_path) -> None: @@ -1505,7 +2654,7 @@ def test_gap_fill_banks_per_target_via_608_store(tmp_path) -> None: tmp_path / "survey", identity=identity, ), - "acs_housing_to_asec": AcsTransferTargetBankStore( + "asec_housing_to_acs": AcsTransferTargetBankStore( tmp_path / "housing", identity=identity, ), @@ -1527,7 +2676,7 @@ def test_gap_fill_banks_per_target_via_608_store(tmp_path) -> None: tmp_path / "survey", identity=identity, ), - "acs_housing_to_asec": AcsTransferTargetBankStore( + "asec_housing_to_acs": AcsTransferTargetBankStore( tmp_path / "housing", identity=identity, ), @@ -1772,6 +2921,63 @@ def test_clone_attachment_manifest_mutation_fails_closed() -> None: validate_puf_clone_attachment(tampered, boundary="tampered attachment") +def test_stacked_validator_rejects_an_unreceipted_clone_role() -> None: + attached = clone_us_frame_for_puf_support( + _stacked_gap_fixture(), + clone_attachment_fraction=1.0, + clone_attachment_seed=578, + ) + tables = {entity: attached.table(entity).copy() for entity in attached.entities} + for entity, table in tables.items(): + clone_column = support_clone_index_column(entity) + table.loc[table[clone_column].eq(1), clone_column] = 3 + relabeled = Frame( + tables, + attached.schema, + {entity: attached.weights_for(entity) for entity in attached.weighted_entities}, + attached.strata, + mass_log=attached.mass_log, + metadata=attached.metadata, + ) + + with pytest.raises(ValueError, match="unreceipted stacked clone roles.*3"): + validate_stacked_spine_frame( + relabeled, + boundary="unreceipted clone-role fixture", + ) + + +def test_stacked_validator_rejects_a_partial_attachment_receipt() -> None: + attached = clone_us_frame_for_puf_support( + _stacked_gap_fixture(), + clone_attachment_fraction=0.5, + clone_attachment_seed=578, + ) + canonical = attached.metadata[PUF_CLONE_ATTACHMENT_MANIFEST_KEY] + malformed = Frame( + {entity: attached.table(entity) for entity in attached.entities}, + attached.schema, + {entity: attached.weights_for(entity) for entity in attached.weighted_entities}, + attached.strata, + mass_log=attached.mass_log, + metadata={ + **attached.metadata, + PUF_CLONE_ATTACHMENT_MANIFEST_KEY: { + "realized_household_count": canonical["realized_household_count"], + "selected_household_source_ids_sha256": canonical[ + "selected_household_source_ids_sha256" + ], + }, + }, + ) + + with pytest.raises(ValueError, match="clone attachment manifest is malformed"): + validate_stacked_spine_frame( + malformed, + boundary="partial attachment receipt fixture", + ) + + def test_run_stacked_puf_pass_imputes_only_the_attached_arm() -> None: gap_filled = _gap_fill_with_test_authority( _stacked_gap_fixture(), @@ -1811,6 +3017,10 @@ def test_run_stacked_puf_pass_imputes_only_the_attached_arm() -> None: assert set(by_origin) == {"asec", "acs"} assert all(count > 0 for count in by_origin.values()) assert result.receipt["doctrines"]["absent_cells"] == "preserve_nulls" + universe = result.receipt["primary_puf_qrf"]["recipient_predictor_universe"] + assert universe["rules"]["employment_income_before_lsr"]["source_column"] == "WAGP" + assert universe["raw_pums_source_cells_mutated"] is False + assert len(universe["sha256"]) == 64 with pytest.raises(ValueError, match="clone attachment"): stacked_spine_module._run_stacked_puf_pass_without_tail_for_test( @@ -1821,6 +3031,78 @@ def test_run_stacked_puf_pass_imputes_only_the_attached_arm() -> None: ) +def test_run_stacked_puf_pass_receipts_raw_child_universe_application() -> None: + gap_filled = _gap_fill_with_test_authority( + _stacked_gap_fixture(), + plan=_GAP_FILL_TEST_PLAN, + seed=578, + n_estimators=10, + ).frame + person = gap_filled.table("person") + acs = person[support_channel_column("person")].eq("acs") + groups = list( + person.loc[acs].groupby("person_tax_unit_id", sort=False).groups.values() + ) + mixed_group = list(next(group for group in groups if len(group) > 1)) + all_child_group = list(next(group for group in groups if len(group) == 1)) + structural_rows = [mixed_group[0], *all_child_group] + person.loc[structural_rows, "age"] = 12.0 + for column in ( + "employment_income_before_lsr", + "self_employment_income_before_lsr", + "WAGP", + "SEMP", + ): + person.loc[structural_rows, column] = np.nan + donor = pd.DataFrame( + { + "employment_income": [45_000.0, 8_000.0, 70_000.0, 22_000.0], + "self_employment_income": [1_000.0, 0.0, 5_000.0, 200.0], + "taxable_interest_income": [120.0, 30.0, 900.0, 0.0], + "weight": [1.0, 1.0, 1.0, 1.0], + } + ) + + result = stacked_spine_module._run_stacked_puf_pass_without_tail_for_test( + gap_filled, + donor, + clone_attachment_fraction=1.0, + clone_attachment_seed=578, + predictors=( + "puf_predictor_employment_income", + "puf_predictor_self_employment_income", + ), + person_outputs=("taxable_interest_income",), + tax_unit_outputs=(), + seed=578, + n_estimators=10, + ) + + receipt = result.receipt["acs_earnings_universe_application"] + assert receipt["structurally_absent_person_rows"] == 2 + assert receipt["affected_tax_unit_rows"] == 2 + assert receipt["mixed_universe_tax_unit_rows"] == 1 + assert receipt["empty_universe_tax_unit_rows"] == 1 + assert receipt["mapped_universe_zero_cells"] == 4 + assert receipt["raw_pums_source_cells_mutated"] is False + assert receipt["mapped_person_cells_materialized"] is True + output_person = result.frame.table("person") + output_child = output_person[support_channel_column("person")].eq( + "acs" + ) & output_person["age"].lt(15) + assert int(output_child.sum()) == 4 + assert ( + output_person.loc[ + output_child, + ["employment_income_before_lsr", "self_employment_income_before_lsr"], + ] + .eq(0.0) + .all() + .all() + ) + assert output_person.loc[output_child, ["WAGP", "SEMP"]].isna().all().all() + + def test_run_stacked_puf_pass_fraction_one_receipts_out_of_frame_identity() -> None: gap_filled = _gap_fill_with_test_authority( _stacked_gap_fixture(), @@ -2109,12 +3391,18 @@ def overlapping_source(stratum: str, income_shift: float) -> Frame: 100_000.0 + income_shift, len(person), ) + person["WAGP"] = person["employment_income_before_lsr"] + person["self_employment_income_before_lsr"] = 0.0 + person["SEMP"] = 0.0 for column in PUF_CAPITAL_GAINS_TAIL_PERSON_COLUMNS: person[column] = 0.0 tax_unit = tables["tax_unit"] tax_unit["filing_status_input"] = "SINGLE" for column in PUF_CAPITAL_GAINS_TAIL_TAX_UNIT_COLUMNS: tax_unit[column] = 0.0 + household = tables["household"] + household["TYPEHUGQ"] = 1 + household["tenure_type"] = "RENTED" return Frame( tables, US_SCHEMA, @@ -2228,7 +3516,109 @@ def test_completeness_gate_passes_on_filled_and_proven_surface() -> None: statuses = { label: receipt["status"] for label, receipt in result.details["targets"].items() } - assert set(statuses.values()) == {"complete"} + assert set(statuses.values()) == {"complete", "proven_absent"} + rent = result.details["targets"]["person/housing/pre_subsidy_rent"] + assert rent["recipient_absence_authority"]["rows"] == 1 + assert rent["proven"]["acs/clone_0"]["structural_absence_rule_id"] == ( + "acs_native_group_quarters_without_housing_unit" + ) + + +def test_structural_rent_absence_covers_every_clone_role_and_battery_scope() -> None: + gap_filled = _gap_fill_with_test_authority( + _stacked_gap_fixture(), + plan=_GAP_FILL_TEST_PLAN, + seed=578, + n_estimators=10, + ).frame + attached = clone_us_frame_for_puf_support( + gap_filled, + clone_attachment_fraction=1.0, + clone_attachment_seed=578, + ) + surface = {"person": {"housing": ("pre_subsidy_rent",)}} + housing_plan = (_GAP_FILL_TEST_PLAN[1],) + completeness = _completeness_with_test_authority( + attached, + declared_surface=surface, + declared_gap_fill_plan=housing_plan, + ) + assert completeness.passed, completeness.failures + target = completeness.details["targets"]["person/housing/pre_subsidy_rent"] + assert set(target["proven"]) == {"acs/clone_0", "acs/clone_1"} + assert target["recipient_absence_authority"]["by_origin_role"] == { + "acs/clone_0": 1, + "acs/clone_1": 1, + } + + authority = stacked_spine_module._make_test_stacked_authority( + declared_surface=surface, + gap_fill_plan=housing_plan, + metric_registry={ + ( + "person", + "housing", + "pre_subsidy_rent", + 0, + ): "monetary_sign_separated" + }, + ) + battery = stacked_spine_module._by_origin_battery_with_test_authority( + attached, + authority=authority, + ) + comparison = battery.details["comparisons"][ + "person/housing/pre_subsidy_rent[clone_0]" + ] + assert comparison["status"] != "null_in_scope" + assert comparison["recipient_absence_authority"]["rows_excluded_from_scope"] == 1 + assert not any( + "exact structural-absence" in failure for failure in battery.failures + ) + + +def test_structural_rent_absence_rejects_a_non_gq_recipient_null() -> None: + gap_filled = _gap_fill_with_test_authority( + _stacked_gap_fixture(), + plan=_GAP_FILL_TEST_PLAN, + seed=578, + n_estimators=10, + ).frame + person = gap_filled.table("person").copy() + channel = person[support_channel_column("person")].astype(str) + household = gap_filled.table("household") + gq_households = set( + household.loc[ + pd.to_numeric(household["TYPEHUGQ"], errors="coerce").isin((2, 3)), + "household_id", + ] + ) + non_gq = channel.eq("acs") & ~person["person_household_id"].isin(gq_households) + person.loc[person.index[non_gq][0], "pre_subsidy_rent"] = np.nan + tables = {entity: gap_filled.table(entity) for entity in gap_filled.entities} + tables["person"] = person + corrupted = Frame( + tables, + gap_filled.schema, + { + entity: gap_filled.weights_for(entity) + for entity in gap_filled.weighted_entities + }, + gap_filled.strata, + mass_log=gap_filled.mass_log, + metadata=gap_filled.metadata, + ) + result = _completeness_with_test_authority( + corrupted, + declared_surface={"person": {"housing": ("pre_subsidy_rent",)}}, + declared_gap_fill_plan=(_GAP_FILL_TEST_PLAN[1],), + ) + assert not result.passed + assert any( + "exact structural-absence equation failed" in failure + and "unexpected_null_rows=1" in failure + for failure in result.failures + ) def test_completeness_gate_names_a_silently_missing_family() -> None: @@ -2386,6 +3776,38 @@ def test_completeness_gate_empty_plan_cannot_launder_canonical_target() -> None: ) +def test_completeness_gate_cannot_launder_post_puf_transfer_target() -> None: + frame = _post_puf_transfer_fixture() + surface = {"person": {"model_required_boolean": ("is_pregnant",)}} + authority = stacked_spine_module._make_test_stacked_authority( + declared_surface=surface, + gap_fill_plan=(), + post_puf_transfer_surface=surface, + ) + result = stacked_spine_module._stacked_completeness_gate_with_test_authority( + frame, + authority=authority, + absence_proofs=tuple( + AbsenceProof( + entity="person", + column="is_pregnant", + channel="*", + clone_index=clone_index, + reason="POST-PUF WILDCARD LAUNDER", + ) + for clone_index in (0, 1) + ), + ) + + assert not result.passed + assert any( + "person/model_required_boolean/is_pregnant" in failure + and "zero-residual post-PUF transfer contract" in failure + and "absence proofs are forbidden" in failure + for failure in result.failures + ) + + def test_completeness_gate_empty_surface_is_terminal() -> None: authority = stacked_spine_module._make_test_stacked_authority( declared_surface={}, @@ -2488,10 +3910,12 @@ def test_completeness_receipts_bind_live_authority_per_target() -> None: == authority ) plan_sha256 = authority["components"]["gap_fill_plan"]["sha256"] + post_puf_sha256 = authority["components"]["post_puf_transfer_surface"]["sha256"] surface_sha256 = authority["components"]["declared_surface"]["sha256"] for receipt in canonical.details["targets"].values(): assert receipt["authority_form"] == "observed_complete" assert receipt["plan_sha256"] == plan_sha256 + assert receipt["post_puf_surface_sha256"] == post_puf_sha256 assert receipt["surface_sha256"] == surface_sha256 stacked = _stacked_gap_fixture() @@ -2579,6 +4003,9 @@ def test_self_digested_partial_authority_cannot_forge_production_identity() -> N authority_id="us_stacked_spine_authority", version=1, gap_fill_plan=(), + post_puf_transfer_surface={}, + post_puf_puf_producer_surface={}, + post_puf_source_producer_surface={}, declared_surface=surface, metric_registry={ ("person", "test_only", "unemployment_compensation", 0): ( @@ -2606,6 +4033,41 @@ def test_self_digested_partial_authority_cannot_forge_production_identity() -> N GateReport((result,)).to_manifest() +@pytest.mark.parametrize("stale_version", (1, 2, 3, 4, 5)) +def test_self_consistent_stale_stacked_authority_versions_are_rejected( + stale_version: int, +) -> None: + canonical = stacked_spine_module._production_stacked_authority() + stale = stacked_spine_module._make_stacked_authority( + authority_id=canonical.authority_id, + version=stale_version, + gap_fill_plan=canonical.gap_fill_plan, + post_puf_transfer_surface=canonical.post_puf_transfer_surface, + post_puf_puf_producer_surface=canonical.post_puf_puf_producer_surface, + post_puf_source_producer_surface=canonical.post_puf_source_producer_surface, + declared_surface=canonical.declared_surface, + metric_registry=canonical.metric_registry, + joint_metric_registry=canonical.joint_metric_registry, + support_profile=canonical.support_profile, + declared_form="CANONICAL", + ) + stale_receipt = stacked_spine_module._authority_receipt(stale) + + assert stacked_spine_module.stacked_spine_authority_receipt()["version"] == 6 + assert stale_receipt["version"] == stale_version + assert stale_receipt["integrity_valid"] is True + assert stale_receipt["digest_matches_declared"] is True + assert stale_receipt["canonical_content"] is False + with pytest.raises( + ValueError, + match="non-canonical stacked authority is forbidden", + ): + stacked_spine_module._validate_production_authority_receipt( + stale_receipt, + boundary=f"stale authority v{stale_version}", + ) + + def test_rebound_anchor_aliases_cannot_replace_captured_canonical_authority( monkeypatch: pytest.MonkeyPatch, ) -> None: @@ -2614,6 +4076,9 @@ def test_rebound_anchor_aliases_cannot_replace_captured_canonical_authority( authority_id="us_stacked_spine_authority", version=1, gap_fill_plan=(), + post_puf_transfer_surface={}, + post_puf_puf_producer_surface={}, + post_puf_source_producer_surface={}, declared_surface=surface, metric_registry={ ("person", "test_only", "unemployment_compensation", 0): ( @@ -2626,11 +4091,15 @@ def test_rebound_anchor_aliases_cannot_replace_captured_canonical_authority( rebound = { "_CANONICAL_STACKED_DECLARED_SURFACE_ANCHOR": forged.declared_surface, "_CANONICAL_STACKED_GAP_FILL_PLAN_ANCHOR": forged.gap_fill_plan, + "_CANONICAL_STACKED_POST_PUF_TRANSFER_SURFACE_ANCHOR": ( + forged.post_puf_transfer_surface + ), "_CANONICAL_ORIGIN_BATTERY_METRIC_REGISTRY_ANCHOR": forged.metric_registry, "_CANONICAL_ORIGIN_BATTERY_SUPPORT_PROFILE_ANCHOR": (forged.support_profile), "_CANONICAL_STACKED_AUTHORITY_ANCHOR": forged, "_STACKED_DECLARED_SURFACE": forged.declared_surface, "_STACKED_GAP_FILL_PLAN": forged.gap_fill_plan, + "_STACKED_POST_PUF_TRANSFER_SURFACE": forged.post_puf_transfer_surface, "_BATTERY_METRIC_REGISTRY": forged.metric_registry, "_BATTERY_SUPPORT_PROFILE": forged.support_profile, } @@ -2654,6 +4123,7 @@ def test_rebound_anchor_aliases_cannot_replace_captured_canonical_authority( "support_threshold", "surface_count", "direction_count", + "post_puf_target_count", "target_binding", ), ) @@ -2677,6 +4147,8 @@ def test_stacked_manifest_rejects_pre_emission_nested_receipt_mutation( authority["components"]["declared_surface"]["target_count"] = 0 elif mutation == "direction_count": authority["components"]["gap_fill_plan"]["direction_count"] = 0 + elif mutation == "post_puf_target_count": + authority["components"]["post_puf_transfer_surface"]["target_count"] = 0 else: target = next(iter(result.details["targets"].values())) target["authority_form"] = "wildcard_no_declared_donor_plan" @@ -2783,6 +4255,9 @@ def test_fresh_gate_result_cannot_forge_a_donor_origin_proof() -> None: binding = { "authority_sha256": authority["sha256"], "plan_sha256": authority["components"]["gap_fill_plan"]["sha256"], + "post_puf_surface_sha256": authority["components"]["post_puf_transfer_surface"][ + "sha256" + ], "surface_sha256": authority["components"]["declared_surface"]["sha256"], } label = "person/puf_tax_itemization/taxable_interest_income" @@ -2813,6 +4288,65 @@ def test_fresh_gate_result_cannot_forge_a_donor_origin_proof() -> None: GateReport((forged,)).to_manifest() +def test_fresh_gate_result_cannot_forge_structural_rent_absence() -> None: + frame = _battery_frame( + { + "taxable_interest_income": ( + np.asarray([100.0] * 8), + np.asarray([100.0] * 11), + ) + } + ) + canonical = stacked_completeness_gate(frame) + assert canonical.passed, canonical.failures + details = deepcopy(dict(canonical.details)) + rent = details["targets"]["person/housing/pre_subsidy_rent"] + rent["recipient_absence_authority"]["reason"] = "FORGED STRUCTURAL ABSENCE" + forged = replace(canonical, details=details) + + with pytest.raises( + ValueError, + match="structural-absence doctrine mismatch.*emission is forbidden", + ): + GateReport((forged,)).to_manifest() + + +@pytest.mark.parametrize( + "evaluate", + (stacked_completeness_gate, by_origin_battery), + ids=("completeness", "battery"), +) +def test_canonical_structural_receipt_cannot_be_grafted_between_evaluations( + evaluate: object, +) -> None: + columns = { + "taxable_interest_income": ( + np.asarray([100.0] * 8), + np.asarray([100.0] * 11), + ) + } + no_gq = evaluate(_battery_frame(columns)) + one_gq = evaluate(_battery_frame(columns, acs_group_quarters=True)) + assert no_gq.passed, no_gq.failures + assert one_gq.passed, one_gq.failures + + # Every field in the replacement receipt is canonical: it came from a + # real evaluation of the same gate and authority. It still cannot stand + # in for the evaluator's evidence about another frame. + with pytest.raises(TypeError, match="init=False"): + replace( + no_gq, + details=deepcopy(dict(one_gq.details)), + _stacked_authority_seal=b"caller-recomputed-seal", + ) + forged = replace(no_gq, details=deepcopy(dict(one_gq.details))) + with pytest.raises( + ValueError, + match="was not sealed by its evaluator.*emission is forbidden", + ): + GateReport((forged,)).to_manifest() + + def test_fresh_battery_result_cannot_forge_canonical_coverage_receipts() -> None: frame = _battery_frame( { @@ -2841,6 +4375,29 @@ def test_fresh_battery_result_cannot_forge_canonical_coverage_receipts() -> None GateReport((forged,)).to_manifest() +def test_fresh_battery_result_requires_structural_scope_receipt() -> None: + frame = _battery_frame( + { + "taxable_interest_income": ( + np.asarray([100.0] * 8), + np.asarray([100.0] * 11), + ) + } + ) + canonical = by_origin_battery(frame) + assert canonical.passed, canonical.failures + details = deepcopy(dict(canonical.details)) + rent_label = "person/housing/pre_subsidy_rent[clone_0]" + details["comparisons"][rent_label].pop("recipient_absence_authority") + forged = replace(canonical, details=details) + + with pytest.raises( + ValueError, + match="must carry canonical recipient-absence authority.*emission is forbidden", + ): + GateReport((forged,)).to_manifest() + + def test_fresh_battery_result_cannot_relabel_a_canonical_metric() -> None: frame = _battery_frame( { @@ -2909,10 +4466,17 @@ def test_stripped_noncanonical_receipt_cannot_escape_under_a_renamed_gate( GateReport((stripped,)).to_manifest() -def test_stripped_five_component_authority_cannot_escape_under_a_renamed_gate() -> None: +def test_stripped_six_component_authority_cannot_escape_under_a_renamed_gate() -> None: authority = stacked_spine_module.stacked_spine_authority_receipt() components = deepcopy(dict(authority["components"])) - assert "joint_metric_registry" in components + assert set(components) == { + "gap_fill_plan", + "post_puf_transfer_surface", + "declared_surface", + "metric_registry", + "joint_metric_registry", + "support_profile", + } stripped = GateResult( name="renamed_stacked_battery", passed=True, @@ -3043,7 +4607,11 @@ def _with_declared_battery_defaults( ) -def _battery_frame(columns: dict[str, tuple[np.ndarray, np.ndarray]]) -> Frame: +def _battery_frame( + columns: dict[str, tuple[np.ndarray, np.ndarray]], + *, + acs_group_quarters: bool = False, +) -> Frame: """A stacked frame with hand-set asec/acs person columns. ``columns`` maps a column name to its (asec values, acs values) pair. @@ -3068,8 +4636,17 @@ def with_columns(frame: Frame, position: int) -> Frame: person = frame.table("person").copy() for column, values in columns.items(): person[column] = values[position] + if position == 1 and acs_group_quarters: + person.loc[person.index[0], "pre_subsidy_rent"] = np.nan tables = {entity: frame.table(entity) for entity in frame.entities} tables["person"] = person + household = tables["household"].copy() + household["TYPEHUGQ"] = np.ones(len(household), dtype=np.int64) + household["tenure_type"] = "RENTED" + if position == 1 and acs_group_quarters: + household.loc[household.index[0], "TYPEHUGQ"] = 2 + household.loc[household.index[0], "tenure_type"] = np.nan + tables["household"] = household return Frame( tables, US_SCHEMA, @@ -3565,11 +5142,13 @@ def _asec_e2e_source() -> Frame: person["is_female"] = index % 2 == 0 person["is_household_head"] = True person["employment_income_before_lsr"] = 20_000.0 + 1_500.0 * index - person["unemployment_compensation"] = np.where(index % 4 == 0, 2_400.0, 0.0) + person["self_employment_income_before_lsr"] = 0.0 # Structurally learnable from a REQUIRED predictor with a stable share - # under any household subsample, so the gap-fill QRF reproduces the - # incidence on the seeded ACS sample without sampling-skew noise. + # under any household subsample, so the gap-fill QRF reproduces both + # incidences on the seeded ACS sample without sampling-skew noise. + person["unemployment_compensation"] = np.where(person["is_female"], 2_400.0, 0.0) person["is_disabled"] = person["is_female"].to_numpy() + person["pre_subsidy_rent"] = np.where(index % 2 == 1, 11_000.0 + 150.0 * index, 0.0) interest = np.where(index % 2 == 0, 1_200.0 + 40.0 * index, 0.0) person["taxable_interest_income"] = interest for column in ( @@ -3581,6 +5160,7 @@ def _asec_e2e_source() -> Frame: ): person[column] = 0.0 household = frame.table("household").copy() + household["TYPEHUGQ"] = np.ones(40, dtype=np.int64) household["tenure_type"] = pd.Series( ["RENTED" if position % 2 else "OWNED_WITH_MORTGAGE" for position in range(40)], dtype=object, @@ -3615,11 +5195,14 @@ def _acs_e2e_source() -> Frame: person["is_female"] = index % 2 == 1 person["is_household_head"] = True person["employment_income_before_lsr"] = 21_000.0 + 1_450.0 * index + person["WAGP"] = person["employment_income_before_lsr"] + person["self_employment_income_before_lsr"] = 0.0 + person["SEMP"] = 0.0 person["acs_interest_dividend_rental_income"] = np.where( index % 2 == 0, 1_250.0 + 42.0 * index, 0.0 ) - person["pre_subsidy_rent"] = np.where(index % 2 == 1, 11_000.0 + 150.0 * index, 0.0) household = frame.table("household").copy() + household["TYPEHUGQ"] = np.ones(40, dtype=np.int64) household["tenure_type"] = pd.Series( ["RENTED" if position % 2 else "OWNED_OUTRIGHT" for position in range(40)], dtype=object, @@ -3655,10 +5238,11 @@ def _acs_e2e_source() -> Frame: }, ), GapFillDirection( - name="acs_housing_to_asec", - recipient_channel="asec", - donor_channel="acs", + name="asec_housing_to_acs", + recipient_channel="acs", + donor_channel="asec", target_families={"person": {"housing": ("pre_subsidy_rent",)}}, + recipient_absence_rules=(_CANONICAL_RENT_ABSENCE_RULE,), ), ) diff --git a/tools/build_us_multispine_pool.py b/tools/build_us_multispine_pool.py index f7510f28..5a21f818 100644 --- a/tools/build_us_multispine_pool.py +++ b/tools/build_us_multispine_pool.py @@ -70,6 +70,9 @@ ) from microcosm.build.logbook import record_build_attempt from microcosm.build.serialization_dtypes import canonicalize_frame_string_dtypes +from microcosm.build.us_runtime.acs_income_universe import ( + acs_pums_earnings_universe_contract_identity, +) from microcosm.build.us_runtime.acs_inputs import map_acs_native_inputs from microcosm.build.us_runtime.acs_pums import ( AcsPumsSource, @@ -131,6 +134,7 @@ ) from microcosm.build.us_runtime.puf_donor_io import load_puf_tax_unit_donor from microcosm.build.us_runtime.puf_qrf_chain import ( + PRIMARY_QRF_CHECKPOINT_SCHEMA_VERSION, PRIMARY_QRF_MANIFEST_FILENAME, PRIMARY_QRF_TARGET_ORDER, finalize_primary_puf_qrf_chain, @@ -141,7 +145,14 @@ PUF_SUPPORT_MAX_CLONE_SAFE_SOURCE_ID, US_PUF_SUPPORT_FIT_NAME, ) +from microcosm.build.us_runtime.qbi_inputs import ( + us_qbi_post_reconciliation_person_columns, + us_qbi_reconciliation_contract_identity, + validate_us_qbi_reconciliation_live_output, +) from microcosm.build.us_runtime.stacked_spine import ( + CANONICAL_STACKED_GAP_FILL_SURFACE, + CANONICAL_STACKED_POST_PUF_TRANSFER_SURFACE, assemble_stacked_spine, assert_stacked_tail_cells_preserved, by_origin_battery, @@ -150,7 +161,10 @@ run_stacked_puf_pass, stacked_completeness_gate, stacked_gap_fill_plan, + stacked_gap_fill_producer_schedule_receipt, stacked_spine_authority_receipt, + transfer_stacked_post_puf_inputs, + validate_stacked_post_puf_transfer_receipt, validate_stacked_spine_frame, ) from microcosm.build.us_runtime.support_provenance import ( @@ -231,9 +245,10 @@ 1.00: "f100", } _STACKED_PIPELINE = "us-stacked-pool" -# Version 2 binds the complete clone-2 tail provenance and attachment-descendant -# metadata. Version-1 stacked checkpoints predate that state and must rebuild. -_STACKED_CHECKPOINT_MATERIALIZER_VERSION = 2 +# Version 6 binds the ACS earnings-universe and whole-pool QBI reconciliation +# contracts, plus primary-QRF schema 6. Earlier checkpoints can carry stale +# recipient and mutation semantics and must rebuild. +_STACKED_CHECKPOINT_MATERIALIZER_VERSION = 6 _STACKED_RELEASE_ID_PATTERN = re.compile( r"^populace-us-2024-stacked-f(?:001|010|100)-s[0-9]+-" r"asec[0-9]+-acs[0-9]+-[0-9]{8}T[0-9]{6}Z-[0-9a-f]{8}$" @@ -265,6 +280,7 @@ class StackedPoolBuildResult: stage_receipts: Mapping[str, Mapping[str, object]] terminal_gates: tuple[GateResult, GateResult] release_id: str + qbi_transition_authority_sha256: str | None = None @property def simulation_ready(self) -> bool: @@ -1012,6 +1028,7 @@ def _stacked_checkpoint_base_identity( "gap_fill_stacked_spine", "run_stacked_puf_pass", "complete_multispine_source_inputs", + "transfer_stacked_post_puf_inputs", "prepare_stacked_tail_derivation", "derive_multispine_pool_inputs", "seed_multispine_pool_inputs", @@ -1022,11 +1039,23 @@ def _stacked_checkpoint_base_identity( "pre_clone_source_operator_order": list( POOL_PRE_CLONE_SOURCE_OPERATOR_ORDER ), + "gap_fill_producer_schedule": ( + stacked_gap_fill_producer_schedule_receipt() + ), "post_clone_source_operator_order": list( POOL_POST_CLONE_SOURCE_OPERATOR_ORDER ), "derive_operator_order": list(POOL_DERIVE_OPERATOR_ORDER), "primary_qrf_target_order": list(PRIMARY_QRF_TARGET_ORDER), + "primary_qrf_checkpoint_schema_version": ( + PRIMARY_QRF_CHECKPOINT_SCHEMA_VERSION + ), + "acs_pums_earnings_universe_contract": ( + acs_pums_earnings_universe_contract_identity() + ), + "us_qbi_reconciliation_contract": ( + us_qbi_reconciliation_contract_identity() + ), "take_up_contract": take_up_contract_identity(), "primary_qrf_n_estimators": _PRIMARY_QRF_N_ESTIMATORS, "acs_transfer_n_estimators": _ACS_TRANSFER_N_ESTIMATORS, @@ -1317,6 +1346,17 @@ def write(self, checkpoint: MultispinePoolCheckpoint) -> None: checkpoint.frame, boundary=f"pool {stage} checkpoint write", ) + qbi_route = _checkpoint_qbi_route(self._base_identity) + if stage == "simulated" and qbi_route is not None: + _validate_qbi_stage_receipt( + persistent_frame, + checkpoint.stage_receipts, + route=qbi_route, + boundary=f"pool {stage} durable checkpoint write", + transition_authority_sha256=( + checkpoint.qbi_transition_authority_sha256 + ), + ) stored_frame = persistent_frame if stage == "simulated": if checkpoint.simulation_frame is None: @@ -1376,6 +1416,10 @@ def write(self, checkpoint: MultispinePoolCheckpoint) -> None: else None ), } + if checkpoint.qbi_transition_authority_sha256 is not None: + metadata["qbi_transition_authority_sha256"] = ( + checkpoint.qbi_transition_authority_sha256 + ) path = self.checkpoint_path(stage) started_at = time.perf_counter() write_frame_checkpoint(path, stored_frame, metadata=metadata) @@ -1399,6 +1443,15 @@ def write(self, checkpoint: MultispinePoolCheckpoint) -> None: "row_counts": _frame_row_counts(stored_frame), "frame_schema": _frame_schema_payload(stored_frame), "frame_metadata": metadata["frame_metadata"], + **( + { + "qbi_transition_authority_sha256": metadata[ + "qbi_transition_authority_sha256" + ] + } + if "qbi_transition_authority_sha256" in metadata + else {} + ), }, ) receipts_record: dict[str, object] = { @@ -1645,7 +1698,12 @@ def _load(self, stage: str) -> MultispinePoolCheckpoint | None: expected_identity=expected_identity, expected_identity_sha256=expected_identity_sha256, ) - for key in ("row_counts", "frame_schema", "frame_metadata"): + for key in ( + "row_counts", + "frame_schema", + "frame_metadata", + "qbi_transition_authority_sha256", + ): if metadata.get(key) != manifest.get(key): raise ValueError( f"{stage} checkpoint {key} differs from its sidecar" @@ -1669,6 +1727,9 @@ def _load(self, stage: str) -> MultispinePoolCheckpoint | None: ) assembly_receipt = metadata.get("assembly_receipt") stage_receipts = metadata.get("stage_receipts") + qbi_transition_authority_sha256 = metadata.get( + "qbi_transition_authority_sha256" + ) input_receipts = metadata.get("input_receipts") if not isinstance(assembly_receipt, Mapping): raise ValueError( @@ -1709,12 +1770,23 @@ def _load(self, stage: str) -> MultispinePoolCheckpoint | None: simulation_frame = frame persistent_frame = _without_simulation_output(frame) + qbi_route = _checkpoint_qbi_route(self._base_identity) + if stage == "simulated" and qbi_route is not None: + _validate_qbi_stage_receipt( + persistent_frame, + restored_stage_receipts, + route=qbi_route, + boundary="pool simulated durable checkpoint load", + transition_authority_sha256=(qbi_transition_authority_sha256), + ) + checkpoint = MultispinePoolCheckpoint( stage=stage, frame=persistent_frame, assembly_receipt=dict(assembly_receipt), stage_receipts=restored_stage_receipts, simulation_frame=simulation_frame, + qbi_transition_authority_sha256=(qbi_transition_authority_sha256), ) self.bind_input_receipts(input_receipts) self._attempts[stage] = { @@ -2406,6 +2478,16 @@ def _stacked_direction_bank_identity( } +def _stacked_post_puf_bank_identity( + checkpoint_identity: Mapping[str, object], +) -> dict[str, object]: + return { + **_pool_checkpoint_stage_identity(checkpoint_identity, "transferred"), + "stacked_transfer_stage": "post_puf_source_completion", + "stacked_transfer_name": "asec_clone_1_to_missing", + } + + def _stacked_tail_manifest( stage_receipts: Mapping[str, Mapping[str, object]], ) -> Mapping[str, object]: @@ -2421,6 +2503,100 @@ def _stacked_tail_manifest( return manifest +def _validate_stacked_post_puf_stage_receipt( + stage_receipts: Mapping[str, Mapping[str, object]], + *, + boundary: str, +) -> None: + """Require the nested late-transfer receipt to carry canonical authority.""" + + impute = stage_receipts.get("impute") + if not isinstance(impute, Mapping): + raise ValueError( + f"{boundary}: stacked transferred receipts have no impute object." + ) + transfer_receipt = impute.get("stacked_post_puf_transfer") + if not isinstance(transfer_receipt, Mapping): + raise ValueError( + f"{boundary}: stacked transferred receipts have no post-PUF " + "transfer object." + ) + validate_stacked_post_puf_transfer_receipt( + transfer_receipt, + boundary=boundary, + ) + + +def _qbi_receipt_from_stage_receipts( + stage_receipts: Mapping[str, Mapping[str, object]], + *, + route: str, + boundary: str, +) -> Mapping[str, object]: + """Resolve the exact legacy or stacked QBI receipt path without fallback.""" + + derive = stage_receipts.get("derive") + if not isinstance(derive, Mapping): + raise ValueError(f"{boundary}: stage receipts have no derive object.") + if route == "legacy": + if "pool_derivation" in derive: + raise ValueError( + f"{boundary}: legacy QBI receipt used the stacked derive route." + ) + receipt = derive.get("qbi_input_reconciliation") + elif route == "stacked": + pool_derivation = derive.get("pool_derivation") + if not isinstance(pool_derivation, Mapping): + raise ValueError( + f"{boundary}: stacked derive receipts have no pool_derivation object." + ) + if "qbi_input_reconciliation" in derive: + raise ValueError( + f"{boundary}: stacked QBI receipt also appears at the legacy route." + ) + receipt = pool_derivation.get("qbi_input_reconciliation") + else: # pragma: no cover - internal callers pass a literal + raise ValueError(f"{boundary}: unknown QBI receipt route {route!r}.") + if not isinstance(receipt, Mapping): + raise ValueError( + f"{boundary}: stage receipts have no QBI reconciliation object." + ) + return receipt + + +def _validate_qbi_stage_receipt( + frame: Frame, + stage_receipts: Mapping[str, Mapping[str, object]], + *, + route: str, + boundary: str, + transition_authority_sha256: str | None, +) -> None: + receipt = _qbi_receipt_from_stage_receipts( + stage_receipts, + route=route, + boundary=boundary, + ) + validate_us_qbi_reconciliation_live_output( + frame, + receipt, + boundary=boundary, + expected_transition_authority_sha256=transition_authority_sha256, + allowed_post_reconciliation_person_columns=( + us_qbi_post_reconciliation_person_columns(stage_receipts.get("seed")) + ), + ) + + +def _checkpoint_qbi_route(base_identity: Mapping[str, object]) -> str | None: + artifact_kind = base_identity.get("artifact_kind") + if artifact_kind == "populace_us_multispine_pool_checkpoint_identity": + return "legacy" + if artifact_kind == "populace_us_stacked_pool_checkpoint_identity": + return "stacked" + return None + + def _emit_stacked_checkpoint( callback: Callable[[MultispinePoolCheckpoint], None] | None, *, @@ -2429,7 +2605,21 @@ def _emit_stacked_checkpoint( assembly_receipt: Mapping[str, object], stage_receipts: Mapping[str, Mapping[str, object]], simulation_frame: Frame | None = None, + qbi_transition_authority_sha256: str | None = None, ) -> None: + if stage in {"transferred", "simulated"}: + _validate_stacked_post_puf_stage_receipt( + stage_receipts, + boundary=f"stacked {stage} checkpoint emission", + ) + if stage == "simulated": + _validate_qbi_stage_receipt( + frame, + stage_receipts, + route="stacked", + boundary="stacked simulated checkpoint emission", + transition_authority_sha256=qbi_transition_authority_sha256, + ) if callback is None: return callback( @@ -2441,6 +2631,7 @@ def _emit_stacked_checkpoint( name: dict(receipt) for name, receipt in stage_receipts.items() }, simulation_frame=simulation_frame, + qbi_transition_authority_sha256=(qbi_transition_authority_sha256), ) ) @@ -2485,6 +2676,7 @@ def mark_phase(name: str) -> None: ) assembly_receipt = current.metadata[SPINE_ASSEMBLY_MANIFEST_KEY] receipts: dict[str, Mapping[str, object]] = {} + qbi_transition_authority_sha256: str | None = None resume_stage: str | None = None _emit_stacked_checkpoint( checkpoint, @@ -2516,7 +2708,21 @@ def mark_phase(name: str) -> None: receipts = { name: dict(receipt) for name, receipt in resume.stage_receipts.items() } + qbi_transition_authority_sha256 = resume.qbi_transition_authority_sha256 resume_stage = resume.stage + if resume_stage in {"transferred", "simulated"}: + _validate_stacked_post_puf_stage_receipt( + receipts, + boundary=f"stacked {resume_stage} checkpoint resume", + ) + if resume_stage == "simulated": + _validate_qbi_stage_receipt( + current, + receipts, + route="stacked", + boundary="stacked simulated checkpoint resume", + transition_authority_sha256=(qbi_transition_authority_sha256), + ) for completed_phase in ( "assembled", *( @@ -2598,12 +2804,6 @@ def mark_phase(name: str) -> None: primary_qrf_checkpoint_dir=primary_qrf_checkpoint_dir, ) mark_phase("puf_passed") - weights_audit = weights_audit_gate(fit_records) - if not weights_audit.passed: - raise ValueError( - "Stacked imputation weights audit failed:\n " - + "\n ".join(weights_audit.failures) - ) puf_receipt = dict(puf_result.receipt) primary_qrf_receipt = puf_receipt.pop("primary_puf_qrf") @@ -2642,8 +2842,34 @@ def mark_phase(name: str) -> None: source_completion.frame, tail_manifest, ) - current = canonicalize_frame_string_dtypes( + post_puf_target_bank = AcsTransferTargetBankStore( + acs_transfer_checkpoint_dir / "post_puf_transfer", + identity=_stacked_post_puf_bank_identity(checkpoint_identity), + ) + post_puf_transfer = transfer_stacked_post_puf_inputs( source_completion.frame, + seed=POOL_RANDOM_SEED, + n_estimators=_ACS_TRANSFER_N_ESTIMATORS, + max_targets_per_fit=DEFAULT_ACS_TRANSFER_MAX_TARGETS_PER_FIT, + target_bank=post_puf_target_bank, + ) + validate_stacked_post_puf_transfer_receipt( + post_puf_transfer.receipt, + boundary="stacked cold-build post-PUF transfer", + ) + fit_records.extend(post_puf_transfer.transfer_result.fit_records) + weights_audit = weights_audit_gate(fit_records) + if not weights_audit.passed: + raise ValueError( + "Stacked imputation weights audit failed:\n " + + "\n ".join(weights_audit.failures) + ) + post_puf_preservation = assert_stacked_tail_cells_preserved( + post_puf_transfer.frame, + tail_manifest, + ) + current = canonicalize_frame_string_dtypes( + post_puf_transfer.frame, boundary="stacked pool transferred checkpoint", in_place=True, ) @@ -2657,12 +2883,19 @@ def mark_phase(name: str) -> None: "post_primary_completion": dict(source_completion.receipt), }, "stacked_gap_fill": dict(gap_filled.receipt), + "stacked_post_puf_transfer": dict(post_puf_transfer.receipt), "primary_puf_qrf": primary_qrf_receipt, "puf_capital_gains_tail_transfer": dict(tail_manifest), "stacked_puf_pass": puf_receipt, "tail_preservation_after_source_completion": completion_preservation, + "tail_preservation_after_post_puf_transfer": post_puf_preservation, "acs_qrf_transfer": { - "target_families": _json_ready(pool_transfer_target_families()), + "target_families": { + "early_gap_fill": _json_ready(CANONICAL_STACKED_GAP_FILL_SURFACE), + "post_puf_transfer": _json_ready( + CANONICAL_STACKED_POST_PUF_TRANSFER_SURFACE + ), + }, "n_estimators": _ACS_TRANSFER_N_ESTIMATORS, "max_targets_per_fit": DEFAULT_ACS_TRANSFER_MAX_TARGETS_PER_FIT, "target_bank": { @@ -2671,6 +2904,7 @@ def mark_phase(name: str) -> None: name: bank.receipt() for name, bank in sorted(target_banks.items()) }, + "post_puf_transfer": post_puf_target_bank.receipt(), }, }, "weights_audit": GateReport((weights_audit,)).to_manifest(), @@ -2695,6 +2929,7 @@ def mark_phase(name: str) -> None: current ) derived = derive_multispine_pool_inputs(derivation_input) + qbi_transition_authority_sha256 = derived.qbi_transition_authority_sha256 current = canonicalize_frame_string_dtypes( derived.frame, boundary="stacked pool derive output", @@ -2751,6 +2986,7 @@ def mark_phase(name: str) -> None: assembly_receipt=assembly_receipt, stage_receipts=receipts, simulation_frame=simulation_frame, + qbi_transition_authority_sha256=(qbi_transition_authority_sha256), ) mark_phase("simulated") else: @@ -2784,6 +3020,7 @@ def mark_phase(name: str) -> None: stage_receipts=receipts, terminal_gates=(completeness, battery), release_id=release_id, + qbi_transition_authority_sha256=qbi_transition_authority_sha256, ) @@ -2801,6 +3038,13 @@ def _manifest_payload( checkpoint_provenance: Mapping[str, object], publication_run_id: str, ) -> dict[str, object]: + _validate_qbi_stage_receipt( + result.frame, + result.stage_receipts, + route="legacy", + boundary="legacy production manifest", + transition_authority_sha256=(result.qbi_transition_authority_sha256), + ) status = "simulation_ready" if result.simulation_ready else "agreement_failed" puf_donor_receipt = input_receipts.get("puf_donor") if not isinstance(puf_donor_receipt, Mapping): @@ -2888,6 +3132,17 @@ def _stacked_manifest_payload( ) -> dict[str, object]: """Build the stacked-only manifest without changing the legacy envelope.""" + _validate_stacked_post_puf_stage_receipt( + result.stage_receipts, + boundary="stacked production manifest", + ) + _validate_qbi_stage_receipt( + result.frame, + result.stage_receipts, + route="stacked", + boundary="stacked production manifest", + transition_authority_sha256=(result.qbi_transition_authority_sha256), + ) status = "simulation_ready" if result.simulation_ready else "gate_failed" puf_donor_receipt = input_receipts.get("puf_donor") if not isinstance(puf_donor_receipt, Mapping): @@ -2909,6 +3164,7 @@ def _stacked_manifest_payload( "gap_fill_stacked_spine", "run_stacked_puf_pass", "complete_multispine_source_inputs", + "transfer_stacked_post_puf_inputs", "prepare_stacked_tail_derivation", "derive_multispine_pool_inputs", "seed_multispine_pool_inputs", @@ -3066,6 +3322,13 @@ def _write_outputs( input_receipts: Mapping[str, object] | None = None, checkpoint_provenance: Mapping[str, object] | None = None, ) -> None: + _validate_qbi_stage_receipt( + result.frame, + result.stage_receipts, + route="legacy", + boundary="legacy publication entry", + transition_authority_sha256=(result.qbi_transition_authority_sha256), + ) if input_receipts is None: if loaded is None: raise ValueError( @@ -3148,6 +3411,17 @@ def _write_stacked_outputs( ) -> dict[str, object]: """Atomically publish the stacked input-only pool and terminal receipts.""" + _validate_stacked_post_puf_stage_receipt( + result.stage_receipts, + boundary="stacked publication entry", + ) + _validate_qbi_stage_receipt( + result.frame, + result.stage_receipts, + route="stacked", + boundary="stacked publication entry", + transition_authority_sha256=(result.qbi_transition_authority_sha256), + ) publication_run_id = _new_publication_run_id() _atomic_write_json( outputs.manifest,