Skip to content

Defer exact classifier residual deviation until needed - #20

Merged
fitz2882 merged 1 commit into
mainfrom
codex/exact-lazy-classifier
Sep 30, 2026
Merged

fitz2882 merged 1 commit into
mainfrom
codex/exact-lazy-classifier

Conversation

@fitz2882

@fitz2882 fitz2882 commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

When a trajectory already meets an earlier classification gate, the classifier currently computes regression and residual population deviation that cannot affect the verdict. This change evaluates those gates first and defers residual deviation until the oscillation gate requires it. Public extract_features() remains eager and keeps the exact statistics.pstdev calculation and rounding.

The two-point ratio gate, liveness/first-tie semantics, custom threshold boundaries, large integers, and non-finite exception behavior are covered by 27 new regression tests. Only the classifier and those tests change; no thresholds, dependencies, version, or public APIs change.

Validation

  • Integrated against current main at 8c3eec4bd2d55cfa271c281c40638abd2290dd19, the same baseline used by the accepted benchmark. Both modified files are byte-identical to the validated candidate.
  • Python 3.12.8 full suite: 262 passed / 6 optional real-framework module skips, including all six opt-in container tests. Unmodified base: 229 passed / 12 optional skips. The opt-in sandbox suite separately passed all 14 tests using the existing local image, with no provider calls.
  • Rerun exact differential audit against the current unmodified base: 1,778 records, zero differences. Finite floats compared bit-exact; NaNs semantically; states, traces, result fields, and exception types/messages exactly. Includes numeric extremes and baseline-derived boundary thresholds. This reuses the now-exposed frozen audit corpus and is a regression check, not fresh unseen evidence.
  • Ruff and Bandit passed; sdist/wheel build and Twine metadata checks passed; fresh-venv installed-wheel import, behavior, and CLI smoke passed.
  • New fixed three-block serial balanced integration timing smoke: composite runtime reduction 47.64%; feature extraction 1.23% slower; legacy 1.11% faster; traced allocation 3.63% lower. All raw samples retained. Descriptive only, separate from the acceptance campaign, with no new promotion confidence claim.

Original benchmark evidence and limits

The separately versioned accepted Mac campaign completed 24 paired blocks plus fresh bracketing A/A controls: composite runtime reduction 46.628%; conservative simultaneous A/A bias-adjusted lower improvement bound 45.280%, exceeding the unchanged 2% minimum. All timing/allocation guards passed the unchanged 5% regression margin. The tightest adjusted guard was full-feature extraction at -4.443%, a narrow 0.557 percentage-point margin. No approximate variance math or candidate tuning was accepted.

Synthetic GC-controlled CPU overhead does not establish production latency, answer quality, or provider spending savings. Supported-Python/framework coverage is supplied by repository CI; original Mac acceptance was Python 3.12.8. Cross-environment baseline p-value drift, the preserved failed initial A/A prerequisite, the reviewed pre-candidate protocol correction, and durable-completion recovery after disconnects are disclosed in the original evidence. Cloud and Mac results are not pooled; the old incomplete cloud confirmation remains invalid.

Original evidence: LoopGain-Mac-complete-evidence.zip, Library identity libfile_ba56c3e69df48191965c17716d5e08a2, version 3, SHA-256 1f3e07f33be573657ce2631d4cb8f2d7d5108379698dfce050abd85341120901. The interactive report is LoopGain-cloud-and-Mac.html, Library identity libfile_c86395b976e481918c57c3d202292373, version 2. Original artifacts and integration logs remain preserved separately. Draft for review; no merge or deployment performed.

Terminal CI result

CI run 36765995379 completed successfully for commit c25cac78e0a84dfda5c6084fb6e1ee0a06330768; all 11 jobs passed. Python 3.10, 3.11, 3.12, and 3.13 each passed 256 tests with 12 documented optional skips. All six framework jobs executed with zero skips: LangGraph 3, CrewAI 4, AutoGen 3, LangChain 4, OpenAI Agents 4 (including the offline real Runner), and Claude Agent SDK 4 tests. Lint/security/build/metadata/installed-wheel CI passed. All six opt-in container tests passed locally in the 262-test full suite.

@fitz2882
fitz2882 marked this pull request as ready for review September 30, 2026 19:58
@fitz2882
fitz2882 merged commit 94882c6 into main Sep 30, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant