Repository navigation
fix(sleep): close five ways invalid evidence reaches the gate - #310
Open
Bogdan (Dan) Baciu (bogdanbaciu21) wants to merge 5 commits into
Open
Bogdan (Dan) Baciu (bogdanbaciu21) wants to merge 5 commits into
Bogdan (Dan) Baciu (bogdanbaciu21) wants to merge 5 commits into
Conversation
recall_similar blocked tonight's held-out tasks by id only. Ids hash the project with the intent and the recall archive is shared across projects, so another project's copy of tonight's val task was recalled into training under a different id. The leak check in _split was id-based too, so the gate certified the result: in a two-project cycle the verdict flipped from reject to accept_new_best with holdout_leaked=False. - task_content_key(): normalized intent plus context excerpt - recall_similar(exclude_tasks=...) blocks archived tasks whose content key matches a held-out task; dream_consolidate passes tonight's val/test tasks - _split treats a val task with a content twin in train as leaked, so the existing reject_unverified abstention applies on any route Follows up the content-hash hardening listed as a non-goal in microsoft#235. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…l route Tool-loop backends detect real calls (for example from the search shim's call log) and return an empty tools_called when the agent never ran the tool. The rule judge then fell back to a TOOL_CALL marker regex over the response text and passed the check anyway, undoing the verification one step later: a Claude replay whose shim log was empty scored 1.0. - _check / score_rule_judge(_with_feedback) take verified_tools; when set, only measured calls satisfy tool_called - replay_one sets it for tool tasks, which always go through attempt_with_tools; the inherited marker fallback still converts markers into calls there, so single-shot backends are unchanged - optimizer feedback for a self-reported call asks for a real call - unmeasured single-shot judging keeps the documented marker approximation Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The judge prompt asks for a 0..1 score, but CliBackend.judge accepted any float, including NaN and Infinity (json.loads parses both) and replies on another scale. On the default mixed gate metric, one out-of-range soft score let a candidate with flat hard accuracy and a 1.0 -> 0.0 regression read as 0.625 -> 0.750 and be accepted. Scores must now be finite and within [0, 1]; anything else fails closed as judge-score-out-of-range with the raw value in the rationale. Clamping was rejected because it would turn an 8/10 reply into a perfect 1.0. Boolean scores fall through to judge-parse-failed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The learned block is read back one "- " bullet per item, but set_learned
wrote items verbatim. A multi-line item lost every continuation line the
next time any edit was applied; a "- X" add slipped past the duplicate check
for an existing "X"; and lstrip('- ') removed leading dashes such as
"--force". Losing adopted text changes what the gate scores: a helpful edit
was rejected (0.5 vs 0.5) because applying it silently deleted an adopted
rule held on a continuation line.
- _learned_item(): one item is one line (whitespace collapsed) with exactly
one leading bullet marker removed; used for add, replace and set_learned
- current_learned_lines joins continuation lines of blocks written by
earlier versions onto their item instead of dropping them
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
_is_headless_replay recognised the engine's own sessions by static prompt markers written for earlier prompt wording. None of them match the current attempt, judge, reflect or miner templates, and the duration fallback only covers prompts under 200 characters, so with projects="all" every engine session was harvested and could be mined back as a user task. - markers now include the opening line of every registry template, default and active override, so rewording a prompt cannot reopen the gap - a shared phrase of the Claude, Codex, Copilot and OpenCode tool-attempt prompts is matched case-insensitively - the existing static markers stay for transcripts from older versions; the Codex harvester (microsoft#286) is left to its own fix Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What Problem This Solves
The nightly gate is only as trustworthy as the evidence it scores. This PR closes five independent paths by which invalid evidence reaches it on current
main. Each was found by executing a documented claim and watching it fail, and each is reproduced by a regression test that fails onmainand passes here.maindream.py)_split's leak check is id-based too.rejecttoaccept_new_bestwithholdout_leaked=Falsebackend.py,replay.py)tools_called=[], then the rule judge falls back to aTOOL_CALL:regex over the response and passes anywaytool_calledprompts.py)CliBackend.judgeaccepts any float, includingNaN/Infinityand replies on another scalemixedmetric, a candidate with flat hard accuracy and a 1.0 → 0.0 regression reads 0.625 → 0.750 and is acceptedaddthat already exists "is skipped" and adopted rules are kept (memory.py)-line per item: a multi-line item loses its continuation lines on the next edit,- Xpasses the duplicate check forX, andlstrip('- ')eats--forceharvest.py)projects: "all", every engine session is harvested and can be mined back as a user taskWhy This Change Was Made
These are correctness fixes to evidence the project already relies on, not new behavior. Each commit is self-contained and can be reviewed or reverted on its own:
task_content_key()(normalized intent plus context excerpt) gives recall a content check alongside the id check.dream_consolidatepasses tonight's val/test tasks asexclude_tasks, and_splittreats a val task with a content twin in train as leaked, so the existingreject_unverifiedabstention applies on any route. This follows up the content-hash hardening that fix(sleep): honor val_fraction and test_fraction in the nightly cycle #235 listed as a non-goal.score_rule_judge(_with_feedback)takesverified_tools.replay_onesets it for tool tasks, which always go throughattempt_with_tools, and only measured calls then satisfytool_called. The inheritedattempt_with_toolsalready converts markers into calls for backends without a tool loop, so single-shot backends are unchanged. Unmeasured judging keeps the documented marker approximation.[0, 1]; anything else fails closed asjudge-score-out-of-rangewith the raw value in the rationale. Clamping was rejected because it would turn an 8/10 reply into a perfect 1.0.add,replace, andset_learned.current_learned_linesjoins continuation lines of blocks written by earlier versions onto their item instead of dropping them, so already-adopted text survives.Project Fit
User Impact
Operators get fewer certified-but-wrong verdicts and clearer diagnostics:
holdout_leakednow catches content twins, a self-reported tool call reads asfailed: tool_called=search, and an off-scale judge reply reads asjudge-score-out-of-range: 8. Adopted multi-line rules stop disappearing between nights.Proof
All tests are deterministic and offline. Provider boundaries are faked (
subprocess.runor a scripted_call), orMockBackendis used. 44 new tests run against both trees:main(343db22)_split, gate abstention, two-project cycle); 1 passesharvest(scope="all")); 5 controls pass (4 real user prompts kept)The controls that pass on both trees are deliberate. Measured calls still pass. The default marker route still converts markers for single-shot backends. In-range judge scores are unchanged. Real user prompts that resemble engine wording are still harvested.
Testing
Current head
964d9c050935d57b73999901961d6657f25a5788Fork-runner validation asserts
actual_sha == expected_sha == 964d9c050935d57b73999901961d6657f25a5788before testing. These are contributor-run checks, not official upstream CI.tests/test_gate.py)The six Windows failures are
test_home_is_refused_even_when_it_is_a_git_rootand five subcases oftest_cycle_stages_only_documents_that_changed. The same Windows job on unmodifiedmain(343db22) fails the identical six (6 failed, 1,663 passed). The 44-test difference is exactly the new tests, and none of the six touch code changed here. They are reported, not fixed, to keep this PR scoped.Local macOS x86_64 / Python 3.12: the complete suite reported 1,759 passed, 11 skipped, 359 subtests; the five new files reported 44 passed; strict docs and
git diff --checkpassed; ruff findings on the touched modules are identical tomain.This branch and #263 also merge cleanly with each other in both orders, and the combined tree reports 1,821 passed, 11 skipped, 359 subtests.
No paid provider was called. These are contributor-run checks, not official upstream CI.
Limitations & Negative Results
_is_headless_replay. The Codex harvester is skillopt-sleep Codex harvest ingests its own headless replay sessions #286's scope and is untouched.--no-session-persistencewas not added toclaude -p, because older CLI versions would reject the flag.Reproduce It Yourself
Linux or macOS, shell in a fresh working directory: