Repository navigation
Act on a from-assets test level only with a fitted model - #217
Merged
Merged
Conversation
yichao-liang
force-pushed
the
from-assets-fit-gate
branch
from
October 9, 2026 08:38
b10513a to
dde8369
Compare
yichao-liang
force-pushed
the
from-assets-plausible-ranges
branch
from
October 9, 2026 08:38
a5e88e7 to
038db71
Compare
yichao-liang
force-pushed
the
from-assets-plausible-ranges
branch
from
October 10, 2026 13:36
038db71 to
950cd14
Compare
yichao-liang
force-pushed
the
from-assets-fit-gate
branch
from
October 10, 2026 13:36
dde8369 to
247982e
Compare
yichao-liang
changed the base branch from
from-assets-plausible-ranges
to
master
October 10, 2026 14:07
Seed 0 of the Domino round fixes_r2 found its level-1 sim.fit() slow and uninformative (39 minutes, 0.8% lower error), tuned its values by hand and played the test level with an unfitted model: the belief that chose its plan was the prior around those hand-written values. The test-level gate (continual_require_model_on_test) now asks the arm through _fit_readiness once a model loads. EMPIRIC from assets refuses skills_invoke and skills_execute_plan on a test level until sim.fit() has run on the current content of simulator.py, when the model declares parameters; every edit needs a new fit. The supplied-base arms leave fitting to the agent as before, and real-to-sim fits nothing. The system prompt's gate section says which rule applies (model_gate_fit).
yichao-liang
force-pushed
the
from-assets-fit-gate
branch
from
October 10, 2026 14:09
247982e to
c428e1b
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Seed 0 of the Domino round
fixes_r2found its level-1sim.fit()slow and uninformative (39 minutes for a 0.8% lower error), tuned its parameter values by hand from roll curves, and played the test level with an unfitted model.The belief that chose its plan was therefore the prior around those hand-written values, not anything the recordings support.
The test-level gate (
continual_require_model_on_test) only asked for a loadable model withRESIDUAL_FEATURES.#218 makes that fit about 3 times faster (from about 40 minutes to about 12 on seed 0's recording) with identical results, so asking for a fit after each edit is affordable.
What changes
_model_readinessasks the arm through_fit_readinessonce a model with parameters loads.skills_invokeandskills_execute_planon a test level untilsim.fit()has run on the current content ofsimulator.py; every edit needs a new fit.A fit that could not run (nothing explainable) still counts: the agent asked, and the belief then stays the prior.
model_gate_fit).Tests
sim.fit(), opens after the fit, and refuses again after an edit; real-to-sim needs only a loadable model; each arm's system prompt states its gate.🤖 Generated with Claude Code