Skip to content

Act on a from-assets test level only with a fitted model - #217

Merged
yichao-liang merged 1 commit into
masterfrom
from-assets-fit-gate
Oct 10, 2026
Merged

yichao-liang merged 1 commit into
masterfrom
from-assets-fit-gate

Conversation

@yichao-liang

@yichao-liang yichao-liang commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Why

Seed 0 of the Domino round fixes_r2 found its level-1 sim.fit() slow and uninformative (39 minutes for a 0.8% lower error), tuned its parameter values by hand from roll curves, and played the test level with an unfitted model.
The belief that chose its plan was therefore the prior around those hand-written values, not anything the recordings support.
The test-level gate (continual_require_model_on_test) only asked for a loadable model with RESIDUAL_FEATURES.
#218 makes that fit about 3 times faster (from about 40 minutes to about 12 on seed 0's recording) with identical results, so asking for a fit after each edit is affordable.

What changes

  • _model_readiness asks the arm through _fit_readiness once a model with parameters loads.
  • EMPIRIC from assets refuses skills_invoke and skills_execute_plan on a test level until sim.fit() has run on the current content of simulator.py; every edit needs a new fit.
    A fit that could not run (nothing explainable) still counts: the agent asked, and the belief then stays the prior.
  • The supplied-base arms leave fitting to the agent as before, and real-to-sim fits nothing.
  • The system prompt's gate section says which rule applies (model_gate_fit).

Tests

🤖 Generated with Claude Code

@yichao-liang
yichao-liang force-pushed the from-assets-plausible-ranges branch from a5e88e7 to 038db71 Compare October 9, 2026 08:38
@yichao-liang
yichao-liang force-pushed the from-assets-plausible-ranges branch from 038db71 to 950cd14 Compare October 10, 2026 13:36
@yichao-liang
yichao-liang changed the base branch from from-assets-plausible-ranges to master October 10, 2026 14:07
Seed 0 of the Domino round fixes_r2 found its level-1 sim.fit() slow
and uninformative (39 minutes, 0.8% lower error), tuned its values by
hand and played the test level with an unfitted model: the belief that
chose its plan was the prior around those hand-written values.

The test-level gate (continual_require_model_on_test) now asks the arm
through _fit_readiness once a model loads. EMPIRIC from assets refuses
skills_invoke and skills_execute_plan on a test level until sim.fit()
has run on the current content of simulator.py, when the model declares
parameters; every edit needs a new fit. The supplied-base arms leave
fitting to the agent as before, and real-to-sim fits nothing. The
system prompt's gate section says which rule applies (model_gate_fit).
@yichao-liang
yichao-liang merged commit 9a44418 into master Oct 10, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant