Repository navigation
Defer to guessed starting values no more than the prior does, and report the belief honestly - #220
Open
yichao-liang wants to merge 3 commits into
Open
yichao-liang wants to merge 3 commits into
yichao-liang wants to merge 3 commits into
Conversation
The grid sweep treats a candidate within 5% of the model-bias SSE as data-equivalent and keeps the one nearest the anchor, and the anchor ablation pins moved parameters back to their anchors. Both defer to the anchor beyond the fit's prior, which a calibrated baseline earns. Under code_sim_learning_prior_spans_bounds the anchors are the agent's guesses: seed 0 of the Domino round fixes_r3 kept a guessed spinning friction of 0.005 although its best candidate, 0.67 with the world at 0.5, fit 2.8 nats better. The deployed value then sat outside the belief's own 68% interval of 0.39 to 0.71, the report contradicted itself, and the agent replayed its plan at the guess. With guessed starting values the flat band is the likelihood floor alone, the scale the parameter belief scores with, and the anchor ablation does not run. Supplied-base arms keep both.
test_run_rollout_sysid_fit_cache_and_report_isolation checks the trust selection, which runs only without the joint belief, but read whatever config the previous test left. After test_continual_joint_belief, which sets belief_joint_draws, the fit took the joint-belief path and applied its point estimate (gain 1.9993 instead of the trusted 1.5), so the test failed whenever the two landed in that order.
The fit report printed each parameter as "most likely <point estimate>" next to the central 68% interval of its belief factor. The point estimate need not be the factor's mode, and a central interval excludes a mode at a bound, so the report could contradict itself: seed 0 of the Domino round fixes_r3 read "most likely 0.05; 68% posterior interval [0.2694, 0.291]" and stopped trusting the belief. Each line now gives the factor's mode and its 68% highest-density interval, which contains the mode, and names the fit's point estimate when it lies outside that interval. The note under the report says the planning model runs at the point estimate.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Seed 0 of the Domino round
fixes_r3lost its test level after it stopped trusting its fitted belief, for two reasons this PR addresses (#219 fixed a third, refits drifting with the last fit's answer).The grid sweep counts a candidate within 5% of the model-bias SSE as data-equivalent and keeps the one nearest the anchor, and the anchor ablation pins moved parameters back to their anchors.
That suits calibrated baselines (a supplied domain base).
Under
code_sim_learning_prior_spans_boundsthe anchors are the agent's guesses: on seed 0's level-1 recording the guessed spinning friction 0.005 stayed put although its best candidate, 0.67 with the world at 0.5, fit 2.8 nats better (5.6 likelihood floors).It printed "most likely " next to the central 68% interval of the belief factor, which can exclude both the point estimate and a mode at a bound: seed 0 read "most likely 0.05; 68% posterior interval [0.2694, 0.291]".
What changes
SysIdConfig.anchors_are_guesses(fromcode_sim_learning_prior_spans_bounds): the flat band drops its relative term (flat_band_frac), so data-equivalence is the likelihood floor the belief scores with, and the anchor ablation does not run. Supplied-base arms keep both.test_run_rollout_sysid_fit_cache_and_report_isolationstarts from the default config:test_continual_joint_beliefleavesbelief_joint_drawsset, which sent the fit down the joint-belief path whenever the two tests ran in that order.Evidence
The scripted from-assets harness fits seed 0's level-1 recording (7 segments) with and without the change; the world's values are known.
Tests
🤖 Generated with Claude Code