Skip to content

Defer to guessed starting values no more than the prior does, and report the belief honestly - #220

Open
yichao-liang wants to merge 3 commits into
masterfrom
fit-guess-anchors
Open

yichao-liang wants to merge 3 commits into
masterfrom
fit-guess-anchors

Conversation

@yichao-liang

Copy link
Copy Markdown
Collaborator

Why

Seed 0 of the Domino round fixes_r3 lost its test level after it stopped trusting its fitted belief, for two reasons this PR addresses (#219 fixed a third, refits drifting with the last fit's answer).

  1. The fit deferred to the agent's guesses beyond its prior.
    The grid sweep counts a candidate within 5% of the model-bias SSE as data-equivalent and keeps the one nearest the anchor, and the anchor ablation pins moved parameters back to their anchors.
    That suits calibrated baselines (a supplied domain base).
    Under code_sim_learning_prior_spans_bounds the anchors are the agent's guesses: on seed 0's level-1 recording the guessed spinning friction 0.005 stayed put although its best candidate, 0.67 with the world at 0.5, fit 2.8 nats better (5.6 likelihood floors).
  2. The report contradicted itself.
    It printed "most likely " next to the central 68% interval of the belief factor, which can exclude both the point estimate and a mode at a bound: seed 0 read "most likely 0.05; 68% posterior interval [0.2694, 0.291]".

What changes

  • SysIdConfig.anchors_are_guesses (from code_sim_learning_prior_spans_bounds): the flat band drops its relative term (flat_band_frac), so data-equivalence is the likelihood floor the belief scores with, and the anchor ablation does not run. Supplied-base arms keep both.
  • The belief report gives each factor's mode and its 68% highest-density interval, which contains the mode, and names the fit's point estimate when it lies outside that interval; the note under the report says the planning model runs at the point estimate.
  • test_run_rollout_sysid_fit_cache_and_report_isolation starts from the default config: test_continual_joint_belief leaves belief_joint_draws set, which sent the fit down the joint-belief path whenever the two tests ran in that order.

Evidence

The scripted from-assets harness fits seed 0's level-1 recording (7 segments) with and without the change; the world's values are known.

  • Fit time: 230 s with the change, 682 to 907 s without (2,403 s before Make the rollout fit about 3 times faster without changing its results #218); the anchor ablation, 6,293 of 7,735 replays, no longer runs.
  • Replay error: 7.139 with the change, 7.187 without.
  • Domino lateral friction 0.388 (world 0.5, was 0.954), spinning friction 0.165 (0.5, was 0.005), rolling friction 0.0067 (0.006, was 0.0005); table rolling friction 0.0010 (0.001, was 0.0025).
  • Domino restitution moved away from the world: 0.245 (world 0.02, was the guessed 0.05).

Tests

  • New: a calibrated anchor keeps its value when a better candidate lies inside the relative band, a guessed one does not; the ablation runs for calibrated anchors only; the highest-density interval holds a mode at a bound that the central interval excludes; the report names a point estimate outside the belief.
  • Local CI replay: the static checks and the 8 shards.

🤖 Generated with Claude Code

The grid sweep treats a candidate within 5% of the model-bias SSE as
data-equivalent and keeps the one nearest the anchor, and the anchor
ablation pins moved parameters back to their anchors. Both defer to the
anchor beyond the fit's prior, which a calibrated baseline earns. Under
code_sim_learning_prior_spans_bounds the anchors are the agent's
guesses: seed 0 of the Domino round fixes_r3 kept a guessed spinning
friction of 0.005 although its best candidate, 0.67 with the world at
0.5, fit 2.8 nats better. The deployed value then sat outside the
belief's own 68% interval of 0.39 to 0.71, the report contradicted
itself, and the agent replayed its plan at the guess.

With guessed starting values the flat band is the likelihood floor
alone, the scale the parameter belief scores with, and the anchor
ablation does not run. Supplied-base arms keep both.
test_run_rollout_sysid_fit_cache_and_report_isolation checks the trust
selection, which runs only without the joint belief, but read whatever
config the previous test left. After test_continual_joint_belief, which
sets belief_joint_draws, the fit took the joint-belief path and applied
its point estimate (gain 1.9993 instead of the trusted 1.5), so the test
failed whenever the two landed in that order.
The fit report printed each parameter as "most likely <point estimate>"
next to the central 68% interval of its belief factor. The point
estimate need not be the factor's mode, and a central interval excludes
a mode at a bound, so the report could contradict itself: seed 0 of the
Domino round fixes_r3 read "most likely 0.05; 68% posterior interval
[0.2694, 0.291]" and stopped trusting the belief.

Each line now gives the factor's mode and its 68% highest-density
interval, which contains the mode, and names the fit's point estimate
when it lies outside that interval. The note under the report says the
planning model runs at the point estimate.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant