Skip to content

Count a guessed fit's misfit once in the belief's temperature - #222

Open
yichao-liang wants to merge 1 commit into
fit-multistartfrom
belief-coherent-misfit
Open

yichao-liang wants to merge 1 commit into
fit-multistartfrom
belief-coherent-misfit

Conversation

@yichao-liang

Copy link
Copy Markdown
Collaborator

Summary

  • The parameter belief's temperature was lambda = max(1, SSE_min / (N sigma_n^2)), which spreads the fit's misfit over the N residuals as independent errors.
    When the fit has measured its misfit (the SSE its best fit leaves above the declared noise's expected SSE), the temperature now counts it as one error that persists through the recordings: lambda >= misfit / sigma_n^2.
    This is the scale that already weighs the multi-start's basins, so every basin's factor and the basin weights share one target.
  • The misfit is measured by the multi-start, or at the local fit when the multi-start is off; only fits from guessed starting values measure it, so every other arm is unchanged.
    Without a declared noise channel the misfit is unknown: the multi-start keeps one basin and the temperature is unchanged.
  • The fit report says what the temperature means: parameter values whose SSE is within about the misfit of the best fit's are about as likely.
  • docs/uncertainty/principled-belief.md step 1 describes the temperature.

Why

The spec already named the gap: the temperature "does not correct for errors that persist through a segment".
On Domino seed 1 of fixes_r4 the best fit sits 0.084 above the noise's expected SSE of 4.27, and the temperature stayed at 1.
A replay that drifts off its recording, or a contact that resolves differently, moves many residuals together, and a parameter change can trade that error for a lower SSE.
At temperature 1 the lines read SSE jitter of 0.01 between neighbouring probes as 2 nats of evidence.
The lines through the multi-start's basins peaked far from the basins themselves: a basin at spinning friction 0.95 traced to a mode at 0.
The world's own materials, 0.10 worse than the best fit, sat 20 nats below it.

Rehearsing seed 1's losing plan and the oracle-dynamics arm's winning plan on 16 joint draws, after replaying seed 1's level 1 and fitting:

belief seed 1's plan (lost) oracle's plan (won)
point mass at the world's materials 0.12 0.62
fixes_r4 fit 0.75 0.69
multi-start only (base PR) 0.56 0.62
multi-start + this PR (temperature 33.6) 0.25 0.62
this PR without the multi-start (temperature 84) 0.06 0.19

With both changes the belief separates the plans as the world's materials do.
Without the multi-start the belief is wide around the guesses' compensating basin, and no plan rehearses as worth executing, so the two changes go together.

Test plan

  • test_parameter_belief.py: a measured misfit sets the temperature to misfit / sigma_n^2 (widths double against the per-residual estimate for the same 27 sigma_n^2), survives mixing, joining and checkpoints, and a misfit within the noise leaves the ordinary posterior.
  • test_physical_sysid.py: the multi-start returns the misfit it weighs the basins with; no declared noise means one basin and no misfit; with the multi-start off a guessed fit measures its misfit at the local fit.
  • test_orchestrator.py: every basin's factor takes the misfit's temperature.
  • End to end on seed 1's recorded level 1, as in the table.
  • Local CI replay on this commit: yapf, isort 5.10.1, docformatter, mypy, pylint and the 8 pytest shards pass.

🤖 Generated with Claude Code

The belief's temperature spread the fit's misfit over every residual as
independent errors, so on Domino seed 1 of fixes_r4 (a best fit 0.084
above the noise's expected SSE of 4.27) it stayed at 1. Each line then
read SSE differences of 0.01, the size of the jitter a contact that
resolves differently leaves, as 2 nats of evidence. The lines through
the multi-start's basins peaked far from the basins themselves (a basin
at spinning friction 0.95 traced to a mode at 0), and the world's own
materials, 0.10 worse than the best fit, sat 20 nats below it.

Replay errors persist through a recording, and a parameter change can
trade one for a lower SSE, so SSE differences up to the misfit are no
evidence. When the multi-start has measured the misfit, the temperature
is at least misfit / sigma_n^2: the scale that already weighs the basins.
Calibrated starting values keep the per-residual estimate.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant