Repository navigation
Keep the from-assets arm's untested physics uncertain - #214
Merged
Merged
Conversation
added 2 commits
October 9, 2026 03:25
Seed 1 of the Domino fixes_r1 round lost its test level on two engine materials its scene never declared: the dominoes' spinning and rolling friction ran at PyBullet's 0 (the world: 0.5 and 0.006), and every rehearsal predicted a cascade that the world did not produce. Its level-1 recording, a straight chain, could not identify either value, and declaring them would not have helped: the fit kept them at their starting values with a narrow belief, because its prior is 0.75 times the starting value (a spinning friction started at 0.05 came out 0.057-0.12). - code_sim_learning_prior_spans_bounds widens the fit's prior on every parameter to at least its declared range, so a parameter the recordings do not constrain keeps that whole range in the belief. The from_assets_opus menu entry sets it, and the arm refuses to start without it; the supplied-base arms keep the anchored prior, whose starting values are calibrated defaults. - SceneBase.sampled_material_specs offers every material the scene does not declare as a parameter over a plausible range around the scene's own value (SAMPLED_MATERIALS), and accepts a value for it, as a rehearsal draw sets one. The from-assets arm joins these prior-only factors to its belief (join_beliefs), so joint draws and physics sweeps vary them; they are never fitted and cost no replays. - The from-assets prompts say so: undeclared materials are sampled, and a declared parameter's prior spans its declared range.
The fit pins every parameter it does not estimate to its default. For an undeclared material that default is read from its group's first body, so applying it set every body of the group to that value: the fit worlds of seed 1's Domino scene ran its tables at the ground plane's friction, and the fit started from three times the error on the same data. A default now restores each body's own value; a rehearsal draw's value still sets the whole group.
yichao-liang
force-pushed
the
from-assets-honest-belief
branch
from
October 9, 2026 07:25
b32d3e8 to
d753bd3
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #210 (from-assets-prompt-fixes); review that one first.
Why
Seed 1 of the Domino
fixes_r1round (EMPIRIC from assets) lost its test level.Replaying its recorded test episode in the real environment shows the failing push needs the world's spinning friction (0.5) and rolling friction (0.006) on the dominoes, together with the gripper's follow-through touching the first bridge.
The agent never declared either friction, so its model ran both at PyBullet's 0 and predicted a cascade in every rehearsal.
Its level-1 recording, a straight chain, could not identify either value.
Declaring them would not have helped yet: the fit kept a parameter the data did not constrain at its starting value with a narrow belief, because its prior is 0.75 times the starting value.
A spinning friction started at 0.05 came out with a 68% interval of 0.057 to 0.12.
What changes
Under
code_sim_learning_prior_spans_bounds, no parameter's prior is narrower than its declared range in fit space, so a parameter the recordings do not constrain keeps that whole range in the belief.The
from_assets_opusmenu entry sets it, and the arm refuses to start without it.The supplied-base arms keep the anchored prior, since their starting values are the base's calibrated defaults.
SceneBase.sampled_material_specsoffers every engine material the scene does not declare as a parameter over a plausible range around the scene's own value (SAMPLED_MATERIALS: spinning friction 0 to 1, rolling friction 0 to 0.02, mass within a factor of 3, ...).The from-assets arm joins these prior-only factors to its belief (
join_beliefs), so joint draws and physics sweeps vary them.They are never fitted and cost no replays.
A rollout world applies a draw's value to the whole group; a fit that pins the material to its default gives each body its own value back, so a group the scene made different (the ground and the tables) stays different.
The from-assets system prompt and scene contract say that undeclared materials are sampled, how to fit or fix one, and that a declared parameter's prior spans its declared range.
docs/amps/empiric-from-assets.mdrecords the design.Evidence
A scripted run of the from-assets harness on seed 1's round: level 1 is seed 1's own
simulator.pyand recorded actions.On level 2 it fits, then rehearses two plans from the level's start on 16 joint draws each: seed 1's chain (3 blues, the 46-degree first hit, lost in the world) and the oracle-dynamics arm's arc (2 blues, won in the world).
The sweep now names the two materials the chain hangs on.
The draws do not yet separate the two plans: the agent's own scene still predicts that seed 1's chain cascades at the world's material values, where the world's chain slides.
Real-environment replays rule out the materials, friction anchors, continuous collision detection, damping and the dominoes' massless top link as that difference.
The likely remaining cause is the push: the model replans it from its noisy estimate of the scene, and whether the gripper's follow-through catches a bridge 8.7 cm away decides the outcome.
Domino round
fixes_r2ran this head on 2 seeds and solved 1 of 2.Seed 0 lost its test level because it declared both frictions, with ranges that exclude the world's values (spinning friction 0 to 0.05, rolling friction 0 to 0.002; the world has 0.5 and 0.006).
Neither change here reaches a declared range, so its 16 joint draws all cascaded.
A real-environment replay of its episode stalls at the world's values and cascades with either friction anywhere inside its ranges.
Covering declared materials whose range the recordings do not test is a follow-up.
The first acceptance run also caught a bug in the first version of the sampling: pinning a material's default homogenized the support group to the ground plane's friction, and the fit started at three times the error on the same data.
The per-body restore fixes it.
Tests
Shard 3's one failure (
test_preflight_audit.py::test_shadow_tools_pair_outcomes_without_changing_execution, "Not connected to physics server") is the registry flake that Resolve a registry name to its own class, never a subclass that inherits it #213 fixes; master has it too.This head with Resolve a registry name to its own class, never a subclass that inherits it #213 applied passes the static checks and all 8 shards.
static-type-checkingfails on every branch until Keep mypy passing on tenacity 9.2's typed retry decorator #212 (tenacity 9.2.1) merges.🤖 Generated with Claude Code