Skip to content

Make the rollout fit about 3 times faster without changing its results - #218

Merged
yichao-liang merged 3 commits into
masterfrom
fit-batch-evaluation
Oct 10, 2026
Merged

yichao-liang merged 3 commits into
masterfrom
fit-batch-evaluation

Conversation

@yichao-liang

@yichao-liang yichao-liang commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Why

Seed 0 of the Domino round fixes_r2 called sim.fit() once on its level-1 recording.
The fit took 39 minutes and lowered the error by 0.8%, so the agent tuned its parameters by hand and played the test level unfitted.
#217 makes a fit a precondition for acting on a from-assets test level, so the fit has to be cheap enough to call after every model edit.

Where the time went

A rollout fit replays recorded segments (7 of 22 to 99 steps for seed 0's level 1) at thousands of parameter points.

  • Every point forked one wave of children, one per segment, and waited for its slowest segment before the next point started, so most of the 6 workers sat idle; the explainability sweep scored its segments serially.
  • Every replay built a fresh world: 0.34 to 0.39 s, four loadURDF calls and two loadTexture calls, against 0.11 s of stepping for a 22-step segment.
  • A wave started its jobs in segment order, so the 99-step replay waited for a worker behind short ones, and the parent noticed each result up to 50 ms late.
  • Stepping itself is 93% stepSimulation, which no change here touches.

What changes

  1. Score independent rollouts together (trajectory_terms_by_point): every (point, segment) pair goes to forked children in one wave.
    The explainability sweep and the grid pools score their candidates through it.
    solve_lm takes a batch scorer and computes least_squares' own 2-point Jacobian with its column points scored together: approx_derivative runs once to record the points and once on their cached residuals.
    The belief traces its lines in lockstep, one batch per round.
  2. Replay forked rollouts on copies of one pre-built world (fork_template): a fit builds one world it never steps, and each forked child replays on its copy-on-write copy of it.
    A copy of a never-stepped world is a fresh world, so determinism holds; a rollout the parent runs itself still builds a fresh world, and nothing is built when waves do not fork.
  3. Schedule each wave better: jobs carry costs (the segment lengths) and the costliest start first; the parent waits on the result pipe instead of polling, and starts the next job without waiting for the finished child to exit.

None of these changes a number.

Evidence

sim.fit() on seed 0's level-1 simulator.py and recording, in the scripted from-assets harness on #216's tree (its wider material boxes give this fit 7,735 replays), one Slurm job per variant with 8 CPUs and 6 workers:

  • master's fit code: 2,403 s
  • scoring independent rollouts together: 1,587 s
  • plus one pre-built world per fit: 907 s
  • plus the wave scheduling: 682 s

The last job ran on a faster node, where batching alone took 1,460 s: the world and the scheduling make the fit 2.1 times faster than batching alone, and all three changes about 3.2 times faster than master.

Every variant printed the same fit report, the same published values (SSE 7.186619535806188) and the same belief intervals, digit for digit.
With all three changes the fit ran in 562 waves; the 409 waves that score a single point take 0.70 s, against 0.6 s for its longest segment alone.

Most of what remains is the anchor ablation: 6,293 of the 7,735 replays, 7 rounds that each refit two parameters with a full LM.
Making it cheaper changes what it computes, so this PR leaves it as it is.

On seed 0's own scene, forked replays of its 22-step and 99-step segments on template copies end in the same state, bit for bit, as replays on fresh builds; freeing a copy costs a child 7 ms.

Tests

  • New: batched (point, trajectory) terms equal the point-by-point terms with and without workers; the rollout LM fit returns the same MAP and Jacobian with and without workers; solve_lm's batched Jacobian equals least_squares' own; lines traced in lockstep equal lines traced one by one; a row of dominoes in real PyBullet contact physics gives bit-identical terms on template copies and fresh worlds, no child builds a world, a template around two waves serves both, and a rollout in the building process still gets a fresh world; one worker runs jobs costliest first with results at their job's index.
  • Local CI replay: the static checks and the 8 shards (with Resolve a registry name to its own class, never a subclass that inherits it #213 applied).

🤖 Generated with Claude Code

A rollout fit scores many parameter points whose rollouts do not depend
on each other: the explainability sweep's candidates, each grid pool,
each LM Jacobian's column points and the parameter belief's line
points. The objective scored them one point at a time. Each point forked
one wave of children, one per segment, and waited on its slowest segment
before the next point started, and the explainability sweep scored its
segments serially. One sim.fit of seed 0 of the Domino round fixes_r2
took 39 minutes.

trajectory_terms_by_point scores every (point, trajectory) pair in one
fork wave, with exactly the serial path's terms (per-step scoring with a
factory env; interval scoring and a shared env keep the point-by-point
path). The explainability sweep and the grid pools score their
candidates through it. solve_lm takes a batch scorer and computes
least_squares' own 2-point Jacobian with the column points scored
together: approx_derivative runs once to record the points and once on
their cached residuals. The belief traces its lines in lockstep, one
batch per round. Every value is the one the point-by-point path
computes; the new tests check it bit for bit.
Building the agent's world took 0.39 s of a 0.51 s segment replay on the
from-assets Domino scene. A forked child's copy of a world the parent
built and never stepped is a fresh world, so a fit (and any other fork
wave of the rollout objective) builds one template world and each child
replays on its copy instead of building its own. A rollout the parent
runs itself still builds a fresh world, and nothing is built when waves
do not fork.

The new test replays a row of dominoes in real PyBullet contact physics
both ways: the terms are bit-identical, and no child built a world.
A single-point wave of the fit scores 7 segments of 22 to 99 steps on 6
workers. The dispatcher started jobs in segment order, so the 99-step
replay waited for a worker behind short ones; it polled for results
every 50 ms; and it joined each finished child before starting the next
job. Jobs can now carry costs (the rollout objective passes segment
lengths) and start costliest first. The parent waits on the result pipe,
takes each result the moment it arrives, starts the next job at once
and reaps exited children without blocking. Results stay at their job's
index, so only the wall time changes.
@yichao-liang
yichao-liang changed the base branch from from-assets-fit-gate to master October 10, 2026 14:09
@yichao-liang
yichao-liang merged commit a38dc96 into master Oct 10, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant