Skip to content

Keep the whole-box search inside the fit's time budget - #224

Open
yichao-liang wants to merge 1 commit into
fit-hidden-state-segmentsfrom
fit-time-budget
Open

yichao-liang wants to merge 1 commit into
fit-hidden-state-segmentsfrom
fit-time-budget

Conversation

@yichao-liang

Copy link
Copy Markdown
Collaborator

Summary

  • The whole-box search (Search the whole box when the starting values are guesses #221) improves on a fit that already exists, so it now yields to the sim.fit tool's time budget (agent_sdk_fit_call_timeout, 3600 s).
    The design batch and each LM start run only if they are projected, timed by the fit so far, to end within half of the time the fit had left when it started.
    The other half is the report's and the belief's, which trace every basin.
  • A search that cannot afford the design and one LM start keeps the fit from the starting values, still measures its misfit for the belief's temperature (Count a guessed fit's misfit once in the belief's temperature #222), and mixes no basins.
    A search cut short says how far it got, in a line under the belief in the fit report.
  • The tool passes its deadline down through a new runtime field, SysIdConfig.deadline.
    A fit that still runs out of time now says so, that nothing was applied, and what makes fits long.
    The watchdog's bare ProbeBudgetExceeded used to print an empty "rollout system-ID fit failed:".

Why

Bridge seed 0 of the from-assets round fixes_r5 wrote a model with hidden state, so each candidate replays its whole 1661-step recording.
With the whole-box search on top, its fit ran past the 3600 s limit twice.
Each time the agent saw an empty error, nothing was applied, and under the from-assets fit gate it could not act on its test level at all; it retried the same fit.

#223 shortens such replays when the scored features come to rest, but a model that scores a moving robot still replays whole recordings, so the search itself has to fit the budget.

Test plan

  • test_physical_sysid.py: with a fake clock and a 10 s LM, half of a 50 s budget runs two of three starts and reports it; half of a 10 s budget runs none, keeps the guesses' fit with its misfit and no basins; without a limit the whole search runs with no note.
  • test_agent_continual_approach.py: a real play round whose fit is stopped by the watchdog shows the agent the overrun message, and the tool passed the fit a deadline.
  • Local CI replay on this commit: yapf, isort 5.10.1, docformatter, mypy, pylint and the 8 pytest shards pass.

🤖 Generated with Claude Code

Bridge seed 0 of the from-assets round fixes_r5 wrote a model with
hidden state, so each candidate replays its whole 1661-step recording.
With the whole-box search on top, its fit ran past the 3600 s fit limit
twice: the watchdog's bare ProbeBudgetExceeded printed an empty "fit
failed:", nothing was applied, and under the from-assets fit gate the
agent could not act on its test level at all.

The search improves on a fit that already exists, so it now yields to
the budget: the design and each LM start run only if they are projected,
timed by the fit so far, to end within half of the time the fit had left
when it started (the rest is the report's and the belief's, which trace
every basin). A search that cannot afford the design and one start keeps
the fit from the starting values, still measures its misfit, and mixes
no basins; one cut short says how far it got in the fit report. The
sim.fit tool passes its deadline down through SysIdConfig.deadline, and
a fit that still runs out of time now says so and what makes fits long.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant