Skip to content

[Feature Request] Native recovery from failed exploration tasks with validated prefix frames and compatible ratio policies #1981

Description

@SchrodingersCattt

Summary

Please add native, opt-in recovery semantics for failed or incomplete model-deviation MD tasks: keep the trajectory marked as failed/incomplete, but allow validated pre-failure frames to contribute to labeling and subsequent learning, without blocking the entire workflow when a configured failure budget permits continuation.

This should integrate explicitly with task-completion/unfinished ratios and frame/label failure ratios, rather than requiring manual edits to record.dpgen, task directories, or completion tags.

Detailed Description

Problem

An MD task can terminate with an error such as lost atoms or non-finite values after producing usable earlier frames. Sibling tasks may finish normally. In an active-learning workflow, this is useful evidence about the explored configuration space, not necessarily a reason to discard every sibling result or repeatedly execute the same failing restart.

However, there are distinct facts that must not be conflated:

  1. The MD task did not reach its requested physical duration.
  2. Some of its saved configurations may still be valid inputs for first-principles labeling.
  3. A dispatcher may have finished monitoring all tasks, including terminal failures.
  4. FP labels are usable only after their own convergence/parsing/quality checks.

A failed MD trajectory must never acquire a success/completion status merely because its valid prefix was harvested. Conversely, a terminal MD failure within an explicitly allowed budget should not unconditionally prevent the rest of the active-learning workflow from proceeding.

Generic scenario

This is a synthetic, system-independent example:

  • Submit four exploration tasks.
  • Three finish and produce all declared outputs.
  • One terminates before its requested end, with a partial dump and model-deviation file.
  • The last deviation row may have no corresponding complete dump frame.
  • Restart attempts may append duplicate timesteps.

With the current default path, task failure can abort the exploration stage. Merely allowing a nonzero unfinished ratio is not a complete solution: it may stop still-running work, omit incomplete tasks from downloads, or leave DP-GEN's task.* discovery expecting files that were not retrieved. Missing-output errors can then reappear during selection. These details depend on the dispatcher version/backend and should be covered by an explicit integration contract.

Source observations

DP-GEN upstream master inspected at commit 0b9acec837474d2a93fa010598472feb8f69f82a:

  • run_md_model_devi() calls submission.run_submission() without a DP-GEN-level incomplete-exploration policy.
  • post_model_devi() is currently a no-op.
  • _read_model_devi_file() expects model-deviation output; accepting a failure count does not itself validate or reconcile matching trajectory frames.
  • ratio_failed is used by FP post-processing, so it is not interchangeable with an exploration task-failure allowance.

Some dispatcher versions can monitor siblings through terminal failures and report failures after result collection. DP-GEN still needs to interpret that outcome explicitly; swallowing every dispatcher exception would be unsafe.

Requested behavior

1. Explicit task states and failure classification

Distinguish completed MD, terminal failed MD, intentionally stopped MD, missing/corrupt output, and transient infrastructure failures. Allow bounded infrastructure recovery without resubmitting already completed work. Do not redefine a scientific MD failure as success.

2. Native frame salvage and validation

Provide an opt-in policy to collect diagnostic/partial outputs and select only eligible pre-failure configurations. At minimum:

  • verify complete frames, expected atom count and consistent ID/type mapping;
  • require finite coordinates/forces and a valid simulation cell;
  • reconcile deviation timesteps with actual complete dump frames;
  • handle restart duplicate timesteps deterministically and check that retained deviations correspond to the retained configurations;
  • exclude unmatched rows and corrupt/invalid frames;
  • make relative-deviation normalization use a documented, reproducible eligible-frame set rather than inadvertently including duplicated or invalid records.

Numerical eligibility is not proof of chemical validity or potential accuracy. Existing scientific selection and FP validation remain necessary.

3. Compatible, clearly separated ratio policies

Document separate denominators and ordering for:

  • exploration tasks that failed or were intentionally left unfinished;
  • usable/excluded frames within those tasks;
  • candidate/accurate/failed uncertainty classifications;
  • successfully parsed FP labels versus failed FP labels.

Specify whether an unfinished-task policy means early termination or acceptance of terminal failures. Reaching one threshold must not mask a different threshold violation or fabricate missing outputs. Training and model export should remain strict by default; a permissive exploration policy must not silently propagate to those stages.

4. Native, auditable continuation

When enough eligible data remain and the configured budgets permit it, allow selection, labeling and subsequent training to proceed without manual task-directory surgery or record edits. Record task outcome, valid physical-time coverage, excluded frames, retry history, accepted labels and reasons for continuation in a machine-readable report.

If a failed exploration condition needs to be repeated with an updated model, the repeat should be explicit workflow state, not a renamed old trajectory. Preserve the failure status and do not claim the original requested duration was completed.

Stop with an actionable report if the failure budget is exceeded or there are no eligible frames/new labels; never loop indefinitely just to keep a process alive.

Acceptance tests

  • Three completed tasks plus one truncated task: keep all original outcomes, preserve completed outputs, salvage only matched valid prefix frames when enabled, and advance only if the policy permits.
  • A task missing all outputs cannot be counted as a successful task or a usable frame source.
  • A last deviation row without a dump frame and duplicate restart records cannot create fictitious FP inputs or duplicate labels.
  • Both grouped and one-task-per-job dispatch must behave consistently at task granularity.
  • Resuming the same submission must not rerun completed jobs or reset retry budgets without an explicit policy.
  • Test unfinished-ratio early termination and terminal-failure continuation separately, then their interaction with FP ratio_failed.
  • All-invalid/all-failed batches halt clearly; model export/training cannot pass with missing models.

Further Information, Files, and Links

This is a request for a supported upstream policy/schema and end-to-end regression coverage, not a request to ignore lost atoms or to introduce ad-hoc recovery scripts. The example above intentionally contains no material, dataset, infrastructure, account, or project-specific information.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions