Skip to content

fix(auto): warn when unavailable accelerators fall back to CPU - #3742

Merged
LauraGPT merged 1 commit into
mainfrom
codex/device-fallback-warning-20260930
Sep 30, 2026
Merged

LauraGPT merged 1 commit into
mainfrom
codex/device-fallback-warning-20260930

Conversation

@LauraGPT

@LauraGPT LauraGPT commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Follow-up to the reporter's request in #3738 after installing a CUDA-enabled PyTorch build resolved their immediate problem.

  • Log a warning when AutoModel falls back from an unavailable CUDA, XPU, MPS, or NPU device to CPU, including the model and requested device plus a PyTorch/accelerator configuration hint.
  • Keep the existing fallback conditions, resolved device, and fallback batch size unchanged. The implicit CUDA default also warns; explicit CPU use and ngpu=0 remain quiet.
  • Reuse ordinary logging, respecting application handlers and levels. Inherited submodels receive the resolved CPU device without duplicate warnings; an independently configured accelerator submodel reports its own fallback.

Type of change

  • Bug fix

Validation

  • Native CPU-only tests against the real AutoModel construction path: 59 passed, 1 deselected. Files: tests/test_submodel_device.py, tests/test_auto_model_ncpu.py, tests/test_auto_model.py, and tests/test_amp_device_type.py. The deselected speaker-clustering test downloads a real model and was deliberately not run.
  • RED before the production change: 5 failed, 17 passed, with all five failures specifically caused by the missing fallback warning.
  • A fresh Python process using default logging visibly emitted exactly one fallback warning; a tiny real torch.nn.Module had CPU parameters after construction. No model download or inference was performed.
  • Additional coverage verifies explicit punctuation-model warning identity, caller-owned configuration preservation, and suppression under an application's ERROR logging level.
  • git diff --check passed. Independent static review found no P1/P2 issues at the tested file hashes.

Accelerator availability is mocked in the configuration tests; the stand-in models retain CPU weights. This is not GPU/MPS/XPU/NPU hardware validation, not a full FunASR suite run, and not a claim that the existing hosted workflows execute these new tests. No dependencies, model assets, CI settings, or inference defaults changed.

User impact

Users expecting accelerator execution can see when initialization has selected CPU instead. Applications can continue controlling verbosity through normal logging configuration.

Notes for reviewers

Related to #3738; the reporter already closed that issue after resolving their environment problem.

Integration

Merged as 7b098acefc22999f1be483b7b6fca62db8c8ad39. Both exact-head hosted checks passed: MOSS adapter and KWS output. These workflows do not replace the 59-test native CPU run above. Verified merge parents, signed commit, source bytes, and exact tree equality with the tested head.

Merged through the ordinary SHA-guarded API under the repository's existing admin exemption. No branch protection changes, fabricated review approval, or workflow approval was made. This is a source-main change, not a PyPI release or production deployment.

Post-merge checks at the exact merge SHA also completed successfully: KWS output, MOSS adapter, and API documentation.

Signed-off-by: LauraGPT <18321252+LauraGPT@users.noreply.github.com>
@LauraGPT
LauraGPT merged commit 7b098ac into main Sep 30, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant