Skip to content

feat(scale-down): sweep stopped warm instances - #5494

Open
Brend-Smits wants to merge 2 commits into
feat/warm-pool-03-scale-up-activationfrom
feat/warm-pool-04-scale-down-sweep
Open

Brend-Smits wants to merge 2 commits into
feat/warm-pool-03-scale-up-activationfrom
feat/warm-pool-04-scale-down-sweep

Conversation

@Brend-Smits

@Brend-Smits Brend-Smits commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Warm pool stack (review in order; each PR is based on the previous one):

  1. feat(compute-providers): add EC2 standby primitives for warm pools #5491 feat(compute-providers): add EC2 standby primitives for warm pools
  2. feat(pool): keep a warm pool of stopped instances #5492 feat(pool): keep a warm pool of stopped instances
  3. feat(scale-up): start warm instances before launching cold runners #5493 feat(scale-up): start warm instances before launching cold runners
  4. feat(scale-down): sweep stopped warm instances #5494 feat(scale-down): sweep stopped warm instances (this PR)
  5. feat(runners): boot modes for warm pool instances #5495 feat(runners): boot modes for warm pool instances
  6. feat(runners): add warm_pool option to the root module #5496 feat(runners): add warm_pool option to the root module
  7. feat(multi-runner): support warm_pool in runner configs #5497 feat(multi-runner): support warm_pool in runner configs
  8. docs(warm-pool): add warm pool guide, ADR and examples #5498 docs(warm-pool): add warm pool guide, ADR and examples
  9. fix(compute-providers): add warm pool boot modes to the EC2 template #5499 fix(compute-providers): add warm pool boot modes to the EC2 template
  10. feat(runners): warm pool support for Windows runners #5500 feat(runners): warm pool support for Windows runners
  11. fix(scale-down): do not sweep warm instances that are being activated #5501 fix(scale-down): do not sweep warm instances that are being activated

Description

Scale-down now destroys stopped warm instances every cycle:

  • activated instances that stopped (for example a job powered the machine off), which would otherwise stay stopped forever
  • instances past their ghr:warm-expires-at time, so standby instances are cleaned up even after warm mode is disabled and the pool lambda no longer evicts them

Destroy cancels a persistent spot request before terminating. Errors are logged and never fail the scale-down run.

Test Plan

  • Extended scale-down.test.ts in both packages and registry.test.ts.
  • vitest, tsc, eslint and prettier for the affected lambda packages.

Related Issues

Supersedes #5204.

@github-actions

Copy link
Copy Markdown
Contributor

Dependency Review

✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.

Scanned Files

None

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The sweep can terminate an instance concurrently being activated, leaving a job without a viable runner or cold fallback.

Review effort: Balanced
Findings: 1 High severity

Open (1)
What changed in this PR

Adds provider-level cleanup of stopped warm EC2 instances during every scale-down cycle.

Changes:

  • Introduces the optional sweepStandby capability.
  • Removes activated or expired stopped warm instances with best-effort error handling.
  • Adds registry, provider, and scale-down tests.
File Description
lambdas/​libs/​compute-providers/​registry.test.ts Verifies EC2 exposes standby sweeping.
lambdas/​libs/​compute-providers/​core/​index.ts Extends the scale-down provider contract.
lambdas/​libs/​compute-providers/​aws/​ec2/​src/​control-plane/​scale-down.ts Implements stopped-instance sweeping.
lambdas/​libs/​compute-providers/​aws/​ec2/​src/​control-plane/​scale-down.test.ts Tests selection and failure handling.
lambdas/​libs/​compute-providers/​aws/​ec2/​control-plane.ts Connects standby operations to scale-down.
lambdas/​functions/​control-plane/​src/​scale-runners/​scale-down.ts Runs the sweep during scale-down.
lambdas/​functions/​control-plane/​src/​scale-runners/​scale-down.test.ts Tests orchestration and error isolation.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Scale-down now destroys two kinds of stopped warm instances every
cycle:

- activated instances that stopped, for example after a job powered
  the machine off, which would otherwise stay stopped forever
- instances past their `ghr:warm-expires-at` time, so standby
  instances are cleaned up even after warm mode is disabled and the
  pool lambda no longer evicts them

Destroy cancels a persistent spot request before terminating. Errors
are logged and never fail the scale-down run.

Signed-off-by: Brend Smits <brend.smits@philips.com>
Scale-down sweeps stopped warm instances past ghr:warm-expires-at. A
warm instance claimed shortly before its expiry could be destroyed by
the sweep before scale-up started it. Scale-up now skips warm
instances that expire within the activation grace period.

Signed-off-by: Brend Smits <brend.smits@philips.com>
@Brend-Smits
Brend-Smits force-pushed the feat/warm-pool-04-scale-down-sweep branch from 71db0c7 to 26def47 Compare October 1, 2026 12:27
@Brend-Smits
Brend-Smits force-pushed the feat/warm-pool-03-scale-up-activation branch from b40a741 to 04b7e7f Compare October 1, 2026 12:27
@Brend-Smits
Brend-Smits marked this pull request as ready for review October 1, 2026 13:58
@Brend-Smits
Brend-Smits requested a review from a team as a code owner October 1, 2026 13:58

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants