Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
48 changes: 48 additions & 0 deletions .github/workflows/lambda.yml
Original file line number Diff line number Diff line change
Expand Up @@ -32,21 +32,69 @@ jobs:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false

- name: Install dependencies
run: yarn install --frozen-lockfile

- name: Run prettier
run: yarn format-check

- name: Run linter
run: yarn lint

- name: Run tests
id: test
run: yarn test

- name: Build distribution
run: yarn build

- name: Upload coverage report
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
if: ${{ failure() }}
with:
name: coverage-reports
path: ./**/coverage
retention-days: 5

scale-set-container:
name: Build scale-set service container
runs-on: ubuntu-latest
steps:
- name: Harden the runner (Audit all outbound calls)
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit

- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false

- name: Set up QEMU
uses: docker/setup-qemu-action@96fe6ef7f33517b61c61be40b68a1882f3264fb8 # v4.2.0

- name: Set up Docker Buildx
uses: docker/setup-buildx-action@37fe631027851001ddb9b187196cc803df7f5f0e # v4.3.0

- name: Build scale-set service image
uses: docker/build-push-action@53b7df96c91f9c12dcc8a07bcb9ccacbed38856a # v7.3.0
with:
context: .
file: ./lambdas/services/scale-set/Dockerfile
platforms: linux/amd64,linux/arm64
push: false
cache-from: type=gha,scope=scale-set-service
cache-to: type=gha,mode=max,scope=scale-set-service

- name: Build scale-set service image for smoke test
uses: docker/build-push-action@53b7df96c91f9c12dcc8a07bcb9ccacbed38856a # v7.3.0
with:
context: .
file: ./lambdas/services/scale-set/Dockerfile
platforms: linux/amd64
load: true
tags: scale-set-service:smoke-test
cache-from: type=gha,scope=scale-set-service

- name: Run scale-set service image smoke test
run: ./tests/scale-set-container-smoke-test.sh
2 changes: 2 additions & 0 deletions .github/workflows/smoke-tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -93,6 +93,8 @@ jobs:
image: ghcr.io/ministackorg/ministack:1.5.12@sha256:41fe1ce2e666c6cc410c6047a9db8bf1df69cd0028ebc0a6c6e5517c3a83d6e0
ports:
- 4566:4566
# MiniStack launches nested service containers for the integration smoke test.
# This digest-pinned, non-fork workflow is the narrowly scoped Docker-daemon exception.
options: >-
--add-host=host.docker.internal:host-gateway
--volume /var/run/docker.sock:/var/run/docker.sock
Expand Down
2 changes: 0 additions & 2 deletions MAINTAINERS.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,8 +48,6 @@ The following steps needs to be applied to test a PR
3. Apply the PR to the deployment. Check output for breaking changes such as destroying resources containing state.
4. Test the PR by running a workflow

Some PR tests can be run against [MiniStack](tests/ministack/README.md), including the example deployments and webhook and runner lifecycle [smoke tests](tests/ministack/README.md#webhook-and-runner-lifecycle-smoke-test). Use these tests where applicable during PR review. MiniStack test coverage is still being expanded, with additional cases in progress.

### Security

Act on security issues as soon as possible. If a security issue is reported.
69 changes: 69 additions & 0 deletions docs/adr/0003-scale-set-resource-ownership.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
# ADR-003: Scale-set Resource Ownership

## Status

Accepted

## Date

2026-09-23

## Context

GitHub Actions runner scale sets are GitHub-side resources with their own
identity and permissions. The TypeScript controller is the component that
configures GitHub through the scale-set API. Terraform only provisions the
AWS controller substrate and supplies the controller with names and secret
references; Terraform never configures the GitHub scale set itself.

The controller resolves and manages the runtime relationship with a scale set
and runner group from the configured GitHub scope and names. The GitHub API
lifecycle is controller-owned at runtime, not implemented by Terraform.

## Decision

The scale-set orchestration provider resolves GitHub scale sets by name.

- `scale_set.name` identifies the GitHub scale set to reconcile.
- `runner.group_name` identifies the runner group used when the scale set is
resolved by name.
- Terraform creates and manages the ECS controller, IAM roles, networking,
logs, and SSM configuration required to run the reconciler. It never calls
the GitHub scale-set API.
- The controller may discover the scale-set and runner-group IDs at runtime,
but it does not write those discovered IDs back to SSM. Operators may
pre-populate optional cache parameters when they want read-side caching.
- If the named scale set is absent, the controller may register it in the
resolved runner group and reconcile its system labels. It does not delete
scale sets.
- The controller currently registers a missing scale set and reconciles its
system labels. Terraform destroy removes the AWS controller and stops future
reconciliation, but it cannot delete or rename the GitHub resources because
Terraform never configures them.

The configured GitHub scope and scale-set name must be unique across controller
groups. A single runner configuration must not be selected by both webhook and
scale-set orchestration.

## Consequences

This keeps ownership boundaries explicit: the TypeScript controller is the sole
GitHub API owner, while Terraform owns only the AWS deployment and its input
references. Destroying Terraform stops the controller but does not issue a
GitHub delete. Operators must authorize the configured GitHub scope before
applying the AWS controller configuration and must handle renames as an
explicit migration. The controller task role can remain read-only for SSM
discovery and credential reads, reducing its blast radius.

## Alternatives considered

- **Model GitHub scale sets as Terraform resources:** not selected for the
current implementation. The TypeScript controller already owns all runtime
GitHub API operations (lookup, registration, and label reconciliation), while
Terraform owns the AWS substrate. Adding a second Terraform owner would
create competing GitHub API lifecycles and require an explicit import,
update, and destroy contract.
- **Persist every discovered ID from the controller:** rejected for now
because it would require explicit, caller-visible SSM write targets and an
expanded IAM contract. Read-side caching remains possible through
pre-provisioned parameters.
75 changes: 75 additions & 0 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -296,6 +296,81 @@ In case the setup does not work as intended, trace the events through this seque

## Experimental features

### GitHub Actions runner scale-set orchestration

Scale-set orchestration is an experimental multi-runner v2 provider for
workloads that should use GitHub's runner scale-set message protocol instead of
webhook-driven Lambda scaling. Select it inside the lane's
`multi_runner_config.<name>.orchestration_provider` block and pair it with a
compute provider that implements the scale-set capability contract. The
controller runs as one ECS Fargate task per resolved controller group and
reconciles the configured scale sets continuously.

Use scale-set orchestration when the GitHub scale-set API and a long-lived
controller are the desired ownership model. Continue using webhook
orchestration when the existing `workflow_job` event, SQS, and Lambda lifecycle
are the better fit. The two modes must not manage the same runner lane.

The scale-set module resolves GitHub scale sets by their configured name. When
a named scale set is absent, the controller registers it in the resolved
runner group and reconciles its system labels at runtime. The TypeScript
controller is the component that configures GitHub; Terraform never calls the
GitHub scale-set API. Terraform destroy removes the AWS controller and stops
reconciliation, but does not issue a GitHub delete. The module requires an
explicit controller image, preferably an immutable digest. The task reads GitHub App credentials
and optional discovery-cache values from SSM but does not write discovered IDs
back to SSM. See the [scale-set provider reference](https://github.com/github-aws-runners/terraform-aws-github-runner/blob/main/modules/orchestration-providers/scale-set/README.md)
for the complete input schema.

#### Scale-set options and defaults

Set these values under
`global_config_orchestration_provider.scale_set`. Per-lane scale-set values
under `multi_runner_config.<lane>.orchestration_provider.scale_set` override
the corresponding lane settings. The controller image is represented as
optional in the Terraform type for normalization, but validation requires a
non-empty value; use an immutable digest.

| Option | Default | Purpose |
| --- | --- | --- |
| `grouping.strategy` | `compute_provider` | Pack reconcilers by compute-provider type; use `runner_config` or `custom` to create narrower task/IAM boundaries. |
| `container.image` | none; required | Controller image reference. Prefer a release digest. |
| `container.user` | `10001:10001` | Numeric non-root UID/GID used by the application container. |
| `container.health_port` | `8080` | ECS health-check port. |
| `container.health_path` | `/healthz` | ECS liveness endpoint; `/readyz` is an application readiness signal. |
| `container.health_check_command` | `null` | Use the image health check unless an explicit ECS command is required. |
| `container.health_check_interval` / `timeout` / `retries` | `30` / `5` / `3` | ECS container health-check timing. |
| `container.health_check_start_period` | `30` | Startup grace period for the ECS health check. |
| `container.health_stale_after_seconds` | `180` | Controller health staleness threshold. |
| `container.shutdown_timeout_seconds` | `110` | Controller shutdown grace period. |
| `container.session_close_timeout_seconds` | `10` | Message-session close timeout. |
| `container.reconnect_initial_backoff_seconds` / `max` | `1` / `30` | Bounds for reconnect backoff. |
| `container.stop_timeout_seconds` | `120` | ECS container stop timeout. |
| `config_store.path_prefix` / `tier` | derived / `Standard` | SSM path prefix and parameter tier for non-secret reconciler configuration. |
| `ecs.cluster.mode` | `managed` | Create a cluster or use an external cluster. |
| `ecs.cluster.container_insights` | `true` | Enable ECS container insights on a managed cluster. |
| `ecs.task.cpu` / `memory` | `512` / `1024` | Fargate task CPU units and memory MiB. |
| `ecs.task.cpu_architecture` | `X86_64` | Fargate task architecture. |
| `ecs.task.ephemeral_storage` | `null` | Use the Fargate platform default unless a size is supplied. |
| `ecs.service.platform_version` | `LATEST` | ECS Fargate platform version. |
| `ecs.iam.path` / `permissions_boundary` | `/` / `null` | IAM role path and optional permissions boundary. |
| `network.vpc_id` / `subnet_ids` | required | Private subnets in which the controller service runs. |
| `network.https_egress.ipv4_cidrs` | `0.0.0.0/0` | Default HTTPS reachability; restrict through GitHub Meta API ranges, NAT, firewall, or proxy as required. |
| `network.https_egress.ipv6_cidrs` | `[]` | IPv6 HTTPS egress destinations. |
| `logging.retention_in_days` / `kms_key_id` | `180` / `null` | CloudWatch log retention and optional customer-managed KMS key ID, alias, or ARN. |
| `logging.log_group_class` | `STANDARD` | CloudWatch log-group class. |
| `tags` | `{}` | Tags applied to scale-set resources. |

The scale-set lane itself defaults to `runner.group_name = "Default"`,
`runner.min_runners = 0`, `runner.max_runners = 10`, and
`runner.boot_time_in_minutes = 10`. Configure the GitHub scope, scale-set
name, runner owner, and GitHub App SSM references in the lane; credential values
are not placed in the controller manifest.

The [multi-runner scale-set example](multi-runner-scale-set.md) shows how these
provider-specific settings coexist with webhook lanes in the same v2
`multi_runner_config` map.

### macOS Runners

This feature is in early stage and should be considered experimental. The module supports macOS-based GitHub Actions self-hosted runners on AWS EC2 Mac instances (`mac1.metal`, `mac2.metal`, `mac2-m2.metal`). macOS runners require dedicated hosts due to Apple's licensing requirements and have longer boot times (6–20 minutes). Set `runner_os = "osx"` and `use_dedicated_host = true` to enable. See the full [macOS Runners documentation](mac-runners.md) for details.
Expand Down
1 change: 1 addition & 0 deletions docs/examples/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ Examples are located in the [examples](https://github.com/github-aws-runners/ter
- _[Ephemeral](ephemeral.md)_: Example usages of ephemeral runners based on the default example.
- _[Multi Runner](multi-runner.md)_ : Example usage of creating a multi runner which creates multiple runners/ configurations with a single deployment. The examples including: "arm64", "windows", and "ubuntu" runners.
- _[Multi Runner v2](multi-runner-v2.md)_ : Example usage of the experimental v2 multi-runner configuration interface with shared defaults and per-lane overrides.
- _[Multi Runner scale-set](multi-runner-scale-set.md)_ : Example usage of a v2 deployment combining webhook lanes with an experimental GitHub Actions scale-set lane.
- _[Permissions boundary](permissions-boundary.md)_: Example usages of permissions boundaries.
- _[Prebuilt Images](prebuilt.md)_: Example usages of deploying runners with a custom prebuilt image.
- _[Termination watcher](termination-watcher.md)_: Example usages of termination watcher.
Expand Down
14 changes: 14 additions & 0 deletions docs/examples/multi-runner-scale-set.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
# Multi-runner scale-set example

This example combines ordinary webhook-managed lanes with one experimental
GitHub Actions runner scale-set lane. It demonstrates that v2 keeps the
deployment-wide defaults in `global_config*` and places orchestration and
compute-provider settings inside each `multi_runner_config` lane.

The source example is available at
[examples/multi-runner-scale-set](https://github.com/github-aws-runners/terraform-aws-github-runner/tree/main/examples/multi-runner-scale-set).
Read its README before applying: the GitHub App values are sensitive, the
scale-set controller image must be supplied explicitly, and the GitHub scale
set/runner group must be authorized for the selected GitHub scope.

--8<-- "examples/multi-runner-scale-set/README.md"
64 changes: 64 additions & 0 deletions docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,70 @@ For ephemeral runners a pool can be configured. The pool maintains a minimum num

For non ephemeral runners with the idle config the module will avoid scaling down back to zero. Instead it will maintain a minimum number of runners based on a schedule. This avoids the need to scale up when a new workflow is triggered.

### Scale-set orchestration (experimental)

Multi-runner v2 can select the experimental scale-set orchestration provider for a
runner lane. Scale-set orchestration uses a long-running ECS Fargate controller
instead of webhook events and Lambda scale-up/scale-down handlers. The
controller maintains a message session with GitHub, reconciles the desired
capacity reported by the scale-set API, and delegates runner provisioning to the
selected compute provider.

The controller flow is:

1. Terraform creates one ECS service, task definition, task role, execution
role, log group, security group, and SSM configuration set for each resolved
controller group.
2. The controller loads non-secret configuration from its task manifest or SSM
group path. GitHub App values remain in caller-managed SSM parameters and
only their names are passed to the task.
3. The controller resolves the configured GitHub scale set by name, registers
it when absent, opens its message session, and reconciles capacity through
the compute provider. The TypeScript controller owns those runtime API
operations; Terraform only provisions AWS and never calls the GitHub
scale-set API. Terraform destroy stops reconciliation without issuing a
GitHub delete.
4. The compute provider creates, refreshes, and removes runner capacity. For
EC2, JIT configuration and instance lifecycle remain provider-owned.

The request and lifecycle path is:

```mermaid
flowchart LR
TF[Terraform\nmulti-runner v2] --> ECS[ECS Fargate\nscale-set controller]
TF --> SSM[(SSM\nconfig and secret references)]
ECS -->|GitHub App auth| API[GitHub Actions\nscale-set APIs]
API -->|desired capacity and jobs| ECS
ECS -->|JIT configuration| EC2[EC2 compute provider]
EC2 -->|runner registration| API
ECS -->|scale-up / scale-down| EC2
EC2 -->|terminate owned capacity| EC2
```

The controller is long-lived and does not receive `workflow_job` webhooks. It
opens a GitHub scale-set message session, acknowledges and processes messages,
then passes desired capacity and busy-runner information to the compute
provider. The EC2 provider publishes JIT configuration through SSM, launches
the runner instance, and later removes only its owned runner capacity when the
scale-set session reports that it is safe to scale down.

The service is compatible with the wire behavior implemented by the upstream
[GitHub Actions scale-set client](https://github.com/actions/scaleset). Its
HTTP paths, message-session behavior, statistics handling, and User-Agent
requirements were implemented by reverse-engineering that Go client and its
protocol behavior. This is a compatibility boundary rather than a promise that
the upstream internal API is stable; validate changes in `actions/scaleset`
before upgrading the controller image.

Controller groups can be formed per compute-provider type, per runner config,
or through explicit custom membership. Grouping shares a task and IAM policy,
so split groups when blast radius or policy size must be reduced. This provider
is experimental: callers must supply an explicit controller image, use the
scale-set configuration contract, and verify the current provider and GitHub
API limitations before production use. See the [scale-set module
documentation](https://github.com/github-aws-runners/terraform-aws-github-runner/blob/main/modules/orchestration-providers/scale-set/README.md) and the
[v1-to-v2 configuration guide](multi-runner-v1-to-v2-configuration.md).


## Detailed design

Expand Down
Loading
Loading