Skip to content

fix(hcloud): back off on rate limits and report failures on the service - #37

Open
lexfrei wants to merge 4 commits into
masterfrom
fix/hcloud-rate-limit
Open

lexfrei wants to merge 4 commits into
masterfrom
fix/hcloud-rate-limit

Conversation

@lexfrei

@lexfrei lexfrei commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Once the Hetzner Cloud API budget runs out, robotlb retries every service every 30 seconds, and each retry spends the budget it is waiting for. The failure is also hard to see from the cluster.

The budget is per project, so after a 429 all services now pause together. The pause starts at one minute and doubles up to 16 minutes while the limit keeps being hit. A 429 from a call that was already in flight does not lengthen it, and adding targets stops at the first 429 instead of calling the API for every remaining node. Services wake up spread over up to a quarter of the pause, so they don't all hit the API at the moment it ends. RateLimit-Reset is not used: the generated hcloud client drops response headers, so the header never reaches robotlb.

A failed reconciliation now leaves a SyncLoadBalancerFailed warning event on the service, the same reason the upstream service controller uses, so kubectl describe shows the last error. A service that wakes up to the paused gate gets one too. Errors are logged on one line with the HTTP status and the code and message Hetzner reported. The API token, which Hetzner quotes in some error messages, is redacted from logs and events. The chart role gets permission to create events.

Charts that override serviceAccount.permissions replace the default list, so they need create on events.k8s.io events added by hand. Without it, a failure only logs that the event could not be published.

Other errors are still retried every 30 seconds, see #46, and kube 0.96 does not aggregate events, see #45.

Closes #36

Once the Hetzner Cloud API budget ran out, every service was retried
every 30 seconds, and each retry spent the budget being waited for. The
budget is per project, so after a 429 all services now pause together,
starting at one minute and doubling up to 16 minutes while the limit
keeps being hit. Adding targets stops at the first 429 instead of
calling the API for every remaining node. The generated client drops
response headers, so RateLimit-Reset cannot be used.

A failed reconciliation now leaves a SyncLoadBalancerFailed warning
event on the service, so kubectl describe shows the last error. Errors
are logged on one line with the status and the message Hetzner reported,
and the API token, which Hetzner quotes in some messages, is redacted
from logs and events. The chart role gains permission to create events.

Assisted-by: LLM
Signed-off-by: Aleksei Sviridkin <f@lex.la>
Services paused by the rate limit gate were all requeued for the moment
it reopened, so they hit the API in one burst while the budget was
still nearly empty. Each wait now grows by up to a quarter, derived
from the service name, so services wake up spread over that window.

Assisted-by: LLM
Signed-off-by: Aleksei Sviridkin <f@lex.la>
With wakeups spread by service name, the same service wakes first after
every pause, hits the limit and closes the gate again. Every other
service only ever saw the closed gate, which published no event and
logged at debug level, so after the event TTL they showed nothing at
all. A service that wakes up to a closed gate now gets an event and an
info line.

Assisted-by: LLM
Signed-off-by: Aleksei Sviridkin <f@lex.la>
The comment claimed the default grants all permissions, while it lists
only what robotlb needs.

Assisted-by: LLM
Signed-off-by: Aleksei Sviridkin <f@lex.la>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Operator retries every 30s forever when the hcloud rate limit is hit, and the state is hard to inspect

1 participant