Skip to content

fix(controller): back off per service on repeated failures - #55

Open
lexfrei wants to merge 1 commit into
fix/hcloud-rate-limitfrom
fix/error-backoff
Open

lexfrei wants to merge 1 commit into
fix/hcloud-rate-limitfrom
fix/error-backoff

Conversation

@lexfrei

@lexfrei lexfrei commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator

Failed reconciles now back off per service instead of retrying every 30 seconds forever. Stacked on #37, so the base is fix/hcloud-rate-limit.

Today every error except a rate limit is retried after a fixed 30 seconds, however long it lasts. When the failing step is a Hetzner API call, a service stuck on a permanent error repeats that call twice a minute. That spends the per-project request budget, and every other service shares it.

The delay now starts at 5 seconds and doubles up to 5 minutes. It's the same per-item backoff the upstream cloud-provider service controller uses (NewTypedItemExponentialFailureRateLimiter(5s, 300s) in controllers/service/controller.go).

  • State is a map keyed by namespace/name in CurrentContext, next to the rate limit gate.
  • A successful reconcile drops the entry. A skip drops it too, which covers a service that stopped being a robotlb load balancer.
  • A wait at the rate limit gate doesn't lengthen or reset the delay. Rate limit handling itself is unchanged.
  • A service deleted before it got the robotlb finalizer is not reconciled again, so its entry can't be dropped on success. Entries that haven't failed for an hour are pruned on the next failure of any service.

Watch events still trigger a reconcile right away. The backoff only spaces out retries of a service where nothing changed.

The first retries now come sooner than before: 5, 10 and 20 seconds instead of 30 each time. Since every failed attempt publishes a warning event, the first minute of a failure shows more events than it used to.

The delay has no jitter, same as upstream, so services that fail together also retry together. The existing spread could fix that later if it turns out to matter. A service held at the rate limit gate for more than an hour loses its entry and starts again at 5 seconds.

The FORGET_AFTER boundary could use a test of its own (a failure right at the limit keeps the entry, one second later drops it). a_long_failing_service_is_not_forgotten would pass with a much shorter window.

The key is namespace/name, not the UID, same as the upstream workqueue. A service recreated under the same name within that hour inherits the old delay until its first success.

README doesn't mention the retry delay, so no docs change.

Closes #46

A reconcile that failed for any reason other than a rate limit was
retried after a fixed 30 seconds, however long the error persisted.
When the failing step is a Hetzner API call, a service stuck on a
permanent error repeats that call twice a minute and spends the
per-project request budget that every other service shares.

Retry each service after 5 seconds, doubling up to 5 minutes, the
per-item backoff the upstream cloud-provider service controller uses.
A successful reconcile, or a skip once the service is no longer a
robotlb load balancer, starts the delay short again. A wait at the
rate limit gate neither lengthens nor resets it. Entries that have
not failed for an hour are dropped, so services deleted before they
got the finalizer do not accumulate.

Assisted-by: LLM
Signed-off-by: Aleksei Sviridkin <f@lex.la>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant