Conversation
A reconcile that failed for any reason other than a rate limit was retried after a fixed 30 seconds, however long the error persisted. When the failing step is a Hetzner API call, a service stuck on a permanent error repeats that call twice a minute and spends the per-project request budget that every other service shares. Retry each service after 5 seconds, doubling up to 5 minutes, the per-item backoff the upstream cloud-provider service controller uses. A successful reconcile, or a skip once the service is no longer a robotlb load balancer, starts the delay short again. A wait at the rate limit gate neither lengthens nor resets it. Entries that have not failed for an hour are dropped, so services deleted before they got the finalizer do not accumulate. Assisted-by: LLM Signed-off-by: Aleksei Sviridkin <f@lex.la>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Failed reconciles now back off per service instead of retrying every 30 seconds forever. Stacked on #37, so the base is
fix/hcloud-rate-limit.Today every error except a rate limit is retried after a fixed 30 seconds, however long it lasts. When the failing step is a Hetzner API call, a service stuck on a permanent error repeats that call twice a minute. That spends the per-project request budget, and every other service shares it.
The delay now starts at 5 seconds and doubles up to 5 minutes. It's the same per-item backoff the upstream cloud-provider service controller uses (
NewTypedItemExponentialFailureRateLimiter(5s, 300s)incontrollers/service/controller.go).namespace/nameinCurrentContext, next to the rate limit gate.Watch events still trigger a reconcile right away. The backoff only spaces out retries of a service where nothing changed.
The first retries now come sooner than before: 5, 10 and 20 seconds instead of 30 each time. Since every failed attempt publishes a warning event, the first minute of a failure shows more events than it used to.
The delay has no jitter, same as upstream, so services that fail together also retry together. The existing
spreadcould fix that later if it turns out to matter. A service held at the rate limit gate for more than an hour loses its entry and starts again at 5 seconds.The
FORGET_AFTERboundary could use a test of its own (a failure right at the limit keeps the entry, one second later drops it).a_long_failing_service_is_not_forgottenwould pass with a much shorter window.The key is
namespace/name, not the UID, same as the upstream workqueue. A service recreated under the same name within that hour inherits the old delay until its first success.README doesn't mention the retry delay, so no docs change.
Closes #46