Skip to content

fix(controller): skip Hetzner calls when nothing robotlb acts on changed - #72

Open
lexfrei wants to merge 1 commit into
fix/local-targets-from-slicesfrom
fix/skip-unchanged-reconcile
Open

lexfrei wants to merge 1 commit into
fix/local-targets-from-slicesfrom
fix/skip-unchanged-reconcile

Conversation

@lexfrei

@lexfrei lexfrei commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

robotlb now skips the Hetzner calls when nothing it acts on has changed since the last successful reconcile. The README now explains how to keep another LoadBalancer controller away from robotlb's Services.

Two controllers that both take Services without a class keep undoing each other's status.loadBalancer. MetalLB started without --lb-class is one example. Every status write is a watch event and starts a reconcile, and every reconcile talks to Hetzner. The loop runs until the project hits the rate limit. Services have no metadata.generation, so a generation predicate does not filter these events. A filter on the main stream needs Controller::for_stream, which is still unstable.

After each successful reconcile, robotlb keeps a hash per Service UID. It covers the spec, the robotlb/ annotations, and the targets and ports robotlb computed. The next reconcile computes targets and ports as before, which costs only Kubernetes reads. If the hash is the same and the next check is not due yet, it stops before the first Hetzner call and does not write the status. It requeues for the time left until that check. The next check is the resync, or the 30-second retry while Hetzner refuses some targets. Node and EndpointSlice changes still get through, because they change the targets.

The hash is dropped after a failed reconcile, and when the Service is released or deleted. It is also dropped when the Service has no port to expose, because that path clears the status and the status has to come back with the ports.

The README now says that robotlb handles Services with no loadBalancerClass or with robotlb. Setting loadBalancerClass: robotlb keeps the other controller away. The field can only be set when the Service is created or its type is changed to LoadBalancer. Recreating a Service deletes its balancer, and the new one gets a new public IP, so the README says to set the class at creation.

There are two trade-offs:

  • While the hash matches, a status cleared by another controller stays empty until the next check. This change stops the quota burn, and only the robotlb class ends the fight over the status.
  • A change made to the balancer directly in Hetzner is not noticed until the next check either. Before, any watch event on the Service caught it.

Some smaller limits I left alone. If a Service is deleted without robotlb seeing its deletion (for example after someone removes the finalizer by hand), its hash stays in memory until the process restarts. If the balancer has no public IP yet on the first reconcile, the status is written at the next check, not at the next Service event. A reconcile that waits at the rate limit gate also counts as failed and drops the hash, so every Service does a full reconcile once the gate opens. Keeping the hash there would save requests right when the budget is lowest, but I kept "any failure drops the hash" as the simpler rule.

Unit tests cover the skip decision before and after the check is due, and with the same or a different hash. They check that the hash changes with a spec field, a robotlb annotation, the targets and a port. They also check that it stays the same when only the status, resourceVersion or a foreign annotation changes, and that the hash is dropped on failure, release and deletion. I broke the code under each test and saw it go red: without the targets, without the annotations, with the expiry ignored, with a fixed expiry instead of the one recorded, and with the hash kept after a failure. Two mutations survive. One records the hash before the Hetzner calls succeed, the other bypasses the check. Both sit in reconcile_load_balancer, and catching them needs a mocked Kubernetes API and Hetzner API. A failed reconcile drops the hash anyway, so the first one only changes how long the hash stays valid when Hetzner refuses some targets.

Stacked on #56, only the last commit is new.

Closes #69

Every write to a Service starts a reconcile, including status writes
by another LoadBalancer controller that also takes Services without a
class. When that controller clears status.loadBalancer, robotlb writes
it back, the other controller clears it again, and each round costs
Hetzner requests until the project hits the rate limit. Services carry
no metadata.generation, so a generation predicate cannot filter these
events, and filtering the main stream needs an unstable kube-runtime
feature.

robotlb now remembers, per Service UID, a hash of the spec, the
robotlb/ annotations and the computed targets and ports after each
successful reconcile. A reconcile that finds the same hash before the
next check is due makes no Hetzner requests and writes no
status. Node and EndpointSlice changes still get through because they
change the targets. An entry is valid until the requeue of the
reconcile that recorded it, so while Hetzner refuses some targets
events are skipped only until the 30-second retry. The entry is
dropped after a failed reconcile, when the Service is released or
deleted, and when it has no port to expose.

The README now says which classes robotlb handles, that setting
loadBalancerClass: robotlb keeps other controllers away, and that
recreating a Service to add the class costs it its balancer and
public IP.

Assisted-by: LLM
Signed-off-by: Aleksei Sviridkin <f@lex.la>
@lexfrei

lexfrei commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

Checked on a live cluster with a build of this branch and debug logging, on a LoadBalancer Service with a nodePort:

  1. Adding a foreign annotation (example.com/skip-test=1) gave one reconcile that logged "Nothing robotlb acts on changed since the last reconcile. Skipping...". The health check interval in Hetzner stayed at 15.
  2. Setting robotlb/lb-check-interval=16 gave a full reconcile, and the interval in Hetzner became 16.
  3. Removing the foreign annotation was skipped again.
  4. Removing robotlb/lb-check-interval gave a full reconcile, and the interval went back to 15.

The API budget can't show the difference over a window this short, since it refills by one request per second. The skip line and the shorter log per reconcile are the evidence.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant