Repository navigation
Support running without a DaemonSet (if desired) #525
Description
Activity
nodeSelectors could be used to restrict the amount of nodes they are deployed to.
I'm unsure if that then makes the service unavailable to workloads off those nodes.In any case, having to option to deploy as a Deployment would be nice.
Easy solution to this might be: #688
I looked at this issue today and don't think we need much more than what's already here.
Trying to sum it up:- We want to allow two deployment models: Deployment and DaemonSet
- To define how to select between those two. I think best would be an explicit CRD choice between those two -> needs a decision
- When using Deployment don't set
internalTrafficPolicyat all - K8s has
PreferSameNodeas of 1.35, see below for details
If you believe implementing this (see below for the optimisation bit) will take more than one week let me know. I'm not talking about elapsed time (waiting for decision etc.) but about the actual work effort.
PreferSameNode
@Techassi recently added feature-gate detection to operator-rs. When I looked at this for OpenShift detection ~a year ago, my understanding was that feature gates are removed some releases after a feature goes GA (https://kubernetes.io/docs/reference/command-line-tools-reference/feature-gates-removed/). The gate here is (PreferSameTrafficDistribution) is already GA so it will be removed in a few releases, which means we can't rely on feature-gate detection alone. We'd need "k8s ≥ 1.35 OR the gate is enabled", and operator-rs has no k8s version detection yet.
Open question: does a preferSameNodeAvailable helper belong in operator-rs or here?This is my understanding of the fields we'd need to set:
Model k8s version internalTrafficPolicy trafficDistribution DaemonSet < 1.35 Local unset DaemonSet >= 1.35 unset PreferSameNode Deployment < 1.35 unset unset Deployment >= 1.35 unset PreferSameNode It'd be really nice to have this optimisation as part of this issue.
But I only want this optimisation if it takes at most two days to implement. If you think it'll take longer skip it.- We want to allow two deployment models: Deployment and DaemonSet
and operator-rs has no k8s version detection yet.
And it doesn't need it, because that functionality is already in
kube, see https://docs.rs/kube/latest/kube/struct.Client.html#method.apiserver_version. Only feature gate detection is currently not available in upstreamkubeand that's the reason why I added it tooperator-rsin stackabletech/operator-rs#1207.Reacted by Lars FranckeUpdate on the
PreferSameNodesituation. @Maleware convinced me to NOT do any of that.So, the table would be:
Model internalTrafficPolicy DaemonSet Local Deployment unset (Defaults to Cluster) Which means we need none of the feature detection stuff.
Reacted by Maximilian Wittich and Sebastian Bernauer- moved this from Selected for Development to In Refinement in Stackable End-to-End Coordination
on Jul 28, 2026 Outline
Due to problems with OPA being purely a DS (reported and mitigated by customers in various ways) we want to come up with the possibility to deploy with either a DaemonSet or a Deployment. The option should enable the customer to choose it's deploy mode tailored to their use case.
Goals
- OPA should be deployable as DaemonSet and Deployment.
Non-Goals
- Deployment specific automatism like auto-scaling
- Multi svc's. One service type per deployment mode. E.g. deployment should only come with one svc and
internalTrafficPolicy: Clusterby default. Supersedes Add internal trafific policy 'Cluster' role service and discovery CM #688
Proposal
OpenPolicyAgent is an AuthZ component and thus naturally needs to answer frequent requests across the platform products such as Kafka, HBase and Trino. Due to variance in workload and infrastructure and thus request amount to OPA, a user should be able to choose what is needed.
Option A
spec: clusterConfig: deploymentMode: DaemonSet # or Deployment; default DaemonSet servers: roleGroups: default: replicas: 3 # Field already exists today. Keep ignoring it in DaemonSet mode.
Advantage: Non-breaking change due to
deploymentMode: DaemonSetas a default, gives the user a choice to opt-in the deployment mode. Cluster wide config visible top level.
Disadvantage: If we add new roles (other then servers), deploy them differently might be difficult.Option B
spec: servers: roleConfig: deploymentMode: DaemonSet # or Deployment; default DaemonSet roleGroups: default: replicas: 3 # Field already exists today. Keep ignoring it in DaemonSet mode.
Advantage: Non-breaking change due to
deploymentMode: DaemonSetas a default, gives the user a choice to opt-in the deployment mode.replicasshould be ignored (or emit a warning) in DaemonSet case
Disadvantage: Confusing if servers.roleGroups.default.replicas has a value while DaemonSet is enabled.Option C
spec: clusterConfig: deploymentMode: DaemonSet # default for rolegroups that don't override servers: roleGroups: latency-critical: # inherits deploymentMode: DaemonSet clusterConfig # replicas is ignored for DaemonSet high-throughput: deploymentMode: Deployment replicas: 3
Advantage: Multiple deployment variants in one crd, fits current Stackable approach on roleGroups.
Disadvantage: This does not work with our service topology. The role level service
<cluster-name>selects every pod of theserversrole and can carry exactly oneinternalTrafficPolicy. With a DaemonSet role group and a Deployment role group behind it, either value is wrong for one half:Localmakes the Deployment pods unreachable from nodes they don't run on,Clusterdrops the node-local routing the DaemonSet exists for. Resolving that requires per-role-group services with differing policies, which contradicts the multi-svc Non-Goal above. Changing the name or the semantics of the existing recommended<cluster-name>service is not an option, since products consume it via the discovery ConfigMap and we would break existing installations.Non-options
- nodeSelector/preferedAffinities to restrict DS: currently this would break installations since
internalTrafficPolicy: Localis set. Products on nodes not propagated by OPA will fail to reach it.
Silent Assumptions
Deployment modes need to assume some service architecture to function properly.
DaemonSet
Assumes all nodes are covered via the DaemonSet and thus a svc with
internalTrafficPolicy: Localroutes every request to the local (present on node) pod to avoid network penalties.Deployment
Assumes a fixed replica count on arbitrary nodes. Therefore a svc with
internalTrafficPolicy: Clusteris required to ensure all products from any node can reach a OPA pod. This is the equivalent of not settinginternalTrafficPolicyat all as agreed. The following assumes we leaveinternalTrafficPolicyunset when statinginternalTrafficPolicy: Cluster.Proposal for internalTrafficPolicy (extends on what was agreed on before, to be discussed)
The assumptions above aren't obvious by default and, although sensible, might not fit every use-case. This needs to be considered since in field, we were mitigating performance issues imposed by local traffic via a custom svc with
internalTrafficPolicy: Cluster. Parallelism has shown to outperform network penalties by amplitude. Since evidence is restricted to specific setups this might not be usable for everyone and we shouldn't force anyone into a approach which might not be suited. Thus we should allow the user to choose his own strategy (based on option A)spec: servers: roleConfig: deploymentMode: DaemonSet # or Deployment; default DaemonSet internalTrafficPolicy: Cluster roleGroups: default: replicas: 3 # Field already exists today. Keep ignoring it in DaemonSet mode.
Which would allow the user to configure the strategy on his own, while when unset defaults to the assumptions stated above. This should be non-breaking as currently we deploy a DaemonSet with
internalTrafficPolicy: Local.For completeness, overriding
internalTrafficPolicycan be achieved already today byobjectOverridesand wouldn't need a official implementation to have a workaround present.Recommendation
I personally would go for
Option Bplus explicitly allowing to configure internalTrafficPolicy on roleConfig level. It provides the most flexibility, holds onto Stackable typical structures and is (with sensible defaults) non-breaking. Also to move from option B to option C looks like a non-breaking change and can be done if demand is seen.ToDo's (estimate)
- Implement CRD field (Decision first)
- Parallel deployment mode (DaemonSet/Deployment)
- Deployment equivalent of DaemonSetConditionBuilder (opa_controller.rs:125)
- Service policies
- default soft per-node anti-affinity (crd/mod.rs:309-311)
- update-strategy decision (DaemonSet e.g. doesn't need PDBs a deployment might)
- ClusterResources::delete_orphaned_resources (operator-rs cluster_resources.rs:667-678) should also clean up Deployments as customer can switch from Deployment -> DaemonSet
- deploymentMode dimension in tests/test-definition.yaml.
- Documentation
Looking good to me. I'll leave the choice of which option you'll pick to a decision. I think I'd prefer B myself but no strong opinion.
Maybe if your recommendation is B you could update the example forinternalTrafficPolicyto use the B model as well. I was confused for a minute.Reacted by Maximilian WittichCleaning up deployments change in
operator-rs: stackabletech/operator-rs#1256@Maleware I just had a brain fart. So far we didn't set any affinity between OPA clients and OPA Pods, as the OPA Pods where running everywhere.
Now that there is a (real) chance that the OPA Pods are only running on some nodes, we should go through all operators and add the according affinities.
Glade fully this is pretty easy, I did it as a blueprint for Trino: stackabletech/trino-operator#924- added a commit that references this issue
on Aug 11, 2026 WIP: #873
- moved this from In Refinement to In Progress in Stackable End-to-End Coordination
on Aug 17, 2026 - added a commit that references this issue
on Sep 2, 2026 Done with: #873
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsDevelopment: Done
- StatusShow more project fieldsDone
On Clusters with many nodes you can not easily deploy a DaemonSet. It might be a solution to deploy a Deployment and accept the performance hit (we need measurements!) or get away with it by adding preferedAffinities to all OPA-consuming services.
For one customer this is a show-stopper to using our opa-operator, simply because they have so many k8s nodes that they can not afford the resources.
They hand-roll their own Deployment instead.
TODOs: (Omitted during development, taken as a follow up task)
Go through all other operators and set up needed affinities to the OPA Pods. stackabletech/trino-operator#924 serves as a blueprint
---> Rollout of affinities superseded by stackabletech/issues#895