Skip to content

Support running without a DaemonSet (if desired) #525

Description

@sbernauer

On Clusters with many nodes you can not easily deploy a DaemonSet. It might be a solution to deploy a Deployment and accept the performance hit (we need measurements!) or get away with it by adding preferedAffinities to all OPA-consuming services.

For one customer this is a show-stopper to using our opa-operator, simply because they have so many k8s nodes that they can not afford the resources.
They hand-roll their own Deployment instead.

TODOs: (Omitted during development, taken as a follow up task)
Go through all other operators and set up needed affinities to the OPA Pods. stackabletech/trino-operator#924 serves as a blueprint

---> Rollout of affinities superseded by stackabletech/issues#895

Activity

  1. NickLarsenNZ commented on Apr 30, 2025

    @NickLarsenNZ
    Member

    nodeSelectors could be used to restrict the amount of nodes they are deployed to.
    I'm unsure if that then makes the service unavailable to workloads off those nodes.

    In any case, having to option to deploy as a Deployment would be nice.

  2. Maleware commented on Mar 18, 2026

    @Maleware
    Member

    Easy solution to this might be: #688

  3. moved this to Selected for Development in Stackable End-to-End Coordinationon Jul 10, 2026
  4. lfrancke commented on Jul 21, 2026

    @lfrancke
    Member

    I looked at this issue today and don't think we need much more than what's already here.
    Trying to sum it up:

    • We want to allow two deployment models: Deployment and DaemonSet
      • To define how to select between those two. I think best would be an explicit CRD choice between those two -> needs a decision
    • When using Deployment don't set internalTrafficPolicyat all
    • K8s has PreferSameNodeas of 1.35, see below for details

    If you believe implementing this (see below for the optimisation bit) will take more than one week let me know. I'm not talking about elapsed time (waiting for decision etc.) but about the actual work effort.

    PreferSameNode

    @Techassi recently added feature-gate detection to operator-rs. When I looked at this for OpenShift detection ~a year ago, my understanding was that feature gates are removed some releases after a feature goes GA (https://kubernetes.io/docs/reference/command-line-tools-reference/feature-gates-removed/). The gate here is (PreferSameTrafficDistribution) is already GA so it will be removed in a few releases, which means we can't rely on feature-gate detection alone. We'd need "k8s ≥ 1.35 OR the gate is enabled", and operator-rs has no k8s version detection yet.
    Open question: does a preferSameNodeAvailable helper belong in operator-rs or here?

    This is my understanding of the fields we'd need to set:

    Model k8s version internalTrafficPolicy trafficDistribution
    DaemonSet < 1.35 Local unset
    DaemonSet >= 1.35 unset PreferSameNode
    Deployment < 1.35 unset unset
    Deployment >= 1.35 unset PreferSameNode

    It'd be really nice to have this optimisation as part of this issue.
    But I only want this optimisation if it takes at most two days to implement. If you think it'll take longer skip it.

  5. Techassi commented on Jul 23, 2026

    @Techassi
    Member

    and operator-rs has no k8s version detection yet.

    And it doesn't need it, because that functionality is already in kube, see https://docs.rs/kube/latest/kube/struct.Client.html#method.apiserver_version. Only feature gate detection is currently not available in upstream kube and that's the reason why I added it to operator-rs in stackabletech/operator-rs#1207.

  6. self-assigned this
    on Jul 27, 2026
  7. lfrancke commented on Jul 27, 2026

    @lfrancke
    Member

    Update on the PreferSameNode situation. @Maleware convinced me to NOT do any of that.

    So, the table would be:

    Model internalTrafficPolicy
    DaemonSet Local
    Deployment unset (Defaults to Cluster)

    Which means we need none of the feature detection stuff.

  8. moved this from Selected for Development to In Refinement in Stackable End-to-End Coordinationon Jul 28, 2026
  9. Maleware commented on Jul 28, 2026

    @Maleware
    Member

    Outline

    Due to problems with OPA being purely a DS (reported and mitigated by customers in various ways) we want to come up with the possibility to deploy with either a DaemonSet or a Deployment. The option should enable the customer to choose it's deploy mode tailored to their use case.

    Goals

    • OPA should be deployable as DaemonSet and Deployment.

    Non-Goals

    Proposal

    OpenPolicyAgent is an AuthZ component and thus naturally needs to answer frequent requests across the platform products such as Kafka, HBase and Trino. Due to variance in workload and infrastructure and thus request amount to OPA, a user should be able to choose what is needed.

    Option A

      spec:
        clusterConfig:
          deploymentMode: DaemonSet  # or Deployment; default DaemonSet 
        servers:
          roleGroups:
            default:
              replicas: 3   # Field already exists today. Keep ignoring it in DaemonSet mode.

    Advantage: Non-breaking change due to deploymentMode: DaemonSet as a default, gives the user a choice to opt-in the deployment mode. Cluster wide config visible top level.
    Disadvantage: If we add new roles (other then servers), deploy them differently might be difficult.

    Option B

      spec:
        servers:
          roleConfig:
            deploymentMode: DaemonSet   # or Deployment; default DaemonSet
          roleGroups:
            default:
              replicas: 3   # Field already exists today. Keep ignoring it in DaemonSet mode.

    Advantage: Non-breaking change due to deploymentMode: DaemonSet as a default, gives the user a choice to opt-in the deployment mode. replicas should be ignored (or emit a warning) in DaemonSet case
    Disadvantage: Confusing if servers.roleGroups.default.replicas has a value while DaemonSet is enabled.

    Option C

      spec:
        clusterConfig:
          deploymentMode: DaemonSet  # default for rolegroups that don't override
        servers:
          roleGroups:
            latency-critical:
              # inherits deploymentMode: DaemonSet clusterConfig
              # replicas is ignored for DaemonSet
            high-throughput:
              deploymentMode: Deployment
              replicas: 3

    Advantage: Multiple deployment variants in one crd, fits current Stackable approach on roleGroups.

    Disadvantage: This does not work with our service topology. The role level service <cluster-name> selects every pod of the servers role and can carry exactly one internalTrafficPolicy. With a DaemonSet role group and a Deployment role group behind it, either value is wrong for one half: Local makes the Deployment pods unreachable from nodes they don't run on, Cluster drops the node-local routing the DaemonSet exists for. Resolving that requires per-role-group services with differing policies, which contradicts the multi-svc Non-Goal above. Changing the name or the semantics of the existing recommended <cluster-name> service is not an option, since products consume it via the discovery ConfigMap and we would break existing installations.

    Non-options

    • nodeSelector/preferedAffinities to restrict DS: currently this would break installations since internalTrafficPolicy: Local is set. Products on nodes not propagated by OPA will fail to reach it.

    Silent Assumptions

    Deployment modes need to assume some service architecture to function properly.

    DaemonSet

    Assumes all nodes are covered via the DaemonSet and thus a svc with internalTrafficPolicy: Local routes every request to the local (present on node) pod to avoid network penalties.

    Deployment

    Assumes a fixed replica count on arbitrary nodes. Therefore a svc with internalTrafficPolicy: Cluster is required to ensure all products from any node can reach a OPA pod. This is the equivalent of not setting internalTrafficPolicy at all as agreed. The following assumes we leave internalTrafficPolicy unset when stating internalTrafficPolicy: Cluster.

    Proposal for internalTrafficPolicy (extends on what was agreed on before, to be discussed)

    The assumptions above aren't obvious by default and, although sensible, might not fit every use-case. This needs to be considered since in field, we were mitigating performance issues imposed by local traffic via a custom svc with internalTrafficPolicy: Cluster. Parallelism has shown to outperform network penalties by amplitude. Since evidence is restricted to specific setups this might not be usable for everyone and we shouldn't force anyone into a approach which might not be suited. Thus we should allow the user to choose his own strategy (based on option A)

      spec:
        servers:
          roleConfig:
            deploymentMode: DaemonSet   # or Deployment; default DaemonSet
            internalTrafficPolicy: Cluster
          roleGroups:
            default:
              replicas: 3   # Field already exists today. Keep ignoring it in DaemonSet mode.

    Which would allow the user to configure the strategy on his own, while when unset defaults to the assumptions stated above. This should be non-breaking as currently we deploy a DaemonSet with internalTrafficPolicy: Local.

    For completeness, overriding internalTrafficPolicy can be achieved already today by objectOverrides and wouldn't need a official implementation to have a workaround present.

    Recommendation

    I personally would go for Option B plus explicitly allowing to configure internalTrafficPolicy on roleConfig level. It provides the most flexibility, holds onto Stackable typical structures and is (with sensible defaults) non-breaking. Also to move from option B to option C looks like a non-breaking change and can be done if demand is seen.

    ToDo's (estimate)

    • Implement CRD field (Decision first)
    • Parallel deployment mode (DaemonSet/Deployment)
    • Deployment equivalent of DaemonSetConditionBuilder (opa_controller.rs:125)
    • Service policies
    • default soft per-node anti-affinity (crd/mod.rs:309-311)
    • update-strategy decision (DaemonSet e.g. doesn't need PDBs a deployment might)
    • ClusterResources::delete_orphaned_resources (operator-rs cluster_resources.rs:667-678) should also clean up Deployments as customer can switch from Deployment -> DaemonSet
    • deploymentMode dimension in tests/test-definition.yaml.
    • Documentation
  10. lfrancke commented on Jul 28, 2026

    @lfrancke
    Member

    Looking good to me. I'll leave the choice of which option you'll pick to a decision. I think I'd prefer B myself but no strong opinion.
    Maybe if your recommendation is B you could update the example for internalTrafficPolicy to use the B model as well. I was confused for a minute.

  11. Maleware commented on Jul 31, 2026

    @Maleware
    Member

    Cleaning up deployments change in operator-rs: stackabletech/operator-rs#1256

  12. sbernauer commented on Aug 7, 2026

    @sbernauer
    MemberAuthor

    @Maleware I just had a brain fart. So far we didn't set any affinity between OPA clients and OPA Pods, as the OPA Pods where running everywhere.
    Now that there is a (real) chance that the OPA Pods are only running on some nodes, we should go through all operators and add the according affinities.
    Glade fully this is pretty easy, I did it as a blueprint for Trino: stackabletech/trino-operator#924

  13. Maleware commented on Aug 11, 2026

    @Maleware
    Member

    WIP: #873

  14. moved this from In Refinement to In Progress in Stackable End-to-End Coordinationon Aug 17, 2026
  15. Maleware commented on Oct 6, 2026

    @Maleware
    Member

    Done with: #873

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions