Describe the bug
GPUCluster.spec.daemonsets.updateStrategy: OnDelete is accepted by the CRD, but the generated DaemonSets use RollingUpdate for the DRA kubelet plugin, DRA validator, DCGM and DCGM exporter. Consequently, the rendered resources allow automatic replacement on a pod-template change despite the requested manual update strategy.
The four templates hard-code RollingUpdate. Their render data already includes Daemonsets.UpdateStrategy, and the GPUCluster apply path does not run the ClusterPolicy strategy transformer.
To Reproduce
At main revision 4fdfb1db7ddb87ba8969c52c634a5aeff104ab5f, use this GPUCluster configuration fragment (enable DCGM to include that optional component):
spec:
daemonsets:
updateStrategy: OnDelete
dcgm:
enabled: true
Render the four component manifests through their existing getManifestObjects helpers. Each DaemonSet has spec.updateStrategy.type: RollingUpdate, although the common configuration requests OnDelete. The DRA plugin and validator also include rollingUpdate.maxUnavailable: "100%".
Expected behavior
All four DaemonSets should honor OnDelete and omit rollingUpdate settings. An omitted or explicit RollingUpdate strategy should preserve the current output, including the DRA components' 100% maximum unavailable setting.
The proposed scope keeps readiness semantics unchanged: healthy old pods do not satisfy the requested configuration, so GPUCluster stays NotReady until an administrator replaces them and the updated pods are available. This follows the existing ClusterPolicy OnDelete revision checks in controllers/object_controls.go. The GPUCluster state manager continues reconciling every component while a preceding one is NotReady, so the other DaemonSets still receive their desired configuration. Maintainer feedback on this contract is welcome.
Environment (please provide the following information):
- GPU Operator: main at
4fdfb1db7ddb87ba8969c52c634a5aeff104ab5f.
- OS/architecture: Linux ARM64; Go 1.27.1.
- Reproduction: real template renderer and typed Kubernetes objects; GPU runtime is not needed for this reproduction.
- Live GPU Operator deployment: not tested. Separate CPU Kubernetes lifecycle validation is recorded with the proposed PR.
Information to attach
The regression covers all four operands with omitted, explicit RollingUpdate, OnDelete, and OnDelete plus leftover rolling-update settings. Readiness cases cover old, partially replaced, updated-but-unavailable and fully updated pods under both strategies. Test results will accompany the proposed PR.
Describe the bug
GPUCluster.spec.daemonsets.updateStrategy: OnDeleteis accepted by the CRD, but the generated DaemonSets useRollingUpdatefor the DRA kubelet plugin, DRA validator, DCGM and DCGM exporter. Consequently, the rendered resources allow automatic replacement on a pod-template change despite the requested manual update strategy.The four templates hard-code
RollingUpdate. Their render data already includesDaemonsets.UpdateStrategy, and the GPUCluster apply path does not run the ClusterPolicy strategy transformer.To Reproduce
At main revision
4fdfb1db7ddb87ba8969c52c634a5aeff104ab5f, use this GPUCluster configuration fragment (enable DCGM to include that optional component):Render the four component manifests through their existing
getManifestObjectshelpers. Each DaemonSet hasspec.updateStrategy.type: RollingUpdate, although the common configuration requestsOnDelete. The DRA plugin and validator also includerollingUpdate.maxUnavailable: "100%".Expected behavior
All four DaemonSets should honor
OnDeleteand omitrollingUpdatesettings. An omitted or explicitRollingUpdatestrategy should preserve the current output, including the DRA components'100%maximum unavailable setting.The proposed scope keeps readiness semantics unchanged: healthy old pods do not satisfy the requested configuration, so GPUCluster stays NotReady until an administrator replaces them and the updated pods are available. This follows the existing ClusterPolicy OnDelete revision checks in
controllers/object_controls.go. The GPUCluster state manager continues reconciling every component while a preceding one is NotReady, so the other DaemonSets still receive their desired configuration. Maintainer feedback on this contract is welcome.Environment (please provide the following information):
4fdfb1db7ddb87ba8969c52c634a5aeff104ab5f.Information to attach
The regression covers all four operands with omitted, explicit RollingUpdate, OnDelete, and OnDelete plus leftover rolling-update settings. Readiness cases cover old, partially replaced, updated-but-unavailable and fully updated pods under both strategies. Test results will accompany the proposed PR.