Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 17 additions & 2 deletions content/en/docs/next/operations/gpu-container-workloads.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,9 +44,20 @@ With `driver.enabled=false` the operator uses the pre-installed host driver at i

## 1. Install the GPU Operator (container variant)

**Do not** add `cozystack.gpu-operator` to `bundles.enabledPackages` for this variant. The `iaas` bundle renders the GPU operator from `bundles.iaas.gpuOperatorVariant`, which only accepts `default` or `vgpu` — any other value, `container` included, makes the platform chart fail the Helm render (`packages/core/platform/templates/bundles/iaas.yaml`). Apply the `Package` CR directly instead; the platform controller installs it without a bundle entry and without the variant restriction.
The platform's `iaas` bundle deploys the gpu-operator Package CR when `cozystack.gpu-operator` is in `bundles.enabledPackages` and not in `bundles.disabledPackages`, with the variant taken from `bundles.iaas.gpuOperatorVariant`. Set it to `container` in the Platform Package values, and add `cozystack.gpu-operator` to your existing `enabledPackages` list rather than replacing it:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[MINOR] No transition note for readers who already applied the Package by hand

The released v1.6 page (and this page before the PR) told the reader to kubectl apply a cozystack.gpu-operator Package and to stay out of enabledPackages. A reader with that object who now follows step 1 ends up with the platform release adopting it: it gains helm.sh/resource-policy: keep and helm ownership annotations, and any spec.components overrides they had set keep applying silently, since the bundle renders none for container. One sentence saying "if you applied the Package by hand on an earlier release, the bundle adopts it; remove or move your spec.components overrides first" would close the gap between the two documented flows.


Apply a `Package` CR with `variant: container`:
```yaml
bundles:
iaas:
enabled: true
gpuOperatorVariant: container
enabledPackages:
- cozystack.gpu-operator
```

For this variant the bundle still renders the `KubeVirt` CR with its base settings, but adds no GPU wiring to it: no `HostDevices` feature gate and no `permittedHostDevices` table. The host driver stays bound, so no GPU on the node can be passed through to a VM.

If you need to override something the bundle does not expose (driver settings, custom node selectors, validator or dcgmExporter tweaks), hand-craft a `Package` CR named `cozystack.gpu-operator` with `variant: container` instead, and leave `cozystack.gpu-operator` out of `bundles.enabledPackages` so the platform release does not also manage that Package and overwrite your changes. Put the overrides under `spec.components.gpu-operator.values.gpu-operator` (the package wraps the upstream chart, so its settings sit under a nested `gpu-operator` key). The platform controller installs it without a bundle entry:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[MINOR] The "overwrite your changes" rationale contradicts vgpu.md and is not what the shipped helm-controller does

The instruction itself (keep cozystack.gpu-operator out of enabledPackages when you hand-craft the Package) is sound, but the reason given is the opposite of virtualization/vgpu.md:112 ("the manual Package-CR override path takes precedence over the bundle render whenever both exist"). Neither mechanism claim is verified, and the one I could trace points a third way: cozystack installs helm-controller v1.5.0 (internal/fluxinstall/manifests/fluxcd.yaml:8097), which sets upgrade.TakeOwnership unless the HelmRelease disables it (internal/action/upgrade.go:72), so an existing Package of that name is adopted into the platform release rather than fought over, and a spec.components key the bundle never renders (it renders none for container, iaas.yaml:196-210) is left in place by both apply methods. Suggest keeping the instruction and dropping the mechanism, or stating one mechanism and making vgpu.md:112 say the same.


```yaml
apiVersion: cozystack.io/v1alpha1
Expand All @@ -55,6 +66,10 @@ metadata:
name: cozystack.gpu-operator
spec:
variant: container
components:
gpu-operator:
values:
gpu-operator: {} # upstream gpu-operator chart overrides
```

```bash
Expand Down
2 changes: 1 addition & 1 deletion content/en/docs/next/virtualization/gpu.md
Original file line number Diff line number Diff line change
Expand Up @@ -102,7 +102,7 @@ For example, the database entry for A10 reads `2236 GA102GL [A10]`, which resul

## 2. KubeVirt is wired automatically

When `cozystack.gpu-operator` is in `bundles.enabledPackages`, Cozystack mirrors the chosen GPU variant into the `KubeVirt` Custom Resource for you. There is no `kubectl edit kubevirt` step.
When `cozystack.gpu-operator` is in `bundles.enabledPackages` (and not in `bundles.disabledPackages`) and `bundles.iaas.gpuOperatorVariant` is `default` (the package default) or `vgpu`, Cozystack mirrors the chosen GPU variant into the `KubeVirt` Custom Resource for you. There is no `kubectl edit kubevirt` step. The `container` variant gets none of this wiring: it keeps the host driver bound, so no GPU can reach a VM (see [containerized GPU workloads](/docs/next/operations/gpu-container-workloads/)).

Specifically, the platform injects:

Expand Down
2 changes: 1 addition & 1 deletion content/en/docs/next/virtualization/vgpu.md
Original file line number Diff line number Diff line change
Expand Up @@ -108,7 +108,7 @@ For Pascal to Ampere GPUs (V100, T4, A100, A30) the mdev model still applies. Fl

## KubeVirt configuration

When `cozystack.gpu-operator` is in `bundles.enabledPackages` (and not also in `bundles.disabledPackages`), the platform mirrors the chosen GPU variant into the `KubeVirt` CR automatically. There is no manual `kubectl patch` step.
When `cozystack.gpu-operator` is in `bundles.enabledPackages` (and not also in `bundles.disabledPackages`) and `bundles.iaas.gpuOperatorVariant` is `default` or `vgpu`, the platform mirrors the chosen GPU variant into the `KubeVirt` CR automatically. There is no manual `kubectl patch` step. The `container` variant gets no KubeVirt wiring at all: it keeps the host driver bound, so no GPU can reach a VM.

If you opt out of bundle management and hand-craft a `cozystack.gpu-operator` Package CR directly — typically to apply overrides the bundle does not expose — the platform does **not** auto-wire `HostDevices` or `permittedHostDevices` into the KubeVirt CR. In that flow you also hand-craft a `cozystack.kubevirt` Package CR with `components.kubevirt.values.extraFeatureGates: [HostDevices]` and the appropriate `permittedHostDevices` block. The escape-hatch values shape under `.gpu` below is documented for the bundle-managed flow only; the manual Package-CR override path takes precedence over the bundle render whenever both exist.

Expand Down
Loading