Project reference ↗

Installing a GPU in a worker does not by itself make that GPU usable inside a Kubernetes container. NVIDIA GPU Operator manages the software layers involved, including drivers, container integration and components that expose GPU resources to workloads. It helps platform teams configure compatible GPU workers consistently instead of assembling the stack on every machine. The supported hardware, operating system and allocation method still depend on the chosen operator release and its documented compatibility requirements.

Its supported configuration depends on drivers, hardware, node operating systems and Kubernetes versions; installing a chart cannot make arbitrary combinations compatible.

Choose one device-management path

Traditional device-plugin allocation and Dynamic Resource Allocation have different prerequisites. Kubernetes core DRA reached general availability in 1.34, but each driver and optional capability still has its own support requirements. Use NVIDIA's instructions for the exact DRA or device-plugin deployment you intend to operate.

Node and workload checks

Validate device discovery, driver loading, scheduling, isolation and recovery after a node reboot. Test the intended sharing or partitioning mechanism against actual workload memory and performance requirements. Keep a recovery plan for a driver or node-image change that leaves GPUs unavailable.

Read the scheduling guide before treating an advertised GPU count as usable inference capacity.

Sources & further reading

  1. NVIDIA DRA installation guidance
  2. Kubernetes DRA release context

Spotted something that needs another look?

Help improve this page →