You have GPU machines, but your training job or model server still will not start. The hardware being present is only the first step: the host must have a working driver, Kubernetes must know how to allocate the device, and the application must fit the available GPU memory.
A device plugin advertises devices to Kubernetes so Pods can request them. Dynamic Resource Allocation (DRA) provides another device-request and allocation system using Kubernetes resources and a vendor driver. Kueue decides when queued workloads are allowed to use limited capacity; the Kubernetes scheduler still chooses where their Pods run. These components solve related problems at different stages.
Use the allocation method supported by your GPU vendor and cluster. Add Kueue when teams need to share capacity through queues and quotas. For troubleshooting, first distinguish a Pod waiting to be scheduled from a running application that runs out of GPU memory. This guide follows those separate paths; Kubedex has not exercised GPU hardware, DRA or Kueue for the article.
Build a compatibility record before installing controllers
Record the accelerator model, node OS and kernel, host driver, container runtime integration, Kubernetes release, device plugin or DRA driver, and workload image. Add the sharing mode and relevant feature gates. For a managed service, check which components the provider owns before installing an overlapping operator.
Use the vendor's supported combination rather than treating a running DaemonSet as proof of compatibility. NVIDIA's GPU Operator DRA instructions, for example, distinguish DRA components and their prerequisites. Other hardware has its own driver and runtime requirements.
Choose how a Pod requests a GPU
Device plugins advertise vendor-specific extended resources. A workload requests the advertised name, such as nvidia.com/gpu. For the conventional device-plugin GPU path, set a GPU limit; if a request is also specified it must match the limit. The unit is an allocation exposed by the plugin, so confirm whether that means a whole device, a hardware partition or a shared slot. Follow the Kubernetes GPU scheduling guide.
DRA describes devices through resources such as DeviceClass, ResourceSlice and ResourceClaim, with driver-specific allocation and preparation. Its core became stable in Kubernetes 1.34. Optional capabilities have independent feature states and driver requirements; a stable core API does not make every advanced allocation mode universally available. See the DRA concepts and 1.34 graduation notice.
Recent releases extend that model. Kubernetes 1.37 includes beta fractional capacity ranges while device compatibility groups remain separately gated. Treat these as capabilities to verify against the exact cluster and driver, not as instructions to express every GPU request as a fraction. The 1.37 DRA update lists the individual states.
Use Kueue to decide which waiting jobs may start
Kueue decides when a workload may consume quota and start; the Kubernetes scheduler still places its Pods. A LocalQueue provides a namespace-facing entry point, while ClusterQueue policies describe capacity and sharing. ResourceFlavors distinguish eligible resource configurations. This is useful when research, training or batch teams compete for a limited GPU fleet. It does not replace the serving router's request queue. See the Kueue overview.
Define which work may borrow quota, which work may be preempted and how checkpointing recovers it. For example, reserve the capacity required to meet interactive serving objectives before lending spare resources to restartable batch jobs. That is a proposed operating policy, not a claim that every Kueue integration implements the same behavior.
With DRA, verify Kueue's supported claim path, quota unit and feature gates. Admission is not proof that the scheduler can allocate a matching physical device. Current integration documentation also identifies restrictions around claim references, topology and shared-capacity accounting. Read the Kueue DRA integration for the installed release rather than copying a feature gate from an older example.
Find where the workload is waiting
Start with the owning workload. If it is suspended or waiting on queue admission, there may be no Pod for the scheduler to place yet. Inspect its queue, quota, admission checks and Kueue Workload conditions. If it is admitted and Pods are Pending, move to scheduling constraints. If a Pod is assigned but containers cannot start, inspect device preparation, runtime, image and storage failures. The Kueue troubleshooting guide follows that distinction.
These read-only commands assume a known context, namespace and Pod. Replace the names first. Run the Kueue query only where its CRDs are installed; a missing API is useful evidence of an installation mismatch, not proof that GPUs are unavailable.
gpu_context=your-reviewed-context
gpu_namespace=your-workload-namespace
gpu_pod=your-pending-pod
kubectl --context "$gpu_context" -n "$gpu_namespace" get pods -o wide
kubectl --context "$gpu_context" -n "$gpu_namespace" describe pod "$gpu_pod"
kubectl --context "$gpu_context" -n "$gpu_namespace" get events --field-selector "involvedObject.name=$gpu_pod" --sort-by=.metadata.creationTimestamp
kubectl --context "$gpu_context" -n "$gpu_namespace" get workloads.kueue.x-k8s.io
Read the exact event message. An untolerated taint, unmatched node selector, unavailable volume and insufficient device resource are separate constraints. Adding GPU nodes of the wrong type does not resolve an architecture or topology mismatch. Compare the workload's CPU, memory and storage needs too; free GPU capacity alone does not establish that the complete Pod fits.
Inspect the selected allocation path
For a device plugin, inspect node capacity and allocatable resources, the plugin's health and the advertised resource name. Compare already allocated resources rather than assuming a physical GPU count equals free capacity. A shared-resource configuration can advertise more slots than physical devices.
For DRA, inspect the DeviceClass and ResourceSlices visible to the scheduler, then the workload's ResourceClaims and their allocation state. After placement, inspect node-side preparation errors. Do not use the absence of nvidia.com/gpu in a node's allocatable list as the sole test for every DRA configuration: its device inventory and request path may be different.
For a DRA workload, the following queries continue the read-only diagnosis. Reuse the context and namespace above, then replace gpu_claim with the claim associated with the affected Pod. DeviceClasses and ResourceSlices are cluster-scoped and require suitable read permissions; ResourceClaims are namespaced.
gpu_claim=your-resource-claim-name
kubectl --context "$gpu_context" get deviceclasses.resource.k8s.io
kubectl --context "$gpu_context" get resourceslices.resource.k8s.io
kubectl --context "$gpu_context" -n "$gpu_namespace" get resourceclaims.resource.k8s.io
kubectl --context "$gpu_context" -n "$gpu_namespace" describe resourceclaims.resource.k8s.io "$gpu_claim"
Trace the Pod's claim reference to the concrete claim. When the Pod uses a ResourceClaimTemplate, inspect the generated claim rather than treating the template itself as an allocation. Compare the requested class and selectors with the driver's published inventory and the claim's allocation status. A permission error or unavailable API needs its own resolution; it does not establish a physical-device shortage. These relationships are described in the DRA resource model.
If the device cannot be used after allocation, move down to host and runtime checks. Have the node owner verify driver health, then run the vendor's small supported device test in an isolated workload before loading a large model. Keep the test image and configuration fixed while diagnosing the driver; otherwise two changing variables can conceal the fault.
Choose sharing by isolation and latency requirements
NVIDIA time-slicing exposes multiple allocations backed by the same physical GPU. It does not provide memory or fault isolation between those replicas. MIG partitions supported hardware and provides a different isolation model, with fixed profiles and capacity tradeoffs. The vendor sharing guide explains the distinction. Neither mechanism turns a busy device into several full-speed devices.
Evaluate whole-device allocation, partitions and sharing against a representative workload mix. Include one memory-heavy neighbor and one sustained compute-heavy neighbor, then compare latency tails, errors and completion time. If isolation is a requirement, verify the property provided by the exact hardware and mode rather than inferring it from a resource name.
For inference, account for model weights, runtime overhead and KV-cache growth. A model that loads successfully at startup can still fail with longer contexts or more simultaneous requests. Capacity planning therefore needs the actual prompt and output distributions described in the serving-path guide.
Plan multi-device placement and recovery together
A request for several GPUs is not just an arithmetic total across the cluster. Establish whether the application requires devices on one node or across nodes, which interconnect it expects, and what happens if one worker fails. Verify the workload controller's supported grouping and topology features before adding scheduler extensions.
For preemptible jobs, record checkpoint location, checkpoint frequency and restart behavior. For interactive serving, preserve enough available capacity during a drain to meet the service objective. Queue admission, Pod disruption and request draining operate at different layers; test the transitions between them.
Rehearse a driver or operator upgrade on a small, dedicated pool. Drain according to the workload's recovery procedure, change one layer, verify device discovery and allocation, then run the same application check and representative load. Retain a known-good pool until the new one passes. A rollback plan may require replacing nodes rather than changing only a container image.
Measure useful work and idle capacity
Separate admitted time, Pending time, model-loading time, useful execution and idle reservation. Track completed work meeting its latency or deadline objective against billed device time. High GPU utilization can coexist with poor user latency, while a deliberately warm replica may be necessary capacity rather than waste.
Before calling the setup ready, record one successful allocation, a deliberate capacity shortage, a failed-device or worker scenario and recovery. Include the exact versions, quota settings and sharing mode. Until those checks run on the intended hardware, this remains a sourced operating plan with unmeasured compatibility and performance.
Locate the wait before changing capacity
Choose columns
| Observed state | First place to inspect | What success establishes |
|---|---|---|
| Job waiting or suspended | Queue, quota, admission checks and workload controller | Permission to start consuming the declared quota |
| Pod Pending and unassigned | Scheduler events, node constraints and claims | A node and device allocation can satisfy the whole Pod |
| Pod assigned but cannot start | Device preparation, runtime, image and storage | The container can start with its required device |
| Container running but model fails | Runtime compatibility, model memory and serving configuration | The selected model can execute the declared workload |
4 rows
Sources & further reading
Spotted something that needs another look?
Help improve this page →