A Kubernetes machine looks nearly idle, but Karpenter keeps it running. Before deleting it, check whether its applications can move somewhere else without breaking their availability rules. Low measured CPU use does not tell you whether that move is possible.

Karpenter creates nodes to fit waiting workloads and can remove or replace nodes when capacity can be used more efficiently. That removal and replacement process is called consolidation. It must consider the resources Pods request, where their storage can attach and the rules protecting running applications.

Use this checklist to find the reason the node remains. Start with Karpenter's events and the relevant NodePool, the configuration that describes allowed nodes and disruption settings. Then change the specific blocker only if the application can tolerate the result.

Work through the blockers

  1. Identify the action. Is the controller considering consolidation, drift, expiration or an interruption? Different disruption paths have different behavior.
  2. Check workload protection. Inspect relevant PodDisruptionBudgets, availability and disruption annotations. A single unavailable replica can change what is permitted.
  3. Check NodePool policy. Review disruption budgets, schedules and consolidation policy for the exact controller release.
  4. Check alternate placement. Requests, node affinity, topology spread, taints, architecture and volume zones can prevent movement even when aggregate CPU looks free.
  5. Check capacity and price conditions. An allowed replacement must actually be available and meet the controller's requirements.

Use requests, not a utilization screenshot

Scheduling uses declared requests and constraints. A node with low observed CPU can still host Pods whose requests cannot fit elsewhere. Conversely, lowering requests solely to force consolidation can create memory or CPU contention later. Correct requests from workload evidence rather than the desired node count.

Make one reversible change

Reproduce the limiting condition in a non-production workload when possible. Change one policy, then compare events and placement. Do not remove all PDBs or force-delete production nodes to prove that consolidation could happen. Those actions bypass the availability requirement rather than solve the constraint.

When to stop

If consolidation causes churn, latency or repeated evictions, halt further voluntary changes and restore stable capacity while inspecting the policy. Keep enough capacity to undo the experiment. This is an evidence-driven diagnostic checklist; it does not contain fabricated before/after controller output or claim a reproduced cluster incident.

Sources & further reading

  1. Karpenter disruption and consolidation
  2. Kubernetes disruption budgets

Spotted something that needs another look?

Help improve this page →