Your applications need more room during a traffic spike, but keeping the same machines running all night can waste money. A node autoscaler adds and removes the machines that run Kubernetes Pods. On Amazon EKS, Cluster Autoscaler, Karpenter and EKS Auto Mode offer different ways to do this.

Cluster Autoscaler grows and shrinks node groups you configure in advance. Karpenter chooses and creates machines that fit waiting workloads, using the limits you set. EKS Auto Mode includes AWS-managed, Karpenter-based node provisioning along with other managed infrastructure features. Comparing Karpenter versus Cluster Autoscaler versus Auto Mode means choosing both how machines are selected and how much of their management you keep.

Choose Cluster Autoscaler when existing node groups already fit your applications. Choose self-managed Karpenter when you need more flexible machine selection or node configuration and can maintain its controller. Choose Auto Mode when its supported configuration fits and the AWS management charge is worth the work it takes off your team. HPA or KEDA still handles how many application copies should run; node scaling gives those copies somewhere to run.

How each option creates and manages nodes

Cluster Autoscaler changes the size of configured node groups. With EKS managed node groups, AWS provides the node-group service while you configure the groups and operate the autoscaler. Mixed instance policies are possible: AWS recommends equivalent CPU, memory and GPU shapes because the first instance type is used for scheduling simulation. A smaller alternate shape can leave Pods unschedulable. Match the Cluster Autoscaler version to the Kubernetes version, as described in AWS's Cluster Autoscaler guidance.

Karpenter works from workload constraints and NodePools to select capacity. You manage the controller, its IAM permissions, interruption integration and upgrades. The AWS provider's EC2NodeClass controls details such as AMI selection and subnet/security-group discovery. That flexibility is useful when your node image or rollout process is part of the platform contract; it also leaves that contract with your team.

Auto Mode includes Karpenter-based provisioning alongside managed networking, load-balancing and storage capabilities. Its managed controllers run outside your workload cluster, and nodes use the AWS-managed Bottlerocket model. Custom AMIs are not supported. You still own application configuration and the cluster-version decision. Read the Auto Mode operating responsibilities before treating it as simply a hosted copy of your current Karpenter installation.

Reject incompatible options before comparing price

Start with one representative difficult workload: a stateful service, a GPU job, an agent that needs host access, or an application with a fixed volume zone. Record its operating system, image requirements, architecture, devices, DaemonSets, networking and shutdown behavior. Then check the target's supported configuration. A platform that runs your stateless demo but cannot support a required security agent is not ready for a general migration.

Auto Mode also changes the node-lifecycle contract: AWS sets a maximum node lifetime of 21 days, which you can shorten, and provides no direct SSH or SSM access to those nodes. That maximum is not a guaranteed uninterrupted run; consolidation or interruption can replace capacity sooner. Validate checkpointing, restart and shutdown behavior for long-running jobs, and establish diagnostics that do not depend on logging into a node. AWS also notes that blocking PDBs or other settings can require intervention before the 21-day limit. See the documented Auto Mode lifecycle and access model.

The API names also matter. Both self-managed Karpenter and Auto Mode use NodePools, but their AWS node classes differ: EC2NodeClass in karpenter.k8s.aws versus NodeClass in eks.amazonaws.com. Auto Mode has its own EC2-related labels. Copying a NodePool without translating its class reference and requirements can leave it unusable. Compare against the Auto Mode NodePool specification.

Calculate cost from a workload, not a node count

A smaller fleet is not automatically cheaper: instances differ in price, and a busy replacement process can overlap old and new capacity. Use the same workload, requests, availability constraints and measurement interval. Include steady operation, a representative burst and the quiet period after the burst. Record completed work and user-visible latency alongside the bill.

AWS charges Auto Mode management fees in addition to EC2, based on instance type and runtime. Those fees are independent of the EC2 purchase option. EKS cluster fees, storage, load balancing and networking also remain relevant; a Spot or Savings Plan discount is not a blanket discount on the whole architecture.

Build two estimates: the AWS bill and the team's operating cost. For the bill, sum the measured runtime of each instance type at its applicable rate, then add the corresponding management fees and other services. For operating cost, record actual time spent on controller upgrades, AMI maintenance and incidents. Keep a one-off migration estimate separate from monthly running cost.

For example, if your own measured comparison predicts an extra $300 per month in management fees and $100 in extra infrastructure, it needs more than $400 in monthly savings elsewhere to reduce total cost. At an assumed internal rate of $100 per hour, four saved hours only break even, before migration effort. These numbers illustrate the calculation; they are not AWS prices or Kubedex benchmark results. A team may still choose the managed option for reliability or staffing reasons, but should state that decision explicitly.

Consolidation is a trade-off with disruption

Consolidation can remove empty nodes or move workloads onto fewer or cheaper nodes. A low CPU graph does not show whether the Pods can move: requests, affinity, topology spread, volume placement and available capacity constrain the answer. If nodes remain after a burst, work through the consolidation diagnostic guide before changing policy.

Keep three controls separate. A PodDisruptionBudget protects workload availability during supported voluntary evictions. A NodePool disruption budget limits the start of particular node disruptions. Node expiration and interruption have their own behavior. A zero voluntary-disruption budget is not a promise that a node cannot disappear, and do-not-disrupt is not protection against every termination path. Read the installed version's disruption rules, including termination grace periods, before using an annotation to protect a long-running job.

Set resource limits and instance-family constraints deliberately. The Auto Mode cost guidance notes that built-in pools have no resource limits and custom pools do not inherit all built-in restrictions. For self-managed Karpenter, NodePool limit checking is eventually consistent and rapid scale-out can overrun a limit. Resource limits constrain capacity; they are not an exact currency budget. Alert on spending separately and decide whether reaching a capacity limit should leave work queued.

The EKS 1.37 consolidation default needs an explicit decision

The October 1 Auto Mode release notes change the default for newly created Auto Mode NodePools on EKS 1.37 to Balanced. This policy weighs savings against disruption instead of taking every qualifying saving. It is an Auto Mode policy; do not copy it into self-managed Karpenter merely because the resource is also called NodePool.

For custom Auto Mode pools, an omitted field can acquire a new value when GitOps deletes and recreates the object. Set spec.disruption.consolidationPolicy explicitly when you need stable intent across recreation. The built-in general-purpose and system pools are reconciled by AWS and use Balanced on 1.37; a workload needing another policy belongs on an appropriate custom pool. Fewer evictions is the design goal, not a measured savings or performance result for your application.

Move one workload cohort and keep capacity to return

A migration changes scheduling and infrastructure ownership, so start with selected workloads and distinct capacity. Preserve enough old capacity for rollback; a manifest reversal cannot resurrect a terminated instance or its local data.

  1. Record the existing controller versions, NodePools or node groups, AMI selection, labels, taints, requests and disruption settings. Identify the add-ons that must remain available to create replacement capacity.
  2. Choose a small workload cohort with observable success criteria: startup time, Pending duration, completed jobs, request errors and a permitted interruption rate. Include restart and scale-down behavior, not only a successful deployment.
  3. Make placement deliberate with selectors and tolerations. A toleration permits placement but does not require it. Check where every Pod actually lands.
  4. Exercise a representative burst and a quiet period. Confirm node registration, storage attachment, network access, Pod identity and workload recovery. Inspect the resulting bill and capacity overlap.
  5. Expand only after the cohort meets its criteria. Keep the old provisioner and capacity until its workloads have moved and rollback no longer depends on it.

AWS's Karpenter-to-Auto Mode procedure supports side-by-side migration using a tainted Auto Mode pool and workload selectors. It requires Karpenter 1.1 or later and warns against changing the shared NodePool/NodeClaim CRDs during the transition. It also delays enabling the broad general-purpose pool. Do not uninstall Karpenter immediately after enabling Auto Mode if the existing fleet still depends on it.

For Cluster Autoscaler-to-Karpenter changes, follow the upstream migration sequence and retain a supported place for the Karpenter controller to run. Two provisioners reacting to an unintended common workload set can obscure the experiment; define which capacity each owns.

Diagnose the layer that failed

If a Pod is Pending, inspect its scheduling events and requirements before assuming more nodes will solve it. If capacity was requested but a node never registered, inspect the NodeClaim and the node class's readiness, identity and network dependencies. If nodes register but workloads fail, move to the application, storage, network and identity checks. If the service works but cost remains high, inspect requests, blocked consolidation and replacement overlap.

Auto Mode diagnostics differ from a self-managed controller: its controller is not a Deployment whose logs you can simply tail in your cluster. Start with resource conditions and events, then the AWS troubleshooting workflow and support path. Preserve the timing and affected resource names so a report explains the failure stage.

This comparison is sourced guidance reviewed on 9 October 2026. Kubedex has not run an AWS cost benchmark or executed these EKS migrations. The useful outcome is a documented choice with measured workload results, explicit ownership and a reversible first rollout.

Which EKS node-scaling model fits?

Choose columns
Visible columns
Which EKS node-scaling model fits?
DecisionCluster AutoscalerSelf-managed KarpenterEKS Auto Mode
Capacity modelResize configured node groupsProvision from NodePool/workload constraintsAWS-managed provisioning using Auto Mode NodePools
Controller operationsYou install and maintain the autoscalerYou install and maintain KarpenterAWS operates the managed controllers
Node configurationNode-group and launch-template modelEC2NodeClass with provider-specific configurationAuto Mode NodeClass and supported AWS node model
Cost questionGroup shape, idle capacity and operating timeCapacity selection, disruption and operating timeEC2 plus management fees versus transferred operational work
Good starting fitEstablished groups that meet workload needsFlexible provisioning with a team able to own itSupported workloads where managed operations justify the fee
Migration evidenceCompatible autoscaler and realistic group shapesController availability and tested replacement behaviorWorkload compatibility and staged placement before decommissioning old capacity

6 rows

Sources & further reading

  1. AWS Cluster Autoscaler configuration and mixed instance policies
  2. EKS Auto Mode operating responsibilities
  3. Karpenter AWS node classes and AMI selection
  4. Auto Mode NodePool configuration and supported labels
  5. Current EKS pricing structure
  6. Karpenter disruption controls
  7. Karpenter NodePool limits
  8. Auto Mode cost controls
  9. Auto Mode release notes: EKS 1.37 Balanced default
  10. Migrate Karpenter workloads to Auto Mode
  11. Migrate from Cluster Autoscaler to Karpenter
  12. Auto Mode troubleshooting
  13. Auto Mode maximum node lifetime and node access model

Spotted something that needs another look?

Help improve this page →