You want to use the CPU you pay for without making the application slow. But “90% CPU” is not a useful target until you know what the percentage measures. It might mean 90% of a whole machine, 90% of a container's requested CPU or 90% of its limit.

A CPU request is the amount Kubernetes uses when finding room for a container. A CPU limit can restrict how much CPU time it receives. A CPU-based Horizontal Pod Autoscaler (HPA) normally compares measured use with requests to decide how many application copies to run. For example, 450 millicores of use against a 500-millicore request is 90%, even on a machine with several CPU cores.

Choose the target by measuring application response time, errors and how quickly extra capacity starts. A batch job can use CPU differently from an interactive website, so the same percentage need not suit both. The original article at this address was not recovered; this is newly written guidance, not a recreation of its experiment.

Find the actual bottleneck

A workload can miss its latency objective at modest CPU utilization because it is waiting on storage, a database, a lock or a constrained dependency. Conversely, a batch worker may use most of its allocated CPU productively while meeting its deadline. Separate interactive traffic from throughput-oriented work and measure the user-visible result.

  1. Record each container's requests and limits, including sidecars. Check whether the autoscaler can obtain the required resource metrics.
  2. Graph request rate, queue delay, response percentiles, errors, CPU usage and throttling over the same interval.
  3. Observe the time from increased demand to ready capacity. Include pod startup and node provisioning where relevant.
  4. Run bounded load steps in a safe environment and identify the point where latency or error budgets deteriorate.

Choose headroom deliberately

A high target can leave little room for bursts while new replicas start. A low target can increase resource use without fixing a dependency bottleneck. Select a target from the workload's measured response curve, startup time and failure tolerance. Also reserve enough cluster capacity for disruption and node loss; an HPA cannot schedule pods onto capacity that does not exist.

Change one layer at a time

Changing CPU requests changes the utilization ratio and can affect scheduling. Changing limits can affect throttling. Changing replica targets changes concurrency against downstream services. Record the initial settings and alter one of these decisions at a time so a regression can be attributed and reversed.

Revert the configuration if user-visible latency or errors cross the agreed threshold, then investigate the actual bottleneck. Kubedex has not benchmarked your application or reproduced the missing original article's test. Continue with the autoscaling decision guide for the relationship between workload scaling and node capacity.

Sources & further reading

  1. Resource requests and limits
  2. Horizontal Pod Autoscaling

Spotted something that needs another look?

Help improve this page →