Project reference ↗

KEDA connects workload demand signals to Kubernetes application scaling. A queue of waiting jobs or a stream of unprocessed events can tell it when more worker replicas are needed, even when CPU usage is a poor guide. It is useful for services whose demand is visible in an external system. Additional replicas still need machines to run on and enough downstream capacity, so KEDA is one part of the scaling design.

Choose a meaningful signal

Define the unit of work, acceptable queue delay, processing time and maximum safe concurrency. A larger replica count can overload a database or external API without making the end-to-end system faster. Check the selected scaler's authentication and failure behavior in the version you install.

Prove the whole path

Test scale-up from the expected idle state, draining during scale-down, missing metrics and source unavailability. Account for pod startup time and whether nodes have capacity for the additional pods. Agree on a fallback capacity and alert when the autoscaler cannot observe demand.

See the autoscaling guide to distinguish workload replicas, resource sizing and node provisioning. KEDA does not replace the node capacity layer.

Sources & further reading

  1. Official deployment guidance
  2. Kubernetes HPA behavior

Spotted something that needs another look?

Help improve this page →