k8s-prom-hpa is a tutorial showing how Kubernetes can add or remove application replicas using measurements collected by Prometheus. An adapter exposes the chosen measurements to the Horizontal Pod Autoscaler, which decides the desired replica count. It is useful for understanding that connection between monitoring and scaling. The repository's example versions are historical, so use the concept while selecting current APIs and testing whether the signal reflects useful application capacity.
Current guidance
The stefanprodan/k8s-prom-hpa repository demonstrates resource and custom-metric autoscaling using a sample application, Metrics Server, Prometheus and an adapter API. Its instructions target early Kubernetes versions and include an autoscaling/v2beta2 example. Preserve the tutorial’s conceptual distinction between node capacity and workload replicas, but update the APIs and components deliberately.
Current Kubernetes documentation describes autoscaling/v2 behavior and the resource, custom and external metrics interfaces. Prometheus scraping alone does not make a metric available to HPA; verify the adapter’s discovery rules, metric naming, labels and the API response for the target workload. Resource-utilization targets also depend on meaningful resource requests.
Test scaling with a controlled load and observe the metric timestamp, desired replicas, readiness and stabilization behavior. Include a missing or stale metric and a backend that cannot support additional replicas. Check that the chosen signal correlates with useful capacity rather than merely growing when the application is already failing. Pin the adapter and monitoring versions separately from the application and avoid running two controllers against the same replica target.
Historical upstream link check · 2026-10-09
The recorded upstream address responded successfully (HTTP 200) on 2026-10-09. GitHub does not mark stefanprodan/k8s-prom-hpa archived or disabled; this does not establish active maintenance, support or compatibility. Link availability does not certify the historical installation instructions or current security support.
Website availability is separate from project, chart and image support. Use the current guidance and primary sources on this page to assess the distribution.
Historical Kubedex content
Original publication: 2018-09-27T09:40:30+00:00. Preserved for context. Commands, versions, prices and results below reflect the original research.
Autoscaling is an approach to automatically scale up or down workloads based on the resource usage. Autoscaling in Kubernetes has two dimensions: the Cluster Autoscaler that deals with node scaling operations and the Horizontal Pod Autoscaler that automatically scales the number of pods in a deployment or replica set. The Cluster Autoscaling together with Horizontal Pod Autoscaler can be used to dynamically adjust the computing power as well as the level of parallelism that your system needs to meet SLAs. While the Cluster Autoscaler is highly dependent on the underling capabilities of the cloud provider that’s hosting your cluster, the HPA can operate independently of your IaaS/PaaS provider.
The Horizontal Pod Autoscaler feature was first introduced in Kubernetes v1.1 and has evolved a lot since then. Version 1 of the HPA scaled pods based on observed CPU utilization and later on based on memory usage. In Kubernetes 1.6 a new API Custom Metrics API was introduced that enables HPA access to arbitrary metrics. And Kubernetes 1.7 introduced the aggregation layer that allows 3rd party applications to extend the Kubernetes API by registering themselves as API add-ons. The Custom Metrics API along with the aggregation layer made it possible for monitoring systems like Prometheus to expose application-specific metrics to the HPA controller.
The Horizontal Pod Autoscaler is implemented as a control loop that periodically queries the Resource Metrics API for core metrics like CPU/memory and the Custom Metrics API for application-specific metrics.
The post k8s-prom-hpa appeared first on kubedex.com.
Sources & further reading
Spotted something that needs another look?
Help improve this page →