A cluster can report healthy machines while a user still cannot start an application or reach a service. Kuberhealthy checks those workflows directly by running small, temporary test workloads and collecting their results. For example, a check can try to resolve a network name or create and contact a service. This makes it useful alongside resource monitoring: it asks whether an operation works, rather than only whether the components appear to be running.
Current guidance
Kuberhealthy runs short-lived check pods that report workflow results to the controller, which exposes status through its interfaces and Prometheus metrics. Current upstream documentation uses HealthCheck custom resources. Old examples using a different resource name should be treated as version-specific rather than pasted into a current installation.
Design checks around a concrete user or platform workflow, such as DNS resolution or the ability to create and reach a small service. Keep their privileges and test data narrowly scoped, and make cleanup reliable after both success and failure. A check that cannot remove its resources can gradually exhaust the capacity it is supposed to monitor.
Define timeouts, failure thresholds and an actionable response for each check. Protect any result-reporting endpoint and verify that stale results or a failed checker are distinguishable from a healthy run. An in-cluster checker also needs an external observer if total cluster unavailability must generate an alert. This review identifies current component and API boundaries; it does not establish an end-to-end availability guarantee for a particular check suite.
Historical upstream link check · 2026-10-09
The recorded upstream address responded successfully (HTTP 200) on 2026-10-09. GitHub confirms that helm/charts is archived: this is a historical chart distribution, not evidence that the application itself is retired. Link availability does not certify the historical installation instructions or current security support.
Website availability is separate from project, chart and image support. Use the current guidance and primary sources on this page to assess the distribution.
Historical Kubedex content
Original publication: 2018-12-22T08:21:40+00:00. Preserved for context. Commands, versions, prices and results below reflect the original research.
Easy synthetic testing for Kubernetes clusters. Supplements other solutions like Prometheus nicely.
Kuberhealthy performs stynthetic tests from within Kubernetes clusters in order to catch issues that would otherwise go unnoticed. Instead of trying to identify all the things that could potentially go wrong, Kuberhealthy replicates real workflow and watches carefully for the expected Kubernetes behavior to occur. Kuberhealthy serves both a JSON status page and a Prometheus metrics endpoint for integration into your choice of alerting solution. More checks will be added in future versions to better cover service provisioning, DNS resolution, disk provisioning, and more.
Some examples of errors Kuberhealthy has detected in production:
- Nodes where new pods get stuck in
Terminating
due to CNI communication failures - Nodes where new pods get stuck in
ContainerCreating
due to disk scheduler errors - Nodes where new pods get stuck in
Pending
due to Docker daemon errors - Nodes where Docker or Kubelet crashes or has restarted
- A node that can not provision or terminate pods quickly enough due to high IO wait
- A pod in the
kube-system
namespace that is restarting too quickly - A Kubernetes component that is in a non-ready state
- Intermittent failures to access or create custom resources
- Kubernetes system services remaining technically “healthy” while their underlying pods are crashing too much
- kube-scheduler
- kube-apiserver
- kube-dns
The post kuberhealthy appeared first on kubedex.com.
Sources & further reading
Spotted something that needs another look?
Help improve this page →