Project reference ↗

A worker machine can have disk, operating-system or container-runtime problems even while its applications appear to be running. Node Problem Detector collects evidence from system logs and health checks, then reports problems through Kubernetes events and node conditions. This makes machine-level faults visible alongside workload health and available to other automation. It is a reporting component: responding to an alert, repairing a machine or replacing a worker still requires an operator or another system.

Current guidance

The upstream Kubernetes project runs node monitors and exports findings through Events, NodeConditions and optional metrics integrations. Different monitors cover system logs, health checks, statistics or custom scripts. The original chart’s broad fault list should not imply that every monitor is enabled or compatible with every node operating system.

Check whether the Kubernetes provider already installs a managed detector before deploying another instance. Match configuration to actual log locations, service names and container runtime, and restrict host access to what those monitors need. A missing log path can make a running DaemonSet appear healthy while providing little useful detection.

Define how each reported condition leads to an alert or a repair action. Detecting a fault does not automatically replace the node, and aggressive remediation can amplify false positives. Test a harmless synthetic condition and verify the event, metric and intended response end to end. Keep custom scripts bounded by timeouts and resource limits. This review describes the reporting boundary, not an executed hardware-failure test or a guarantee that all node problems are detected.

Historical upstream link check · 2026-10-09

The recorded upstream address responded successfully (HTTP 200) on 2026-10-09. GitHub confirms that helm/charts is archived: this is a historical chart distribution, not evidence that the application itself is retired. Link availability does not certify the historical installation instructions or current security support.

Source for this check ↗

Website availability is separate from project, chart and image support. Use the current guidance and primary sources on this page to assess the distribution.

The original record

Historical Kubedex content

Original publication: 2018-11-04T05:46:53+00:00. Preserved for context. Commands, versions, prices and results below reflect the original research.

node-problem-detector aims to make various node problems visible to the upstream layers in cluster management stack. It is a daemon which runs on each node, detects node problems and reports them to apiserver. node-problem-detector can either run as a DaemonSet or run standalone. It also runs as a Kubernetes Addon enabled by default in the GKE cluster.

Background

There are tons of node problems could possibly affect the pods running on the node such as:

  • Infrastructure daemon issues: ntp service down;
  • Hardware issues: Bad cpu, memory or disk, ntp service down;
  • Kernel issues: Kernel deadlock, corrupted file system;
  • Container runtime issues: Unresponsive runtime daemon;
  • …

Currently these problems are invisible to the upstream layers in cluster management stack, so Kubernetes will continue scheduling pods to the bad nodes.

To solve this problem, we introduced this new daemon node-problem-detector to collect node problems from various daemons and make them visible to the upstream layers. Once upstream layers have the visibility to those problems, we can discuss the remedy system.

Problem API

node-problem-detector uses

  • Event: Temporary problem that has limited impact on pod but is informative should be reported as Event.

Problem Daemon

A problem daemon is a sub-daemon of node-problem-detector. It monitors a specific kind of node problems and reports them to node-problem-detector.

A problem daemon could be:

  • A tiny daemon designed for dedicated usecase of Kubernetes.
  • An existing node health monitoring daemon integrated with node-problem-detector.

Currently, a problem daemon is running as a goroutine in the node-problem-detector binary. In the future, we’ll separate node-problem-detector and problem daemons into different containers, and compose them with pod specification.

List of supported problem daemons:

Problem Daemon NodeCondition Description
KernelMonitor KernelDeadlock A system log monitor monitors kernel log and reports problem according to predefined rules.
AbrtAdaptor None Monitor ABRT log messages and report them further. ABRT (Automatic Bug Report Tool) is health monitoring daemon able to catch kernel problems as well as application crashes of various kinds occurred on the host. For more information visit the link.
CustomPluginMonitor On-demand(According to users configuration) A custom plugin monitor for node-problem-detector to invoke and check various node problems with user defined check scripts. See proposal here.

The post node-problem-detector appeared first on kubedex.com.

Sources & further reading

  1. Node Problem Detector monitors and exporters
  2. Recovered historical source

Spotted something that needs another look?

Help improve this page →