Operating-system updates sometimes leave a Kubernetes worker needing a restart, but rebooting several workers together can interrupt applications. Kured coordinates those restarts after the operating-system workflow signals they are needed. It marks a node unavailable for new work, attempts to move existing workloads away, then reboots it and returns it to service. It helps automate maintenance coordination; installing updates and deciding whether applications can tolerate the disruption remain separate responsibilities.
Current guidance
Kured now lives under kubereboot, following the verified relocation of the original repository. It detects a configured sentinel file or command, takes a Kubernetes-backed lock and coordinates cordon, drain, reboot and uncordon. Package installation and the decision that a reboot is needed remain responsibilities of the node operating-system workflow.
Review maintenance windows, time zones, lock behavior and the signals that block a reboot. The current configuration includes a force-reboot option that can proceed after a failed or timed-out drain; that is a deliberate disruption tradeoff, not a harmless recovery default. Check pod disruption budgets, local state and workloads that cannot move.
Test one disposable node through the complete sequence, including a stuck drain and an unexpected controller restart. Confirm that monitoring and the configured blocking-pod or alert conditions actually delay reboot as intended, and define how an abandoned lock is investigated. Limit host privileges to the documented need. This source review explains the coordination model and risky options without claiming an executed reboot test or universal application safety.
Historical upstream link check · 2026-10-09
The recorded upstream address redirects to https://github.com/kubereboot/kured and returned HTTP 200 on 2026-10-09. GitHub does not mark kubereboot/kured archived or disabled; this does not establish active maintenance, support or compatibility. GitHub resolves the old repository identity to kubereboot/kured. Link availability does not certify the historical installation instructions or current security support.
Website availability is separate from project, chart and image support. Use the current guidance and primary sources on this page to assess the distribution.
Historical Kubedex content
Original publication: 2018-09-13T20:20:27+00:00. Preserved for context. Commands, versions, prices and results below reflect the original research.
Kured (KUbernetes REboot Daemon) is a Kubernetes daemonset that performs safe automatic node reboots when the need to do so is indicated by the package management system of the underlying OS.
- Watches for the presence of a reboot sentinel e.g. /var/run/reboot-required
- Utilises a lock in the API server to ensure only one node reboots at a time
- Optionally defers reboots in the presence of active Prometheus alerts
- Cordons & drains worker nodes before a reboot, uncordoning them after
Kubernetes & OS Compatibility
The daemon image contains versions of k8s.io/client-go and the kubectl binary for the purposes of maintaining the lock and draining worker nodes. See the release notes for specific version compatibility information.
Additionally, the image contains a systemctl binary from Ubuntu 16.04 in order to command reboots. Again, although this has not been tested against other systems distributions there is a good chance that it will work.
The Reboot Problem
At Weaveworks the development and production clusters underpinning Weave Cloud are orchestrated with Kubernetes running on EC2, maintained with Terraform and Ansible.
The EC2 instances run Ubuntu 16.04 with unattended-upgrades enabled, so the machines need to be rebooted periodically (mainly in response to kernel upgrades). If they aren’t, the clusters are at risk from security vulnerabilities, and eventually, run out of disk space as the OS is unable to remove older kernels and modules.
The first attempt
Our initial approach to this problem was to trigger a Prometheus alert whenever the /var/run/reboot-required file appeared on any of the nodes. We tried coupling it with a manual process that entailed waiting for a safe moment – defined as no active alerts – before draining the application pods and then rebooting each node in turn.
Automation makes everything better
Whilst this worked in practice, the frequency of OS updates coupled with the number of nodes drove us eventually to an automated solution. And so for the past six months, all reboots have been conducted safely and automatically by kured, our Kubernetes reboot daemon.
During this time kured has affected hundreds of node reboots in our dev and prod clusters without human intervention – in fact, until the relatively recent addition of Slack notifications, we were mostly unaware that it was happening at all.
The post Kured appeared first on kubedex.com.
Sources & further reading
Spotted something that needs another look?
Help improve this page →