Software is useful only while people can rely on it. Site reliability engineering (SRE) applies engineering work to keeping services dependable: understanding user needs, measuring failures, reducing repetitive work and improving recovery.
Start with one service and a clear question, such as whether users can complete a request when one instance fails. Monitoring, deployment automation and controlled failure exercises help answer that question. The historical outline below groups the supporting subjects; it does not provide a current hosted lab or assessment.
Practice reliability as an engineering task
Choose a service and define what users need from it. Add useful signals, a recovery procedure and a controlled failure exercise. Learn to distinguish an availability symptom from its cause and to make an improvement that prevents recurrence.
Kubernetes, Go, CI/CD and observability are tools in that work. Distributed-systems concepts, capacity planning and incident communication connect them. Chaos experiments belong in a bounded environment with an abort condition and recovery owner. The restored outline does not imply that a current hosted lab or assessment is available.
Historical Kubedex content
Preserved for context. Commands, versions, prices and results below reflect the original research.
Course Content
Sources & further reading
Spotted something that needs another look?
Help improve this page →