Software is useful only while people can rely on it. Site reliability engineering (SRE) applies engineering work to keeping services dependable: understanding user needs, measuring failures, reducing repetitive work and improving recovery.

Start with one service and a clear question, such as whether users can complete a request when one instance fails. Monitoring, deployment automation and controlled failure exercises help answer that question. The historical outline below groups the supporting subjects; it does not provide a current hosted lab or assessment.

Practice reliability as an engineering task

Choose a service and define what users need from it. Add useful signals, a recovery procedure and a controlled failure exercise. Learn to distinguish an availability symptom from its cause and to make an improvement that prevents recurrence.

Kubernetes, Go, CI/CD and observability are tools in that work. Distributed-systems concepts, capacity planning and incident communication connect them. Chaos experiments belong in a bounded environment with an abort condition and recovery owner. The restored outline does not imply that a current hosted lab or assessment is available.

The original record

Historical Kubedex content

Preserved for context. Commands, versions, prices and results below reflect the original research.

Sources & further reading

  1. Google SRE books
  2. Kubernetes workload documentation
  3. Recovered historical source (Common Crawl index)

Spotted something that needs another look?

Help improve this page →