An application has stopped working, and several components could be responsible. Troubleshooting is the process of narrowing that uncertainty using observations: what failed, who is affected, when it started and what changed.

Begin with the user's symptom, then follow the request or workload until you find where actual behavior differs from the expected result. Form one explanation and choose a small check that could disprove it. Restarting everything may hide the evidence and create more problems, so capture useful logs and state first.

Try it in a lab

Break one dependency in a disposable application. Collect timestamps, relevant events and component logs, form a hypothesis, change one variable and record the result. Restore the original configuration afterward.

Check your understanding

Write a short incident note that another engineer can follow without knowing the answer in advance. Include evidence that rules out a plausible alternative cause.

Before you start

Restarting everything can erase evidence and add new failures. Capture the useful state before making a broad change.

Read the official guide

Use the project documentation for version-specific commands and prerequisites. Record the versions and results of your own exercise. This page proposes a learning activity; it does not report a Kubedex test.

This is a newly written study reference at an address from the original Kubedex course outline. The original lesson was not recovered. It does not include course enrolment, progress tracking or a certificate.

Sources & further reading

  1. Primary learning documentation

Spotted something that needs another look?

Help improve this page →