Lab 1 · Failure Triage#
Break it, diagnose it blind, fix the declaration — Kubernetes From The Ground Up#
David Chan, Claude Opus 4.8 AI-Symbiosis Research · July 2026
Lab 0 was the map: Kubernetes is a reconcile loop, a pod is born in stages, and the status string localizes a failure to a stage before you run a single command. This lab is that map, drilled on a live cluster. Four failures, each broken on purpose, each diagnosed blind — no peeking at what changed, just the three lenses:
get → what does the cluster think the state is?
describe → what did the machinery DO, step by step, and where did it stall?
logs → what did the app itself say as it died?The whole discipline: don’t guess — make the system tell you where it hurts.
The diagnostic map#
Print this on the inside of your skull. Each failure is stuck at a different birth-stage, and the stage decides which lens holds the evidence:
| Status | Stage that failed | Restarts | Winning lens |
|---|---|---|---|
CrashLoopBackOff | container ran, then died | climbing | logs --previous |
ImagePullBackOff | image fetch (never started) | 0 | describe Events (kubelet) |
Pending | scheduling (never placed) | 0 | describe Events (scheduler) |
| CoreDNS outage | a shared cluster service | — (app is healthy!) | kube-system + a DNS test |
Notice the Restarts column already disambiguates the first three: a climbing count means the container started and keeps dying; 0 means it never started at all. One column, half the diagnosis.
Failure 1 — CrashLoopBackOff#
(Walked in full detail in Lab 0’s worked solution — here’s the compact version.)
Symptom. A pod won’t stay up; restarts climbing.
signal-api-5d8cfcc999-ggs52 0/1 CrashLoopBackOff 6 (17s ago)Diagnose. get → CrashLoopBackOff (container starts then exits — worker-side, post-startup). describe → Events cycle Pulled → Created → Started → BackOff, and under Last State: Terminated, Exit Code: 1 (the app chose to die — a normal error exit, not 137 OOMKill or 139 segfault). The current container is a fresh restart with empty logs, so read the corpse:
kubectl logs signal-api-5d8cfcc999-ggs52 --previous
# → AttributeError: attribute 'app_handler' not found in module 'main'Root cause. The deployment’s command: overrode the image’s correct default (uvicorn main:app) with main:app_handler — a name the code never defines (main.py has app = FastAPI(...)). The bug was in the desired state, not the app.
Fix. Correct the declaration: main:app_handler → main:app.
Postmortem. Symptom:
CrashLoopBackOff, restarts climbing. Diagnosis:get→describe(exit 1, started-then-died) →logs --previous(app_handler not found). Root cause:command:overrode the image’s default with a bad module attribute. Fix:main:app. Verify: new ReplicaSet rolled out1/1; broken one scaled to0.
Failure 2 — ImagePullBackOff#
Symptom. New pod won’t start — and note the tell: 0 restarts.
signal-api-6d6c88bc48-mcmjm 0/1 ImagePullBackOff 0Zero restarts, because the container never started — there’s nothing to restart. The failure is upstream of the process even existing.
Diagnose. logs is useless here (no container ever ran):
kubectl logs signal-api-6d6c88bc48-mcmjm
# → Error from server (BadRequest): ... is waiting to start: trying and failing to pull imageThat’s logs refusing you — proof there’s nothing to read. So all evidence is in describe Events (the kubelet’s pull attempts). The key move is to read the exact error word, because it forks the fix:
| Error word | Meaning | Fix lives in |
|---|---|---|
unauthorized / denied | credentials | imagePullSecrets |
not found / manifest unknown | image or tag doesn’t exist | the image reference |
Here it’s not found. Diff the broken pod’s image against the healthy pod’s:
kubectl get pod <healthy> -o jsonpath='{.spec.containers[0].image}' # ...:latest
kubectl get pod <broken> -o jsonpath='{.spec.containers[0].image}' # ...:v1.4.2 ← ghost tagRoot cause. The deployment pointed at image tag :v1.4.2, which doesn’t exist. Not a credentials problem — a nonexistent tag. (No secret on earth fixes a tag that isn’t there.)
Fix. Point the image back at the tag that exists (:latest).
Postmortem. Symptom:
ImagePullBackOff,0restarts. Diagnosis:logsempty (never started) →describeEvents:Failed to pull ... not found→ diffed broken vs healthy image → tag mismatch. Root cause: bad image tag:v1.4.2. Fix: corrected to:latest. Verify: deployment reconciled to the existing healthy ReplicaSet.
The trap: not found vs unauthorized are two different bugs with two different fixes. Don’t guess between them — make describe say the word.
Failure 3 — Pending#
Symptom. The earliest failure of all — 0 restarts, and status Pending.
signal-api-d456bf64b-zj5bt 0/1 Pending 0Pending ≠ ImagePullBackOff. ImagePull means the pod got scheduled to a node but couldn’t fetch its image. Pending means it never got a node at all — a manager-side scheduling failure. No node → no container → no image → no logs.
Diagnose. logs useless again, but for a different reason (never placed vs never pulled). All evidence is in describe Events — and this time the messages come from the scheduler:
kubectl describe pod signal-api-d456bf64b-zj5bt
# Events: 0/1 nodes are available: 1 Insufficient memory.Insufficient memory has two completely different fixes — add hardware, or fix an over-inflated request — so before buying a bigger node, compare what the pod asks for against what it could possibly need:
kubectl get pod ... -o jsonpath='{.spec.containers[0].resources.requests}' # memory: 500Gi (!)
kubectl get node minikube -o jsonpath='{.status.allocatable.memory}' # ~28GiA tiny FastAPI app requesting 500 GiB — more RAM than exists in a small datacenter. The datacenter isn’t the bug.
Root cause. A fat-fingered memory request of 500Gi in the deployment — larger than any node. A bad declaration, not a hardware shortage.
Fix. Lower requests.memory to something sane (128Mi). Mind the units — this failure is a units bug:
| You write | Means | For |
|---|---|---|
128Mi | 128 mebibytes | ✅ memory |
128m | 128 millicpu (0.128 core) | CPU only — on memory this is ~0.128 of a byte |
Fix. kubectl set resources deployment/signal-api --requests=memory=128Mi --limits=memory=256Mi.
Postmortem. Symptom:
Pending,0restarts, never scheduled. Diagnosis:logsuseless →describeEvents: schedulerInsufficient memory→ request (500Gi) vs node (28Gi) vs need (128Mi). Root cause: fat-fingered500Gimemory request. Fix: lowered to128Mi. Verify: pod scheduled1/1.
Failure 4 — CoreDNS: when the pod is green but nothing works#
The finale is a different muscle. The first three were bugs in your deployment. This one isn’t — and it’s invisible to a casual glance.
Symptom. kubectl get pods shows your app 1/1 Running, no restarts, no errors. Green. Yet calls to other services by name start failing with Name or service not known.
Diagnose. Confirm it’s not your app (it’s healthy). Then prove DNS is broken — resolve a name from inside a pod:
kubectl exec deploy/signal-api -- python -c "import socket; print(socket.gethostbyname('kubernetes.default'))"
# → socket.gaierror: [Errno -3] Try againThat gaierror isn’t a crash or a refused connection — it’s a name-resolution failure. The pod asked “what’s the IP for kubernetes.default?” and nobody answered. So who answers DNS? A shared cluster service — in a namespace you haven’t looked at:
kubectl get pods -n kube-system
# etcd ... apiserver ... scheduler ... kube-proxy ... storage-provisioner ...
# — every core service Running, but CoreDNS is MISSING ENTIRELY.Root cause. The coredns deployment was scaled to 0. No CoreDNS pod → nothing answers DNS → every service-by-name call in every pod fails — while each individual pod stays perfectly healthy.
Fix. It’s not in your deployment at all — it’s a cluster service in another namespace:
kubectl scale deployment coredns -n kube-system --replicas=2
kubectl exec deploy/signal-api -- python -c "import socket; print(socket.gethostbyname('kubernetes.default'))"
# → 10.96.0.1 ← the phone book is open againPostmortem. Symptom: app
1/1 Running, but name lookups fail. Diagnosis:gethostbyname→gaierror(resolution, not connection) →kube-systemmissingcoredns. Root cause: CoreDNS scaled to0— a shared service, not your app. Fix: scaled it back up. Verify: lookup returns an IP.
The namespace lesson#
kubernetes.default reads right-to-left as <service>.<namespace> — the service kubernetes in the namespace default. Namespaces are partitions of the cluster, and they’re the reason this failure hid in plain sight:
YOUR WORKLOADS (namespace: default) ← the "leaf" — your app; where kubectl looks by default
│ depends on ↓
SHARED FOUNDATION (namespace: kube-system) ← DNS, API server, scheduler, networkingkubectl get pods defaults to default — so all lab, you were looking at your layer, and it was green. Failures 1–3 lived there. Failure 4 lived a layer down, in the shared foundation your app stands on. The heuristic:
When your layer looks clean but things still break, drop down to the shared layer it depends on — but only for cross-cutting symptoms (naming, networking, scheduling, node health). A
NullPointerExceptionis still your layer.
The meta-pattern#
Four failures, one spine:
- The status string localized the stage before any command —
CrashLoop(died),ImagePull(fetch),Pending(schedule), green-but-broken (foundation). - The stage chose the lens —
logs --previousfor the crash;describeEvents for the two that never started;kube-systemfor the outage. - Three of four bugs lived in the declaration — a bad command, a ghost image tag, an absurd resource request. Kubernetes was doing exactly what the YAML told it to. The fourth lived one layer down.
That’s the operator’s shift: from “debug my app” to “read the cluster.” You reason from symptom → stage → lens → root cause, and you fix the desired state — not by guessing.
Next up — Lab 2 · Observability End-to-End: stop reacting to failures and start seeing them coming — a custom metric → recording rule → Grafana panel → a tripped alert.