Lab 1 · Failure Triage
#

Break it, diagnose it blind, fix the declaration — Kubernetes From The Ground Up
#

David Chan, Claude Opus 4.8 AI-Symbiosis Research · July 2026


Lab 0 was the map: Kubernetes is a reconcile loop, a pod is born in stages, and the status string localizes a failure to a stage before you run a single command. This lab is that map, drilled on a live cluster. Four failures, each broken on purpose, each diagnosed blind — no peeking at what changed, just the three lenses:

get       →  what does the cluster think the state is?
describe  →  what did the machinery DO, step by step, and where did it stall?
logs      →  what did the app itself say as it died?

The whole discipline: don’t guess — make the system tell you where it hurts.

The diagnostic map
#

Print this on the inside of your skull. Each failure is stuck at a different birth-stage, and the stage decides which lens holds the evidence:

StatusStage that failedRestartsWinning lens
CrashLoopBackOffcontainer ran, then diedclimbinglogs --previous
ImagePullBackOffimage fetch (never started)0describe Events (kubelet)
Pendingscheduling (never placed)0describe Events (scheduler)
CoreDNS outagea shared cluster service— (app is healthy!)kube-system + a DNS test

Notice the Restarts column already disambiguates the first three: a climbing count means the container started and keeps dying; 0 means it never started at all. One column, half the diagnosis.


Failure 1 — CrashLoopBackOff
#

(Walked in full detail in Lab 0’s worked solution — here’s the compact version.)

Symptom. A pod won’t stay up; restarts climbing.

signal-api-5d8cfcc999-ggs52   0/1   CrashLoopBackOff   6 (17s ago)

Diagnose. get → CrashLoopBackOff (container starts then exits — worker-side, post-startup). describe → Events cycle Pulled → Created → Started → BackOff, and under Last State: Terminated, Exit Code: 1 (the app chose to die — a normal error exit, not 137 OOMKill or 139 segfault). The current container is a fresh restart with empty logs, so read the corpse:

kubectl logs signal-api-5d8cfcc999-ggs52 --previous
# → AttributeError: attribute 'app_handler' not found in module 'main'

Root cause. The deployment’s command: overrode the image’s correct default (uvicorn main:app) with main:app_handler — a name the code never defines (main.py has app = FastAPI(...)). The bug was in the desired state, not the app.

Fix. Correct the declaration: main:app_handler → main:app.

Postmortem. Symptom: CrashLoopBackOff, restarts climbing. Diagnosis: get → describe (exit 1, started-then-died) → logs --previous (app_handler not found). Root cause: command: overrode the image’s default with a bad module attribute. Fix: main:app. Verify: new ReplicaSet rolled out 1/1; broken one scaled to 0.


Failure 2 — ImagePullBackOff
#

Symptom. New pod won’t start — and note the tell: 0 restarts.

signal-api-6d6c88bc48-mcmjm   0/1   ImagePullBackOff   0

Zero restarts, because the container never started — there’s nothing to restart. The failure is upstream of the process even existing.

Diagnose. logs is useless here (no container ever ran):

kubectl logs signal-api-6d6c88bc48-mcmjm
# → Error from server (BadRequest): ... is waiting to start: trying and failing to pull image

That’s logs refusing you — proof there’s nothing to read. So all evidence is in describe Events (the kubelet’s pull attempts). The key move is to read the exact error word, because it forks the fix:

Error wordMeaningFix lives in
unauthorized / deniedcredentialsimagePullSecrets
not found / manifest unknownimage or tag doesn’t existthe image reference

Here it’s not found. Diff the broken pod’s image against the healthy pod’s:

kubectl get pod <healthy> -o jsonpath='{.spec.containers[0].image}'   # ...:latest
kubectl get pod <broken>  -o jsonpath='{.spec.containers[0].image}'   # ...:v1.4.2  ← ghost tag

Root cause. The deployment pointed at image tag :v1.4.2, which doesn’t exist. Not a credentials problem — a nonexistent tag. (No secret on earth fixes a tag that isn’t there.)

Fix. Point the image back at the tag that exists (:latest).

Postmortem. Symptom: ImagePullBackOff, 0 restarts. Diagnosis: logs empty (never started) → describe Events: Failed to pull ... not found → diffed broken vs healthy image → tag mismatch. Root cause: bad image tag :v1.4.2. Fix: corrected to :latest. Verify: deployment reconciled to the existing healthy ReplicaSet.

The trap: not found vs unauthorized are two different bugs with two different fixes. Don’t guess between them — make describe say the word.


Failure 3 — Pending
#

Symptom. The earliest failure of all — 0 restarts, and status Pending.

signal-api-d456bf64b-zj5bt   0/1   Pending   0

Pending ≠ ImagePullBackOff. ImagePull means the pod got scheduled to a node but couldn’t fetch its image. Pending means it never got a node at all — a manager-side scheduling failure. No node → no container → no image → no logs.

Diagnose. logs useless again, but for a different reason (never placed vs never pulled). All evidence is in describe Events — and this time the messages come from the scheduler:

kubectl describe pod signal-api-d456bf64b-zj5bt
# Events: 0/1 nodes are available: 1 Insufficient memory.

Insufficient memory has two completely different fixes — add hardware, or fix an over-inflated request — so before buying a bigger node, compare what the pod asks for against what it could possibly need:

kubectl get pod ... -o jsonpath='{.spec.containers[0].resources.requests}'   # memory: 500Gi (!)
kubectl get node minikube -o jsonpath='{.status.allocatable.memory}'          # ~28Gi

A tiny FastAPI app requesting 500 GiB — more RAM than exists in a small datacenter. The datacenter isn’t the bug.

Root cause. A fat-fingered memory request of 500Gi in the deployment — larger than any node. A bad declaration, not a hardware shortage.

Fix. Lower requests.memory to something sane (128Mi). Mind the units — this failure is a units bug:

You writeMeansFor
128Mi128 mebibytes✅ memory
128m128 millicpu (0.128 core)CPU only — on memory this is ~0.128 of a byte

Fix. kubectl set resources deployment/signal-api --requests=memory=128Mi --limits=memory=256Mi.

Postmortem. Symptom: Pending, 0 restarts, never scheduled. Diagnosis: logs useless → describe Events: scheduler Insufficient memory → request (500Gi) vs node (28Gi) vs need (128Mi). Root cause: fat-fingered 500Gi memory request. Fix: lowered to 128Mi. Verify: pod scheduled 1/1.


Failure 4 — CoreDNS: when the pod is green but nothing works
#

The finale is a different muscle. The first three were bugs in your deployment. This one isn’t — and it’s invisible to a casual glance.

Symptom. kubectl get pods shows your app 1/1 Running, no restarts, no errors. Green. Yet calls to other services by name start failing with Name or service not known.

Diagnose. Confirm it’s not your app (it’s healthy). Then prove DNS is broken — resolve a name from inside a pod:

kubectl exec deploy/signal-api -- python -c "import socket; print(socket.gethostbyname('kubernetes.default'))"
# → socket.gaierror: [Errno -3] Try again

That gaierror isn’t a crash or a refused connection — it’s a name-resolution failure. The pod asked “what’s the IP for kubernetes.default?” and nobody answered. So who answers DNS? A shared cluster service — in a namespace you haven’t looked at:

kubectl get pods -n kube-system
# etcd ... apiserver ... scheduler ... kube-proxy ... storage-provisioner ...
# — every core service Running, but CoreDNS is MISSING ENTIRELY.

Root cause. The coredns deployment was scaled to 0. No CoreDNS pod → nothing answers DNS → every service-by-name call in every pod fails — while each individual pod stays perfectly healthy.

Fix. It’s not in your deployment at all — it’s a cluster service in another namespace:

kubectl scale deployment coredns -n kube-system --replicas=2
kubectl exec deploy/signal-api -- python -c "import socket; print(socket.gethostbyname('kubernetes.default'))"
# → 10.96.0.1     ← the phone book is open again

Postmortem. Symptom: app 1/1 Running, but name lookups fail. Diagnosis: gethostbyname → gaierror (resolution, not connection) → kube-system missing coredns. Root cause: CoreDNS scaled to 0 — a shared service, not your app. Fix: scaled it back up. Verify: lookup returns an IP.

The namespace lesson
#

kubernetes.default reads right-to-left as <service>.<namespace> — the service kubernetes in the namespace default. Namespaces are partitions of the cluster, and they’re the reason this failure hid in plain sight:

   YOUR WORKLOADS  (namespace: default)      ← the "leaf" — your app; where kubectl looks by default
        │ depends on ↓
   SHARED FOUNDATION (namespace: kube-system) ← DNS, API server, scheduler, networking

kubectl get pods defaults to default — so all lab, you were looking at your layer, and it was green. Failures 1–3 lived there. Failure 4 lived a layer down, in the shared foundation your app stands on. The heuristic:

When your layer looks clean but things still break, drop down to the shared layer it depends on — but only for cross-cutting symptoms (naming, networking, scheduling, node health). A NullPointerException is still your layer.


The meta-pattern
#

Four failures, one spine:

  1. The status string localized the stage before any command — CrashLoop (died), ImagePull (fetch), Pending (schedule), green-but-broken (foundation).
  2. The stage chose the lens — logs --previous for the crash; describe Events for the two that never started; kube-system for the outage.
  3. Three of four bugs lived in the declaration — a bad command, a ghost image tag, an absurd resource request. Kubernetes was doing exactly what the YAML told it to. The fourth lived one layer down.

That’s the operator’s shift: from “debug my app” to “read the cluster.” You reason from symptom → stage → lens → root cause, and you fix the desired state — not by guessing.

Next up — Lab 2 · Observability End-to-End: stop reacting to failures and start seeing them coming — a custom metric → recording rule → Grafana panel → a tripped alert.