Lab 0 · The Core 5 & 5#
The one idea that unlocks the rest — Kubernetes From The Ground Up#
David Chan, Claude Opus 4.8 AI-Symbiosis Research · July 2026
Most Kubernetes tutorials hand you a glossary of forty nouns and hope a mental model precipitates out. It doesn’t. So this series goes the other way: one load-bearing idea first, then the handful of concepts and commands that idea makes inevitable — and then we go break a real cluster and diagnose it blind.
This is Lab 0: the spine. It’s the map I wish I’d had before my first kubectl describe — the foundation the hands-on labs build on.
The ONE idea: Kubernetes is a reconciliation loop#
Forget everything else for a moment. Kubernetes is a control loop, like a thermostat.
YOU declare CONTROLLERS
"desired state" ──▶ compare desired vs actual ──▶ act to close the gap
(a YAML object) (the reconcile loop) (create / kill / reschedule)
▲ │
└────────── observe reality ◀──────┘You never say “start this container.” You declare “I want one replica of signal-api that passes its /health check,” and a controller notices that reality doesn’t match and works — forever — to fix it.
That single idea explains ~80% of Kubernetes behavior:
- Self-healing is just the loop: kill a pod and it comes back, because actual ≠ desired.
CrashLoopBackOffis the loop trying so hard it has to back off — an exponential delay so it doesn’t hammer a container that keeps dying.- Rolling updates, autoscaling, rollbacks — all the same loop, different controllers.
Internalize this and the rest of Kubernetes stops being magic and becomes bookkeeping.
Your app is a stack of objects#
When you kubectl apply a Deployment, you don’t create one thing — you trigger a chain, each layer owning the one below it:
Deployment ← YOU write this. "Keep N healthy copies; roll updates this way."
└─ ReplicaSet ← auto-created. Its ONE job: keep exactly N pods alive.
└─ Pod ← the unit K8s schedules. 1+ containers sharing an IP + lifecycle.
└─ Container ← your actual process, from the image.Why it matters: a Pod is the smallest thing you debug. A rolling update is nothing more exotic than a new ReplicaSet scaling up while the old one scales down — two ReplicaSets coexisting for a few seconds.
Brain vs. muscle#
The cluster splits cleanly into a control plane (decides what should be true) and nodes (make it true, locally).
CONTROL PLANE (the brain) NODE (the muscle)
┌─────────────────────────────────────────┐ ┌────────────────────────────────────┐
│ kube-apiserver the front door; every │◀─ kubelet ─▶│ kubelet the node's agent; runs │
│ request goes through it │ (reports) │ pods, reports status │
│ etcd the database; desired + │ │ runtime containerd: pulls images,│
│ actual state live here │ │ starts containers │
│ scheduler picks WHICH node a pod │ │ kube-proxy programs Service routing│
│ runs on │ └────────────────────────────────────┘
│ controller-mgr runs the reconcile loops │
└─────────────────────────────────────────┘One immediate payoff: kubectl only ever talks to the apiserver. So when the apiserver is down, every command returns connection refused — not because your app is broken, but because the brain’s front door never opened. Knowing the topology tells you which failure you’re even looking at.
Manager + worker — and the debugging tax#
Step back and the shape is a manager / worker split. Traditional software is one thing: a process that decides and does. Kubernetes cleaves those apart — the control plane decides (what should be true), the nodes do (make it true, locally).
That split is the whole reason Kubernetes exists. Because the manager doesn’t do the work itself, it can add workers, kill them, reschedule them, replace a dead one — all without touching the brain. That decoupling is self-healing and autoscaling. You get elasticity for free, because “the thing that decides” was pried apart from “the thing that runs.”
But there’s a tax, and it’s worth naming honestly: more moving parts means more places to fail. A monolith can’t have a “the scheduler couldn’t place me” failure — Kubernetes can. Distributing the work distributes the failure surface too.
Here’s the twist, though — and it’s the point of this whole series. Kubernetes doesn’t leave you blind to which box broke. The status string localizes the failure to a layer before you run a single command:
Pending → MANAGER side (scheduler couldn't place the pod)
ImagePullBackOff → WORKER side (kubelet couldn't fetch the image)
CrashLoopBackOff → WORKER side (container ran, then died)
routes to nothing → the GLUE (labels/selectors — neither box)So the honest trade is: you pay a more complex failure surface, and in return the architecture hands you a map to navigate it. Learning to read that map is the skill. The rest of these labs are that map, drilled on a live cluster.
🧠 The 5 concepts that carry their weight#
1. Desired state + reconciliation — the concept. You declare what should be true; controllers loop forever to make reality match. Everything below is a consequence.
2. Pod — the atom. The smallest thing K8s schedules and debugs. One or more containers sharing an IP, storage, and lifecycle. Pods are cattle, not pets: disposable, replaceable, never patched in place — killed and reborn.
3. Deployment (which owns a ReplicaSet) — the manager. “Keep N healthy copies, and roll updates this way.” This is what gives you self-healing and zero-downtime deploys.
4. Service — the stable front door. Pods are mortal; their IPs churn constantly. A Service is a permanent virtual IP + DNS name that load-balances across whatever pods currently match — so callers never chase moving targets.
5. Labels & Selectors — the glue.
There are no hard wires in Kubernetes. A Service finds its pods by matching labels (app: signal-api), not by ID; a Deployment owns its pods the same way. This loose coupling is powerful — and a top source of silent bugs: wrong label → Service routes to nothing, with no error.
These five snap together: a Deployment keeps Pods alive → a Service fronts them → labels wire Service ↔ Pods → the reconcile loop enforces all of it.
⌨️ The 5 commands that are your everyday hands#
| # | Command | What it’s for | The mental hook |
|---|---|---|---|
| 1 | kubectl get <res> | State at a glance | “What does the cluster think is true?” — add -o wide, -w (watch), -A (all namespaces) |
| 2 | kubectl describe <res> <name> | Full narrative + Events | “What did the machinery DO, step by step, and where did it stall?” — 80% of triage lives here |
| 3 | kubectl logs <pod> [--previous] | The app’s own stdout | “What did the process say as it died?” — --previous reads the dead container on a crash loop |
| 4 | kubectl apply -f <file> | Declare / change desired state | “Make reality match this file.” — idempotent; re-run it safely |
| 5 | kubectl exec -it <pod> -- sh | Step inside a running container | “Let me poke around from the inside” — check env, curl localhost, test DNS |
The three that solve most incidents#
get → describe → logs. Outside-in: the cluster’s opinion → the machinery’s actions → the app’s confession. Each taps a different layer of the loop, so together they triangulate.
Two bonus verbs you’ll reach for fast#
kubectl rollout status/undo deployment/<name>— watch a deploy land, or instantly roll back a bad one.kubectl delete pod <name>— kill a pod and watch the ReplicaSet resurrect it: the reconcile loop, made visible in one command.
Where pod failures live in the lifecycle#
Here’s the architectural kicker that turns this from vocab into a debugging superpower. A pod is born in stages, and each classic failure is stuck at a different stage — which tells you which lens holds the evidence:
Scheduler finds a node → kubelet pulls image → container starts → probes pass → Running
│ │ │ │
▼ ▼ ▼ ▼
PENDING ImagePullBackOff CrashLoopBackOff readiness fail →
(no node fits: (bad image name / (process starts never Ready; old
resources, taints) no pull secret) then exits) pod stays up)
CoreDNS failure = the cluster's phone book is down → pods can't resolve each other by name- Pending → never scheduled → evidence is in
describeEvents (scheduler messages). Logs are useless — no container exists yet. - ImagePullBackOff → node can’t fetch the image →
describeEvents again (kubelet pull errors). Still no logs. - CrashLoopBackOff → image ran, process died →
describegives the exit code, but the smoking gun is inlogs --previous. - CoreDNS failure → a networking-layer problem, not your pod → you debug a different component entirely.
The meta-skill: the status string tells you which stage failed, which tells you which lens holds the evidence. That’s the whole game.
Setting up your lab cluster#
You don’t need a cloud bill to do these labs. A local single-node cluster on your laptop behaves like the real thing for everything that matters here. I use minikube on Docker — and critically, I do not break my live showcase cluster; the whole point of a sandbox is that you can wreck it freely.
1. Start a cluster.
minikube start --driver=docker --cpus=2 --memory=38002. Build your app image straight into the cluster. This sidesteps needing a private registry pull secret — the image resolves locally:
# from your app directory (the one with the Dockerfile)
minikube image build -t ghcr.io/you/your-app:latest ./app3. Deploy a working baseline first. You want to see healthy → broken → fixed, so start from healthy. A trimmed Deployment + Service is enough (set imagePullPolicy: IfNotPresent so it uses the local image):
kubectl apply -f your-app.yaml
kubectl rollout status deployment/your-app
kubectl get pods -o wide # expect: 1/1 RunningNow you have a green cluster to break in Lab 1.
The setup itself is the first triage lesson#
Standing this up, I hit two failures before the cluster was even healthy — and both are exactly the kind of thing these labs teach. Worth calling out so you recognize them:
apiserver process never appeared. The cluster wouldn’t come up: the API server and etcd were crash-looping (attempt 14 and climbing). The cause was a stale minikube profile left over from months earlier — corrupted etcd state from an old boot.kubectl get nodesjust returnedconnection refused(remember: the brain’s front door never opened). Fix: wipe the rotten profile and start clean.minikube delete && minikube start --driver=docker --cpus=2 --memory=3800Note it was not a resource problem — the host had 20+ GB free. “apiserver never appeared” reads like a scary control-plane bug; it was just stale state.
A boot that looked hung but wasn’t. The next start sat silent for minutes with no output. It wasn’t frozen — it was downloading a 500 MB Kubernetes preload at a crawl. Running the start in the foreground (instead of piping it somewhere) showed the progress bar and the truth. Lesson: “silent” and “stuck” are not the same thing — get the process to tell you what it’s doing before you assume it’s dead.
Both of these are the same muscle Lab 1 trains: don’t guess — make the system tell you where it hurts.
Solution: a worked triage — CrashLoopBackOff#
Theory sticks once you’ve walked one failure end to end. Here’s the whole method applied to a single real bug — the exact path from “something’s wrong” to “fixed and verified.”
The situation. You deploy signal-api and a pod won’t stay up:
NAME READY STATUS RESTARTS
signal-api-5d8cfcc999-ggs52 0/1 CrashLoopBackOff 6 (17s ago)Work the three lenses, outside-in.
Lens 1 — get: what does the cluster think? CrashLoopBackOff — the reconcile loop is restarting a container that keeps dying, backing off exponentially so it doesn’t hammer. The status string already localizes it: the container starts, then exits — a worker-side, post-startup failure. Not Pending (scheduler), not ImagePullBackOff (image fetch).
Lens 2 — describe: what did the machinery do?
kubectl describe pod signal-api-5d8cfcc999-ggs52Two things to read. The Events show the loop — Pulled → Created → Started → BackOff, over and over. Note Container image ... already present on machine: the image is fine, it starts. That rules out an image problem. Then, under Last State: Terminated, the Exit Code: 1. That number is a classifier before you read a single log line: 1 = the application chose to die (a normal error exit); 137 would be OOMKilled; 139 a segfault. Exit 1 says: ask the app why.
Lens 3 — logs --previous: the app’s dying words. The current container is a fresh restart with empty logs — the evidence is in the corpse:
kubectl logs signal-api-5d8cfcc999-ggs52 --previousAttributeError: attribute 'app_handler' not found in module 'main'There’s the smoking gun. But here’s the trap — and the real lesson.
The bug is not in the app. The instinct is “add app_handler to the code.” Wrong question. Ask instead: who is telling the app to load app_handler? Look at the deployment’s start command:
kubectl get deployment signal-api -o yaml | grep -A4 "command:"command:
- uvicorn
- main:app_handler # ← the culprit
- --host
- 0.0.0.0main:app_handler is uvicorn’s module:attribute syntax — “import main.py, then load the object app_handler from it.” But the code defines no such object:
# main.py
app = FastAPI(title="DevOps Showcase API") # ← the object is named `app`app isn’t a FastAPI rule — it’s just the variable name. If the code said api = FastAPI(), the correct reference would be main:api. The name is whatever’s on the left of the =. So how do you know the right name cold? Two sources of truth: the code (grep "= *FastAPI(" main.py), and — the tell you’ll use in an interview — the image’s own Dockerfile CMD, which shipped a correct default:
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"]The image was built to run main:app. The deployment’s command: overrode that good default with a bad one. That’s the deeper takeaway: an image ships with a working default command; a command: in a manifest overrides it. Whenever you see command:/args: in a Deployment, ask “is this clobbering a default that already worked?”
The fix — correct the declaration, not the app:
kubectl edit deployment signal-api # change main:app_handler → main:app
kubectl rollout status deployment/signal-apiVerify — watch the reconcile loop converge:
$ kubectl get rs -l app=signal-api
NAME DESIRED CURRENT READY
signal-api-57878dc598 0 0 0 # original
signal-api-5d8cfcc999 0 0 0 # broken (app_handler)
signal-api-5d8dc5bc7b 1 1 1 # fixed (app) ← only this one wantedThree ReplicaSets, one alive. That table is the reconcile loop’s history — the new template scaled up healthy, the broken one scaled to 0. Kubernetes healed itself the moment the desired state was correct.
Postmortem (five lines). Symptom: pod in
CrashLoopBackOff, restarts climbing. Diagnosis:get→describe(exit code1; container started then died — not image/scheduling) →logs --previous(app_handler not found). Root cause: deploymentcommand:overrode the image’s correct default (main:app) withmain:app_handler, a name the code never defines. The bug was in the desired state, not the app. Fix: correctedcommandtomain:app. Verify: new ReplicaSet rolled out1/1; broken ReplicaSet scaled to0.
Symptom → tool → root cause → fix → verification. That’s the shape of every triage in this series.
What’s next in this series#
Reading about Kubernetes and owning it are different sports. Lab 0 was the map; the rest of this series is hands-on break → diagnose → fix on a real cluster — no answers handed over, just the three lenses and the lifecycle map above:
- Lab 1 · Failure triage — force
CrashLoopBackOff,ImagePullBackOff,Pending, and a CoreDNS outage; diagnose each blind. - Lab 2 · Observability end-to-end — a custom metric → recording rule → Grafana panel → a tripped alert.
- Lab 3 · SLOs & error budgets — define one, burn it in a simulated incident, read the burn-rate.
- Lab 4 · GitOps — put the cluster under Argo CD; change by PR, force drift, watch it reconcile.
- Lab 5 · Blast radius — bad deploy → rollback; failing readiness probe; a deliberate OOMKill — with a five-line postmortem for each.
The loop is the idea. The five concepts are the nouns. The five commands are the hands. Everything after this is just practice.