Lab 5 · Blast Radius
#

Fail small, fail contained, fail reversible — Kubernetes From The Ground Up
#

David Chan, Claude Opus 4.8 AI-Symbiosis Research · July 2026


Every change you ship has a blast radius — the amount of damage if it goes wrong. You will ship bad things; that’s guaranteed. Platform engineering isn’t about preventing every failure — it’s about making sure a bad thing stays small, caught, and reversible instead of cascading into an outage. This finale practices the three safety mechanisms that do that, and there’s one theme running through all of them: Kubernetes never routes traffic to a pod that isn’t Ready.

1. A bad deploy that can’t hurt you
#

Ship a “new version” with a broken health check — the readiness probe now points at a path that doesn’t exist:

kubectl patch deployment signal-api --type=json \
  -p='[{"op":"replace","path":"/spec/template/spec/containers/0/readinessProbe/httpGet/path","value":"/healthz"}]'
kubectl rollout status deployment/signal-api
# → Waiting for deployment "signal-api" rollout to finish: 1 old replicas are pending termination...  (STUCK)

The rollout refuses to finish. Look at why:

signal-api-78f48ff458-wjsck   1/1  Running   ← OLD pod: Ready, still serving
signal-api-8f55cb49b-24dzg    0/1  Running   ← NEW pod: Running but NEVER Ready

The new pod runs, but its readiness probe fails forever → it never goes Ready → Kubernetes won’t kill the old one. Prove the app never blinked:

kubectl get endpoints signal-api -o jsonpath='{.subsets[*].addresses[*].ip}'
# → 10.244.0.54   (ONLY the old pod — the new one is excluded)

# from inside the cluster:
curl -s -o /dev/null -w '%{http_code}\n' http://signal-api.default.svc/health
# → 200   200   200

The Service routes only to ready=true pods, so the broken new pod (ready=false) got zero traffic. No user ever saw it. Then the escape hatch:

kubectl rollout undo deployment/signal-api    # revert to the last-good ReplicaSet, one command

Postmortem. Symptom: deploy hung on “old replicas pending termination.” Diagnosis: new pod Running but 0/1; endpoints showed only the old pod; curl stayed 200. Root cause: new readiness probe pointed at a nonexistent path. Why no outage: the probe kept the broken pod out of the Service. Fix: rollout undo.

2. The readiness probe is the hero
#

Scenario 1 already showed it, so name it explicitly — because it’s the single most underrated safety mechanism in Kubernetes:

  • A readiness probe answers “should this pod receive traffic?” A failing readiness probe pulls the pod out of the Service’s endpoints — but does not kill it.
  • A liveness probe answers “is this pod wedged and needs a restart?” A failing liveness probe restarts the container.

The distinction is everything. During a bad rollout, readiness failing = “hold traffic back, keep the old version live.” That’s why a broken deploy is a non-event instead of an outage — the probe quarantines the sick pod at the exact layer (Service routing) where users would have been hurt.

3. OOMKill — contained to one pod
#

Set a memory limit far below what the app needs, and the kernel’s out-of-memory killer does the rest:

kubectl patch deployment signal-api --type=json \
  -p='[{"op":"replace","path":"/spec/template/spec/containers/0/resources/limits/memory","value":"12Mi"}]'

The new pod can’t even boot inside 12Mi → CrashLoopBackOff. But this looks like Lab 1’s crash loop, and the exit code is what tells them apart:

kubectl describe pod <new-pod> | grep -A6 "Last State"
#   Last State:  Terminated
#   Reason:      OOMKilled
#   Exit Code:   137
Exit codeMeaning
1 (Lab 1)the app chose to die — a normal error exit
137128 + 9 = SIGKILL — the kernel’s OOM killer

137 + Reason: OOMKilled is an unmistakable signature: the container exceeded its memory limit and was killed. The fix is about resources, not code — raise the limit or shrink the footprint. And once again the pod never went Ready, so the old pod served 100% of traffic — the OOMKill was contained to one pod, not the node.

Postmortem. Symptom: CrashLoopBackOff, restarts climbing. Diagnosis: describe → OOMKilled, exit 137. Root cause: memory limit below the app’s real footprint. Why no outage: never Ready → kept out of the Service. Fix: rollout undo / raise the limit.

The pattern
#

Three different explosions — a broken deploy, a broken health check, an out-of-memory kill — and not one caused an outage. The mechanism was the same every time:

new/sick pod starts  →  readiness probe fails  →  pod excluded from Service endpoints
                                                →  old healthy pod serves 100%
                                                →  rollout undo reverts in one command

That’s the discipline: contain the blast radius so failure is a shrug, not a page. Bad deploys can’t go live, sick pods get no traffic, greedy pods get killed alone, and every mistake is one rollout undo from gone.


The series, in one map
#

This closes Kubernetes From The Ground Up. The whole arc, one reconcile loop seen from six angles:

LabThe one idea
0The Core 5 & 5Kubernetes is a reconcile loop; status → stage → lens
1Failure Triagediagnose blind: the status string localizes the failure
2Observabilitya dashboard is a hospital vitals monitor for a service
3SLOs & Error Budgetsship or freeze, decided by burn rate not opinion
4GitOps with Argo CDthe cluster becomes a mirror of a Git repo
5Blast Radiusfail small, contained, reversible

From diagnosing failures, to seeing them, to promising against them, to declaring the whole cluster, to containing the damage when something breaks anyway. Six posts, one mental model — and a cluster I can now own, break, and defend cold.