Lab 3 · SLOs & Error Budgets
#

Ship or freeze — decided by math, not opinion — Kubernetes From The Ground Up
#

David Chan, Claude Opus 4.8 AI-Symbiosis Research · July 2026


Lab 2 gave you the vital signs — Rate, Errors, Duration. This lab turns them into a promise, and turns the eternal “should we ship this risky feature or stop and fix reliability?” argument into a single number that decides for you.

The reframe: reliability is a budget, not a wall
#

Chasing 100% reliability is a trap — it’s impossible, and every extra nine costs exponentially more. The SRE move is to flip it: don’t promise “never fail.” Promise “fail no more than X,” and treat that X as a budget you’re allowed to spend.

TermWhat it isExample
SLI — Indicatorthe number you measuresuccess ratio = good ÷ total
SLO — Objectivethe target you promise“99% of requests succeed over 30 days”
Error budgetthe flip side = 100% − SLO99% SLO → you’re allowed 1% failure
Burn ratehow fast you’re spending itactual error ratio ÷ budgeted error ratio

Why the error budget is genius: it turns a political fight into arithmetic. Dev wants to ship features (risky); Ops wants stability. The budget referees:

Budget REMAINING  →  within your promise  →  SHIP freely, take risks 🚀
Budget EXHAUSTED  →  spent your allowed failure  →  FREEZE features, fix reliability 🛑

No arguing. The number decides.

Define the SLI
#

Availability = the fraction of requests that aren’t errors. Against our signal-api (from Lab 2), in Prometheus:

1 - (
  sum(rate(api_requests_total{status="404"}[5m]))
  /
  sum(rate(api_requests_total[5m]))
)

“1 minus (error rate ÷ total rate).” With a healthy baseline it reads ~0.993 (99.3%).

Nuance worth stating: a real availability SLO usually counts 5xx server errors and excludes 4xx (a 404 is arguably the client’s fault). Our demo app has no 5xx, so we use the 404 as a stand-in “failure.” The SLO math is identical either way.

Read the number as a decision
#

SLI = 99.3%. SLO = 99%. So:

QuestionAnswer
Meeting the SLO?Yes — 99.3% > 99% ✅
Error budget (allowed failure)1%
Actually failing~0.7%
Burn rate = actual ÷ budget0.7% ÷ 1% = ~0.7

Burn rate < 1 → sustainable. You’re spending budget slower than the window allows — you’d finish the month with budget to spare. Compute it directly (budget = 0.01):

(
  sum(rate(api_requests_total{status="404"}[5m]))
  /
  sum(rate(api_requests_total[5m]))
) / 0.01

The beauty of burn rate: it’s dimensionless and universal. 1 = exactly on pace to spend the whole budget. 10 = you’ll blow the entire month’s budget in ~3 days. The threshold is the same regardless of your SLO — which is why modern alerting fires on burn rate, not raw error counts.

Burn it — a live incident
#

Crank the error traffic to ~12/s and re-run the burn-rate query (switch to a [1m] window so it responds fast):

burn rate  0.5   →  healthy, ship freely 🟢
           1     →  spending exactly on pace
           38    →  🔥 3800% of sustainable — see below
           ~63   →  where it settles under the flood

Translate burn rate to time — that’s the gut-punch. At burn rate 38, your entire 30-day budget is gone in 30 ÷ 38 ≈ 0.8 days — under a day. A whole month of allowed failure, incinerated before tomorrow. That is why anything past ~10 is a hard freeze: at this rate you break your monthly promise by lunch. Cut the incident, and the burn rate crashes back under 1 — the freeze lifts, ship again. You just watched the ship-or-freeze decision get made, twice, in opposite directions, purely by a number.

The production alert: multi-window burn rate
#

A naïve burn alert picks one window, and both choices are bad:

  • Short window (5m) — detects fast, but flaps: a 30-second blip pages someone at 3am for nothing.
  • Long window (1h) — stable, no false pages, but slow to detect and slow to reset — it keeps paging you for an hour after you’ve already fixed it.

The multi-window pattern (Google SRE Workbook) ands them together:

( burn_rate over [5m] > 14.4 )     # LONG window — entry gate: is the burn REAL & sustained?
        and
( burn_rate over [1m] > 14.4 )     # SHORT window — exit gate: is it STILL happening NOW?
  • The long window is the entry gate — it refuses to page until the burn is confirmed sustained → kills false alarms.
  • The short window is the exit gate — it clears the page the instant the problem actually stops → fast auto-resolve.

Threshold 14.4 = burning ~2% of a 30-day budget in a single hour (page-worthy). The full rule (windows compressed to 5m/1m so it trips in a lab session; production uses 1h/5m):

- alert: ErrorBudgetFastBurn
  expr: |-
    (sum(rate(api_requests_total{status="404"}[5m])) / sum(rate(api_requests_total[5m])) / 0.01 > 14.4)
    and
    (sum(rate(api_requests_total{status="404"}[1m])) / sum(rate(api_requests_total[1m])) / 0.01 > 14.4)
  for: 1m
  labels: { severity: critical }
  annotations:
    summary: "signal-api is burning its error budget fast"
    description: "Multi-window burn rate > 14.4 — ~2% of the 30-day budget gone in an hour. FREEZE and investigate."

Trip it (sustained ~12 errors/s): both windows climb past 14.4 → and true → FIRING.

Then the payoff — watch it resolve. Cut the flood and observe the exact moment of recovery:

burn5m = 49.4   ← STILL 3× over threshold (a 5m average takes minutes to forget)
burn1m = 0.3    ← already crashed (the flood stopped)
alert  = INACTIVE ✅  ← RESOLVED

The alert cleared while the 5-minute burn rate was still 49. With a single 5m window you’d keep getting paged for ~4 more minutes after the fix — the 1m window cleared it instantly. No false pages on the way in, no lingering pages on the way out.

What you built
#

metric → SLI (availability) → SLO (99%) → error budget (1%) → burn rate
                                                                   ↘ multi-window alert → FREEZE 🛑 → recover → SHIP 🚀

The shift this lab teaches is cultural as much as technical: reliability stops being a vibe and becomes a number everyone agreed to in advance. “Should we ship?” is answered by the burn rate, not the loudest voice in the room. That’s the SRE discipline in one lab.


This closes the observability arc — Labs 1–3 took you from diagnosing failures, to seeing them, to promising against them. From here the series turns to delivery and blast-radius: GitOps and safe rollbacks.