Lab 2 · Observability End-to-End#
Stop reacting. Start seeing. — Kubernetes From The Ground Up#
David Chan, Claude Opus 4.8 AI-Symbiosis Research · July 2026
Lab 1 was reactive: something broke, you diagnosed it. This lab is proactive — you instrument the system so you can watch its vital signs and get paged before a user files a ticket. By the end you’ll have scraped a real app, learned the PromQL that matters, built a RED dashboard by hand, and wired an alert that fires on a simulated incident and then resolves — the full loop.
There’s a mental model that makes all of this click, and it’s worth stating up front because it’s not a metaphor — it’s an isomorphism:
A monitoring dashboard is a hospital vitals monitor pointed at a different patient. Heart-rate / BP / O₂ are a body’s vital signs; Rate / Errors / Duration are a service’s. The beep when a vital crosses a threshold is an alert firing. You can’t see inside a living system directly, so you watch a few well-chosen numbers and alarm when one goes bad. Bodies, services, aircraft, reactors — same shape.
The stack, and the one idea behind it#
We use the kube-prometheus-stack (Prometheus + Grafana + Alertmanager, driven by the Prometheus Operator):
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm install monitoring prometheus-community/kube-prometheus-stack \
-n monitoring --create-namespace \
--set grafana.adminPassword=showcase \
--set prometheus.prometheusSpec.serviceMonitorSelectorNilUsesHelmValues=falseThe one idea: Prometheus is a pull-based metrics database. On a schedule it scrapes a /metrics endpoint, stores every number as a time series, and lets you query them with PromQL. Grafana draws them; Alertmanager watches them and fires. Nothing pushes to Prometheus — it goes and gets the data itself.
Our app (signal-api) is already instrumented — a FastAPI service exposing:
api_requests_total{endpoint,status}— a Counterapi_request_duration_seconds{endpoint}— a Histogram
So the metrics exist; the lab is about scraping, querying, visualizing, and alerting on them.
Scraping: the ServiceMonitor (and a nasty label collision)#
With the Operator, you don’t hand-edit Prometheus config — you declare a scrape target with a ServiceMonitor CRD: “scrape any Service labeled app=signal-api, on the port named http, path /metrics.”
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata: { name: signal-api, labels: { app: signal-api } }
spec:
selector: { matchLabels: { app: signal-api } }
endpoints: [{ port: http, path: /metrics, interval: 15s }]Target goes UP — and then the first real gotcha appears. Query api_requests_total and every series has endpoint="http" — not the URL path. Your real paths are hiding under exported_endpoint.
Why: a label collision. Your app labels each metric with endpoint (the URL path). The Operator also attaches a label called endpoint — set to the scrape port name (http). Two labels, same name. Prometheus’s default rule (honor_labels: false) lets the scraper’s label win and renames the app’s to exported_endpoint.
The fix — tell the ServiceMonitor to let the app’s own labels win:
endpoints: [{ port: http, path: /metrics, honorLabels: true }]Now endpoint is your path again. Bank the word honorLabels — most people don’t learn it until it bites them in production, and saying it in an interview signals you’ve actually operated Prometheus.
The PromQL core#
Three moves cover most of what you’ll ever write.
1. rate() — because counters only go up. A Counter is monotonic; graphing the raw number is useless. You want how fast it’s climbing:
rate(api_requests_total[1m]) # per-second rate, per label set2. sum by (…) — aggregate away labels you don’t care about:
sum(rate(api_requests_total[1m])) # one line: total req/s
sum by (endpoint) (rate(api_requests_total[1m])) # one line per endpoint3. {…} — filter to specific label values. This is the distinction that trips everyone up:
| Syntax | Job | Where the label name goes |
|---|---|---|
by (endpoint) | group → one line per value | keep endpoint literally |
{endpoint="/health"} | filter → keep only matches | the actual value in quotes |
rate(api_requests_total{status="404"}[1m]) # ONE line: just the error trafficA habit worth keeping: sanity-check a new metric against something you know. Our error rate read ~0.9/s. Traffic mix:
/nopeis 1 of 4 equally-likely paths at ~3.5 req/s total → expected ~0.9/s. It matched, so the metric is trustworthy. If it had read 0 or 50/s, the query or labels were wrong — before you built an alert on it.
The RED dashboard, built by hand#
RED = Rate, Errors, Duration — the golden signals for any request-driven service. Three Grafana panels, each the same flow (New visualization → Prometheus → Code mode → paste), unit requests/sec or seconds:
R — request rate by endpoint
sum by (endpoint) (rate(api_requests_total[1m]))E — error rate
sum(rate(api_requests_total{status="404"}[1m]))D — p95 latency, and why averages lie. Latency isn’t a count — it’s a distribution, and the average smothers the tail:
A statistician drowned crossing a river of average depth three feet.
If the average response is 50ms but 5% of users wait 2s, the average says “all good” while 1 in 20 suffers. So you use percentiles: p50 (typical), p95 (the health number), p99 (the tail). That’s why the app emits a Histogram — it buckets each request by how long it took (le = less-than-or-equal), and histogram_quantile reconstructs the percentile from the buckets:
histogram_quantile(0.95, sum by (le) (rate(api_request_duration_seconds_bucket[5m])))Reads inside-out: rate of each bucket → sum keeping the le boundaries → “the value 95% of requests fall under.” Add a second query at 0.50 and you can watch the gap between p50 and p95 — that gap ballooning is a slow tail forming before errors appear. Latency is a leading indicator.
The alert: the beep, in YAML#
A PrometheusRule turns a query into a pager. Anatomy is four parts:
- alert: HighErrorRate
expr: sum(rate(api_requests_total{status="404"}[1m])) > 2 # the CONDITION
for: 1m # must HOLD this long before firing
labels: { severity: warning } # ROUTING (Alertmanager uses this)
annotations: { summary: "...", description: "404 rate is {{ $value }}..." } # the human message| Field | Job | Hospital-monitor analogy |
|---|---|---|
expr | the PromQL that must be true | “heart rate > 120” |
for: 1m | anti-flap — must hold before firing | ignore one weird beat; alarm on a sustained one |
labels.severity | routing — critical→page, warning→Slack | which staff get paged |
annotations | the human-readable text | what the alarm display says |
And the state machine — the whole point of for::
inactive (green) ──expr true──▶ pending (yellow) ──held 1m──▶ firing (red) 🔔 ──recovers──▶ inactive
▲ │
└──────── expr false (a blip) ───┘ (a spike never reaches "firing")Trip it. Crank /nope traffic to ~10/s and watch: within ~15s the 1-minute rate crosses 2 → pending; after it holds for the full for: 1m → FIRING. Cut the flood, and as the rate decays under 2 → back to inactive (Alertmanager sends a “resolved”). You’ve just watched a complete incident lifecycle — vitals normal → spike → alarm → intervention → recover → all-clear. In production the firing alert routes to Slack or a pager via that severity label.
What you built#
app Counter/Histogram → ServiceMonitor scrape (15s) → PromQL → Grafana RED dashboard
↘ PrometheusRule → FIRING 🔔 → Alertmanager → SlackEvery link is something you wired by hand: the scrape, the label-collision fix, the three golden-signal panels, the alert rule, the threshold, the trip, the resolve. That’s the shift from “debug my app when it breaks” to “watch my service’s vitals and get paged before it does.”
Next — Lab 3 · SLOs & Error Budgets: turn these signals into a promise. Define “99% of requests under 200ms,” burn the budget in a simulated incident, and read the burn-rate — the math that decides whether you ship or freeze.