Prometheus is great until someone adds user_id to a metric’s labels. Then every request from every user becomes its own permanent time series, the TSDB index grows without bound, and a few weeks later Prometheus falls over from memory pressure — usually during an incident, when you actually need it.
This is a cardinality explosion, and it’s one of the most common ways observability systems quietly poison themselves. The fix is simple in principle (“don’t put unbounded values in labels”), but in practice it’s a rule that lives in someone’s head, gets violated the first time a new engineer adds a metric, and isn’t caught until production starts paging.
So I built a gate for it: Cardinality Guard, a custom GitHub Action that statically scans Python source for Prometheus metric definitions and fails the build if any label name matches a blocklist of known-unbounded patterns — user_id, session_id, email, uuid, request_id, and similar.
Where it sits in the pipeline#
This is the first job in the deploy workflow for devops-showcase, a small FastAPI service running on a self-managed k3s cluster (Terraform-provisioned on Vultr):
cardinality-guard → build-and-push (GHCR) → kubectl rollout (k3s)jobs:
cardinality-guard:
name: Check metric cardinality labels
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: ./.github/actions/cardinality-guard
build-and-push:
needs: cardinality-guard
...If the guard fails, the image never gets built. The cost of a bad metric definition is “fix a line and re-push,” not “debug an OOMKilled Prometheus pod at 2am.”
How the scan works#
The action parses every .py file into an AST and walks it looking for calls to Counter, Gauge, Histogram, or Summary — the four prometheus_client metric types. For each one, it pulls the label names (whether passed positionally or as labelnames=[...]) and checks them against a list of regex patterns:
class CardinalityChecker(ast.NodeVisitor):
def __init__(self):
self.violations = []
def visit_Call(self, node):
func_name = (
getattr(node.func, "id", None) or
getattr(node.func, "attr", None)
)
if func_name not in {"Counter", "Gauge", "Histogram", "Summary"}:
self.generic_visit(node)
return
label_lists = []
if len(node.args) >= 3 and isinstance(node.args[2], ast.List):
label_lists.append(node.args[2])
for kw in node.keywords:
if kw.arg == "labelnames" and isinstance(kw.value, ast.List):
label_lists.append(kw.value)
for label_list in label_lists:
for elt in label_list.elts:
if isinstance(elt, ast.Constant):
label = str(elt.value)
for pattern in blocked_patterns:
if pattern.search(label):
self.violations.append(
(node.lineno, func_name, label, pattern.pattern)
)
self.generic_visit(node)The blocklist itself is just a config file, so it’s tunable per project without touching the action:
blocked_label_patterns:
- "user_?id"
- "request_?id"
- "session_?id"
- "trace_?id"
- "order_?id"
- "transaction_?id"
- "ip_?address"
- "email"
- "token"
- "uuid"
- "filename"
- "query"A failing run looks like this:
Cardinality Guard scanned 3 file(s).
FAILED — high-cardinality labels detected:
app/main.py:30 — 'Counter' uses label 'user_id' (matches blocked pattern: user_?id)
These labels produce one time series per unique value, causing Prometheus
memory exhaustion at scale. Move the value to a log line or structured event instead.That last line matters as much as the failure itself — the guard doesn’t just say “no,” it tells you where the data should go. High-cardinality fields like user_id or trace_id belong in logs, not metric labels — which is exactly what the rest of the stack is built for.
The bigger picture#
Cardinality Guard is one gate in a larger setup: Terraform provisions a Vultr VPS and bootstraps k3s, GitHub Actions builds and rolls out the FastAPI service, and a full PLG-plus-traces stack (Prometheus, Loki/Promtail, Tempo, Grafana, Alertmanager) gives RED dashboards, an SLO error-budget view, synthetic uptime checks via Blackbox Exporter, and one-click trace ↔ log ↔ metric correlation.
The guard is a small piece, but it’s the kind of small piece that determines whether the rest of that stack stays usable six months from now. Static checks like this are cheap insurance against the failure mode where your observability tooling becomes the outage.
Repo: github.com/DeArchiTech/devops-showcase · The live demo has been decommissioned — the VPS was torn down once the writeup was done, rather than paying to keep a finished demo running. Everything it served is reproducible from the Terraform in the repo.