SLOs & burn-rate alerts
The lab’s defined objectives, how they are measured, and how the alerts are structured. They are engineering targets and rule-as-code examples, not an externally backed SLA or evidence that a remote environment is currently meeting them. Rules as code live in k8s/monitoring/slo-rules.yaml.
Published SLOs
config/observability/slos.yaml
is the versioned source of truth. It declares the owner, tier, environment scope, 30-day window,
traffic floor, good/total queries, burn policy, and release policy for each row below.
| Service | SLI | Objective | Window and minimum traffic |
|---|---|---|---|
| gateway-api | non-error SERVER spans / all SERVER spans | 99.5% | 30d, 100 events |
| gateway-api | SERVER latency ≤500ms / all SERVER spans | 99.0% | 30d, 100 events |
| order-api | non-error SERVER spans / all SERVER spans | 99.5% | 30d, 100 events |
| order-api | SERVER latency ≤300ms / all SERVER spans | 99.0% | 30d, 100 events |
| notification-svc | terminal ACK/duplicate outcomes / terminal ACK, duplicate, or DLQ outcomes | 99.5% | 30d, 20 events |
| notification-svc | CONSUMER latency ≤5s / all ended CONSUMER spans | 99.0% | 30d, 20 events |
| otel-frontend (browser) | non-error CLIENT spans / all CLIENT spans | 99.5% | 30d, 20 events |
| otel-frontend (browser) | CLIENT latency ≤1000ms / all ended CLIENT spans | 99.0% | 30d, 20 events |
The frontend rows are a proxy, not a literal page-load/JS-error SLI: Faro’s TracingInstrumentation
instruments the browser’s own outbound fetch/XHR calls, and those spans flow through the same
spanmetrics connector as backend traces (service_name="otel-frontend", span_kind="SPAN_KIND_CLIENT"),
which is the only 30-day, continuously-recordable Mimir signal the browser produces today. The
release gate’s frontend-rum check (below) verifies the literal page-load-success/no-JS-error signal
per release instead, from the Playwright journey’s own result plus Faro’s Loki log intake.
The notification read API is deliberately absent. Its real user-impacting work completes on the RabbitMQ consumer. The messaging SLO additionally requires p99 message age ≤5 minutes and queue backlog ≤100. Missing either input blocks the release gate instead of reporting invented health.
The 99.5% error budget is 0.5% of eligible events; the 99.0% latency budget is 1%. The release
policy warns below 25% remaining and blocks an exhausted PROD budget unless an approved, expiring
slo-error-budget-exhausted remediation waiver exists.
Where the numbers come from
All SLIs are computed from metrics the spanmetrics connector generates from app traces:
traces_spanmetrics_calls_total{service_name, deployment_environment, span_kind, status_code}— one counter per spantraces_spanmetrics_duration_milliseconds_bucket{service_name, deployment_environment, span_kind, le}— latency histogram
The connector runs in both modes (local: hand-authored River; cloud:
applicationObservability.connectors.spanMetrics in values-cloud.yaml.tmpl). namespace is
pinned to traces.spanmetrics and the histogram unit to ms so the series names are stable across
modes; scripts/ci/validate_observability.py asserts the two connectors stay in lockstep.
Availability is scoped to inbound requests — span_kind="SPAN_KIND_SERVER". The earlier rules
summed every span per service (client, DB, gRPC-client, RabbitMQ producer/consumer, internal),
which diluted the SLI. gRPC server handlers are SERVER spans and are included; client spans are not.
notification-svc’s async work is a SPAN_KIND_CONSUMER span and terminal outcome counter, not
folded into HTTP availability. Its signal-forge.messaging group derives good/total from
notifications_processed_total{status=~"acknowledged|duplicate|failed"}; a requeued transient
attempt is not a terminal event.
Recording rules in slo-rules.yaml
pre-compute:
slo:<id>:good_events:30d = sum(increase(<versioned good-event query>[30d])) by (service_name, deployment_environment)
slo:<id>:total_events:30d = sum(increase(<versioned total-event query>[30d])) by (service_name, deployment_environment)
slo:<id>:compliance:30d = good_events / total_events, only when the configured traffic floor is met
slo:<id>:budget_consumed:30d = (1 - compliance) / (1 - objective)
slo:<id>:budget_remaining:30d = 1 - budget_consumed
No-traffic and low-traffic behaviour: a 30-day compliance/budget record is absent until its configured event floor is met; it is never treated as 100% healthy. A candidate still has its own freshness/evidence floor, and the gate warns when it cannot read an operational 30-day budget.
Rates are also computed at 30m and 6h windows — those feed the multi-window burn alerts. Every
operational aggregation retains deployment_environment; the physical tenant is verified as the
target environment before the rules load, and Alertmanager also matches/groups/inhibits on that
label. The signal-forge.pipeline group detects a missing application series by comparing it with
the per-environment Alloy self-scrape, so a firing absence alert retains its environment label.
Multi-window burn-rate alerts
Following the Google SRE workbook approach: compute the rate at which the error budget is burning, over two windows simultaneously, and alert when both windows cross their threshold. This gives you:
- Fast-burn (page) — 14.4× budget burn rate. At this rate, 2% of the 30-day budget is consumed in 1 hour. Requires a short window (5m) to be current AND a longer window (30m) to confirm it’s not a transient spike.
- Slow-burn (ticket) — 6× budget burn rate. 5% of the budget in 6 hours. Detected by 30m + 6h windows both exceeding the threshold.
Threshold math for a 99.5% SLO: target error budget = 0.005, so the fast-burn alert fires when
error_ratio > 14.4 × 0.005 = 0.072 (= 7.2% errors) in BOTH windows.
Alert summary
Every alert in k8s/monitoring/slo-rules.yaml — scripts/ci/validate_observability.py fails the
build if this table and the rule file drift apart.
| Alert | Group | Severity | for | Trigger |
|---|---|---|---|---|
SignalForgeAvailabilityFastBurn | availability | page | 2m | error_ratio > 7.2% in 5m AND 30m |
SignalForgeAvailabilitySlowBurn | availability | ticket | 15m | error_ratio > 3% in 30m AND 6h |
SignalForgeLatencyFastBurn | latency | page | 2m | latency bad-event ratio >14.4× 1% in 5m AND 30m |
SignalForgeLatencySlowBurn | latency | ticket | 15m | latency bad-event ratio >6× 1% in 30m AND 6h |
SignalForgeGatewayLatencyHigh | latency | ticket | 10m | gateway-api p99 > 500ms (diagnostic only) |
SignalForgeDownstreamLatencyHigh | latency | ticket | 10m | order-api p99 > 300ms (diagnostic only) |
SignalForgeNotificationConsumerErrors | messaging | ticket | 15m | notification-svc consumer error ratio > 5% / 30m |
SignalForgeNotificationConsumerFastBurn | messaging | page | 2m | terminal outcome bad ratio >14.4× 0.5% in 5m AND 30m |
SignalForgeNotificationConsumerSlowBurn | messaging | ticket | 15m | terminal outcome bad ratio >6× 0.5% in 30m AND 6h |
SignalForgeSloBudgetWarning | release | ticket | 15m | a 30-day SLO has <25% error budget remaining |
SignalForgeSloBudgetExhausted | release | page | 5m | a 30-day SLO has exhausted its error budget |
SignalForgeCardinalityBudgetWarning | cardinality | ticket | 10m | active series or sample rate > 80% of an instrument budget |
SignalForgeCardinalityBudgetExceeded | cardinality | page | 5m | active series or sample rate >= 100% of an instrument budget |
SignalForgeSpanMetricsAbsent | pipeline | page | 10m | no SERVER span-metrics for gateway-api or order-api |
SignalForgeNotificationSpanMetricsAbsent | pipeline | ticket | 15m | no span-metrics for notification-svc |
SignalForgeCollectorExportFailing | pipeline | page | 10m | Alloy dropping > 1% of spans to Tempo |
SignalForgeReceiverRefusing | pipeline | page | 10m | otelcol_receiver_refused_spans_total rising |
SignalForgeRemoteWriteFailing | pipeline | page | 15m | remote-write to Mimir dropping/failing samples |
SignalForgeCollectorQueueSaturated | pipeline | ticket | 10m | an exporter sending queue > 80% full |
AlloyReceiverDown | infra | page | 5m | per-environment receiver up is 0 |
DatastoreDown | infra | page | 3m | a single-replica datastore pod not Ready 3m |
The pipeline and infra collector alerts need the chart’s integrations.alloy self-scrape
(job="integrations/alloy") — until a cluster runs it, those series do not exist and the alerts
stay inactive.
Why two-window + for clause
The for: 2m on the fast-burn alert is not redundant with the multi-window AND — it’s a scheduling
dampener. Without it, a single scrape interval of bad data could fire. 2m is short enough to still
page within the SLO’s fast-burn target (we want to know within ~10m that the budget is on fire) and
long enough to absorb individual scrape failures.
The slow-burn alert has for: 15m because its whole purpose is non-urgent — 15 minutes of delay is
cheap and kills noise.
Where the alerts are evaluated
k8s/monitoring/slo-rules.yaml is a bare Prometheus/Mimir rule file (native
groups: format, not a PrometheusRule CRD) — one file, no Prometheus-Operator dependency,
consumed differently per monitoring.mode:
| Mode | Evaluator | How rules land |
|---|---|---|
local | In-cluster vanilla Prometheus | Automatic. deploy-local.sh’s apply_local_slo_rules() generates a prometheus-slo-rules ConfigMap from this file and Prometheus loads it via rule_files: — no kube-prometheus-stack needed. |
cloud (promotion) | Grafana Cloud Mimir (Ruler) | Automatic in CD. The apply-observability-rules phase runs mimirtool rules load (via scripts/ci/lib/mimir.sh) into the target environment’s tenant before the cluster is touched, alongside mimirtool alertmanager load for k8s/monitoring/alerting/<env>.yaml.tmpl. |
cloud (manual) | Grafana Cloud Mimir (Ruler) | ./scripts/push-slo-rules-to-mimir.sh (--dry-run for a read-only mimirtool rules diff; --alertmanager <env> to also push routing). For local dev / ad-hoc pushes outside the promotion flow. |
If a cluster does have kube-prometheus-stack installed, deploy-local.sh’s apply_slo_rules()
wraps the same file’s groups: into a PrometheusRule on the fly (no separate CRD-shaped copy is
maintained anywhere).
Local mode has no Alertmanager. In
monitoring.mode: localthe rules load into vanilla in-cluster Prometheus, which evaluates them but has nowhere to route a firing alert — there is no Alertmanager in the local stack, andk8s/monitoring/alerting/<env>.yaml.tmpltargets Grafana Cloud only. Local alerts are visible on Prometheus/alerts; delivery (Slack / PagerDuty) exists only on thecloudpromotion path. Wiring a local Alertmanager is a known follow-up (ADR-011).
observability.slo_rules.enabled in conf.yml (default true) is the master switch for both the
local ConfigMap load and the kube-prometheus-stack fallback; it doesn’t gate the cloud script, which
you run explicitly when you want rules live in Grafana Cloud.
Release-gate boundary
Operational rules and release evidence deliberately answer different questions:
slo-rules.yamlevaluates the complete service population in one environment. Its long-lived recording rules must never containservice_version; partitioning an operational SLO by every release would reset its history and grow cardinality without bound.scripts/ci/observability_gate.pyevaluates only the proposed release. It queries rawtraces_spanmetrics_*series withservice_name, the full 40-characterservice_version, and the exactdeployment_environment. It never consumes unversioned operational recording rules to decide candidate health.
The candidate check requires every declared resource attribute to be promoted for every candidate
service (for example, service.name becomes service_name), plus at least 100 SERVER spans for gateway-api and order-api, and 20
CONSUMER spans for notification-svc, during the configured evidence window. It computes the 30m
error ratio and 5m p99 latency directly from that candidate selector. Missing counters,
insufficient traffic, missing histogram buckets, error ratio above 3%, gateway p99 above 500ms, or
downstream p99 above 300ms all fail closed.
The heavyweight E2E workflow proves the boundary against real Mimir data: bad candidate beside a healthy old revision, healthy candidate beside a bad old revision, healthy DEV beside bad QA, and candidate counters with the histogram connector disabled. The test restores the DEV collector and application identity before continuing.
Frontend RUM evidence is a separate frontend-rum gate check, active only when a browser
journey supplies a trace ID (the E2E frontend-journey) phase in scripts/ci/e2e/run.sh, driven by
src/frontend/e2e/observability.spec.ts via a pinned Playwright/Chromium build). It blocks on: the
Playwright run itself reporting a failed page load or JS error; the browser trace failing to
propagate a traceparent into a gateway-api/order-api span (Tempo); a missing Faro session ID
(fails closed — see below); missing Faro log intake in Loki for that exact browser session; or Faro
independently reporting a JS exception (kind="exception") for that same session — a server-side
confirmation distinct from the Playwright script’s own self-reported result. It does not run as
part of an ordinary dev/qa/prod promotion (deploy-environment.sh does not yet invoke the browser
journey) — today it is E2E-only release evidence, not a per-environment gate.
The Faro receiver these Loki queries hit is k8s/monitoring/e2e/alloy-faro.yaml — a dedicated,
E2E-only Alloy instance. The vendored grafana/k8s-monitoring chart has no Faro/RUM receiver in
any version (verified against its own feature-application-observability templates through
v4.5.2: only _receiver_{otlp,zipkin,jaeger}.tpl exist), and the bespoke faro.receiver in
k8s/monitoring/grafana/local/configmap.yaml.tmpl is monitoring.mode: local only —
monitoring.manifests.cloud: [] in conf.yml means cloud mode (what E2E runs) never applies it.
alloy-faro terminates the browser’s Faro traffic, relays traces into the chart’s own
grafana-k8s-alloy-receiver OTLP endpoint (so frontend spans get the same spanmetrics/resource
promotion as backend traces, from one already-validated pipeline), and writes logs straight to E2E
Loki with extra_log_labels promoting only bounded-cardinality fields (app_name, kind,
app_environment) as real stream labels. The gate’s log/exception queries then filter at query time
on session_id (| logfmt | session_id="...", Faro’s own per-browser-session identifier) — not
service.version/git SHA, which is constant for an entire E2E job and so cannot tell this browser
run’s Faro output apart from a different run’s output for the same candidate emitted minutes
earlier in the same job (the negative fixtures run several journeys back-to-back). session_id is
also deliberately not a stream label itself — even higher cardinality than a release SHA.
Do not describe an ordinary CI run as SLO-gated. The evaluator polls both warn and block until
the evidence becomes pass or its deadline expires. A deadline block fails DEV/QA/PROD; a
deadline warn is informational in DEV/QA and release-blocking in PROD. Operational rules may
still be warming without weakening the candidate decision. See
Immutable CI/CD Promotion,
ADR-011, and
OTel Signal Contracts.
Runbook links
Every alert’s runbook_url points at a section in
docs/operations/runbooks.md (via the published
docs site, e.g. .../operations/runbooks/#availability-burn-rate-alert-firing). Each section gives
the confirming query, likely causes, first response, and escalation. Keep the heading text and the
anchor in the rule annotation in sync when adding an alert.
Tuning
- Bump targets: edit
config/observability/slos.yaml, not an alert expression. CI requires the 30-day records and burn expressions to match the policy’s objective and multipliers. - Different sample windows: the recording rules produce 5m, 30m, and 6h ratios. To add a 1h
window, add a new
- record:entry and a new alert that combines 1h with another window. - Per-service objectives: add a complete SLO record in
slos.yaml; validation requires all five 30-day records, an environment-scoped query, and matching owner/tier before it can ship.
What this doesn’t cover
- Synthetic monitoring. The k6 load-test Job in
k8s/loadtest/ is manual.
For continuous synthetic traffic, convert to a
CronJobthat runs every 5 minutes and points at/healthz. - A true client-latency SLI from web-vitals. The
otel-frontendSLOs above are a CLIENT-span proxy (fetch/XHR call success and latency), not Faro’s own page-load/web-vitals measurements — those land as LOG events in Loki, and there is currently no Loki→Mimir recording-rule bridge to turn them into a 30-day good/total series. A real web-vitals SLI would need that bridge (or a Loki ruler metric query) first. - Multi-window burn for latency. Latency alerts are single-window (5m > threshold for 10m). A multi-window burn for latency is possible (see Google SRE workbook), but we picked single-window for simplicity — latency deserves different tuning than availability.