SLOs & burn-rate alerts

SignalForge's published SLOs, how their SLIs are computed from span metrics, and how multi-window burn-rate alerts are structured.

Updated September 12, 2026
On this page
Navigation

SLOs & burn-rate alerts

The lab’s defined objectives, how they are measured, and how the alerts are structured. They are engineering targets and rule-as-code examples, not an externally backed SLA or evidence that a remote environment is currently meeting them. Rules as code live in k8s/monitoring/slo-rules.yaml.

Published SLOs

config/observability/slos.yaml is the versioned source of truth. It declares the owner, tier, environment scope, 30-day window, traffic floor, good/total queries, burn policy, and release policy for each row below.

ServiceSLIObjectiveWindow and minimum traffic
gateway-apinon-error SERVER spans / all SERVER spans99.5%30d, 100 events
gateway-apiSERVER latency ≤500ms / all SERVER spans99.0%30d, 100 events
order-apinon-error SERVER spans / all SERVER spans99.5%30d, 100 events
order-apiSERVER latency ≤300ms / all SERVER spans99.0%30d, 100 events
notification-svcterminal ACK/duplicate outcomes / terminal ACK, duplicate, or DLQ outcomes99.5%30d, 20 events
notification-svcCONSUMER latency ≤5s / all ended CONSUMER spans99.0%30d, 20 events
otel-frontend (browser)non-error CLIENT spans / all CLIENT spans99.5%30d, 20 events
otel-frontend (browser)CLIENT latency ≤1000ms / all ended CLIENT spans99.0%30d, 20 events

The frontend rows are a proxy, not a literal page-load/JS-error SLI: Faro’s TracingInstrumentation instruments the browser’s own outbound fetch/XHR calls, and those spans flow through the same spanmetrics connector as backend traces (service_name="otel-frontend", span_kind="SPAN_KIND_CLIENT"), which is the only 30-day, continuously-recordable Mimir signal the browser produces today. The release gate’s frontend-rum check (below) verifies the literal page-load-success/no-JS-error signal per release instead, from the Playwright journey’s own result plus Faro’s Loki log intake.

The notification read API is deliberately absent. Its real user-impacting work completes on the RabbitMQ consumer. The messaging SLO additionally requires p99 message age ≤5 minutes and queue backlog ≤100. Missing either input blocks the release gate instead of reporting invented health.

The 99.5% error budget is 0.5% of eligible events; the 99.0% latency budget is 1%. The release policy warns below 25% remaining and blocks an exhausted PROD budget unless an approved, expiring slo-error-budget-exhausted remediation waiver exists.

Where the numbers come from

All SLIs are computed from metrics the spanmetrics connector generates from app traces:

  • traces_spanmetrics_calls_total{service_name, deployment_environment, span_kind, status_code} — one counter per span
  • traces_spanmetrics_duration_milliseconds_bucket{service_name, deployment_environment, span_kind, le} — latency histogram

The connector runs in both modes (local: hand-authored River; cloud: applicationObservability.connectors.spanMetrics in values-cloud.yaml.tmpl). namespace is pinned to traces.spanmetrics and the histogram unit to ms so the series names are stable across modes; scripts/ci/validate_observability.py asserts the two connectors stay in lockstep.

Availability is scoped to inbound requests — span_kind="SPAN_KIND_SERVER". The earlier rules summed every span per service (client, DB, gRPC-client, RabbitMQ producer/consumer, internal), which diluted the SLI. gRPC server handlers are SERVER spans and are included; client spans are not. notification-svc’s async work is a SPAN_KIND_CONSUMER span and terminal outcome counter, not folded into HTTP availability. Its signal-forge.messaging group derives good/total from notifications_processed_total{status=~"acknowledged|duplicate|failed"}; a requeued transient attempt is not a terminal event.

Recording rules in slo-rules.yaml pre-compute:

slo:<id>:good_events:30d      = sum(increase(<versioned good-event query>[30d])) by (service_name, deployment_environment)
slo:<id>:total_events:30d     = sum(increase(<versioned total-event query>[30d])) by (service_name, deployment_environment)
slo:<id>:compliance:30d       = good_events / total_events, only when the configured traffic floor is met
slo:<id>:budget_consumed:30d  = (1 - compliance) / (1 - objective)
slo:<id>:budget_remaining:30d = 1 - budget_consumed

No-traffic and low-traffic behaviour: a 30-day compliance/budget record is absent until its configured event floor is met; it is never treated as 100% healthy. A candidate still has its own freshness/evidence floor, and the gate warns when it cannot read an operational 30-day budget.

Rates are also computed at 30m and 6h windows — those feed the multi-window burn alerts. Every operational aggregation retains deployment_environment; the physical tenant is verified as the target environment before the rules load, and Alertmanager also matches/groups/inhibits on that label. The signal-forge.pipeline group detects a missing application series by comparing it with the per-environment Alloy self-scrape, so a firing absence alert retains its environment label.

Multi-window burn-rate alerts

Following the Google SRE workbook approach: compute the rate at which the error budget is burning, over two windows simultaneously, and alert when both windows cross their threshold. This gives you:

  • Fast-burn (page) — 14.4× budget burn rate. At this rate, 2% of the 30-day budget is consumed in 1 hour. Requires a short window (5m) to be current AND a longer window (30m) to confirm it’s not a transient spike.
  • Slow-burn (ticket) — 6× budget burn rate. 5% of the budget in 6 hours. Detected by 30m + 6h windows both exceeding the threshold.

Threshold math for a 99.5% SLO: target error budget = 0.005, so the fast-burn alert fires when error_ratio > 14.4 × 0.005 = 0.072 (= 7.2% errors) in BOTH windows.

Alert summary

Every alert in k8s/monitoring/slo-rules.yaml — scripts/ci/validate_observability.py fails the build if this table and the rule file drift apart.

AlertGroupSeverityforTrigger
SignalForgeAvailabilityFastBurnavailabilitypage2merror_ratio > 7.2% in 5m AND 30m
SignalForgeAvailabilitySlowBurnavailabilityticket15merror_ratio > 3% in 30m AND 6h
SignalForgeLatencyFastBurnlatencypage2mlatency bad-event ratio >14.4× 1% in 5m AND 30m
SignalForgeLatencySlowBurnlatencyticket15mlatency bad-event ratio >6× 1% in 30m AND 6h
SignalForgeGatewayLatencyHighlatencyticket10mgateway-api p99 > 500ms (diagnostic only)
SignalForgeDownstreamLatencyHighlatencyticket10morder-api p99 > 300ms (diagnostic only)
SignalForgeNotificationConsumerErrorsmessagingticket15mnotification-svc consumer error ratio > 5% / 30m
SignalForgeNotificationConsumerFastBurnmessagingpage2mterminal outcome bad ratio >14.4× 0.5% in 5m AND 30m
SignalForgeNotificationConsumerSlowBurnmessagingticket15mterminal outcome bad ratio >6× 0.5% in 30m AND 6h
SignalForgeSloBudgetWarningreleaseticket15ma 30-day SLO has <25% error budget remaining
SignalForgeSloBudgetExhaustedreleasepage5ma 30-day SLO has exhausted its error budget
SignalForgeCardinalityBudgetWarningcardinalityticket10mactive series or sample rate > 80% of an instrument budget
SignalForgeCardinalityBudgetExceededcardinalitypage5mactive series or sample rate >= 100% of an instrument budget
SignalForgeSpanMetricsAbsentpipelinepage10mno SERVER span-metrics for gateway-api or order-api
SignalForgeNotificationSpanMetricsAbsentpipelineticket15mno span-metrics for notification-svc
SignalForgeCollectorExportFailingpipelinepage10mAlloy dropping > 1% of spans to Tempo
SignalForgeReceiverRefusingpipelinepage10motelcol_receiver_refused_spans_total rising
SignalForgeRemoteWriteFailingpipelinepage15mremote-write to Mimir dropping/failing samples
SignalForgeCollectorQueueSaturatedpipelineticket10man exporter sending queue > 80% full
AlloyReceiverDowninfrapage5mper-environment receiver up is 0
DatastoreDowninfrapage3ma single-replica datastore pod not Ready 3m

The pipeline and infra collector alerts need the chart’s integrations.alloy self-scrape (job="integrations/alloy") — until a cluster runs it, those series do not exist and the alerts stay inactive.

Why two-window + for clause

The for: 2m on the fast-burn alert is not redundant with the multi-window AND — it’s a scheduling dampener. Without it, a single scrape interval of bad data could fire. 2m is short enough to still page within the SLO’s fast-burn target (we want to know within ~10m that the budget is on fire) and long enough to absorb individual scrape failures.

The slow-burn alert has for: 15m because its whole purpose is non-urgent — 15 minutes of delay is cheap and kills noise.

Where the alerts are evaluated

k8s/monitoring/slo-rules.yaml is a bare Prometheus/Mimir rule file (native groups: format, not a PrometheusRule CRD) — one file, no Prometheus-Operator dependency, consumed differently per monitoring.mode:

ModeEvaluatorHow rules land
localIn-cluster vanilla PrometheusAutomatic. deploy-local.sh’s apply_local_slo_rules() generates a prometheus-slo-rules ConfigMap from this file and Prometheus loads it via rule_files: — no kube-prometheus-stack needed.
cloud (promotion)Grafana Cloud Mimir (Ruler)Automatic in CD. The apply-observability-rules phase runs mimirtool rules load (via scripts/ci/lib/mimir.sh) into the target environment’s tenant before the cluster is touched, alongside mimirtool alertmanager load for k8s/monitoring/alerting/<env>.yaml.tmpl.
cloud (manual)Grafana Cloud Mimir (Ruler)./scripts/push-slo-rules-to-mimir.sh (--dry-run for a read-only mimirtool rules diff; --alertmanager <env> to also push routing). For local dev / ad-hoc pushes outside the promotion flow.

If a cluster does have kube-prometheus-stack installed, deploy-local.sh’s apply_slo_rules() wraps the same file’s groups: into a PrometheusRule on the fly (no separate CRD-shaped copy is maintained anywhere).

Local mode has no Alertmanager. In monitoring.mode: local the rules load into vanilla in-cluster Prometheus, which evaluates them but has nowhere to route a firing alert — there is no Alertmanager in the local stack, and k8s/monitoring/alerting/<env>.yaml.tmpl targets Grafana Cloud only. Local alerts are visible on Prometheus /alerts; delivery (Slack / PagerDuty) exists only on the cloud promotion path. Wiring a local Alertmanager is a known follow-up (ADR-011).

observability.slo_rules.enabled in conf.yml (default true) is the master switch for both the local ConfigMap load and the kube-prometheus-stack fallback; it doesn’t gate the cloud script, which you run explicitly when you want rules live in Grafana Cloud.

Release-gate boundary

Operational rules and release evidence deliberately answer different questions:

  • slo-rules.yaml evaluates the complete service population in one environment. Its long-lived recording rules must never contain service_version; partitioning an operational SLO by every release would reset its history and grow cardinality without bound.
  • scripts/ci/observability_gate.py evaluates only the proposed release. It queries raw traces_spanmetrics_* series with service_name, the full 40-character service_version, and the exact deployment_environment. It never consumes unversioned operational recording rules to decide candidate health.

The candidate check requires every declared resource attribute to be promoted for every candidate service (for example, service.name becomes service_name), plus at least 100 SERVER spans for gateway-api and order-api, and 20 CONSUMER spans for notification-svc, during the configured evidence window. It computes the 30m error ratio and 5m p99 latency directly from that candidate selector. Missing counters, insufficient traffic, missing histogram buckets, error ratio above 3%, gateway p99 above 500ms, or downstream p99 above 300ms all fail closed.

The heavyweight E2E workflow proves the boundary against real Mimir data: bad candidate beside a healthy old revision, healthy candidate beside a bad old revision, healthy DEV beside bad QA, and candidate counters with the histogram connector disabled. The test restores the DEV collector and application identity before continuing.

Frontend RUM evidence is a separate frontend-rum gate check, active only when a browser journey supplies a trace ID (the E2E frontend-journey) phase in scripts/ci/e2e/run.sh, driven by src/frontend/e2e/observability.spec.ts via a pinned Playwright/Chromium build). It blocks on: the Playwright run itself reporting a failed page load or JS error; the browser trace failing to propagate a traceparent into a gateway-api/order-api span (Tempo); a missing Faro session ID (fails closed — see below); missing Faro log intake in Loki for that exact browser session; or Faro independently reporting a JS exception (kind="exception") for that same session — a server-side confirmation distinct from the Playwright script’s own self-reported result. It does not run as part of an ordinary dev/qa/prod promotion (deploy-environment.sh does not yet invoke the browser journey) — today it is E2E-only release evidence, not a per-environment gate.

The Faro receiver these Loki queries hit is k8s/monitoring/e2e/alloy-faro.yaml — a dedicated, E2E-only Alloy instance. The vendored grafana/k8s-monitoring chart has no Faro/RUM receiver in any version (verified against its own feature-application-observability templates through v4.5.2: only _receiver_{otlp,zipkin,jaeger}.tpl exist), and the bespoke faro.receiver in k8s/monitoring/grafana/local/configmap.yaml.tmpl is monitoring.mode: local only — monitoring.manifests.cloud: [] in conf.yml means cloud mode (what E2E runs) never applies it. alloy-faro terminates the browser’s Faro traffic, relays traces into the chart’s own grafana-k8s-alloy-receiver OTLP endpoint (so frontend spans get the same spanmetrics/resource promotion as backend traces, from one already-validated pipeline), and writes logs straight to E2E Loki with extra_log_labels promoting only bounded-cardinality fields (app_name, kind, app_environment) as real stream labels. The gate’s log/exception queries then filter at query time on session_id (| logfmt | session_id="...", Faro’s own per-browser-session identifier) — not service.version/git SHA, which is constant for an entire E2E job and so cannot tell this browser run’s Faro output apart from a different run’s output for the same candidate emitted minutes earlier in the same job (the negative fixtures run several journeys back-to-back). session_id is also deliberately not a stream label itself — even higher cardinality than a release SHA.

Do not describe an ordinary CI run as SLO-gated. The evaluator polls both warn and block until the evidence becomes pass or its deadline expires. A deadline block fails DEV/QA/PROD; a deadline warn is informational in DEV/QA and release-blocking in PROD. Operational rules may still be warming without weakening the candidate decision. See Immutable CI/CD Promotion, ADR-011, and OTel Signal Contracts.

Every alert’s runbook_url points at a section in docs/operations/runbooks.md (via the published docs site, e.g. .../operations/runbooks/#availability-burn-rate-alert-firing). Each section gives the confirming query, likely causes, first response, and escalation. Keep the heading text and the anchor in the rule annotation in sync when adding an alert.

Tuning

  • Bump targets: edit config/observability/slos.yaml, not an alert expression. CI requires the 30-day records and burn expressions to match the policy’s objective and multipliers.
  • Different sample windows: the recording rules produce 5m, 30m, and 6h ratios. To add a 1h window, add a new - record: entry and a new alert that combines 1h with another window.
  • Per-service objectives: add a complete SLO record in slos.yaml; validation requires all five 30-day records, an environment-scoped query, and matching owner/tier before it can ship.

What this doesn’t cover

  • Synthetic monitoring. The k6 load-test Job in k8s/loadtest/ is manual. For continuous synthetic traffic, convert to a CronJob that runs every 5 minutes and points at /healthz.
  • A true client-latency SLI from web-vitals. The otel-frontend SLOs above are a CLIENT-span proxy (fetch/XHR call success and latency), not Faro’s own page-load/web-vitals measurements — those land as LOG events in Loki, and there is currently no Loki→Mimir recording-rule bridge to turn them into a 30-day good/total series. A real web-vitals SLI would need that bridge (or a Loki ruler metric query) first.
  • Multi-window burn for latency. Latency alerts are single-window (5m > threshold for 10m). A multi-window burn for latency is possible (see Google SRE workbook), but we picked single-window for simplicity — latency deserves different tuning than availability.