Signal Forge ADR-011: Observability is a release gate
Status: Accepted
Context
The CD pipeline already had an observability-gate phase, but it was a bare curl to an
OBSERVABILITY_GATE_URL expecting {"status": "...", "summary": "..."}. Nothing implemented that
service; a {"status":"pass"} stub would have satisfied it. Separately, the published SLO design in
SLOs had no data in the default cloud mode: values-cloud.yaml.tmpl
carried no spanmetrics connector, so traces_spanmetrics_* — the series every SLI recording rule
queries — was never produced. And slo-rules.yaml was only ever loaded into Grafana Cloud Mimir by
a manual scripts/push-slo-rules-to-mimir.sh; CI/CD never touched it, and there was no alert
routing anywhere in the repo.
Net effect: a green CI run proved configuration shape, not that telemetry arrived, that alerts could route, or that the candidate’s SLO burn was within budget.
Decision
-
The gate is an in-repo Python evaluator, not a hosted service.
scripts/ci/observability_gate.pyqueries Grafana Cloud Mimir / Loki / Tempo with read-only credentials, applies a release-bound evidence model, and emits the same{status, summary}contract to stdout. Theobservability-gatephase indeploy-environment.shruns it directly with a bounded poll loop. Thevalidatephase now requires theOBS_GATE_*query credentials (GitHub Environment secrets) instead of an external URL;OBSERVABILITY_GATE_ENABLEDstays the master switch.Evidence checked, all scoped to
service_version=<full git sha>+deployment_environment=signal-forge-<env>: candidate span-metric freshness per service; per-signal presence (metric + log + trace); every required resource attribute on every service; collector health (otelcol_exporter_send_failed_*, remote-write failures, receiver refusals, queue depth); candidate error ratio and latency calculated directly from raw span metrics; and, when the synthetic journey ran, that its exact trace threads all required services in Tempo.Operational SLO recording rules are deliberately separate: they measure the complete service in one environment and retain
service_name, deployment_environment, but notservice_version. A release decision never reads those aggregated series as candidate evidence.Decision mapping: missing candidate telemetry / missing resource attributes / collector hard failure / candidate already above the 6× slow-burn (ticket) rate →
block(all environments). Collector self-metrics unavailable or soft queue saturation →warn. The evaluator continues polling warnings until its deadline; the finalwarnis informational in DEV/QA and release-blocking in PROD. The nominal E2E positive case requirespass, neverwarn. -
A tagged synthetic journey runs before the gate.
scripts/ci/synthetic_journey.pyfires one deterministicPOST /api/projects→POST /api/orders→ poll/api/notificationstransaction through the public ingress carrying a minted W3Ctraceparent, then exportsSYNTHETIC_TRACE_IDfor the gate. This guarantees the candidate environment has candidate-revision traffic to find and exercises the full gateway → order → RabbitMQ → notification path the SLOs are about (smoke-testonly touched/and/api/projects/). -
Cloud mode produces RED metrics declaratively.
values-cloud.yaml.tmplnow configuresapplicationObservability.connectors.spanMetrics(namespace pinned totraces.spanmetrics, histogram unitms, dimensions from the sharedspan-dimensions.txt), a/healthzdrop viatraces.filters.span+metrics.filters.datapoint, andintegrations.alloyself-scrape.scripts/ci/validate_observability.pyasserts the local River connector and the cloud chart values stay in lockstep. The chart pin moved from an untested4.3.1back to3.8.12— the latest 3.x, whose v3 values schema is what both templates are actually written for and which already exposes the fullspanMetricsschema. -
SLO rules and alert routing are promotion-controlled. The
apply-observability-rulesCD phase runsmimirtool rules load+mimirtool alertmanager loadper environment, before the cluster is touched, from GitHub Environment secrets. Routing lives ink8s/monitoring/alerting/{dev,qa,prod}.yaml.tmpl(Slack everywhere; PagerDuty forseverity=pagein prod), with receiver secrets templated in.scripts/push-slo-rules-to-mimir.shand the CD phase sharescripts/ci/lib/mimir.sh. DEV, QA, and PROD use separate tenant identities as required by ADR-012, and environment remains part of every operational recording and routing key. -
SLIs are scoped to inbound requests.
slo-rules.yamlrecording rules now filterspan_kind="SPAN_KIND_SERVER"; notification-svc’s async work has its ownsignal-forge.messaginggroup (SPAN_KIND_CONSUMER); a newsignal-forge.pipelinegroup covers telemetry freshness and collector loss. Every alert carries a realrunbook_urlinto operations/runbooks.md. -
The cross-language trace propagation test runs in CI (
src/integration-tests) as a blocking job on push-to-mainandworkflow_dispatch— heavy (6 containers), so not per-PR.
Alternatives considered
- A hosted gate service honouring the HTTP contract as-is. Rejected: it adds a second deploy surface and its own pipeline for a lab, and the evaluation logic is more testable and reviewable as versioned code run by CD.
- Grafana-managed alerting (unified alerting via the provisioning API / Terraform) instead of
Mimir Alertmanager. Rejected for now:
mimirtool alertmanager loadreuses the exact tenant and tooling the rule push already uses. Revisit if the stack does not expose the Mimir Alertmanager. - kube-prometheus-stack in-cluster to evaluate rules and route alerts. Rejected: cloud mode has no in-cluster Prometheus by design (see the one load-bearing knob).
- Migrating the Helm values to the chart-4.x schema as part of this change. Deferred: the 4.x
schema (map
destinations,collectors:block) is a larger, separately-testable change; 3.8.12 delivers everything the gate needs on the schema the files already use. The v4spanMetricsschema is a superset, so the connector config ported forward cleanly when that work happened.
Consequences
- A real deployment now fails closed unless the candidate’s telemetry is demonstrably live and
healthy in Grafana Cloud.
render-onlyenvironments (DEPLOY_ENABLED != true) are unaffected. - New GitHub Environment configuration is required per environment:
OBS_GATE_*query URLs + per-signal users + a read token;GRAFANA_CLOUD_MIMIR_*write URLs + tenant + token;SLACK_WEBHOOK_URL;PAGERDUTY_ROUTING_KEY(prod). Until an environment is enabled and these are set, all of this is exercised only by unit tests and the CI policy job. - The collector-loss alerts and the gate’s collector-health check depend on a cluster actually
running the updated
values-cloud.yaml.tmpl(withintegrations.alloy). Until then the gate emitswarn, notblock, for that check. - Still not covered (follow-ups): full cardinality budgeting beyond the dimension allowlist; frontend RUM release identity; admission-time enforcement (Kyverno / policy-controller); cloud tail sampling; local-mode Alertmanager parity; the chart-4.x values migration.
- The complete immutable release unit and rollback boundary are governed by ADR-013. Application images, collector configuration, rules, routes, dashboards, policy, and same-SHA E2E evidence are not independently promotable in QA/PROD.
Links