Signal Forge ADR-011: Observability is a release gate

Promotes the runtime observability gate from an unimplemented HTTP contract to an in-repo evaluator that queries Grafana Cloud for release-bound telemetry evidence, and moves SLO rules + alert routing into the promotion flow.

Updated September 11, 2026
On this page
Navigation

Signal Forge ADR-011: Observability is a release gate

Status: Accepted

Context

The CD pipeline already had an observability-gate phase, but it was a bare curl to an OBSERVABILITY_GATE_URL expecting {"status": "...", "summary": "..."}. Nothing implemented that service; a {"status":"pass"} stub would have satisfied it. Separately, the published SLO design in SLOs had no data in the default cloud mode: values-cloud.yaml.tmpl carried no spanmetrics connector, so traces_spanmetrics_* — the series every SLI recording rule queries — was never produced. And slo-rules.yaml was only ever loaded into Grafana Cloud Mimir by a manual scripts/push-slo-rules-to-mimir.sh; CI/CD never touched it, and there was no alert routing anywhere in the repo.

Net effect: a green CI run proved configuration shape, not that telemetry arrived, that alerts could route, or that the candidate’s SLO burn was within budget.

Decision

  1. The gate is an in-repo Python evaluator, not a hosted service. scripts/ci/observability_gate.py queries Grafana Cloud Mimir / Loki / Tempo with read-only credentials, applies a release-bound evidence model, and emits the same {status, summary} contract to stdout. The observability-gate phase in deploy-environment.sh runs it directly with a bounded poll loop. The validate phase now requires the OBS_GATE_* query credentials (GitHub Environment secrets) instead of an external URL; OBSERVABILITY_GATE_ENABLED stays the master switch.

    Evidence checked, all scoped to service_version=<full git sha> + deployment_environment=signal-forge-<env>: candidate span-metric freshness per service; per-signal presence (metric + log + trace); every required resource attribute on every service; collector health (otelcol_exporter_send_failed_*, remote-write failures, receiver refusals, queue depth); candidate error ratio and latency calculated directly from raw span metrics; and, when the synthetic journey ran, that its exact trace threads all required services in Tempo.

    Operational SLO recording rules are deliberately separate: they measure the complete service in one environment and retain service_name, deployment_environment, but not service_version. A release decision never reads those aggregated series as candidate evidence.

    Decision mapping: missing candidate telemetry / missing resource attributes / collector hard failure / candidate already above the 6× slow-burn (ticket) rate → block (all environments). Collector self-metrics unavailable or soft queue saturation → warn. The evaluator continues polling warnings until its deadline; the final warn is informational in DEV/QA and release-blocking in PROD. The nominal E2E positive case requires pass, never warn.

  2. A tagged synthetic journey runs before the gate. scripts/ci/synthetic_journey.py fires one deterministic POST /api/projects → POST /api/orders → poll /api/notifications transaction through the public ingress carrying a minted W3C traceparent, then exports SYNTHETIC_TRACE_ID for the gate. This guarantees the candidate environment has candidate-revision traffic to find and exercises the full gateway → order → RabbitMQ → notification path the SLOs are about (smoke-test only touched / and /api/projects/).

  3. Cloud mode produces RED metrics declaratively. values-cloud.yaml.tmpl now configures applicationObservability.connectors.spanMetrics (namespace pinned to traces.spanmetrics, histogram unit ms, dimensions from the shared span-dimensions.txt), a /healthz drop via traces.filters.span + metrics.filters.datapoint, and integrations.alloy self-scrape. scripts/ci/validate_observability.py asserts the local River connector and the cloud chart values stay in lockstep. The chart pin moved from an untested 4.3.1 back to 3.8.12 — the latest 3.x, whose v3 values schema is what both templates are actually written for and which already exposes the full spanMetrics schema.

  4. SLO rules and alert routing are promotion-controlled. The apply-observability-rules CD phase runs mimirtool rules load + mimirtool alertmanager load per environment, before the cluster is touched, from GitHub Environment secrets. Routing lives in k8s/monitoring/alerting/{dev,qa,prod}.yaml.tmpl (Slack everywhere; PagerDuty for severity=page in prod), with receiver secrets templated in. scripts/push-slo-rules-to-mimir.sh and the CD phase share scripts/ci/lib/mimir.sh. DEV, QA, and PROD use separate tenant identities as required by ADR-012, and environment remains part of every operational recording and routing key.

  5. SLIs are scoped to inbound requests. slo-rules.yaml recording rules now filter span_kind="SPAN_KIND_SERVER"; notification-svc’s async work has its own signal-forge.messaging group (SPAN_KIND_CONSUMER); a new signal-forge.pipeline group covers telemetry freshness and collector loss. Every alert carries a real runbook_url into operations/runbooks.md.

  6. The cross-language trace propagation test runs in CI (src/integration-tests) as a blocking job on push-to-main and workflow_dispatch — heavy (6 containers), so not per-PR.

Alternatives considered

  • A hosted gate service honouring the HTTP contract as-is. Rejected: it adds a second deploy surface and its own pipeline for a lab, and the evaluation logic is more testable and reviewable as versioned code run by CD.
  • Grafana-managed alerting (unified alerting via the provisioning API / Terraform) instead of Mimir Alertmanager. Rejected for now: mimirtool alertmanager load reuses the exact tenant and tooling the rule push already uses. Revisit if the stack does not expose the Mimir Alertmanager.
  • kube-prometheus-stack in-cluster to evaluate rules and route alerts. Rejected: cloud mode has no in-cluster Prometheus by design (see the one load-bearing knob).
  • Migrating the Helm values to the chart-4.x schema as part of this change. Deferred: the 4.x schema (map destinations, collectors: block) is a larger, separately-testable change; 3.8.12 delivers everything the gate needs on the schema the files already use. The v4 spanMetrics schema is a superset, so the connector config ported forward cleanly when that work happened.

Consequences

  • A real deployment now fails closed unless the candidate’s telemetry is demonstrably live and healthy in Grafana Cloud. render-only environments (DEPLOY_ENABLED != true) are unaffected.
  • New GitHub Environment configuration is required per environment: OBS_GATE_* query URLs + per-signal users + a read token; GRAFANA_CLOUD_MIMIR_* write URLs + tenant + token; SLACK_WEBHOOK_URL; PAGERDUTY_ROUTING_KEY (prod). Until an environment is enabled and these are set, all of this is exercised only by unit tests and the CI policy job.
  • The collector-loss alerts and the gate’s collector-health check depend on a cluster actually running the updated values-cloud.yaml.tmpl (with integrations.alloy). Until then the gate emits warn, not block, for that check.
  • Still not covered (follow-ups): full cardinality budgeting beyond the dimension allowlist; frontend RUM release identity; admission-time enforcement (Kyverno / policy-controller); cloud tail sampling; local-mode Alertmanager parity; the chart-4.x values migration.
  • The complete immutable release unit and rollback boundary are governed by ADR-013. Application images, collector configuration, rules, routes, dashboards, policy, and same-SHA E2E evidence are not independently promotable in QA/PROD.

Links