Signal Forge ADR-013: Promote observability as part of the release unit
Status: Accepted
Context
Application images already move through promotion as immutable digests. Collector values are validated but deployed through a separate path, cloud dashboards are pushed manually, and the runtime observability E2E workflow is not attached to release eligibility. Source at one Git SHA can therefore describe a different observability system from the one evaluating the candidate.
Decision
The releasable unit contains all of the following evidence:
| Component | Immutable identity |
|---|---|
| Applications | One registry digest for each catalogued deployable |
| Collector | Pinned Helm chart identity, collector image digests, canonical non-secret values digest, and compatibility version |
| Policy | Policy schema/version and canonical service-contract digest |
| SLO rules | Canonical rule-bundle digest |
| Alert routing | Canonical environment-specific routing digest |
| Dashboards | Canonical dashboard-bundle digest and datasource mapping |
| Runtime proof | Successful positive and negative E2E run ID, attempt, full Git SHA, result, and diagnostic artifact digest |
CI signs and attests this evidence with the release manifest. CD rejects absent, unverifiable, different-SHA, failed, cancelled, skipped, stale, or superseded evidence.
Deployment is one ordered transaction:
- Validate signatures, attestations, compatibility, expected tenant identity, and live diffs.
- Snapshot the current collector revision, rules, routes, dashboards, runtime ConfigMaps, and image digests with their hashes.
- Deploy and verify the collector when its desired digest differs from live state.
- Reconcile and verify SLO rules, Alertmanager routes, and dashboards.
- Deploy the immutable application digests and verify workload health and image IDs.
- Generate backend and browser candidate traffic, then evaluate release-bound telemetry.
- Record success, or restore every mutated component in reverse order and verify the restored hashes and health.
No manual collector, rule, route, or dashboard push is an authoritative QA/PROD release path. Local-development helpers remain supported but are explicitly non-promotional.
Candidate versus operational evidence
- Operational SLOs measure the complete service in one environment and do not retain
service.versionin long-lived recording rules. - Release checks query raw span metrics for the exact
service.name, fullservice.version, anddeployment.environment. Healthy traffic from another revision or environment cannot influence the decision. service.versionon candidate span metrics is a deliberate, budgeted exception with bounded retention churn; it is not a general permission to promote high-churn resource attributes.
Consequences
- Releases take longer because runtime negative proof is part of eligibility.
- Observability control-plane changes gain the same audit and rollback expectations as application changes.
- A collector-only change still produces a release manifest and compatibility evidence; unchanged application digests may be reused, but their identity remains recorded.
- A failure to restore any component is a paging condition rather than a partial-success warning.
Links