Immutable CI/CD promotion
The release contract is build once → secure once → attest once → promote the same release unit. CI
builds each of the four application images exactly once; CD never recompiles source or invokes a
container build. A human-friendly commit tag and latest may exist in GHCR, but neither is a
deployment input.
The release unit also binds the collector chart and canonical values digest, policy version, SLO rules, environment-specific Alertmanager routes, dashboards, and successful positive/negative E2E evidence for the same full Git SHA. ADR-013 defines the authoritative artifact identities, deployment order, and complete rollback boundary.
The repository contains a complete promotion workflow, but does not contain GitHub Environment
variables, secrets, cluster credentials, or live Grafana query credentials. Therefore, CD is
render-only by default. A repository maintainer must deliberately configure an Environment and
set DEPLOY_ENABLED=true before it can mutate a cluster. This is a safety boundary, not a claim
that a live DEV, QA, or PROD target currently exists.
CI: validation and release creation
[.github/workflows/ci.yml](https://github.com/shipsolid/signal-forge/blob/main/.github/workflows/ci.yml)
runs on matching pull requests, pushes to main, and manual dispatch.
PR or main
-> Gitleaks + repository policy
-> unit tests, frontend production build, protobuf contract
-> CodeQL + dependency analysis + Trivy IaC scan
-> observability-as-policy + promtool + Alloy validation
-> quality gate
-> build each application image once
-> local image scan + CycloneDX SBOM
-> (main only) push registry digest + keyless sign/attest/verify
-> immutable release-manifest.json
All preceding controls are BLOCKING. A pull request proves that the source can pass the same
validation and image build/scan path but does not receive GHCR credentials and cannot publish a
release. Only a successful trusted main run publishes images and release metadata.
Release artifact
CI creates one metadata record for each required service:
otel-frontendgateway-apiorder-apinotification-svc
release-manifest.json is the sole supported hand-off to CD. It binds the Git commit, CI run and
attempt, SBOM artifact name, Sigstore evidence, and each exact
ghcr.io/<owner>/signal-forge/<service>@sha256:... reference. CI rejects a partial four-service
set before publishing the manifest, so a promotion and a rollback always operate on one complete
release. The extended manifest additionally binds every observability artifact and same-SHA E2E
result listed above; CD treats missing or mismatched evidence as a release-integrity failure.
Security and observability controls
| Control | Classification | What fails the release |
|---|---|---|
| Gitleaks, lint/repository policy, tests, protobuf contract | BLOCKING | Any failed command |
| CodeQL | BLOCKING | Analysis/upload failure |
| SCA | BLOCKING for .NET/Python and runtime-critical frontend findings | Vulnerable dependency policy breach |
| Trivy IaC | BLOCKING for CRITICAL; HIGH is a warning | Critical misconfiguration |
| Trivy image | BLOCKING | HIGH/CRITICAL, fixed CVE in the release image |
| Observability-as-policy | BLOCKING | Invalid telemetry contract, rendered collector configuration, dashboard shape/metric reference/missing $environment scoping, required assets, or forbidden span-metric / instrument dimension |
The static observability policy checks service.name identity, common resource attributes
(service.namespace, service.version, deployment.environment), Alloy/Helm renderability,
Prometheus rule syntax, Grafana dashboard structure plus panel metric references,
canonical/embedded parity, and environment template-variable/$environment scoping on every panel
query, runbook-link and alert-catalog lockstep, the rendered
values-cloud ⇄ Secret key contract, required SLO/runbook assets, and focused denylists of unbounded
span-metric and application-instrument dimensions. It does not prove data reaches a backend,
dashboard queries work, alerts route, SLOs are being met, or a full cardinality/cost budget is
acceptable. Those require environment-specific runtime data.
What runs by default vs. what an Environment adds
The control tables in this document mark gates BLOCKING because that is their classification when they run. It is not a claim that they all run on this repo’s default pipeline. The repository as shipped configures no GitHub Environment, so everything in the right-hand column below is dormant.
| Stage | Every PR / push to main | Only when an Environment sets DEPLOY_ENABLED=true (+ OBSERVABILITY_GATE_ENABLED=true) |
|---|---|---|
Static observability policy — validate_observability.py + unit suite + promtool / mimirtool / alloy validate / helm template / amtool | always | — |
| Image build, scan, CycloneDX SBOM, keyless sign/attest/verify | main only | — |
| Server-side apply + rollout status, exact-digest verification, HTTP/API smoke, synthetic journey | — | per environment |
Runtime observability gate — observability_gate.py: candidate freshness, per-signal presence, resource-attribute promotion, collector health, SLO burn + p99 latency | — | per environment; block fails all, warn fails PROD only |
SLO rule + Alertmanager route load into Grafana Cloud — apply-observability-rules | — | per environment |
Dashboard reconciliation + live verification into Grafana Cloud — apply-dashboards | — | per environment |
A green pipeline on this repo therefore proves the telemetry contract shape (left column), not
that a candidate’s signals are live or within SLO. Each observability_gate.py decision is
regression-covered by scripts/ci/tests/test_observability_gate.py (see that file’s decision-model
map), so a change to a threshold or an outcome is visible without a live environment.
CD: select evidence, then promote it
[.github/workflows/cd.yml](https://github.com/shipsolid/signal-forge/blob/main/.github/workflows/cd.yml)
is manual by design. The operator supplies a successful CI run ID, not a branch, source commit,
or image tag. CD validates through GitHub’s API that the selected run is a successful main CI run
from this repository, then downloads the artifact whose name includes that run’s attempt number.
It validates the manifest again before handing it to every deployment leg.
trusted CI release manifest
-> DEV: verify evidence -> deploy digest -> health/smoke/telemetry/DAST
-> QA: verify the same evidence -> deploy the same digest -> health/smoke/telemetry
-> PROD (explicit selection + Environment approval): same digest -> health/smoke/telemetry
Promotion is globally serialized, and each environment has its own non-cancelling concurrency lock. Two releases cannot interleave between DEV, QA, and PROD, and a new rollout never cancels one that may need rollback.
Environment contract
The reusable [deploy-environment.yml](https://github.com/shipsolid/signal-forge/blob/main/.github/workflows/deploy-environment.yml)
binds the job to the GitHub Environment (dev, qa, or prod). Environment protection rules control
approval and the moment secrets become available. PROD is selected explicitly with
confirm_production and requires QA success; any configured reviewers or wait timers remain outside
workflow-input control.
For a real deployment, the protected environment needs:
- Variables:
DEPLOY_ENABLED=true,ENVIRONMENT_URL,FARO_COLLECTOR_URL,OTEL_EXPORTER_OTLP_ENDPOINT,OBSERVABILITY_GATE_ENABLED=true, and the Grafana Cloud query endpointsOBS_GATE_MIMIR_QUERY_URL,OBS_GATE_LOKI_QUERY_URL,OBS_GATE_TEMPO_QUERY_URL. Repository variableOBSERVABILITY_TENANT_MAPis a JSON registry for all DEV/QA/PROD Mimir, Loki, Tempo, and Alertmanager identities; the job rejects a missing or reused identity. - Secrets:
KUBE_CONFIG,DB_SECRETS_ENV; the in-repo gate’s read credentialsOBS_GATE_MIMIR_USER,OBS_GATE_LOKI_USER,OBS_GATE_TEMPO_USER,OBS_GATE_TOKEN(one access-policy token,metrics:read logs:read traces:read); the rule/route write credentialsGRAFANA_CLOUD_MIMIR_TENANT,GRAFANA_CLOUD_MIMIR_RULER_URL,GRAFANA_CLOUD_ALERTMANAGER_URL,GRAFANA_CLOUD_MIMIR_RW_TOKEN(metrics:write rules:write alerts:write); the dashboard write credentialsGRAFANA_CLOUD_STACK_URL,GRAFANA_CLOUD_DASHBOARD_TOKEN(dashboards:writeonly — deliberately separate from the rule/route token above, least privilege); and the alert receiversSLACK_WEBHOOK_URLand (PROD)PAGERDUTY_ROUTING_KEY;OBSERVABILITY_ROLLBACK_KEYencrypts retained previous Mimir-rule, Alertmanager, and dashboard payloads.
Each Environment also declares the expected observability tenant identity. DEV, QA, and PROD use
separate Grafana stacks or tenants; CD verifies that identity before any rule, route, dashboard, or
collector mutation. Logical deployment.environment labels remain mandatory even with physical
isolation. See ADR-012.
The former OBSERVABILITY_GATE_URL / OBSERVABILITY_GATE_TOKEN are retired — the gate is now
scripts/ci/observability_gate.py, run by CD, not an external service. See
ADR-011.
These values are intentionally not committed. CD creates environment-specific ConfigMaps for
runtime endpoint, browser config, and telemetry resource attributes, while preserving the image
digest. The full Git SHA becomes service.version; the deployment environment becomes
deployment.environment. Database secrets are created separately from the protected Environment
and never enter the uploaded deployment plan.
Deployment gates and rollback
Before writing kubeconfig or applying a manifest, CD verifies that every GHCR digest exists and has a
keyless Cosign signature plus a CycloneDX attestation from the trusted main CI workflow identity.
It renders Kustomize topology with those exact digests, strips placeholder Secrets, and uploads the
secret-free plan for audit.
Before application rollout, the collector is reconciled when its desired digest changes, then
apply-observability-rules verifies the credential identities against the three-environment tenant
registry, captures and checks live rule/routing state, performs read-only diffs, then loads the SLO
recording/alerting rules (k8s/monitoring/slo-rules.yaml) and the per-environment Alertmanager
routing (k8s/monitoring/alerting/<env>.yaml.tmpl) into the tenant’s Mimir Ruler + Alertmanager, so
the candidate is gated against rules that already exist. Dashboards are reconciled in the same
control-plane phase. Every previous payload/revision is captured before mutation; the sensitive
rules/routes snapshot is uploaded only as an AES-256 encrypted 30-day artifact. A failed release
restores collector, rules, routes, dashboards, runtime ConfigMaps, and application image digests —
each restore step is attempted independently (a failure in one, e.g. Mimir unreachable, does not
skip the others), so a partial failure never leaves more of the release un-restored than necessary.
The restore order does not currently mirror strict reverse-mutation order (collector is restored
first in both apply and rollback); nothing in this pipeline depends on cross-component restore
ordering today, since collector, Mimir, Grafana dashboards, and the application cluster are
independent systems, but this is worth revisiting if that ever changes.
| Gate | DEV | QA | PROD | Classification |
|---|---|---|---|---|
| Server-side apply and rollout status | yes | yes | yes | BLOCKING |
| Deployment/pod exact-digest verification | yes | yes | yes | BLOCKING |
| HTTP and API-shape smoke tests | yes | yes | yes | BLOCKING |
| Synthetic journey (tagged gateway → order → RabbitMQ → notification) | yes | yes | yes | BLOCKING |
Observability release gate (scripts/ci/observability_gate.py) | pass/block; deadline warn reports warning | pass/block; deadline warn reports warning | pass only; deadline warn and block fail | BLOCKING; polls until pass/deadline and fails closed on unknown/error |
| OWASP ZAP baseline | yes | no | never | blocking execution/findings; warning-level ZAP rules informational |
If any post-apply gate fails, CD restores the complete previous four-image immutable release and the captured runtime ConfigMaps. It refuses to roll back to a partial or tag-based prior state. If no complete prior digest set exists, rollback fails loudly rather than inventing one.
Checked-in policy vs. retained runtime evidence
Three different things all live under “policy” in this repo, and they answer different questions:
| What | Where | Proves |
|---|---|---|
| Checked-in policy declarations | config/observability/*.yaml, deploy/charts/signal-forge-policy/ | What SHOULD be true — a Kyverno policy’s intended rule, a declared branch-protection or tenant-retention requirement. Reviewable in a PR; enforces nothing by itself until applied. |
| Live cluster/GitHub state | Kyverno admission decisions, GitHub branch/environment protection, Grafana Cloud tenant settings | What IS actually configured right now. A checked-in policy that was never applied, or was applied and later drifted, proves nothing about this. |
| Retained signed evidence | Release-manifest fields (collector, dashboards, tenant_settings, repository_controls) plus their cosign signatures/bundles, the admission E2E run, the break-glass reaper’s pod logs | That a SPECIFIC, named check of the live state happened at a SPECIFIC commit/run and produced a SPECIFIC result — the thing CD and an external reviewer can actually verify without re-running the check themselves. |
A change to the first column is a normal reviewable PR. A change to the second happens outside this
repo (GitHub Settings, the Kyverno admission webhook, Grafana Cloud) and this repo can only audit
it, not edit it directly. The third is what makes the first two trustworthy without re-auditing them
by hand on every release — see scripts/ci/audit_repository_controls.py and
scripts/ci/audit_tenant_settings.py for the two audits that produce it today, and
docs/operations/break-glass.md for the one runtime control (the exception reaper) whose evidence
is a retained log, not a signed manifest field.
Current boundaries and next operational work
- Admission control is now real, not aspirational (OAP-013):
deploy/charts/signal-forge-policyinstalls a pinned Kyverno release with fail-closedValidatingPolicy/ImageValidatingPolicyresources that deny unsigned images, mutable tags, digest-less references, privileged workloads, and workloads missing the approved service identity/observability-config references — enforced in-cluster at admission time, not only checked by CD before apply. Proven against a real signed GHCR digest inadmission-e2e.yml.signal-forge-deployer(ordinary CD’s identity) has no rights to touch Kyverno, its policies, orPolicyException— seedocs/operations/break-glass.mdfor the one path around it and how it’s time-boxed and auto-revoked. - The observability gate (
scripts/ci/observability_gate.py) queries Grafana Cloud directly; its checks and thresholds are code and unit-tested. What is still environment configuration: theOBS_GATE_*query credentials, and whether a cluster runs the chart withintegrations.alloy(until it does, the collector-health check emitswarn, notblock). - SLO rules and Alertmanager routing are validated as code in CI (
promtool/mimirtool rules check/amtool check-config) and loaded per environment byapply-observability-rules. A production promotion is SLO-gated wheneverDEPLOY_ENABLED=trueand the gate’sslo-burncheck has recording-rule data. - The scheduled vulnerability rescan is manual-only today and examines compatibility
latesttags; it is informational and never a CD input. - Frontend RUM release evidence (a real Chromium/Playwright browser journey, gated on Faro
browser-to-backend trace propagation and Loki log intake — the
frontend-rumcheck inscripts/ci/observability_gate.py) exists only in the E2E workflow’sfrontend-journey)/frontend-gate-negative)phases today.deploy-environment.shdoes not yet run a browser journey per dev/qa/prod promotion, so this table’s per-environment gates do not include it — see SLOs & burn-rate alerts.
For local k3d deployment, use Local Deployment. It builds/imports local images for the lab and is intentionally separate from the immutable GHCR promotion path.