Immutable CI/CD promotion

How SignalForge validates, builds, signs, and promotes one immutable four-service release from CI through DEV, QA, and protected PROD.

Updated September 11, 2026
On this page
Navigation

Immutable CI/CD promotion

The release contract is build once → secure once → attest once → promote the same release unit. CI builds each of the four application images exactly once; CD never recompiles source or invokes a container build. A human-friendly commit tag and latest may exist in GHCR, but neither is a deployment input.

The release unit also binds the collector chart and canonical values digest, policy version, SLO rules, environment-specific Alertmanager routes, dashboards, and successful positive/negative E2E evidence for the same full Git SHA. ADR-013 defines the authoritative artifact identities, deployment order, and complete rollback boundary.

The repository contains a complete promotion workflow, but does not contain GitHub Environment variables, secrets, cluster credentials, or live Grafana query credentials. Therefore, CD is render-only by default. A repository maintainer must deliberately configure an Environment and set DEPLOY_ENABLED=true before it can mutate a cluster. This is a safety boundary, not a claim that a live DEV, QA, or PROD target currently exists.

CI: validation and release creation

[.github/workflows/ci.yml](https://github.com/shipsolid/signal-forge/blob/main/.github/workflows/ci.yml) runs on matching pull requests, pushes to main, and manual dispatch.

PR or main
  -> Gitleaks + repository policy
  -> unit tests, frontend production build, protobuf contract
  -> CodeQL + dependency analysis + Trivy IaC scan
  -> observability-as-policy + promtool + Alloy validation
  -> quality gate
  -> build each application image once
  -> local image scan + CycloneDX SBOM
  -> (main only) push registry digest + keyless sign/attest/verify
  -> immutable release-manifest.json

All preceding controls are BLOCKING. A pull request proves that the source can pass the same validation and image build/scan path but does not receive GHCR credentials and cannot publish a release. Only a successful trusted main run publishes images and release metadata.

Release artifact

CI creates one metadata record for each required service:

  • otel-frontend
  • gateway-api
  • order-api
  • notification-svc

release-manifest.json is the sole supported hand-off to CD. It binds the Git commit, CI run and attempt, SBOM artifact name, Sigstore evidence, and each exact ghcr.io/<owner>/signal-forge/<service>@sha256:... reference. CI rejects a partial four-service set before publishing the manifest, so a promotion and a rollback always operate on one complete release. The extended manifest additionally binds every observability artifact and same-SHA E2E result listed above; CD treats missing or mismatched evidence as a release-integrity failure.

Security and observability controls

ControlClassificationWhat fails the release
Gitleaks, lint/repository policy, tests, protobuf contractBLOCKINGAny failed command
CodeQLBLOCKINGAnalysis/upload failure
SCABLOCKING for .NET/Python and runtime-critical frontend findingsVulnerable dependency policy breach
Trivy IaCBLOCKING for CRITICAL; HIGH is a warningCritical misconfiguration
Trivy imageBLOCKINGHIGH/CRITICAL, fixed CVE in the release image
Observability-as-policyBLOCKINGInvalid telemetry contract, rendered collector configuration, dashboard shape/metric reference/missing $environment scoping, required assets, or forbidden span-metric / instrument dimension

The static observability policy checks service.name identity, common resource attributes (service.namespace, service.version, deployment.environment), Alloy/Helm renderability, Prometheus rule syntax, Grafana dashboard structure plus panel metric references, canonical/embedded parity, and environment template-variable/$environment scoping on every panel query, runbook-link and alert-catalog lockstep, the rendered values-cloud ⇄ Secret key contract, required SLO/runbook assets, and focused denylists of unbounded span-metric and application-instrument dimensions. It does not prove data reaches a backend, dashboard queries work, alerts route, SLOs are being met, or a full cardinality/cost budget is acceptable. Those require environment-specific runtime data.

What runs by default vs. what an Environment adds

The control tables in this document mark gates BLOCKING because that is their classification when they run. It is not a claim that they all run on this repo’s default pipeline. The repository as shipped configures no GitHub Environment, so everything in the right-hand column below is dormant.

StageEvery PR / push to mainOnly when an Environment sets DEPLOY_ENABLED=true (+ OBSERVABILITY_GATE_ENABLED=true)
Static observability policy — validate_observability.py + unit suite + promtool / mimirtool / alloy validate / helm template / amtoolalways—
Image build, scan, CycloneDX SBOM, keyless sign/attest/verifymain only—
Server-side apply + rollout status, exact-digest verification, HTTP/API smoke, synthetic journey—per environment
Runtime observability gate — observability_gate.py: candidate freshness, per-signal presence, resource-attribute promotion, collector health, SLO burn + p99 latency—per environment; block fails all, warn fails PROD only
SLO rule + Alertmanager route load into Grafana Cloud — apply-observability-rules—per environment
Dashboard reconciliation + live verification into Grafana Cloud — apply-dashboards—per environment

A green pipeline on this repo therefore proves the telemetry contract shape (left column), not that a candidate’s signals are live or within SLO. Each observability_gate.py decision is regression-covered by scripts/ci/tests/test_observability_gate.py (see that file’s decision-model map), so a change to a threshold or an outcome is visible without a live environment.

CD: select evidence, then promote it

[.github/workflows/cd.yml](https://github.com/shipsolid/signal-forge/blob/main/.github/workflows/cd.yml) is manual by design. The operator supplies a successful CI run ID, not a branch, source commit, or image tag. CD validates through GitHub’s API that the selected run is a successful main CI run from this repository, then downloads the artifact whose name includes that run’s attempt number. It validates the manifest again before handing it to every deployment leg.

trusted CI release manifest
  -> DEV: verify evidence -> deploy digest -> health/smoke/telemetry/DAST
  -> QA:  verify the same evidence -> deploy the same digest -> health/smoke/telemetry
  -> PROD (explicit selection + Environment approval): same digest -> health/smoke/telemetry

Promotion is globally serialized, and each environment has its own non-cancelling concurrency lock. Two releases cannot interleave between DEV, QA, and PROD, and a new rollout never cancels one that may need rollback.

Environment contract

The reusable [deploy-environment.yml](https://github.com/shipsolid/signal-forge/blob/main/.github/workflows/deploy-environment.yml) binds the job to the GitHub Environment (dev, qa, or prod). Environment protection rules control approval and the moment secrets become available. PROD is selected explicitly with confirm_production and requires QA success; any configured reviewers or wait timers remain outside workflow-input control.

For a real deployment, the protected environment needs:

  • Variables: DEPLOY_ENABLED=true, ENVIRONMENT_URL, FARO_COLLECTOR_URL, OTEL_EXPORTER_OTLP_ENDPOINT, OBSERVABILITY_GATE_ENABLED=true, and the Grafana Cloud query endpoints OBS_GATE_MIMIR_QUERY_URL, OBS_GATE_LOKI_QUERY_URL, OBS_GATE_TEMPO_QUERY_URL. Repository variable OBSERVABILITY_TENANT_MAP is a JSON registry for all DEV/QA/PROD Mimir, Loki, Tempo, and Alertmanager identities; the job rejects a missing or reused identity.
  • Secrets: KUBE_CONFIG, DB_SECRETS_ENV; the in-repo gate’s read credentials OBS_GATE_MIMIR_USER, OBS_GATE_LOKI_USER, OBS_GATE_TEMPO_USER, OBS_GATE_TOKEN (one access-policy token, metrics:read logs:read traces:read); the rule/route write credentials GRAFANA_CLOUD_MIMIR_TENANT, GRAFANA_CLOUD_MIMIR_RULER_URL, GRAFANA_CLOUD_ALERTMANAGER_URL, GRAFANA_CLOUD_MIMIR_RW_TOKEN (metrics:write rules:write alerts:write); the dashboard write credentials GRAFANA_CLOUD_STACK_URL, GRAFANA_CLOUD_DASHBOARD_TOKEN (dashboards:write only — deliberately separate from the rule/route token above, least privilege); and the alert receivers SLACK_WEBHOOK_URL and (PROD) PAGERDUTY_ROUTING_KEY; OBSERVABILITY_ROLLBACK_KEY encrypts retained previous Mimir-rule, Alertmanager, and dashboard payloads.

Each Environment also declares the expected observability tenant identity. DEV, QA, and PROD use separate Grafana stacks or tenants; CD verifies that identity before any rule, route, dashboard, or collector mutation. Logical deployment.environment labels remain mandatory even with physical isolation. See ADR-012.

The former OBSERVABILITY_GATE_URL / OBSERVABILITY_GATE_TOKEN are retired — the gate is now scripts/ci/observability_gate.py, run by CD, not an external service. See ADR-011.

These values are intentionally not committed. CD creates environment-specific ConfigMaps for runtime endpoint, browser config, and telemetry resource attributes, while preserving the image digest. The full Git SHA becomes service.version; the deployment environment becomes deployment.environment. Database secrets are created separately from the protected Environment and never enter the uploaded deployment plan.

Deployment gates and rollback

Before writing kubeconfig or applying a manifest, CD verifies that every GHCR digest exists and has a keyless Cosign signature plus a CycloneDX attestation from the trusted main CI workflow identity. It renders Kustomize topology with those exact digests, strips placeholder Secrets, and uploads the secret-free plan for audit.

Before application rollout, the collector is reconciled when its desired digest changes, then apply-observability-rules verifies the credential identities against the three-environment tenant registry, captures and checks live rule/routing state, performs read-only diffs, then loads the SLO recording/alerting rules (k8s/monitoring/slo-rules.yaml) and the per-environment Alertmanager routing (k8s/monitoring/alerting/<env>.yaml.tmpl) into the tenant’s Mimir Ruler + Alertmanager, so the candidate is gated against rules that already exist. Dashboards are reconciled in the same control-plane phase. Every previous payload/revision is captured before mutation; the sensitive rules/routes snapshot is uploaded only as an AES-256 encrypted 30-day artifact. A failed release restores collector, rules, routes, dashboards, runtime ConfigMaps, and application image digests — each restore step is attempted independently (a failure in one, e.g. Mimir unreachable, does not skip the others), so a partial failure never leaves more of the release un-restored than necessary. The restore order does not currently mirror strict reverse-mutation order (collector is restored first in both apply and rollback); nothing in this pipeline depends on cross-component restore ordering today, since collector, Mimir, Grafana dashboards, and the application cluster are independent systems, but this is worth revisiting if that ever changes.

GateDEVQAPRODClassification
Server-side apply and rollout statusyesyesyesBLOCKING
Deployment/pod exact-digest verificationyesyesyesBLOCKING
HTTP and API-shape smoke testsyesyesyesBLOCKING
Synthetic journey (tagged gateway → order → RabbitMQ → notification)yesyesyesBLOCKING
Observability release gate (scripts/ci/observability_gate.py)pass/block; deadline warn reports warningpass/block; deadline warn reports warningpass only; deadline warn and block failBLOCKING; polls until pass/deadline and fails closed on unknown/error
OWASP ZAP baselineyesnoneverblocking execution/findings; warning-level ZAP rules informational

If any post-apply gate fails, CD restores the complete previous four-image immutable release and the captured runtime ConfigMaps. It refuses to roll back to a partial or tag-based prior state. If no complete prior digest set exists, rollback fails loudly rather than inventing one.

Checked-in policy vs. retained runtime evidence

Three different things all live under “policy” in this repo, and they answer different questions:

WhatWhereProves
Checked-in policy declarationsconfig/observability/*.yaml, deploy/charts/signal-forge-policy/What SHOULD be true — a Kyverno policy’s intended rule, a declared branch-protection or tenant-retention requirement. Reviewable in a PR; enforces nothing by itself until applied.
Live cluster/GitHub stateKyverno admission decisions, GitHub branch/environment protection, Grafana Cloud tenant settingsWhat IS actually configured right now. A checked-in policy that was never applied, or was applied and later drifted, proves nothing about this.
Retained signed evidenceRelease-manifest fields (collector, dashboards, tenant_settings, repository_controls) plus their cosign signatures/bundles, the admission E2E run, the break-glass reaper’s pod logsThat a SPECIFIC, named check of the live state happened at a SPECIFIC commit/run and produced a SPECIFIC result — the thing CD and an external reviewer can actually verify without re-running the check themselves.

A change to the first column is a normal reviewable PR. A change to the second happens outside this repo (GitHub Settings, the Kyverno admission webhook, Grafana Cloud) and this repo can only audit it, not edit it directly. The third is what makes the first two trustworthy without re-auditing them by hand on every release — see scripts/ci/audit_repository_controls.py and scripts/ci/audit_tenant_settings.py for the two audits that produce it today, and docs/operations/break-glass.md for the one runtime control (the exception reaper) whose evidence is a retained log, not a signed manifest field.

Current boundaries and next operational work

  • Admission control is now real, not aspirational (OAP-013): deploy/charts/signal-forge-policy installs a pinned Kyverno release with fail-closed ValidatingPolicy/ImageValidatingPolicy resources that deny unsigned images, mutable tags, digest-less references, privileged workloads, and workloads missing the approved service identity/observability-config references — enforced in-cluster at admission time, not only checked by CD before apply. Proven against a real signed GHCR digest in admission-e2e.yml. signal-forge-deployer (ordinary CD’s identity) has no rights to touch Kyverno, its policies, or PolicyException — see docs/operations/break-glass.md for the one path around it and how it’s time-boxed and auto-revoked.
  • The observability gate (scripts/ci/observability_gate.py) queries Grafana Cloud directly; its checks and thresholds are code and unit-tested. What is still environment configuration: the OBS_GATE_* query credentials, and whether a cluster runs the chart with integrations.alloy (until it does, the collector-health check emits warn, not block).
  • SLO rules and Alertmanager routing are validated as code in CI (promtool / mimirtool rules check / amtool check-config) and loaded per environment by apply-observability-rules. A production promotion is SLO-gated whenever DEPLOY_ENABLED=true and the gate’s slo-burn check has recording-rule data.
  • The scheduled vulnerability rescan is manual-only today and examines compatibility latest tags; it is informational and never a CD input.
  • Frontend RUM release evidence (a real Chromium/Playwright browser journey, gated on Faro browser-to-backend trace propagation and Loki log intake — the frontend-rum check in scripts/ci/observability_gate.py) exists only in the E2E workflow’s frontend-journey)/frontend-gate-negative) phases today. deploy-environment.sh does not yet run a browser journey per dev/qa/prod promotion, so this table’s per-environment gates do not include it — see SLOs & burn-rate alerts.

For local k3d deployment, use Local Deployment. It builds/imports local images for the lab and is intentionally separate from the immutable GHCR promotion path.