Signal Forge ADR-004: Helm-managed Alloy stack (grafana/k8s-monitoring)

Standardizes on the grafana/k8s-monitoring Helm chart's five-role Alloy topology, keeping the hand-rolled DaemonSet only as a non-deployed reference, and documents the CD reconciliation + rollback contract that promotes it to real clusters.

Updated September 12, 2026
On this page
Navigation

Signal Forge ADR-004: Helm-managed Alloy stack (grafana/k8s-monitoring)

Status: Accepted

Decision: The production collector stack uses the grafana/k8s-monitoring Helm chart (five specialised Alloy roles), pinned in conf.yml (monitoring.helm.version, currently 3.8.12 on the v3 values schema). The hand-rolled DaemonSet in k8s/monitoring/grafana/ is kept as a reference artifact but is not deployed. The split between cloud and local collector configs within this stack is covered separately in ADR-005.

Rationale:

  • Running two Alloy instances receiving the same OTLP traffic caused duplicate spans, duplicate metric samples, version mismatches, and CrashLoopBackOff.
  • The Helm chart manages RBAC, ServiceAccounts, and River configs with versioned upgrades. The hand-rolled version required manual maintenance of all these.
  • The five-role split (metrics, logs, singleton, receiver, profiles) mirrors production AKS configuration, providing parity for validation.

The five roles:

RoleKindPurpose
alloy-receiverDaemonSetOTLP push receiver — app telemetry
alloy-logsDaemonSetPod + node log tailing → Loki
alloy-metricsStatefulSetkubelet, cAdvisor, KSM → Prometheus
alloy-singletonDeploymentCluster events, KSM API → Loki/Prometheus
alloy-profilesDaemonSetContinuous profiling (disabled locally)

Alternative considered: Single hand-rolled DaemonSet — rejected due to operational complexity and the duplicate-collector problem.

CD reconciliation and rollback contract (OAP-007)

This chart is deployed to real dev/qa/prod clusters only through scripts/ci/deploy-environment.sh’s apply-collector phase — never through deploy-local.sh, which remains restricted to local k3d clusters by its own assert_k3d_context guard. The phase runs in the exact slot ADR-011 declares: after verify-supply-chain (and, necessarily, after configure-cluster materializes kubeconfig), before apply-observability-rules and long before any application rollout. This section is the authoritative contract for what that phase — and its rollback counterpart — do.

Desired state. The chart, chart_sha256, collector_policy_version, and config_sha256 for the target environment are read from the promoted release manifest’s .collector.environments.<env> block (the same shape scripts/ci/validate-release-manifest.sh already enforces) — never recomputed independently at deploy time. apply-collector does re-render the Helm values via collector_policy.py (it needs the actual YAML to install), and asserts that render’s own config_sha256 equals the manifest’s before ever installing it; a mismatch fails closed. The chart archive is downloaded from the manifest’s chart_url and its SHA-256 verified against chart_sha256 before helm upgrade ever runs, then installed from that local, verified archive path — not helm repo add + --version — so the exact pinned bytes are what gets deployed.

Live state. collector_policy.py’s _stamp_collector_digest writes two pod annotations onto every Alloy workload’s podAnnotations: signal-forge.io/collector-config-sha256 and signal-forge.io/collector-policy-version. apply-collector reads both back via kubectl get pods -n <monitoring namespace> -l app.kubernetes.io/name=alloy and reads the current Helm release revision via helm status <release> -n <namespace> -o json. No Alloy pods + no Helm release = a first deploy, allowed. Alloy pods present but disagreeing with each other on either annotation, present with no Helm release installed at all, or a Helm release installed with zero matching Alloy pods (e.g. all crashed or evicted) is an ambiguous state and is rejected (fail closed) rather than guessed at.

Policy-version mismatch. If a live collector is running (not a first deploy) and its live collector-policy-version annotation differs from the manifest’s declared collector_policy_version, apply-collector exits non-zero before ever deciding noop/upgrade. A running collector built under a different policy contract must never silently receive a candidate that assumes the new one.

Noop vs. upgrade. collector_policy.reconciliation_action(desired_config_sha256, live_config_sha256) is the sole decision function (invoked via collector_policy.py --reconcile-action from bash) — a matching live digest is a noop (no Helm call, pods are not restarted); anything else, including no live digest, is an upgrade via helm upgrade --install --atomic --wait. Helm’s own --atomic rolls back that single release on a failed upgrade, but the phase still captures the Helm revision that was current before the upgrade (exported via GITHUB_ENV, then also snapshotted into a capture-previous file for cross-component rollback) — the failure that needs a rollback is often detected downstream of this phase (rules, application rollout, the observability gate), after Helm’s own atomic rollback has already returned success for the collector step in isolation.

“Verified” before application rollout means all three hold, checked immediately after any upgrade and before the phase can succeed:

  1. The live signal-forge.io/collector-config-sha256 annotation, re-read via kubectl, equals the desired config_sha256 consistently across every Alloy pod.
  2. Every Alloy pod is Running with all containers Ready.
  3. scripts/ci/observability_gate.py’s collector-health check passes in its scoped --only-collector-health --strict-collector mode, which also takes --expected-collector-config-sha256 / --live-collector-config-sha256 and reports a digest mismatch as its own distinct blocking reason — separate from “collector unreachable” or “self-metrics unavailable” — before falling through to the existing export-loss / remote-write / receiver-refusal / queue-saturation checks.

Any of the three failing is fail-closed: the phase exits non-zero without proceeding to apply-observability-rules.

Rollback restores exactly this: rollback reads the same captured pre-mutation state (from capture-previous’s snapshot file, or directly from the COLLECTOR_* environment variables apply-collector exported if capture-previous itself never ran because apply-collector failed first). If the collector was not upgraded this run, rollback leaves it untouched. If it was a first-time install (no prior revision exists), rollback runs helm uninstall to restore the true prior state — no collector — rather than guessing a revision to roll back to. Otherwise it runs helm rollback <release> <captured revision> -n <namespace> --wait and re-runs the identical three-part verification above against the prior digest. An incomplete or missing capture for a run that did upgrade is rejected rather than guessed at, mirroring the existing image/ConfigMap rollback phase’s handling of a partial previous-state set.

Out of scope / operational prerequisite. Provisioning the grafana-cloud-secrets Kubernetes Secret (the glc_ access-policy token + per-signal usernames) in each real cluster’s monitoring namespace remains an out-of-band prerequisite, unchanged by this phase — see Grafana Cloud Deployment. apply-collector neither creates nor reads that Secret directly; it only installs Helm values that reference it by name (secret.create: false + secretKeyRef, optional: true). A missing or invalid Secret does not fail helm upgrade itself (pods still start), but does fail the collector-health self-metrics check above, since a collector that cannot export also cannot be scraped for job="integrations/alloy" — the existing fail-closed path already covers this, it does not need a separate check.