Signal Forge ADR-004: Helm-managed Alloy stack (grafana/k8s-monitoring)
Status: Accepted
Decision: The production collector stack uses the grafana/k8s-monitoring Helm chart (five
specialised Alloy roles), pinned in conf.yml (monitoring.helm.version, currently 3.8.12 on the
v3 values schema). The hand-rolled DaemonSet in k8s/monitoring/grafana/ is kept as a reference
artifact but is not deployed. The split between cloud and local collector configs within
this stack is covered separately in
ADR-005.
Rationale:
- Running two Alloy instances receiving the same OTLP traffic caused duplicate spans, duplicate metric samples, version mismatches, and CrashLoopBackOff.
- The Helm chart manages RBAC, ServiceAccounts, and River configs with versioned upgrades. The hand-rolled version required manual maintenance of all these.
- The five-role split (metrics, logs, singleton, receiver, profiles) mirrors production AKS configuration, providing parity for validation.
The five roles:
| Role | Kind | Purpose |
|---|---|---|
alloy-receiver | DaemonSet | OTLP push receiver — app telemetry |
alloy-logs | DaemonSet | Pod + node log tailing → Loki |
alloy-metrics | StatefulSet | kubelet, cAdvisor, KSM → Prometheus |
alloy-singleton | Deployment | Cluster events, KSM API → Loki/Prometheus |
alloy-profiles | DaemonSet | Continuous profiling (disabled locally) |
Alternative considered: Single hand-rolled DaemonSet — rejected due to operational complexity and the duplicate-collector problem.
CD reconciliation and rollback contract (OAP-007)
This chart is deployed to real dev/qa/prod clusters only through
scripts/ci/deploy-environment.sh’s apply-collector phase — never through deploy-local.sh,
which remains restricted to local k3d clusters by its own assert_k3d_context guard. The phase runs
in the exact slot ADR-011 declares: after verify-supply-chain (and, necessarily, after
configure-cluster materializes kubeconfig), before apply-observability-rules and long before any
application rollout. This section is the authoritative contract for what that phase — and its
rollback counterpart — do.
Desired state. The chart, chart_sha256, collector_policy_version, and config_sha256 for the
target environment are read from the promoted release manifest’s
.collector.environments.<env> block (the same shape scripts/ci/validate-release-manifest.sh
already enforces) — never recomputed independently at deploy time. apply-collector does re-render
the Helm values via collector_policy.py (it needs the actual YAML to install), and asserts that
render’s own config_sha256 equals the manifest’s before ever installing it; a mismatch fails
closed. The chart archive is downloaded from the manifest’s chart_url and its SHA-256 verified
against chart_sha256 before helm upgrade ever runs, then installed from that local, verified
archive path — not helm repo add + --version — so the exact pinned bytes are what gets deployed.
Live state. collector_policy.py’s _stamp_collector_digest writes two pod annotations onto
every Alloy workload’s podAnnotations: signal-forge.io/collector-config-sha256 and
signal-forge.io/collector-policy-version. apply-collector reads both back via
kubectl get pods -n <monitoring namespace> -l app.kubernetes.io/name=alloy and reads the current
Helm release revision via helm status <release> -n <namespace> -o json. No Alloy pods + no Helm
release = a first deploy, allowed. Alloy pods present but disagreeing with each other on either
annotation, present with no Helm release installed at all, or a Helm release installed with zero
matching Alloy pods (e.g. all crashed or evicted) is an ambiguous state and is rejected (fail closed)
rather than guessed at.
Policy-version mismatch. If a live collector is running (not a first deploy) and its live
collector-policy-version annotation differs from the manifest’s declared collector_policy_version,
apply-collector exits non-zero before ever deciding noop/upgrade. A running collector built under a
different policy contract must never silently receive a candidate that assumes the new one.
Noop vs. upgrade. collector_policy.reconciliation_action(desired_config_sha256, live_config_sha256)
is the sole decision function (invoked via collector_policy.py --reconcile-action from bash) — a
matching live digest is a noop (no Helm call, pods are not restarted); anything else, including no
live digest, is an upgrade via helm upgrade --install --atomic --wait. Helm’s own --atomic rolls
back that single release on a failed upgrade, but the phase still captures the Helm revision that was
current before the upgrade (exported via GITHUB_ENV, then also snapshotted into a
capture-previous file for cross-component rollback) — the failure that needs a rollback is often
detected downstream of this phase (rules, application rollout, the observability gate), after Helm’s
own atomic rollback has already returned success for the collector step in isolation.
“Verified” before application rollout means all three hold, checked immediately after any upgrade and before the phase can succeed:
- The live
signal-forge.io/collector-config-sha256annotation, re-read viakubectl, equals the desiredconfig_sha256consistently across every Alloy pod. - Every Alloy pod is
Runningwith all containersReady. scripts/ci/observability_gate.py’scollector-healthcheck passes in its scoped--only-collector-health --strict-collectormode, which also takes--expected-collector-config-sha256/--live-collector-config-sha256and reports a digest mismatch as its own distinct blocking reason — separate from “collector unreachable” or “self-metrics unavailable” — before falling through to the existing export-loss / remote-write / receiver-refusal / queue-saturation checks.
Any of the three failing is fail-closed: the phase exits non-zero without proceeding to
apply-observability-rules.
Rollback restores exactly this: rollback reads the same captured pre-mutation state (from
capture-previous’s snapshot file, or directly from the COLLECTOR_* environment variables
apply-collector exported if capture-previous itself never ran because apply-collector failed
first). If the collector was not upgraded this run, rollback leaves it untouched. If it was a
first-time install (no prior revision exists), rollback runs helm uninstall to restore the true
prior state — no collector — rather than guessing a revision to roll back to. Otherwise it runs
helm rollback <release> <captured revision> -n <namespace> --wait and re-runs the identical
three-part verification above against the prior digest. An incomplete or missing capture for a run
that did upgrade is rejected rather than guessed at, mirroring the existing image/ConfigMap rollback
phase’s handling of a partial previous-state set.
Out of scope / operational prerequisite. Provisioning the grafana-cloud-secrets Kubernetes
Secret (the glc_ access-policy token + per-signal usernames) in each real cluster’s monitoring
namespace remains an out-of-band prerequisite, unchanged by this phase — see
Grafana Cloud Deployment. apply-collector neither creates nor
reads that Secret directly; it only installs Helm values that reference it by name (secret.create: false + secretKeyRef, optional: true). A missing or invalid Secret does not fail helm upgrade
itself (pods still start), but does fail the collector-health self-metrics check above, since a
collector that cannot export also cannot be scraped for job="integrations/alloy" — the existing
fail-closed path already covers this, it does not need a separate check.