Signal Forge ADR-003: Span metrics generated before tail sampling
Status: Accepted
Decision: The spanmetrics connector is placed in the pipeline before tail_sampling. See
the instrumentation reference for the full pipeline
walkthrough and PromQL examples built on this ordering.
Rationale:
- If span metrics were generated after sampling, only ~25% of traces would contribute to rate and error counters. A “request rate” metric reading 25% of actual traffic would be operationally useless.
- Placing
spanmetricsbefore sampling means every span contributes to RED metrics, regardless of whether the trace is kept. The sampled traces are for debugging; the span metrics are for SLO dashboards.
Pipeline order:
filter(healthz) → spanmetrics (ALL spans)
↘
tail_sampling (25% + errors + slow)
↓
batch
Scope: This ordering is realized only in the bespoke local-mode Alloy pipeline
(k8s/monitoring/grafana/local/configmap.yaml). In monitoring.mode: cloud (the default) the
pinned grafana/k8s-monitoring chart runs no tail_sampling — 100% of traces reach Tempo and the
spanmetrics connector sees every span regardless. The “before sampling” property still holds
trivially there; the ADR’s concern only bites once cloud-mode tail sampling exists (chart 4.x).
Alternative considered: After sampling — rejected because it produces misleading metrics.