Observability Pipeline
Grafana Alloy is the collector for all signals. The local deployment always applies the bespoke
local Alloy DaemonSet; the grafana/k8s-monitoring Helm chart is opt-in in local mode
(./deploy-local.sh --with-helm) and mandatory in cloud mode. Local and cloud mode are two
structurally different implementations, not two configmaps for the same pipeline:
- Local mode (
monitoring.mode: local) — a hand-authored River pipeline ink8s/monitoring/grafana/local/configmap.yaml, applied directly bydeploy-local.sh. Every stage below (receivers, k8sattributes, env-label, healthz filter, spanmetrics, tail sampling, batch, exporters) is custom code in that file. - Cloud mode (
monitoring.mode: cloud, default) — entirely the Helm chart’s ownapplicationObservabilityfeature, configured declaratively viak8s/monitoring/grafana-helm/values-cloud.yaml.tmpl. There is no equivalent hand-authored configmap — the chart generates its own fixed internal pipeline from those values. See Cloud mode pipeline below; it does not have the same stages as local mode, and that’s a real, documented capability gap (see Known gaps in cloud mode), not a doc-only difference.
Local mode pipeline
Stage 1: Receivers
OTLP receiver — accepts push from all services:
otelcol.receiver.otlp "default" {
grpc { endpoint = "0.0.0.0:4317" }
http { endpoint = "0.0.0.0:4318" }
output {
traces = [otelcol.processor.k8sattributes.default.input]
metrics = [otelcol.processor.k8sattributes.default.input]
logs = [otelcol.processor.k8sattributes.default.input]
}
}
Faro receiver — accepts browser RUM from Angular SPA:
faro.receiver "frontend" {
server {
listen_address = "0.0.0.0"
listen_port = 12347
cors_allowed_origins = ["*"]
}
output {
traces = [otelcol.processor.k8sattributes.default.input]
logs = [loki.write.local.receiver] // logs bypass OTel pipeline → direct to Loki
}
}
Note: Faro logs go directly to loki.write (bypassing the OTel batch processor) because they are
already structured for Loki and do not need OTLP processing.
Stage 2: K8s attribute enrichment
otelcol.processor.k8sattributes "default" {
extract {
metadata = [
"k8s.namespace.name",
"k8s.deployment.name",
"k8s.pod.name",
"k8s.node.name",
"k8s.container.name",
]
label {
from = "pod"
key_regex = "app\\.kubernetes\\.io/.*"
}
}
pod_association {
source { from = "connection" } // resolves pod from OTLP connection source IP
}
}
Attributes added to every span and metric point (regardless of which service sent it):
k8s.namespace.name,k8s.pod.name,k8s.deployment.name,k8s.node.name,k8s.container.name- Any pod label matching
app.kubernetes.io/*(e.g.app.kubernetes.io/name=gateway-api)
Stage 3: Environment label
otelcol.processor.transform "env_label" {
error_mode = "ignore"
trace_statements {
context = "resource"
statements = [
"set(attributes[\"deployment.environment\"], \"signal-forge-dev\") where attributes[\"deployment.environment\"] == nil",
"set(attributes[\"deployment_environment\"], \"signal-forge-dev\") where attributes[\"deployment_environment\"] == nil"
]
}
// metric_statements and log_statements only set the dot key — Prometheus's
// exporter auto-sanitizes dots to underscores, and loki.write's
// external_labels stamps the underscore key separately, so both already
// land as deployment_environment without a second statement. Traces have
// no such sanitization step, so trace_statements sets both keys explicitly.
}
Stamps deployment_environment as a resource attribute on any signal that doesn’t already have it —
identically across metrics, logs, and traces. This enables environment-based filtering in Grafana
without requiring every service to set the attribute explicitly.
Stage 4: Health-check filter (traces only)
otelcol.processor.filter "healthz" {
error_mode = "ignore"
traces {
span = [
"attributes[\"http.route\"] == \"/healthz\"",
"attributes[\"url.path\"] == \"/healthz\"",
]
}
output {
traces = [
otelcol.connector.spanmetrics.default.input,
otelcol.processor.tail_sampling.default.input,
]
}
}
Drops /healthz spans before they reach spanmetrics or the sampler. K8s liveness probes fire every
15 seconds across all pods — without this filter they would dominate trace count and span metrics.
Belt-and-suspenders design: the SDK-level filter
(opts.Filter = ctx => ctx.Request.Path != "/healthz") prevents spans from being created at all for
.NET services. The collector filter catches anything that slips through from Python or future
services.
Stage 5: Span metrics connector
otelcol.connector.spanmetrics "default" {
namespace = "traces.spanmetrics"
// Extra dimensions are the single shared allowlist:
// k8s/monitoring/grafana/shared/span-dimensions.txt
dimension { name = "http.request.method" }
dimension { name = "http.route" }
dimension { name = "http.response.status_code" }
dimension { name = "rpc.method" }
dimension { name = "rpc.service" }
dimension { name = "rpc.grpc.status_code" }
dimension { name = "messaging.operation" }
dimension { name = "messaging.system" }
histogram {
unit = "ms"
explicit {
buckets = ["5ms", "10ms", "25ms", "50ms", "100ms", "250ms", "500ms", "1s", "2.5s", "5s", "10s"]
}
}
exemplars { enabled = true }
output {
metrics = [otelcol.processor.batch.default.input]
}
}
Generates RED (Rate / Error / Duration) metrics — with
exemplars attached — from all spans,
before tail sampling. The namespace + unit pins are enforced (both pipelines) by
scripts/ci/validate_observability.py. Produces:
traces_spanmetrics_calls_total— rate countertraces_spanmetrics_duration_milliseconds_bucket— latency histogram
See Tail-Based Sampling for why this runs before sampling.
Stage 6: Tail sampling
otelcol.processor.tail_sampling "default" {
decision_wait = "10s"
num_traces = 1000
expected_new_traces_per_sec = 100
policy { name = "errors-always"; type = "status_code"; status_code { status_codes = ["ERROR"] } }
policy { name = "slow-requests"; type = "latency"; latency { threshold_ms = 2000 } }
policy { name = "probabilistic-rest"; type = "probabilistic"; probabilistic { sampling_percentage = 25 } }
output { traces = [otelcol.processor.batch.default.input] }
}
See Tail-Based Sampling for policy rationale and validation approach. This processor has no equivalent in cloud mode — see Known gaps in cloud mode.
Stage 7: Batch processor
otelcol.processor.batch "default" {
timeout = "5s"
send_batch_size = 1024
}
Buffers up to 1024 signals or 5 seconds, whichever comes first, before flushing to exporters. Reduces exporter connections and amortises network round trips.
Stage 8: Exporters
// Traces → Jaeger (OTLP gRPC, insecure)
otelcol.exporter.otlp "jaeger_local" {
client {
endpoint = "jaeger.otel-lab.svc.cluster.local:4317"
tls { insecure = true }
}
}
// Metrics → Prometheus (remote-write)
otelcol.exporter.prometheus "local" {
forward_to = [prometheus.remote_write.local.receiver]
}
prometheus.remote_write "local" {
endpoint {
url = "http://prometheus.otel-lab.svc.cluster.local:9090/api/v1/write"
}
}
Log tailing pipeline
This pipeline runs independently from the OTLP pipeline. It tails pod stdout from the otel-lab
namespace.
discovery.kubernetes "pods" {
role = "pod"
namespaces { names = ["otel-lab"] }
}
discovery.relabel "pod_logs" {
targets = discovery.kubernetes.pods.targets
rule { source_labels = ["__meta_kubernetes_namespace"]; target_label = "namespace" }
rule { source_labels = ["__meta_kubernetes_pod_name"]; target_label = "pod" }
rule { source_labels = ["__meta_kubernetes_pod_container_name"]; target_label = "container" }
rule { source_labels = ["__meta_kubernetes_pod_label_app"]; target_label = "app" }
}
loki.source.kubernetes "pod_logs" {
targets = discovery.relabel.pod_logs.output
forward_to = [loki.process.trace_correlation.receiver]
}
The trace_correlation processor extracts trace IDs from JSON log lines of both .NET and Python
services. Its trace_id/span_id extraction + structured-metadata stages are single-sourced at
k8s/monitoring/grafana/shared/trace-correlation-stages.alloy
and spliced into configmap.yaml.tmpl at deploy time (deploy-local.sh’s
render_local_alloy_configmap()) — the same fragment cloud mode uses, see
Log↔trace correlation (cloud mode) below. Only the level
extraction/label stages below are local-only:
loki.process "trace_correlation" {
stage.json {
expressions = {
dotnet_trace = "TraceId", // .NET field name
dotnet_span = "SpanId",
dotnet_level = "Level",
python_trace = "otelTraceID", // Python field name
python_span = "otelSpanID",
python_level = "levelname",
}
}
stage.template { source = "trace_id"; template = "{{ if .dotnet_trace }}{{ .dotnet_trace }}{{ else }}{{ .python_trace }}{{ end }}" }
stage.template { source = "span_id"; template = "{{ if .dotnet_span }}{{ .dotnet_span }}{{ else }}{{ .python_span }}{{ end }}" }
stage.template { source = "level"; template = "{{ if .dotnet_level }}{{ .dotnet_level }}{{ else }}{{ .python_level }}{{ end }}" }
stage.labels { values = { level = "" } }
stage.structured_metadata { values = { trace_id = "trace_id", span_id = "span_id" } }
forward_to = [loki.write.local.receiver]
}
trace_id and span_id are stored as Loki structured metadata (not stream labels) because
trace IDs are high-cardinality — using them as stream labels would cause label explosion in Loki.
Grafana’s Jaeger datasource tracesToLogsV2 config queries Loki for {trace_id="<id>"} to provide
“Logs for this span” in the trace view.
Full local-mode data flow
flowchart LR
otlp_grpc["OTLP gRPC :4317"] --> k8sattr
otlp_http["OTLP HTTP :4318"] --> k8sattr
faro["Faro HTTP :12347"] --> k8sattr
k8sattr["k8sattributes<br/>(pod metadata)"] --> envlabel["env_label<br/>(env stamp)"]
envlabel --> traces["traces"]
envlabel --> metrics["metrics"]
envlabel --> logs["logs"]
traces --> filter["filter"]
filter --> spanmetrics["spanmetrics"]
filter --> tailsampling["tail_sampling"]
spanmetrics --> batch1["batch"] --> exp1["exporters"]
tailsampling --> batch1
metrics --> batch2["batch"] --> exp2["exporters"]
logs --> batch3["batch"] --> exp3["exporters"]
podstdout["Pod stdout"] --> lokisrc["loki.source.kubernetes"] --> tracecorr["trace_correlation"] --> lokiwrite["loki.write"] --> loki["Loki"]
Cloud mode pipeline
Cloud mode has no hand-authored River config at all. deploy-local.sh renders
values-cloud.yaml.tmpl
(substituting ${...} placeholders from conf.yml) and passes it to helm upgrade for the
grafana/k8s-monitoring chart. The chart’s applicationObservability feature generates its own
fixed, templated pipeline from those values — not something this repo controls stage-by-stage
the way local mode’s configmap does:
flowchart LR
grpc["OTLP gRPC :4317"] --> rd["resourcedetection"]
http["OTLP HTTP :4318"] --> rd
rd --> k8s["k8sattributes"] --> tr["transform"] --> filt["filter (/healthz drop)"]
filt --> sm["spanmetrics connector<br/>traces_spanmetrics_*"] --> batch["batch"]
filt --> batch
batch --> dest["destinations<br/>(Mimir / Loki / Tempo)"]
values-cloud.yaml.tmpl configures, under applicationObservability:
-
connectors.spanMetrics— RED metrics from spans,namespace: traces.spanmetrics,histogram.unit: ms, dimensions from the sharedspan-dimensions.txt. This is the only producer oftraces_spanmetrics_*in cloud mode — every SLI recording rule and the release gate’s SLO check depend on it.scripts/ci/validate_observability.pyasserts it stays in lockstep with the local River connector. -
traces.filters.span+metrics.filters.datapoint— drop/healthzfrom the trace path and the span-metric datapoints (mirrors local mode’sfilter "healthz"). -
integrations.alloy.instances— self-scrape the chart’s Alloy pods (job="integrations/alloy") for thesignal-forge.pipelinecollector-loss alerts. -
destinations[2] (grafana-cloud-traces).processors.tailSampling— the same errors-always / slow-requests / 25%-probabilistic policy set as local mode’s hand-authored River (audit finding M4; see sampling.md). This is a different part of the chart’s schema thanapplicationObservabilityabove, which has no sampling hook of its own — sampling is declared per-destination, downstream of the spanmetrics connector, so ADR-003’s “spanmetrics before sampling” ordering holds in cloud mode the same way it does locally. Enabling it provisions a dedicated Alloy instance (the chart’s own sampler collector) that this repository has verified only viahelm template, not yet against a live cluster rollout.
Log↔trace correlation (cloud mode)
podLogs.extraLogProcessingStages in values-cloud.yaml.tmpl carries the same JSON-extraction +
structured-metadata logic as local mode’s trace_correlation stage above — both are rendered from
the single shared fragment,
k8s/monitoring/grafana/shared/trace-correlation-stages.alloy,
spliced in by deploy-local.sh’s render_helm_values(), through the chart’s raw-River-snippet hook
(the same mechanism already used there for ANSI-stripping and kube-system log-level dropping). The
Helm chart runs tpl on this value, so render_helm_values() escapes the fragment’s Go-template
delimiters ({{"{{"}} / {{"}}"}}) before substitution, to survive Helm’s render pass intact.
Known gaps in cloud mode
Tail sampling is declared but not yet live-verified. destinations[2].processors.tailSampling
(audit finding M4, closed) renders correctly through helm template against the pinned chart
version — the expected otelcol.processor.tail_sampling component, exact policy set, and pipeline
ordering all match local mode’s reference implementation (see sampling.md). What remains
unverified is the live behavior of the dedicated Alloy sampler instance and its load-balanced
routing that enabling this provisions, since this repository has no cluster to roll it out against
during development. Confirm on the next real cloud-mode deploy: the sampler pod comes up healthy, and
Tempo search volume for a known route drops to roughly 25% of that route’s
traces_spanmetrics_calls_total rate outside error/slow traffic (the same validation queries
sampling.md documents for local mode).