Runbooks
Troubleshooting playbooks for every known failure mode.
Immutable CD promotion blocked or rolled back
Scope
This runbook applies only when a GitHub Environment has DEPLOY_ENABLED=true. Without that flag,
CD is intentionally render-only and no cluster rollback is expected. The promotion input is a
successful CI run ID and its immutable release manifest — never a tag, a source branch, or a locally
rebuilt image.
First checks
- Open the CD run summary and record the selected CI run ID, release commit, target environment,
and the four
repository@sha256:...references. Do not copy these fromlatest. - Determine the failed boundary: pre-deploy evidence verification, server-side apply/rollout, exact-digest health gate, smoke test, observability gate, or DEV-only ZAP baseline.
- Download the
deployment-plan-<environment>-<commit>artifact. It is intentionally secret-free and shows the rendered digest references, runtime ConfigMaps, ingress host, and labels used for the attempted deployment. - If the failure occurred before
Apply immutable release, no cluster mutation occurred. Fix the environment configuration or the release, then start a new promotion from a successful CI run.
Verify the currently running release
namespace=otel-lab # replace only if the protected Environment uses another namespace
for service in otel-frontend gateway-api order-api notification-svc; do
kubectl -n "$namespace" get deployment "$service" \
-o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
done
Every returned application image must be a ghcr.io/...@sha256:... reference. A mutable tag is not
a valid known-good rollback target. Confirm workload availability and restart counts before declaring
recovery:
kubectl -n "$namespace" get deployments
kubectl -n "$namespace" get pods -l tier=app
Automatic rollback behavior
After an apply, any failed health, smoke, observability, or DEV DAST gate triggers CD to restore the
complete previous four-image immutable set and the previous signal-forge-app-env and
frontend-env-js ConfigMaps. It refuses partial rollback because mixing independent service
versions creates a release that was never tested together.
If the run reports “No complete previous immutable release is available for rollback”, stop.
Do not run kubectl apply -k k8s/overlays/prod, use latest, or rebuild a prior commit. Recover
only with an operator-approved complete digest manifest after investigating why the target had no
captured known-good state.
Observability release gate blocked
The gate is scripts/ci/observability_gate.py, run by the observability-gate CD phase. It polls
Grafana Cloud (Mimir / Loki / Tempo) with the OBS_GATE_* read credentials for the candidate’s
telemetry — keyed on service_version=<release commit> and
deployment_environment=signal-forge-<env> — and prints one JSON line {status, summary, checks}.
It fails closed for block or an unknown/error response; warn is visible in DEV/QA and blocks
PROD.
Read the checks array in the step log. Each entry names the failing check:
| check | meaning | first response |
|---|---|---|
candidate-freshness | no recent traces_spanmetrics_* for service_version=<commit> | confirm the pods run the release digest (verify-deployment passed?); check OTEL_EXPORTER_OTLP_ENDPOINT; check the synthetic-journey step actually completed |
resource-attributes | service_namespace absent on live series | the cloud values lost openTelemetryConversion.resourceToTelemetryConversion: true, or the chart is an old render |
collector-health | otelcol_exporter_send_failed_* / remote-write failing / receiver refusing | see Collector export failures / Remote-write to Mimir failing |
slo-burn | candidate already above the 6× slow-burn ratio or the p99 target | this is a real regression — do not force past it; roll back and investigate the candidate |
synthetic-journey | the synthetic trace is missing spans from a service | context propagation broke on one hop; run the cross-language integration test locally |
warn-only outcomes (collector-health unavailable, slo-burn rules not populated) mean the
supporting infra is not fully in place — the chart’s integrations.alloy self-scrape or the
apply-observability-rules phase. They do not indicate a bad candidate in DEV/QA.
Local repro: python scripts/ci/observability_gate.py --environment dev --git-sha <sha> --once
with the OBS_GATE_* vars exported.
See Immutable CI/CD Promotion for the full gate order and environment contract.
Availability burn-rate alert firing
SignalForgeAvailabilityFastBurn (page) or SignalForgeAvailabilitySlowBurn (ticket) means a
service’s SERVER-span error ratio is burning the 99.5% error budget 14.4× (fast, 5m+30m) or 6×
(slow, 30m+6h) faster than sustainable.
- Confirm scope:
sli:error_ratio:rate5m{service_name="<svc>"}in Grafana Explore (Mimir). A single spiking service points at that service; all three points at a shared dependency (datastore, RabbitMQ) — checkDatastoreDown. - Identify the failing operation:
sum by (http_route, http_response_status_code) (rate(traces_spanmetrics_calls_total{service_name="<svc>",span_kind="SPAN_KIND_SERVER",status_code="STATUS_CODE_ERROR"}[5m])). - Pull an exemplar trace from the error-ratio panel and follow it to the failing downstream span.
- If the burn is release-correlated, roll back (
deployment-planartifact has the prior digests). - Escalation: GitHub issue on the repo with the alert, the top failing route, and one exemplar trace ID.
Latency SLO breach
SignalForgeLatencyFastBurn / SignalForgeLatencySlowBurn are the customer-facing latency SLO
alerts: they use the configured good/total histogram buckets and 1% error budget. The p99 alerts
remain diagnostics: gateway >500ms and order-api >300ms over 5m.
- Confirm the SLO:
sli:latency_bad_ratio:rate5m{service_name="<svc>"}and its 30m companion. Usesli:latency_p99:rate5monly to describe the tail, not eligibility. - Break down by route:
histogram_quantile(0.99, sum by (le, http_route) (rate(traces_spanmetrics_duration_milliseconds_bucket{service_name="<svc>",span_kind="SPAN_KIND_SERVER"}[5m]))). - Use an exemplar to find where the time goes — DB span, gRPC fan-out, RabbitMQ publish.
- Check
SignalForgeCollectorQueueSaturatedand pod CPU throttling before assuming an app cause.
Notification consumer error rate high
SignalForgeNotificationConsumerErrors, SignalForgeNotificationConsumerFastBurn, or
SignalForgeNotificationConsumerSlowBurn apply to terminal outcomes only: ACK/duplicate are good;
DLQ is bad; requeued transient attempts are excluded.
- Check the dead-letter queue depth (see Consumer not processing messages).
kubectl -n otel-lab logs deploy/notification-svc | grep -i 'failed\|traceback'for the failingorder_ids.- Distinguish transient (RabbitMQ / Redis blips — self-heal) from poison messages (same
order_idfailing repeatedly — see the DLQ ADR).
Frontend RUM
SignalForgeSloBudgetWarning/SignalForgeSloBudgetExhausted for frontend-rum-availability or
frontend-rum-latency use a CLIENT-span good/total ratio as a proxy for client-perceived health (see
SLOs & burn-rate alerts) — a failed/slow browser call to the API, not
literally page-load success or a JS error. The release gate’s own frontend-rum check
(scripts/ci/observability_gate.py) verifies that literal signal directly, from a Playwright journey
plus Faro’s own telemetry, and fails closed for a different, more specific set of reasons:
- Playwright journey itself failed — page load or a JS exception the script observed directly.
Re-run
src/frontend/e2e/observability.spec.tsagainst the target environment and read its own console/page-error output first; the CLIENT-span ratio checks below assume the journey ran. - Browser trace missing a backend span —
otel-frontendand at least one backend service must both appear under the same Tempo trace ID. A missing backend span means thetraceparentheader never propagated from Faro’sTracingInstrumentationinto gateway-api/order-api. Check the frontend’s CORS/allowed-origins config for the collector endpoint and confirmFARO_COLLECTOR_URLinfrontend-env-jspoints at the livealloy-faroreceiver. - No Faro log intake in Loki for the session — the collector never received Faro’s payload at
all for that browser run. Confirm the
faro.receiverroute is up (kubectl -n otel-lab logs deploy/otel-frontendfor client-side POST failures to/collect) and that the session’sapp_name/app_environmentlabels match what the query expects. - Faro reported a JS exception for the session — a real client-side error occurred even though
the Playwright script’s own assertions passed. Query
{app_name="otel-frontend",kind="exception"} | logfmt | session_id="<id>"in Loki for the stack.
For the operational SLO alerts specifically (not the release gate), treat them the same as any other 30-day good/total SLO — see SLO error budget low or exhausted.
SLO error budget low or exhausted
SignalForgeSloBudgetWarning means a 30-day good/total SLO has less than 25% budget remaining.
SignalForgeSloBudgetExhausted means its remaining budget is zero or below.
- Identify the exact
service_nameandsloin the alert, then inspectslo:<id>:compliance:30d,:budget_consumed:30d, and:budget_remaining:30din Mimir. - Confirm the configured traffic floor is met; absent 30-day records are unknown, never healthy.
- Stop feature promotion, prioritize rollback/remediation, and attach trace/log evidence to the incident or corrective-action issue.
- For an exhausted PROD budget, only a remediation release may proceed, and only with an approved,
expiring
slo-error-budget-exhaustedwaiver inconfig/observability/waivers.yaml.
Cardinality or ingestion budget warning
Scope
SignalForgeCardinalityBudgetWarning and SignalForgeCardinalityBudgetExceeded identify the
specific service_name and cardinality_budget_id that crossed its declared active-series or
delivered-samples-per-second threshold. The source of truth is
config/observability/cardinality-budgets.yaml.
First checks
- Confirm the target environment and inspect the exact budget id in the alert labels.
- In Grafana Explore, group the matching metric family by
__name__,service_name, and each declared dimension; do not add pod, user, request, or trace identifiers as labels while debugging. - Compare the resulting series and sample rate to the ledger’s measured value and 20% headroom.
- Roll back the release that added the dimension or instrument if the owner cannot reduce it safely. An increase to the ledger requires a reviewed pricing/capacity update, not a silence.
Span metrics absent
SignalForgeSpanMetricsAbsent (page, gateway-api/order-api) or
SignalForgeNotificationSpanMetricsAbsent (ticket) — traces_spanmetrics_calls_total has produced
no SERVER samples for 10–15m. Every SLO and the release gate’s SLO check has no input.
up{job=~".*alloy.*"}— is the receiver alive? See Alloy receiver down.- cloud mode: confirm the running Helm release still has
applicationObservability.connectors.spanMetrics.enabled: trueandnamespace: traces.spanmetrics—helm get values <release> -n monitoring. An old render (or a revert to chart 4.x without the connector) is the usual cause. - local mode:
kubectl -n otel-lab get cm alloy-config -o yaml | grep -A2 'connector.spanmetrics'. - App side: is the service emitting traces at all? Check
otelcol_receiver_accepted_spans_total.
Collector export failures
SignalForgeCollectorExportFailing — Alloy is dropping >1% of spans on the way to Tempo.
sum(rate(otelcol_exporter_send_failed_spans_total{job="integrations/alloy"}[10m]))vs..._sent_spans_total— confirm the ratio.- Alloy logs:
kubectl -n monitoring logs -l app.kubernetes.io/name=alloy | grep -i 'export\|tempo\|429\|401'. 401 → wrong token type (glc_access-policy, notglsa_). 429 → Tempo rate limit / quota. SignalForgeCollectorQueueSaturatedalongside → downstream is slow; check Grafana Cloud status../scripts/debug.shruns the mode-aware exporter-counter + reachability probe.
Remote-write to Mimir failing
SignalForgeRemoteWriteFailing — prometheus_remote_storage_samples_{failed,dropped}_total is
increasing; metric samples (including the SLI rule inputs) are not reaching Mimir.
prometheus_remote_storage_samples_failed_totalbyurl— which destination.- Alloy
alloy-metricslogs forremote_writeerrors — 401 (token), 400 (out-of-order / label issues), 429 (ingestion limit). prometheus_remote_storage_shardsvs_shards_max— if pinned at max, Alloy cannot keep up; check CPU and the Grafana Cloud metrics quota.
OTLP receiver refusing spans
SignalForgeReceiverRefusing — otelcol_receiver_refused_spans_total is climbing; app traffic is
turned away at the collector.
- Alloy receiver logs for the refusal reason (memory limiter, payload too large, auth).
SignalForgeCollectorQueueSaturated→ backpressure from a slow exporter is propagating to the receiver; fix the export path first.- Check
alloy-receiverpod memory vs itsmemory_limiterlimit.
Alloy receiver down
AlloyReceiverDown — up{job="integrations/alloy", pod=~".*alloy-receiver.*"} == 0 for 5m. All
app OTLP is being discarded.
kubectl -n monitoring get pods -l app.kubernetes.io/name=alloy— isalloy-receiverRunning/Ready?kubectl -n monitoring describe pod <alloy-receiver-pod>— OOMKilled, image pull, scheduling.kubectl -n monitoring logs <alloy-receiver-pod> --previousif it is crash-looping.- Same symptom set as No traces in Jaeger — that section has the connectivity checks from the app side.
Datastore not Ready
DatastoreDown — a single-replica datastore pod (mysql-*, postgres-*, redis-*, rabbitmq-*)
has been not-Ready for 3m. App-tier requests are failing.
kubectl -n otel-lab describe pod <pod>andlogs --previous.redis/rabbitmqerrors also show up as Redis connection errors / Consumer not processing messages.- Graduating past single-replica: datastore HA.
No traces in Jaeger
Symptoms
- Jaeger UI shows no services
make validatepasses but traces don’t appear
Diagnosis
Step 1: Is alloy-receiver running?
kubectl -n monitoring get pods -l app.kubernetes.io/component=alloy-receiver
# Expected: Running
Step 2: Is the app sending to the right endpoint?
kubectl -n otel-lab exec deploy/gateway-api -- env | grep OTEL
# OTEL_EXPORTER_OTLP_ENDPOINT should be:
# http://grafana-k8s-alloy-receiver.monitoring.svc.cluster.local:4317
Step 3: Can the app reach Alloy?
kubectl -n otel-lab exec deploy/gateway-api -- \
wget -qO- http://grafana-k8s-alloy-receiver.monitoring.svc.cluster.local:4317
# gRPC will return an HTTP 400 (expected — it's not HTTP/1.1) — this confirms connectivity
Step 4: Check Alloy receiver logs
kubectl -n monitoring logs daemonset/grafana-k8s-alloy-receiver --tail=100 \
| grep -E "error|warn|export"
Step 5: Check Alloy pipeline UI
kubectl port-forward svc/grafana-k8s-alloy-receiver 12345 -n monitoring
open http://localhost:12345
# Navigate to Components → otelcol.receiver.otlp.default → check "Data received" counter
Step 6: Is Jaeger accessible?
curl -s http://localhost:16686/api/services
# Should return {"data":["gateway-api","order-api",...]}
Metrics missing from Prometheus
Symptoms
- Prometheus has no metrics from app services
traces_spanmetrics_calls_totalquery returns nothing
Diagnosis
Step 1: Check Prometheus is up
curl -s http://localhost:9090/-/ready
Step 2: Prometheus has remote-write receiver enabled?
kubectl -n otel-lab describe deploy/prometheus | grep -A5 "Command"
# Should include: --web.enable-remote-write-receiver
# and: --enable-feature=exemplar-storage
Step 3: Check Alloy is writing to Prometheus
kubectl port-forward svc/grafana-k8s-alloy-receiver 12345 -n monitoring
# Navigate to Components → prometheus.remote_write.local → check "Samples sent" counter
Step 4: Query Prometheus directly
curl "http://localhost:9090/api/v1/query?query=up" | jq '.data.result'
Async propagation not working
Symptoms
notification.processspan in Jaeger has a differenttraceIdthanorder.publish- SpanLink is missing (no dashed arrow in Jaeger)
Diagnosis
Step 1: Verify traceparent is in the RabbitMQ message
In RabbitMQ Management (http://localhost:15672):
- Go to Queues →
notifications - Click “Get Message(s)”
- Inspect the Properties → Headers
- Should contain key
traceparentwith value00-<32 hex chars>-<16 hex chars>-01
If missing: the order-api publisher is not writing the header. Check OrderPublisher.cs — it sets
props.Headers["traceparent"] from the traceParent string persisted with the outbox row (the
order.create context), as raw UTF-8 bytes. If OutboxMessage.TraceParent was never stored,
OrderGrpcService.cs did not capture Activity.Current?.Id at order-create time.
Step 2: Verify the consumer extracts it correctly
kubectl -n otel-lab logs deploy/notification-svc --tail=50 | grep -i trace
Add temporary debug logging to consumer.py:
logger.debug("headers: %s", properties.headers)
Step 3: Check for pika instrumentation conflict
If opentelemetry-instrumentation-pika is also running, it may overwrite the extracted context.
Verify requirements.txt — opentelemetry-instrumentation-pika should not be present (we use
manual extraction).
Logs not appearing in Loki with trace correlation
Symptoms
- “Logs for this span” in Grafana returns no results
- Loki has logs but they lack
trace_idstructured metadata
Diagnosis
Step 1: Confirm apps write JSON
kubectl -n otel-lab logs deploy/gateway-api --tail=3
# Should be JSON: {"Timestamp":"...","Level":"Information","TraceId":"4bf..."}
# NOT plain text: info: Processing request
Step 2: Check alloy-logs is running
kubectl -n monitoring get pods -l app.kubernetes.io/component=alloy-logs
kubectl -n monitoring logs daemonset/grafana-k8s-alloy-logs --tail=50
Step 3: Query Loki directly
kubectl port-forward svc/loki 3100 -n otel-lab
curl -G "http://localhost:3100/loki/api/v1/query_range" \
--data-urlencode 'query={namespace="otel-lab"}' \
--data-urlencode 'limit=5'
# If logs arrive but lack trace_id, the stage.json field names don't match
Step 4: Check field name mismatch
Alloy’s stage.json extracts:
.TraceIdfor .NET.otelTraceIDfor Python
If a service uses different field names, trace_id will be empty. Check the raw log JSON.
Step 5: Verify structured metadata is enabled in Loki
kubectl -n otel-lab exec statefulset/loki -- cat /etc/loki/config.yaml | grep allow_structured
# Should show: allow_structured_metadata: true
Exemplar dots not showing in Grafana
Symptoms
- Histogram panels show time series but no scatter dots
Diagnosis checklist (must ALL be true)
-
Panel → Edit → Query → “Exemplars” toggle is ON
-
Panel → Options → Data links has an entry pointing to Jaeger datasource, URL field =
${__value.raw} -
Prometheus has
--enable-feature=exemplar-storage:kubectl -n otel-lab describe deploy/prometheus | grep exemplar-storage -
App has
OTEL_METRICS_EXEMPLAR_FILTER=trace_basedin Deployment env:kubectl -n otel-lab exec deploy/gateway-api -- env | grep EXEMPLAR -
The histogram observation happens inside a sampled span. Use
/api/slow(always sampled) to test.
Force an exemplar-generating request:
curl http://localhost:8080/api/slow
# Wait ~5s, then check Grafana panel for new exemplar dot
K8s attributes missing from spans
Symptoms
- Spans in Jaeger lack
k8s.pod.name,k8s.namespace.nameetc.
Diagnosis
Step 1: Check RBAC
kubectl get clusterrolebinding alloy -o yaml
# Should reference ServiceAccount alloy in otel-lab namespace
kubectl auth can-i list pods --as=system:serviceaccount:otel-lab:alloy
# Should return: yes
Step 2: Check k8sattributes processor logs
kubectl -n monitoring logs daemonset/grafana-k8s-alloy-receiver --tail=100 \
| grep -i "k8sattr\|k8s.pod"
Step 3: Verify pod association mode
The configmap uses source { from = "connection" } — it resolves the pod from the OTLP connection
source IP. This works when pods have their own network namespace (standard in k3d). If pods share
the node network namespace, use source { from = "resource_attribute" } instead and set
k8s.pod.name in the app’s OTEL_RESOURCE_ATTRIBUTES.
Grafana Cloud export not working
Symptoms
- Alloy logs show export errors
- Traces/metrics/logs missing in Grafana Cloud
Diagnosis
# Mode-aware triage: conf.yml values, pod state, Alloy exporter counters,
# remote-write reachability probe, alloy-receiver endpoint check — start here.
./scripts/debug.sh
# Check the secret exists and is populated
kubectl -n monitoring get secret grafana-cloud-secrets -o json \
| python3 -c 'import json,sys,base64; d=json.load(sys.stdin)["data"]; [print(f"{k}: {base64.b64decode(v).decode()[:4]}****") for k,v in d.items()]'
# Check Alloy is reading the env vars
kubectl -n monitoring exec daemonset/grafana-k8s-alloy-receiver -- env | grep GRAFANA
# Check Alloy logs for export errors
kubectl -n monitoring logs daemonset/grafana-k8s-alloy-receiver --tail=100 \
| grep -E "grafana_cloud|export.*fail|endpoint.*empty|401|403"
| Error | Cause | Fix |
|---|---|---|
endpoint is empty | conf.yml’s monitoring.grafana_cloud.* is unset, or Alloy wasn’t redeployed after it changed | Populate via ./scripts/fetch-grafana-cloud-conf-from-akv.sh, then ./deploy-local.sh --skip-cluster --skip-build |
401 Unauthorized | Wrong API key or wrong instance ID | Re-check with ./scripts/fetch-grafana-cloud-conf-from-akv.sh --dry-run; verify Grafana Cloud Access Policies |
connection refused | Wrong endpoint format | Tempo must be host:443 (no https://) in conf.yml; the fetch script applies this adjustment automatically |
403 Forbidden | API key lacks scope | Ensure scopes: metrics:write logs:write traces:write |
Prefer
./scripts/fetch-grafana-cloud-conf-from-akv.sh+./deploy-local.shovermake secrets-fetch-akv/make secrets-applyfor this. The Makefile targets are legacy — they write the K8s Secret directly and drive their ownhelm upgrade, bypassingdeploy-local.shentirely, andsecrets-applyin particular is only as correct as whatever you put in.envmanually.secrets-fetch-akvwrites the correct Mimir endpoint format (.../api/prom/push, matching values-cloud.yaml.tmpl’s Prometheus remote_write destination) as of this fix, but the script-based flow remains the canonical path — see docs/deployment/grafana-cloud.md for the full credential model.
Consumer not processing messages
Symptoms
- Messages accumulate in RabbitMQ
notificationsqueue - Notification-svc pods appear Running but notifications don’t appear
Diagnosis
Step 1: Check consumer thread is alive
kubectl -n otel-lab logs deploy/notification-svc --tail=50 | grep -i "consumer\|rabbit"
Step 2: Check for backoff
kubectl -n otel-lab logs deploy/notification-svc --tail=100 | grep "Consumer crashed"
# If present, the consumer is in exponential backoff — check the delay and underlying error
Step 3: Check RabbitMQ connectivity
kubectl -n otel-lab exec deploy/notification-svc -- python3 -c \
"import pika; pika.BlockingConnection(pika.ConnectionParameters('rabbitmq.otel-lab'))"
# Should succeed with no output
Step 4: Check DLQ
In RabbitMQ Management → Queues → notifications.dlq:
- If messages are here, they were NACKed with
requeue=False(unrecoverable errors) - Inspect the message body and headers to understand the failure
Redis connection errors
Symptoms
- Notification-svc logs:
Redis connection lost, reconnecting - Notifications API returns 500
Diagnosis
kubectl -n otel-lab get pod -l app=redis
kubectl -n otel-lab exec deploy/notification-svc -- python3 -c \
"import redis; r=redis.Redis(host='redis.otel-lab'); print(r.ping())"
If Redis has restarted, all notification state is lost (ephemeral Deployment, no PVC). Consumer will re-process messages from RabbitMQ on the next delivery, and new notifications will be stored correctly.
App pods in CrashLoopBackOff
Most common causes and fixes:
| App | Likely cause | Fix |
|---|---|---|
| gateway-api | Missing GATEWAY_DB_CONNECTION secret | kubectl -n otel-lab get secret db-secrets + verify key exists |
| order-api | Missing ORDER_DB_CONNECTION secret or PostgreSQL not ready | Check datastore pod status |
| notification-svc | RabbitMQ not ready | Consumer has backoff — pod stays Running, consumer retries internally |
| Any | Image not imported into k3d | make import |
# Get detailed startup error
kubectl -n otel-lab describe pod <pod-name>
kubectl -n otel-lab logs <pod-name> --previous