Runbooks

Troubleshooting playbooks for every known Signal Forge failure mode, from missing traces to Grafana Cloud export errors.

Updated September 6, 2026

Runbooks

Troubleshooting playbooks for every known failure mode.


Immutable CD promotion blocked or rolled back

Scope

This runbook applies only when a GitHub Environment has DEPLOY_ENABLED=true. Without that flag, CD is intentionally render-only and no cluster rollback is expected. The promotion input is a successful CI run ID and its immutable release manifest — never a tag, a source branch, or a locally rebuilt image.

First checks

  1. Open the CD run summary and record the selected CI run ID, release commit, target environment, and the four repository@sha256:... references. Do not copy these from latest.
  2. Determine the failed boundary: pre-deploy evidence verification, server-side apply/rollout, exact-digest health gate, smoke test, observability gate, or DEV-only ZAP baseline.
  3. Download the deployment-plan-<environment>-<commit> artifact. It is intentionally secret-free and shows the rendered digest references, runtime ConfigMaps, ingress host, and labels used for the attempted deployment.
  4. If the failure occurred before Apply immutable release, no cluster mutation occurred. Fix the environment configuration or the release, then start a new promotion from a successful CI run.

Verify the currently running release

namespace=otel-lab  # replace only if the protected Environment uses another namespace
for service in otel-frontend gateway-api order-api notification-svc; do
  kubectl -n "$namespace" get deployment "$service" \
    -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
done

Every returned application image must be a ghcr.io/...@sha256:... reference. A mutable tag is not a valid known-good rollback target. Confirm workload availability and restart counts before declaring recovery:

kubectl -n "$namespace" get deployments
kubectl -n "$namespace" get pods -l tier=app

Automatic rollback behavior

After an apply, any failed health, smoke, observability, or DEV DAST gate triggers CD to restore the complete previous four-image immutable set and the previous signal-forge-app-env and frontend-env-js ConfigMaps. It refuses partial rollback because mixing independent service versions creates a release that was never tested together.

If the run reports “No complete previous immutable release is available for rollback”, stop. Do not run kubectl apply -k k8s/overlays/prod, use latest, or rebuild a prior commit. Recover only with an operator-approved complete digest manifest after investigating why the target had no captured known-good state.

Observability release gate blocked

The gate is scripts/ci/observability_gate.py, run by the observability-gate CD phase. It polls Grafana Cloud (Mimir / Loki / Tempo) with the OBS_GATE_* read credentials for the candidate’s telemetry — keyed on service_version=<release commit> and deployment_environment=signal-forge-<env> — and prints one JSON line {status, summary, checks}. It fails closed for block or an unknown/error response; warn is visible in DEV/QA and blocks PROD.

Read the checks array in the step log. Each entry names the failing check:

checkmeaningfirst response
candidate-freshnessno recent traces_spanmetrics_* for service_version=<commit>confirm the pods run the release digest (verify-deployment passed?); check OTEL_EXPORTER_OTLP_ENDPOINT; check the synthetic-journey step actually completed
resource-attributesservice_namespace absent on live seriesthe cloud values lost openTelemetryConversion.resourceToTelemetryConversion: true, or the chart is an old render
collector-healthotelcol_exporter_send_failed_* / remote-write failing / receiver refusingsee Collector export failures / Remote-write to Mimir failing
slo-burncandidate already above the 6× slow-burn ratio or the p99 targetthis is a real regression — do not force past it; roll back and investigate the candidate
synthetic-journeythe synthetic trace is missing spans from a servicecontext propagation broke on one hop; run the cross-language integration test locally

warn-only outcomes (collector-health unavailable, slo-burn rules not populated) mean the supporting infra is not fully in place — the chart’s integrations.alloy self-scrape or the apply-observability-rules phase. They do not indicate a bad candidate in DEV/QA.

Local repro: python scripts/ci/observability_gate.py --environment dev --git-sha <sha> --once with the OBS_GATE_* vars exported.

See Immutable CI/CD Promotion for the full gate order and environment contract.


Availability burn-rate alert firing

SignalForgeAvailabilityFastBurn (page) or SignalForgeAvailabilitySlowBurn (ticket) means a service’s SERVER-span error ratio is burning the 99.5% error budget 14.4× (fast, 5m+30m) or 6× (slow, 30m+6h) faster than sustainable.

  1. Confirm scope: sli:error_ratio:rate5m{service_name="<svc>"} in Grafana Explore (Mimir). A single spiking service points at that service; all three points at a shared dependency (datastore, RabbitMQ) — check DatastoreDown.
  2. Identify the failing operation: sum by (http_route, http_response_status_code) (rate(traces_spanmetrics_calls_total{service_name="<svc>",span_kind="SPAN_KIND_SERVER",status_code="STATUS_CODE_ERROR"}[5m])).
  3. Pull an exemplar trace from the error-ratio panel and follow it to the failing downstream span.
  4. If the burn is release-correlated, roll back (deployment-plan artifact has the prior digests).
  5. Escalation: GitHub issue on the repo with the alert, the top failing route, and one exemplar trace ID.

Latency SLO breach

SignalForgeLatencyFastBurn / SignalForgeLatencySlowBurn are the customer-facing latency SLO alerts: they use the configured good/total histogram buckets and 1% error budget. The p99 alerts remain diagnostics: gateway >500ms and order-api >300ms over 5m.

  1. Confirm the SLO: sli:latency_bad_ratio:rate5m{service_name="<svc>"} and its 30m companion. Use sli:latency_p99:rate5m only to describe the tail, not eligibility.
  2. Break down by route: histogram_quantile(0.99, sum by (le, http_route) (rate(traces_spanmetrics_duration_milliseconds_bucket{service_name="<svc>",span_kind="SPAN_KIND_SERVER"}[5m]))).
  3. Use an exemplar to find where the time goes — DB span, gRPC fan-out, RabbitMQ publish.
  4. Check SignalForgeCollectorQueueSaturated and pod CPU throttling before assuming an app cause.

Notification consumer error rate high

SignalForgeNotificationConsumerErrors, SignalForgeNotificationConsumerFastBurn, or SignalForgeNotificationConsumerSlowBurn apply to terminal outcomes only: ACK/duplicate are good; DLQ is bad; requeued transient attempts are excluded.

  1. Check the dead-letter queue depth (see Consumer not processing messages).
  2. kubectl -n otel-lab logs deploy/notification-svc | grep -i 'failed\|traceback' for the failing order_ids.
  3. Distinguish transient (RabbitMQ / Redis blips — self-heal) from poison messages (same order_id failing repeatedly — see the DLQ ADR).

Frontend RUM

SignalForgeSloBudgetWarning/SignalForgeSloBudgetExhausted for frontend-rum-availability or frontend-rum-latency use a CLIENT-span good/total ratio as a proxy for client-perceived health (see SLOs & burn-rate alerts) — a failed/slow browser call to the API, not literally page-load success or a JS error. The release gate’s own frontend-rum check (scripts/ci/observability_gate.py) verifies that literal signal directly, from a Playwright journey plus Faro’s own telemetry, and fails closed for a different, more specific set of reasons:

  1. Playwright journey itself failed — page load or a JS exception the script observed directly. Re-run src/frontend/e2e/observability.spec.ts against the target environment and read its own console/page-error output first; the CLIENT-span ratio checks below assume the journey ran.
  2. Browser trace missing a backend span — otel-frontend and at least one backend service must both appear under the same Tempo trace ID. A missing backend span means the traceparent header never propagated from Faro’s TracingInstrumentation into gateway-api/order-api. Check the frontend’s CORS/allowed-origins config for the collector endpoint and confirm FARO_COLLECTOR_URL in frontend-env-js points at the live alloy-faro receiver.
  3. No Faro log intake in Loki for the session — the collector never received Faro’s payload at all for that browser run. Confirm the faro.receiver route is up (kubectl -n otel-lab logs deploy/otel-frontend for client-side POST failures to /collect) and that the session’s app_name/app_environment labels match what the query expects.
  4. Faro reported a JS exception for the session — a real client-side error occurred even though the Playwright script’s own assertions passed. Query {app_name="otel-frontend",kind="exception"} | logfmt | session_id="<id>" in Loki for the stack.

For the operational SLO alerts specifically (not the release gate), treat them the same as any other 30-day good/total SLO — see SLO error budget low or exhausted.

SLO error budget low or exhausted

SignalForgeSloBudgetWarning means a 30-day good/total SLO has less than 25% budget remaining. SignalForgeSloBudgetExhausted means its remaining budget is zero or below.

  1. Identify the exact service_name and slo in the alert, then inspect slo:<id>:compliance:30d, :budget_consumed:30d, and :budget_remaining:30d in Mimir.
  2. Confirm the configured traffic floor is met; absent 30-day records are unknown, never healthy.
  3. Stop feature promotion, prioritize rollback/remediation, and attach trace/log evidence to the incident or corrective-action issue.
  4. For an exhausted PROD budget, only a remediation release may proceed, and only with an approved, expiring slo-error-budget-exhausted waiver in config/observability/waivers.yaml.

Cardinality or ingestion budget warning

Scope

SignalForgeCardinalityBudgetWarning and SignalForgeCardinalityBudgetExceeded identify the specific service_name and cardinality_budget_id that crossed its declared active-series or delivered-samples-per-second threshold. The source of truth is config/observability/cardinality-budgets.yaml.

First checks

  1. Confirm the target environment and inspect the exact budget id in the alert labels.
  2. In Grafana Explore, group the matching metric family by __name__, service_name, and each declared dimension; do not add pod, user, request, or trace identifiers as labels while debugging.
  3. Compare the resulting series and sample rate to the ledger’s measured value and 20% headroom.
  4. Roll back the release that added the dimension or instrument if the owner cannot reduce it safely. An increase to the ledger requires a reviewed pricing/capacity update, not a silence.

Span metrics absent

SignalForgeSpanMetricsAbsent (page, gateway-api/order-api) or SignalForgeNotificationSpanMetricsAbsent (ticket) — traces_spanmetrics_calls_total has produced no SERVER samples for 10–15m. Every SLO and the release gate’s SLO check has no input.

  1. up{job=~".*alloy.*"} — is the receiver alive? See Alloy receiver down.
  2. cloud mode: confirm the running Helm release still has applicationObservability.connectors.spanMetrics.enabled: true and namespace: traces.spanmetrics — helm get values <release> -n monitoring. An old render (or a revert to chart 4.x without the connector) is the usual cause.
  3. local mode: kubectl -n otel-lab get cm alloy-config -o yaml | grep -A2 'connector.spanmetrics'.
  4. App side: is the service emitting traces at all? Check otelcol_receiver_accepted_spans_total.

Collector export failures

SignalForgeCollectorExportFailing — Alloy is dropping >1% of spans on the way to Tempo.

  1. sum(rate(otelcol_exporter_send_failed_spans_total{job="integrations/alloy"}[10m])) vs ..._sent_spans_total — confirm the ratio.
  2. Alloy logs: kubectl -n monitoring logs -l app.kubernetes.io/name=alloy | grep -i 'export\|tempo\|429\|401'. 401 → wrong token type (glc_ access-policy, not glsa_). 429 → Tempo rate limit / quota.
  3. SignalForgeCollectorQueueSaturated alongside → downstream is slow; check Grafana Cloud status.
  4. ./scripts/debug.sh runs the mode-aware exporter-counter + reachability probe.

Remote-write to Mimir failing

SignalForgeRemoteWriteFailing — prometheus_remote_storage_samples_{failed,dropped}_total is increasing; metric samples (including the SLI rule inputs) are not reaching Mimir.

  1. prometheus_remote_storage_samples_failed_total by url — which destination.
  2. Alloy alloy-metrics logs for remote_write errors — 401 (token), 400 (out-of-order / label issues), 429 (ingestion limit).
  3. prometheus_remote_storage_shards vs _shards_max — if pinned at max, Alloy cannot keep up; check CPU and the Grafana Cloud metrics quota.

OTLP receiver refusing spans

SignalForgeReceiverRefusing — otelcol_receiver_refused_spans_total is climbing; app traffic is turned away at the collector.

  1. Alloy receiver logs for the refusal reason (memory limiter, payload too large, auth).
  2. SignalForgeCollectorQueueSaturated → backpressure from a slow exporter is propagating to the receiver; fix the export path first.
  3. Check alloy-receiver pod memory vs its memory_limiter limit.

Alloy receiver down

AlloyReceiverDown — up{job="integrations/alloy", pod=~".*alloy-receiver.*"} == 0 for 5m. All app OTLP is being discarded.

  1. kubectl -n monitoring get pods -l app.kubernetes.io/name=alloy — is alloy-receiver Running/Ready?
  2. kubectl -n monitoring describe pod <alloy-receiver-pod> — OOMKilled, image pull, scheduling.
  3. kubectl -n monitoring logs <alloy-receiver-pod> --previous if it is crash-looping.
  4. Same symptom set as No traces in Jaeger — that section has the connectivity checks from the app side.

Datastore not Ready

DatastoreDown — a single-replica datastore pod (mysql-*, postgres-*, redis-*, rabbitmq-*) has been not-Ready for 3m. App-tier requests are failing.

  1. kubectl -n otel-lab describe pod <pod> and logs --previous.
  2. redis / rabbitmq errors also show up as Redis connection errors / Consumer not processing messages.
  3. Graduating past single-replica: datastore HA.

No traces in Jaeger

Symptoms

  • Jaeger UI shows no services
  • make validate passes but traces don’t appear

Diagnosis

Step 1: Is alloy-receiver running?

kubectl -n monitoring get pods -l app.kubernetes.io/component=alloy-receiver
# Expected: Running

Step 2: Is the app sending to the right endpoint?

kubectl -n otel-lab exec deploy/gateway-api -- env | grep OTEL
# OTEL_EXPORTER_OTLP_ENDPOINT should be:
# http://grafana-k8s-alloy-receiver.monitoring.svc.cluster.local:4317

Step 3: Can the app reach Alloy?

kubectl -n otel-lab exec deploy/gateway-api -- \
  wget -qO- http://grafana-k8s-alloy-receiver.monitoring.svc.cluster.local:4317
# gRPC will return an HTTP 400 (expected — it's not HTTP/1.1) — this confirms connectivity

Step 4: Check Alloy receiver logs

kubectl -n monitoring logs daemonset/grafana-k8s-alloy-receiver --tail=100 \
  | grep -E "error|warn|export"

Step 5: Check Alloy pipeline UI

kubectl port-forward svc/grafana-k8s-alloy-receiver 12345 -n monitoring
open http://localhost:12345
# Navigate to Components → otelcol.receiver.otlp.default → check "Data received" counter

Step 6: Is Jaeger accessible?

curl -s http://localhost:16686/api/services
# Should return {"data":["gateway-api","order-api",...]}

Metrics missing from Prometheus

Symptoms

  • Prometheus has no metrics from app services
  • traces_spanmetrics_calls_total query returns nothing

Diagnosis

Step 1: Check Prometheus is up

curl -s http://localhost:9090/-/ready

Step 2: Prometheus has remote-write receiver enabled?

kubectl -n otel-lab describe deploy/prometheus | grep -A5 "Command"
# Should include: --web.enable-remote-write-receiver
# and: --enable-feature=exemplar-storage

Step 3: Check Alloy is writing to Prometheus

kubectl port-forward svc/grafana-k8s-alloy-receiver 12345 -n monitoring
# Navigate to Components → prometheus.remote_write.local → check "Samples sent" counter

Step 4: Query Prometheus directly

curl "http://localhost:9090/api/v1/query?query=up" | jq '.data.result'

Async propagation not working

Symptoms

  • notification.process span in Jaeger has a different traceId than order.publish
  • SpanLink is missing (no dashed arrow in Jaeger)

Diagnosis

Step 1: Verify traceparent is in the RabbitMQ message

In RabbitMQ Management (http://localhost:15672):

  1. Go to Queues → notifications
  2. Click “Get Message(s)”
  3. Inspect the Properties → Headers
  4. Should contain key traceparent with value 00-<32 hex chars>-<16 hex chars>-01

If missing: the order-api publisher is not writing the header. Check OrderPublisher.cs — it sets props.Headers["traceparent"] from the traceParent string persisted with the outbox row (the order.create context), as raw UTF-8 bytes. If OutboxMessage.TraceParent was never stored, OrderGrpcService.cs did not capture Activity.Current?.Id at order-create time.

Step 2: Verify the consumer extracts it correctly

kubectl -n otel-lab logs deploy/notification-svc --tail=50 | grep -i trace

Add temporary debug logging to consumer.py:

logger.debug("headers: %s", properties.headers)

Step 3: Check for pika instrumentation conflict

If opentelemetry-instrumentation-pika is also running, it may overwrite the extracted context. Verify requirements.txt — opentelemetry-instrumentation-pika should not be present (we use manual extraction).


Logs not appearing in Loki with trace correlation

Symptoms

  • “Logs for this span” in Grafana returns no results
  • Loki has logs but they lack trace_id structured metadata

Diagnosis

Step 1: Confirm apps write JSON

kubectl -n otel-lab logs deploy/gateway-api --tail=3
# Should be JSON: {"Timestamp":"...","Level":"Information","TraceId":"4bf..."}
# NOT plain text: info: Processing request

Step 2: Check alloy-logs is running

kubectl -n monitoring get pods -l app.kubernetes.io/component=alloy-logs
kubectl -n monitoring logs daemonset/grafana-k8s-alloy-logs --tail=50

Step 3: Query Loki directly

kubectl port-forward svc/loki 3100 -n otel-lab
curl -G "http://localhost:3100/loki/api/v1/query_range" \
  --data-urlencode 'query={namespace="otel-lab"}' \
  --data-urlencode 'limit=5'
# If logs arrive but lack trace_id, the stage.json field names don't match

Step 4: Check field name mismatch

Alloy’s stage.json extracts:

  • .TraceId for .NET
  • .otelTraceID for Python

If a service uses different field names, trace_id will be empty. Check the raw log JSON.

Step 5: Verify structured metadata is enabled in Loki

kubectl -n otel-lab exec statefulset/loki -- cat /etc/loki/config.yaml | grep allow_structured
# Should show: allow_structured_metadata: true

Exemplar dots not showing in Grafana

Symptoms

  • Histogram panels show time series but no scatter dots

Diagnosis checklist (must ALL be true)

  • Panel → Edit → Query → “Exemplars” toggle is ON

  • Panel → Options → Data links has an entry pointing to Jaeger datasource, URL field = ${__value.raw}

  • Prometheus has --enable-feature=exemplar-storage:

    kubectl -n otel-lab describe deploy/prometheus | grep exemplar-storage
  • App has OTEL_METRICS_EXEMPLAR_FILTER=trace_based in Deployment env:

    kubectl -n otel-lab exec deploy/gateway-api -- env | grep EXEMPLAR
  • The histogram observation happens inside a sampled span. Use /api/slow (always sampled) to test.

Force an exemplar-generating request:

curl http://localhost:8080/api/slow
# Wait ~5s, then check Grafana panel for new exemplar dot

K8s attributes missing from spans

Symptoms

  • Spans in Jaeger lack k8s.pod.name, k8s.namespace.name etc.

Diagnosis

Step 1: Check RBAC

kubectl get clusterrolebinding alloy -o yaml
# Should reference ServiceAccount alloy in otel-lab namespace
kubectl auth can-i list pods --as=system:serviceaccount:otel-lab:alloy
# Should return: yes

Step 2: Check k8sattributes processor logs

kubectl -n monitoring logs daemonset/grafana-k8s-alloy-receiver --tail=100 \
  | grep -i "k8sattr\|k8s.pod"

Step 3: Verify pod association mode

The configmap uses source { from = "connection" } — it resolves the pod from the OTLP connection source IP. This works when pods have their own network namespace (standard in k3d). If pods share the node network namespace, use source { from = "resource_attribute" } instead and set k8s.pod.name in the app’s OTEL_RESOURCE_ATTRIBUTES.


Grafana Cloud export not working

Symptoms

  • Alloy logs show export errors
  • Traces/metrics/logs missing in Grafana Cloud

Diagnosis

# Mode-aware triage: conf.yml values, pod state, Alloy exporter counters,
# remote-write reachability probe, alloy-receiver endpoint check — start here.
./scripts/debug.sh

# Check the secret exists and is populated
kubectl -n monitoring get secret grafana-cloud-secrets -o json \
  | python3 -c 'import json,sys,base64; d=json.load(sys.stdin)["data"]; [print(f"{k}: {base64.b64decode(v).decode()[:4]}****") for k,v in d.items()]'

# Check Alloy is reading the env vars
kubectl -n monitoring exec daemonset/grafana-k8s-alloy-receiver -- env | grep GRAFANA

# Check Alloy logs for export errors
kubectl -n monitoring logs daemonset/grafana-k8s-alloy-receiver --tail=100 \
  | grep -E "grafana_cloud|export.*fail|endpoint.*empty|401|403"
ErrorCauseFix
endpoint is emptyconf.yml’s monitoring.grafana_cloud.* is unset, or Alloy wasn’t redeployed after it changedPopulate via ./scripts/fetch-grafana-cloud-conf-from-akv.sh, then ./deploy-local.sh --skip-cluster --skip-build
401 UnauthorizedWrong API key or wrong instance IDRe-check with ./scripts/fetch-grafana-cloud-conf-from-akv.sh --dry-run; verify Grafana Cloud Access Policies
connection refusedWrong endpoint formatTempo must be host:443 (no https://) in conf.yml; the fetch script applies this adjustment automatically
403 ForbiddenAPI key lacks scopeEnsure scopes: metrics:write logs:write traces:write

Prefer ./scripts/fetch-grafana-cloud-conf-from-akv.sh + ./deploy-local.sh over make secrets-fetch-akv / make secrets-apply for this. The Makefile targets are legacy — they write the K8s Secret directly and drive their own helm upgrade, bypassing deploy-local.sh entirely, and secrets-apply in particular is only as correct as whatever you put in .env manually. secrets-fetch-akv writes the correct Mimir endpoint format (.../api/prom/push, matching values-cloud.yaml.tmpl’s Prometheus remote_write destination) as of this fix, but the script-based flow remains the canonical path — see docs/deployment/grafana-cloud.md for the full credential model.


Consumer not processing messages

Symptoms

  • Messages accumulate in RabbitMQ notifications queue
  • Notification-svc pods appear Running but notifications don’t appear

Diagnosis

Step 1: Check consumer thread is alive

kubectl -n otel-lab logs deploy/notification-svc --tail=50 | grep -i "consumer\|rabbit"

Step 2: Check for backoff

kubectl -n otel-lab logs deploy/notification-svc --tail=100 | grep "Consumer crashed"
# If present, the consumer is in exponential backoff — check the delay and underlying error

Step 3: Check RabbitMQ connectivity

kubectl -n otel-lab exec deploy/notification-svc -- python3 -c \
  "import pika; pika.BlockingConnection(pika.ConnectionParameters('rabbitmq.otel-lab'))"
# Should succeed with no output

Step 4: Check DLQ

In RabbitMQ Management → Queues → notifications.dlq:

  • If messages are here, they were NACKed with requeue=False (unrecoverable errors)
  • Inspect the message body and headers to understand the failure

Redis connection errors

Symptoms

  • Notification-svc logs: Redis connection lost, reconnecting
  • Notifications API returns 500

Diagnosis

kubectl -n otel-lab get pod -l app=redis
kubectl -n otel-lab exec deploy/notification-svc -- python3 -c \
  "import redis; r=redis.Redis(host='redis.otel-lab'); print(r.ping())"

If Redis has restarted, all notification state is lost (ephemeral Deployment, no PVC). Consumer will re-process messages from RabbitMQ on the next delivery, and new notifications will be stored correctly.


App pods in CrashLoopBackOff

Most common causes and fixes:

AppLikely causeFix
gateway-apiMissing GATEWAY_DB_CONNECTION secretkubectl -n otel-lab get secret db-secrets + verify key exists
order-apiMissing ORDER_DB_CONNECTION secret or PostgreSQL not readyCheck datastore pod status
notification-svcRabbitMQ not readyConsumer has backoff — pod stays Running, consumer retries internally
AnyImage not imported into k3dmake import
# Get detailed startup error
kubectl -n otel-lab describe pod <pod-name>
kubectl -n otel-lab logs <pod-name> --previous