Project README

The full SignalForge repository README.

On this page
Navigation

End-to-end OpenTelemetry instrumentation lab across .NET 8, Python/FastAPI, and Angular 17, deployed on k3d with Helm-managed Grafana Alloy agents as the collector stack.

What it validates: traces (5-hop cross-language), span metrics, exemplars, async trace propagation via RabbitMQ, frontend RUM with Faro, log-to-trace correlation via Loki, and tail-based sampling.

Work outside-in, purpose → design → implementation:

StepFileWhy
1https://shipsolid.github.io/signal-forge/spec/The “what” — all services, patterns to validate, the validation checklist at §11
2https://shipsolid.github.io/signal-forge/architecture/overview/Topology diagram, signal flow per type, port map
3https://shipsolid.github.io/signal-forge/architecture/adrs/10 ADRs that explain the non-obvious “why” (most important before touching anything)
4conf.ymlThe single control file — every knob deploy-local.sh reads
5https://shipsolid.github.io/signal-forge/observability/pipeline/Alloy River config stage-by-stage; the heart of the lab
6src/order-api/Richest service: gRPC, Outbox, RabbitMQ publish with W3C traceparent injection
7src/notification-svc/Python consumer, SpanLink, cross-language async propagation
8src/gateway-api/.NET BFF, exemplars, UpDownCounter, fan-out pattern
9src/frontend/Angular + Faro RUM — browser-to-backend trace propagation
10k8s/Manifests: infra/ → datastores/ → app/ → monitoring/

For ops understanding: deploy-local.sh → scripts/debug.sh → .github/workflows/ci.yml.


Purpose

SignalForge exists to provide a portable, reproducible environment for validating OpenTelemetry instrumentation patterns across multiple runtimes and communication protocols. It is not a toy: it models production-grade concerns — tail-based sampling, async context propagation, exemplar plumbing, SLO recording rules, and supply-chain controls — in a self-contained k3d cluster that any engineer can spin up on a laptop.

The lab is consumed by engineers who need to test instrumentation changes before they land on production clusters, and by anyone building familiarity with the Grafana Alloy / k8s-monitoring Helm chart. It is a reference implementation, not a template — copy patterns from it, but do not fork it as application scaffolding.

It lives here rather than inside the main monorepo because it has its own k3d cluster lifecycle, separate image builds, and Grafana Cloud credentials that are scoped to a dev stack and should not bleed into production pipelines.


Architecture

%%{init: {'theme':'base','themeVariables':{'background':'#0f172a','primaryColor':'#bae6fd','primaryTextColor':'#0f172a','primaryBorderColor':'#7dd3fc','secondaryColor':'#bbf7d0','tertiaryColor':'#fde68a','lineColor':'#cbd5e1','clusterBkg':'#1e293b','clusterBorder':'#94a3b8','titleColor':'#f8fafc','edgeLabelBackground':'#1e293b'}}}%%
flowchart TD
    Browser["Browser (Faro RUM)"] --> Gateway["gateway-api (.NET 8, :5000)"]
    Gateway --> MySQL["MySQL 8 (EF Core)"]
    Gateway --> OrderAPI["order-api (.NET 8, gRPC :5002)"]
    Gateway -->|HTTP| Notification["notification-svc (Python, :8000)"]
    OrderAPI --> Postgres["PostgreSQL 16 (Npgsql)"]
    OrderAPI -->|outbox relay| RabbitMQ["RabbitMQ"]
    RabbitMQ --> Notification
    Notification --> Redis["Redis 7"]
    classDef app fill:#bae6fd,stroke:#7dd3fc,color:#0f172a;
    classDef data fill:#bbf7d0,stroke:#4ade80,color:#0f172a;
    class Browser,Gateway,OrderAPI,Notification app;
    class MySQL,Postgres,RabbitMQ,Redis data;

All services push OTLP to alloy-receiver (Helm-managed DaemonSet, monitoring namespace):

%%{init: {'theme':'base','themeVariables':{'background':'#0f172a','primaryColor':'#bae6fd','primaryTextColor':'#0f172a','primaryBorderColor':'#7dd3fc','secondaryColor':'#bbf7d0','tertiaryColor':'#fde68a','lineColor':'#cbd5e1','clusterBkg':'#1e293b','clusterBorder':'#94a3b8','titleColor':'#f8fafc','edgeLabelBackground':'#1e293b'}}}%%
flowchart TD
    SDK["App SDK"] -->|OTLP gRPC :4317| Receiver[alloy-receiver]
    Receiver --> K8sAttrs[k8sattributes enrichment]
    K8sAttrs --> Transform["transform (stamp deployment.environment)"]
    Transform --> Filter["filter (drop /healthz spans)"]
    Filter --> SpanMetrics[spanmetrics connector]
    SpanMetrics --> RED["RED metrics (before sampling)"]
    Filter --> TailSampling["tail_sampling (errors=100%, slow>2s=100%, rest=25%)"]
    TailSampling --> Batch[batch]
    Batch --> Backend1["Tempo (cloud) | Jaeger (local)"]

    Logs["alloy-logs (DaemonSet)"] --> LogsTail["pod stdout tailing"] --> LogsCorr["trace correlation"] --> Loki[Loki]

    Metrics["alloy-metrics (StatefulSet)"] --> MetricsScrape["kubelet/cAdvisor/KSM"] --> Backend2["Mimir | Prometheus"]
    classDef pipeline fill:#bae6fd,stroke:#7dd3fc,color:#0f172a;
    classDef backend fill:#bbf7d0,stroke:#4ade80,color:#0f172a;
    class SDK,Receiver,K8sAttrs,Transform,Filter,SpanMetrics,RED,TailSampling,Batch,Logs,LogsTail,LogsCorr,Metrics,MetricsScrape pipeline;
    class Backend1,Loki,Backend2 backend;

The single most important configuration knob is monitoring.mode in conf.yml:

monitoring.modeAlloy destinationsIn-cluster backends
cloud (default)Grafana Cloud Tempo / Mimir / Lokinone
localIn-cluster Jaeger / Prometheus / Loki / GrafanaJaeger :16686, Prometheus :9090, Grafana :3000

The two modes are mutually exclusive — there is no dual-export. Any doc saying otherwise is stale.

In cloud mode, the chart’s Alloy agents are the entire pipeline — no in-cluster Jaeger / Prometheus / Loki / Grafana are deployed. In local mode, a bespoke Alloy DaemonSet in k8s/monitoring/grafana/ exports to in-cluster backends. Helm is optional there and is installed only with --with-helm; it is mandatory in cloud mode.

alloy-logs tails pod stdout/stderr with trace-id correlation. alloy-metrics scrapes cluster infra metrics. See docs/observability/pipeline.md for the full signal flow and docs/OTEL-PATTERNS.md for per-runtime instrumentation choices.

Alloy roles (Helm release, monitoring namespace)

Alloy roleKindResponsibilityDestination (cloud)Destination (local)
alloy-metricsStatefulSetScrapes cluster infra metricsGrafana Cloud Mimirin-cluster Prometheus
alloy-singletonDeploymentCluster events, kube-state-metricsCloud Mimir + Lokiin-cluster Prom + Loki
alloy-logsDaemonSetPod + node log tailingGrafana Cloud Lokiin-cluster Loki
alloy-receiverDaemonSetOTLP push receiver (app telemetry)Cloud Tempo + Mimirin-cluster Jaeger + Prom
alloy-profilesDaemonSetDisabled — no Pyroscope——

Values: values-local.yaml.tmpl or values-cloud.yaml.tmpl — both rendered at deploy time from conf.yml.

Trace propagation

A single “Create Order” click produces a 5-hop trace across three runtimes. The RabbitMQ hop uses a SpanLink (not parent-child) because message processing is async; both spans share the same traceId and appear as a dashed arrow in Jaeger.

flowchart LR
    Browser["Browser (Faro)"] --> Gateway[gateway-api]
    Gateway --> OrderAPI[order-api]
    OrderAPI -->|SpanLink| RabbitMQ[RabbitMQ]
    RabbitMQ --> Notification[notification-svc]
    OrderAPI --> Postgres[PostgreSQL]
    Notification --> Redis[Redis]

See docs/architecture/overview.md for the full signal flow diagrams.


Ownership Boundary

DimensionDetail
TeamPersonal lab / portfolio (Amit Singh)
Primary ownerAmit Singh — see GitHub profile / repo issues for contact
On-callNone — lab environment, no production SLA
Escalation pathGitHub issues on this repo

This component does not own anything in shared infrastructure. It creates and manages its own k3d cluster (otel-lab) and its own Kubernetes namespace (otel-lab). The only external dependency with shared ownership is the Grafana Cloud stack (example-org.grafana.net) and the Azure Key Vault (example-org-prd-kv) — those are the parent organization’s platform resources and are consumed read-only by this lab.

The lab does not own the Grafana Cloud instance, the AKV vault, or any network resources outside the k3d cluster. Changes to Grafana Cloud credentials are fetched from AKV; they are never committed as live values.


Deployment Model

StageMethodTriggerOutcome
local lab./deploy-local.shmanualBuilds/imports local images into k3d otel-lab
PR CIGitHub ActionsPR / manualBlocking validation, local build/scan/SBOM; no registry publication
trusted main CIGitHub Actionspush / manualBuilds each image once, publishes signed digest references and a release manifest
CDGitHub Actionsmanual successful CI run selectionPromotes the same four digest references DEV → QA → optional protected PROD

The repository contains no live target credentials or URLs. CD therefore renders a secret-free plan and makes no cluster change unless environment-scoped DEPLOY_ENABLED=true, variables, secrets, and an observability gate are configured. This is distinct from the local k3d lab, which remains ephemeral and is the default development path. See Immutable CI/CD Promotion.

# Full deploy: cluster + Docker builds + manifests + Helm (5-15 min cold)
./deploy-local.sh

# Subsequent iteration, manifests only (<1 min)
./deploy-local.sh --skip-cluster --skip-build

# Local mode: also install the Helm monitoring chart
./deploy-local.sh --with-helm

# Local-only reapply: this is not the CD rollback path.
./deploy-local.sh --skip-cluster --skip-build

# Teardown: delete the k3d cluster entirely
./deploy-local.sh --teardown

Two deployment tools — do not mix them

ToolConfig sourceCredentialsUse
./deploy-local.sh (primary)conf.yml./scripts/fetch-grafana-cloud-conf-from-akv.sh → updates conf.ymlRecommended for all new work
Makefile (legacy).env + hand-edited valuesmake secrets-fetch-akv → writes Secret directlyReference only — prefer deploy-local.sh for new work

make secrets-fetch-akv used to write a stale GRAFANA_CLOUD_MIMIR_ENDPOINT (/api/v1/otlp) that didn’t match the chart’s expected /api/prom/push — fixed, it now writes the correct format. Still secondary/legacy; the script-based flow remains the recommended path.

Credentials (Grafana Cloud, cloud mode only)

# Preview credential diff against AKV
./scripts/fetch-grafana-cloud-conf-from-akv.sh --dry-run

# Pull and write into conf.yml (creates conf.yml.bak)
./scripts/fetch-grafana-cloud-conf-from-akv.sh

# Re-deploy with new credentials (no cluster rebuild)
./deploy-local.sh --skip-cluster --skip-build

Auth: az login first, or export ARM_CLIENT_ID + ARM_CLIENT_SECRET in the shell. See docs/deployment/grafana-cloud.md for the full credential model and rotation procedure.

Helm version and values

deploy-local.sh reads the pinned monitoring.helm.version from conf.yml and renders the cloud values template before invoking Helm. Do not duplicate the chart version in commands or hand-edit a rendered values file; update the source configuration and then re-run the deploy script.

Kustomize overlays

kubectl kustomize k8s/base                  # render full stack
kubectl kustomize k8s/overlays/prod         # render prod overlay (replicas=6, required anti-affinity)
kubectl apply -k k8s/overlays/dev           # apply dev overlay

Dependencies

DependencyTypeRequiredNotes
Docker 24+toolingyesImage builds + k3d node images
k3d v5+toolingyesLocal Kubernetes cluster
kubectl v1.28+toolingyesManifest apply
helm v3.14+toolingyesgrafana/k8s-monitoring chart install
Python 3.9+toolingyesdeploy-local.sh uses Python to parse conf.yml, render templates, run scripts
Azure CLI 2.50+toolingcloud mode only./scripts/fetch-grafana-cloud-conf-from-akv.sh — not needed if credentials are already in conf.yml
Grafana Cloud stack (example-org.grafana.net)upstreamcloud mode onlyTempo, Mimir, Loki endpoints; credentials in AKV
Azure Key Vault (example-org-prd-kv)upstreamcloud mode onlyStores Grafana Cloud API key + endpoint URLs
Zscaler CA (zcert.crt)infracorporate networks onlyPassed to local builds as an optional BuildKit secret; omitted in CI and never copied into a build context or image
grafana/k8s-monitoring Helm chartinfracloud mode (auto-installed)Version is pinned in conf.yml; pulled at deploy time with no local vendored copy
cert-manager (jetstack chart)infrawhen security.tls.enabled: trueVersion is pinned in conf.yml; installs into cert-manager; skip by setting security.tls.enabled: false

Version pins that must not drift:

  • grafana/k8s-monitoring is pinned in conf.yml. Upgrading requires re-validating all Alloy role names and values schema — the chart has breaking changes between minor versions.
  • .NET 8.0 in Dockerfiles — do not bump to .NET 9 without re-testing the OTel SDK compatibility matrix.

Operational Model

Health check:

# Mode-aware triage: pod state, Alloy exporter counters, remote-write probe
./scripts/debug.sh

# Verify all service endpoints are responding
make validate

Logs:

ServiceHow to access
Any app podkubectl -n otel-lab logs deploy/<service-name>
Alloy receiverkubectl -n monitoring logs daemonset/grafana-k8s-alloy-receiver
Loki query (local mode){namespace="otel-lab"} in Grafana Explore or via Loki API on :3100
Loki query (cloud mode){namespace="otel-lab", deployment_environment="signal-forge-dev"} in Grafana Cloud Explore

All services write structured JSON logs. Alloy’s alloy-logs DaemonSet extracts TraceId/SpanId fields and attaches them as Loki structured metadata, enabling “Logs for this span” in Grafana.

Metrics / dashboards:

  • Span-derived RED metrics: traces_spanmetrics_calls_total{service_name}, traces_spanmetrics_duration_milliseconds_bucket{service_name}
  • Cluster infra: standard kubelet/cAdvisor/KSM metrics scraped by alloy-metrics
  • Alloy pipeline UI: kubectl -n monitoring port-forward svc/grafana-k8s-alloy-receiver 12345 → http://localhost:12345

Alerts:

SLO rules live in k8s/monitoring/slo-rules.yaml. The local deploy path loads the bare Prometheus rule groups when observability.slo_rules.enabled: true (the current default); no Prometheus Operator CRD is required. Cloud Mimir publication remains an explicit operation via scripts/push-slo-rules-to-mimir.sh.

AlertSeverityTrigger
SignalForgeAvailabilityFastBurnpageerror_ratio > 7.2% in 5m AND 30m windows
SignalForgeAvailabilitySlowBurnticketerror_ratio > 3% in 30m AND 6h windows
SignalForgeGatewayLatencyHighticketgateway-api p99 > 500ms for 10m
SignalForgeDownstreamLatencyHighticketorder-api or notification-svc p99 > 300ms for 10m
AlloyReceiverDownpageup == 0 for alloy-receiver for 5m
DatastoreDownpageany datastore pod not Ready for 3m

Runbook: docs/operations/runbooks.md — covers no-traces, missing metrics, async propagation failures, log correlation gaps, exemplar troubleshooting, Grafana Cloud export errors.


Quick Start

# Full deploy: cluster + builds + apply + Helm
./deploy-local.sh

# Subsequent iteration (manifests only):
./deploy-local.sh --skip-cluster --skip-build

# Triage after a deploy:
./scripts/debug.sh

# Tear down:
./deploy-local.sh --teardown

First run: ~5-15 minutes (4 Docker builds + k3d create + cert-manager + Helm rollout). --skip-cluster --skip-build runs complete in <1 min.

Prerequisites

ToolMin versionInstall
Docker24+docs.docker.com
k3dv5+curl -s https://raw.githubusercontent.com/k3d-io/k3d/main/install.sh | bash
kubectlv1.28+kubernetes.io/docs
helmv3.14+curl https://raw.githubusercontent.com/helm/helm/main/scripts/get-helm-3 | bash
Azure CLI2.50+needed only for Grafana Cloud credential fetch via AKV
Python 33.9+system package — drives deploy-local.sh

Endpoints (after deploy)

Always available:

URLServiceNotes
http://localhost:8080Angular SPA + Gateway API/api/* → gateway, / → frontend
https://signal-forge.local:8443Same, TLSRequires security.tls.enabled: true + /etc/hosts entry
http://localhost:15672RabbitMQ Managementguest/guest

Local mode only:

URLServiceCredentials
http://localhost:16686Jaeger UI—
http://localhost:3000Grafanaadmin/admin
http://localhost:9090Prometheus—

Cloud mode only: your Grafana Cloud stack (e.g. https://example-org.grafana.net) — Explore for Tempo/Mimir/Loki.


Repository Layout

signal-forge/
├── src/                        # Application source
│   ├── gateway-api/            # .NET 8 Minimal API — BFF, MySQL, gRPC client
│   ├── order-api/              # .NET 8 gRPC — PostgreSQL, RabbitMQ publisher
│   ├── notification-svc/       # Python FastAPI — RabbitMQ consumer, Redis
│   ├── frontend/               # Angular 17 SPA — Faro RUM, nginx
│   └── proto/                  # Shared gRPC protobuf definitions
│
├── k8s/
│   ├── base/                   # Kustomize base (ArgoCD / Flux entrypoint)
│   ├── overlays/{dev,staging,prod}/
│   ├── infra/                  # namespace, secrets, PDB, NetworkPolicies, ingress, cert-manager issuer
│   ├── app/                    # Application Deployments + Services
│   ├── datastores/             # MySQL, PostgreSQL, Redis, RabbitMQ
│   ├── monitoring/
│   │   ├── slo-rules.yaml      #   Prometheus/Mimir rule groups (SLOs + burn-rate alerts)
│   │   ├── grafana/            #   Bespoke Alloy DaemonSet (local mode only)
│   │   ├── grafana-helm/       #   grafana/k8s-monitoring Helm values
│   │   └── local/              #   In-cluster backends: Jaeger, Prometheus, Loki, Grafana
│   └── loadtest/               # k6 load test Job
│
├── conf.yml                    # Single source of truth for all deploy-local.sh knobs
├── deploy-local.sh             # Idempotent local deploy script
├── scripts/                    # AKV fetch, debug, smoke tests
├── Makefile                    # Legacy lifecycle targets
└── docs/
    ├── spec.md                 # OTel validation test scenarios and checklist
    ├── OTEL-PATTERNS.md        # Instrumentation patterns per runtime
    ├── architecture/           # System topology, decisions
    ├── observability/          # SLOs, pipeline, sampling, exemplars, correlation
    ├── infrastructure/         # Hardening, Kustomize, datastores, HA paths
    └── operations/             # Networking, runbooks, supply chain, reliability

Deploy order within k8s/: infra/ → app-env ConfigMap → grafana-cloud-secrets → cert-manager → datastores/ → monitoring/ → app/ → post (ingress). Driven by deploy-local.sh with context-guard and NodePort drift-check.


Services

ServiceStackPort (cluster)DBRole
otel-frontendAngular 17 + nginx80 (host: 8080 via ingress)—SPA + Faro RUM
gateway-api.NET 8 Minimal API5000MySQLBFF, gRPC client
order-api.NET 8 gRPC5001 health; 5002 gRPCPostgreSQLOrder CRUD + transactional-outbox publisher
notification-svcPython/FastAPI8000RedisRabbitMQ consumer + REST

Testing & CI

# .NET — order-api.Tests requires Docker (Testcontainers starts a real postgres:16.4
# for OutboxRelayWorkerTests)
dotnet test src/order-api.Tests/order-api.Tests.csproj --configuration Release
dotnet test src/gateway-api.Tests/gateway-api.Tests.csproj --configuration Release

# Python
python -m pytest src/notification-svc/tests/ -v --tb=short

# Frontend
cd src/frontend && npm ci --legacy-peer-deps && npx jest --config jest.config.js

CI (.github/workflows/ci.yml) gates a release with Gitleaks, repository policy, unit/build/protobuf validation, CodeQL, SCA, Trivy IaC/container scans, and observability-as-policy validation. A trusted main run then builds each service image once, generates a CycloneDX SBOM, signs and attests the GHCR digest with keyless Cosign, and emits the release manifest consumed by CD. See Immutable CI/CD Promotion for the trust and rollback boundaries.


Grafana Cloud Mode

Set monitoring.mode: cloud in conf.yml and the Helm chart’s Alloy agents ship every signal to Grafana Cloud Tempo / Mimir / Loki. The in-cluster Jaeger / Prometheus / Loki / Grafana are not deployed — the cloud backends are the only sink.

Credentials live in Azure Key Vault. The fetch script pulls them, writes them into conf.yml in place (preserving comments), and then ./deploy-local.sh materialises them into the grafana-cloud-secrets Kubernetes Secret that the chart’s destinations reference by name.

Setup

# Azure auth: either an existing `az login` session, or export ARM_CLIENT_ID +
# ARM_CLIENT_SECRET in your shell (no .env loading).

# AKV coordinates live in conf.yml → monitoring.grafana_cloud.akv.{tenant_id,
#   subscription_id, resource_group, vault_name}. Edit these if they change.

# Pull every Grafana Cloud secret and update conf.yml in place:
./scripts/fetch-grafana-cloud-conf-from-akv.sh             # writes conf.yml + conf.yml.bak
./scripts/fetch-grafana-cloud-conf-from-akv.sh --dry-run   # preview diff only

# Re-apply (no cluster rebuild):
./deploy-local.sh --skip-cluster --skip-build

See docs/deployment/grafana-cloud.md for the full credential model and rotation procedure.

Credentials

Credentials are stored in Azure Key Vault (example-org-prd-kv) under the grafana-example-org-* secret prefix. The fetch script writes them into conf.yml at monitoring.grafana_cloud.*; deploy-local.sh materialises them into the grafana-cloud-secrets Kubernetes Secret.

AKV secret nameconf.yml keyNotes
grafana-example-org-alloy-writer-example-org-tokenmonitoring.grafana_cloud.api_keyglc_ access-policy token — scopes: metrics:write logs:write traces:write
grafana-example-org-cloud-tempo-endpointmonitoring.grafana_cloud.tempo.endpointhost only — :443 suffix added by fetch script
grafana-example-org-cloud-tempo-usernamemonitoring.grafana_cloud.tempo.userTempo instance ID
grafana-example-org-cloud-mimir-endpointmonitoring.grafana_cloud.mimir.endpointbase URL → /push suffix added if missing
grafana-example-org-cloud-mimir-usernamemonitoring.grafana_cloud.mimir.userMimir instance ID
grafana-example-org-cloud-loki-endpointmonitoring.grafana_cloud.loki.endpointbase URL → /loki/api/v1/push suffix added if missing
grafana-example-org-cloud-loki-usernamemonitoring.grafana_cloud.loki.userLoki instance ID
grafana-example-org-faro-api-endpointmonitoring.grafana_cloud.faro.endpointbrowser Faro collector (runtime env)
grafana-example-org-faro-sourcemap-tokenmonitoring.grafana_cloud.faro.api_keywebpack source-map upload (build arg)

Production Readiness Controls

For any step beyond local lab, the following controls are already implemented:

  • Container hardening — non-root UIDs per image, readOnlyRootFilesystem, securityContext
  • Kustomize layout — base + overlays for dev/staging/prod
  • Reliability — PodDisruptionBudgets, pod anti-affinity, graceful shutdown
  • Networking & TLS — NetworkPolicies, cert-manager, and the lab’s kube-router enforcement model
  • Supply-chain security — SBOM, keyless signing/attestation, immutable digest verification before deploy
  • Immutable CI/CD promotion — build-once release manifest, DEV → QA → PROD serialization, protected-environment rollback
  • SLOs & burn-rate alerts — validated SLO rules and multi-window burn thresholds; live promotion needs a configured telemetry gate
  • Datastore HA migration — CloudNativePG / RabbitMQ Operator / Redis Sentinel paths

Observability UIs

Always available (both modes):

UIURLNotes
Angular SPAhttp://localhost:8080Frontend entry point
RabbitMQhttp://localhost:15672guest / guest — inspect queues and message headers
Alloy UIkubectl -n monitoring port-forward svc/grafana-k8s-alloy-receiver 12345 then http://localhost:12345Pipeline graph, component status, debug traces

Only in monitoring.mode: local:

UIURLNotes
Grafanahttp://localhost:3000admin / admin
Jaegerhttp://localhost:16686Trace search and waterfall
Prometheushttp://localhost:9090Metric explorer + exemplars

Only in monitoring.mode: cloud: your Grafana Cloud stack’s Explore + dashboards (e.g. https://example-org.grafana.net/explore).


Make Targets Reference

./deploy-local.sh is the sole deploy path. The Makefile no longer deploys anything — it only builds images, runs tests, and fetches/applies Grafana Cloud credentials. Its former deploy/deploy-cloud/deploy-local/full/helm-repo/helm-render/deploy-helm/ deploy-helm-cloud/teardown-helm/full-helm targets (a second, parallel Jinja2-based Helm-values pipeline plus a legacy kubectl apply -f flow, both superseded by deploy-local.sh) were retired; make deploy/deploy-cloud/deploy-local/full now just print a redirect to ./deploy-local.sh and exit non-zero, so old muscle memory fails loudly instead of silently doing the wrong thing.

Cluster lifecycle

TargetDescriptionEquivalent
make cluster-upCreate k3d cluster with port mappings./deploy-local.sh (builds + deploys too)
make cluster-downDelete k3d cluster./deploy-local.sh --teardown
make buildBuild all 4 Docker images locallyimplicit in ./deploy-local.sh
make importBuild + import images into k3dimplicit
make teardownDelete the otel-lab namespaceuse ./deploy-local.sh --teardown to drop the whole cluster

Testing & ops

TargetDescription
make testRun k6 load test Job (generates realistic traffic)
make validateSmoke-test all endpoints with curl
make logsStream logs from all app pods
./scripts/debug.shMode-aware triage — pod state, Alloy exporter counters, remote-write reachability probe
./scripts/smoke-test-conf-updater.shOffline regression test for the conf.yml in-place updater

Grafana Cloud credentials (Azure Key Vault)

PathDescription
./scripts/fetch-grafana-cloud-conf-from-akv.shPreferred. Pulls AKV secrets, writes them into conf.yml in place (preserving comments), supports --dry-run. Auth via existing az login or shell-exported ARM_CLIENT_ID/SECRET.
make secrets-fetch-akvLegacy — writes the K8s Secret directly and also drives its own helm upgrade, bypassing deploy-local.sh. Kept for the manual-.env fallback case; the script-based flow above covers everything else.
make secrets-applyLegacy — applies credentials from .env.
make secrets-showPrint stored Secret values (API key redacted). Still accurate.

Helm monitoring

There’s no separate make *-helm step anymore — ./deploy-local.sh handles the Helm install inline, adding the grafana repo itself and rendering values-cloud.yaml.tmpl directly from conf.yml (see docs/deployment/helm.md). The Jinja2-based render pipeline this table used to document (render.py + config.yaml.j2, real prod Grafana Cloud fingerprints left over from a copy-paste) has been deleted along with the Makefile targets that drove it.


OTel Validation Checklist

See docs/spec.md for the full test scenarios, including:

  • Trace propagation (HTTP, gRPC, RabbitMQ async)
  • Span metrics (RED) + exemplars
  • K8s attribute enrichment via k8sattributes processor
  • Log-to-trace correlation via Loki structured metadata
  • Tail sampling rates (verify 25% sampling of normal traffic)
  • Frontend RUM (Faro) session and page view spans
  • Resilience / negative scenarios (datastore down, consumer crash)

Roadmap

PhaseItemStatusTarget
1Core 4-service OTel stack with cloud + local modedone—
1Kustomize overlays (dev/staging/prod)done—
1Immutable CI/CD: security gates, SBOM, Cosign, digest promotiondoneRender-only until GitHub Environment configuration is supplied
1SLO recording rules + multi-window burn alertsdone—
1Tail-based sampling + spanmetricsdone—
2Faro RUM source-map upload + error groupingactiveTBD — depends on Faro source-map token in AKV
1Local SLO rule loading (slo_rules.enabled: true)doneBare rules load into local Prometheus; cloud Mimir publishing remains an explicit operation
3k6 CronJob for continuous synthetic trafficplannedRequired for client-perceived SLO validation
3Pyroscope continuous profiling (alloy-profiles enable)plannedBlocked on in-cluster Pyroscope backend
3Datastore HA operators (CNPG / RabbitMQ Operator / Redis Sentinel)deferredSee https://shipsolid.github.io/signal-forge/infrastructure/datastore-ha/