Break-glass: bypassing admission policy in an incident

Who can create a PolicyException, what it must contain, how it expires, and how to revoke it manually.

Updated September 13, 2026
On this page
Navigation

Break-glass: bypassing admission policy in an incident

Break-glass is not an application-CD capability. signal-forge-deployer (the identity ordinary CD runs as) has no RBAC rights on PolicyException at all — it cannot create, modify, or delete one, by design (deploy/charts/signal-forge-policy/templates/platform-rbac.yaml). This exists precisely so a compromised or careless CD run cannot grant itself an exception to the admission policies it is otherwise bound by (docs/architecture/adrs/adr-observability-release-gate.md).

A break-glass exception is created by a human, through kubectl, using a platform identity that has its own separate RBAC (outside this chart, provisioned per the platform team’s own cluster-admin process) — never through this repository’s CI/CD.

Who

  • Approver — platform incident commander. Authorizes the exception before it’s created. Named per the active on-call rotation at the time of the incident.
  • Executor — on-call platform owner. Creates the PolicyException, runs the manual revoke command if the reaper doesn’t fire in time, and owns the post-incident review.

What the exception must contain

Enforced at admission time by signal-forge-breakglass-contract (deploy/charts/signal-forge-policy/templates/admission-policies.yaml) — a malformed or over-broad request is rejected before it ever reaches the cluster, not caught after the fact:

  • signal-forge.io/incident-id, signal-forge.io/owner, signal-forge.io/approved-by, signal-forge.io/created-at, signal-forge.io/expires-at, signal-forge.io/post-incident-review — all six annotations are required; any missing is a rejection.
  • expires-at must be no more than one hour after created-at, and must not already be expired at creation time.
  • The exception may reference only signal-forge-workload-contract and signal-forge-image-signatures — it cannot exempt any other policy, including itself (signal-forge-breakglass-contract cannot exempt signal-forge-breakglass-contract).
  • It must name exactly one protected Signal Forge workload by its real Pod-template match condition — a break-glass exception cannot be workload-unscoped.
  • created-at is immutable on update — extending an exception’s life by editing its own creation time is rejected the same way creating an already-expired one is.

Automatic revocation

signal-forge-breakglass-reaper (a Kyverno-adjacent CronJob, deploy/charts/signal-forge-policy/templates/platform-rbac.yaml) runs every five minutes under its own ServiceAccount, scoped by RBAC to exactly get/list/delete on policyexceptions.policies.kyverno.io in otel-lab — nothing else, not even read access to the Kyverno policies themselves. Every run:

  1. Lists all PolicyException objects in otel-lab.
  2. Deletes any whose signal-forge.io/expires-at has passed.
  3. Prints one sanitized JSON line per deletion to its own pod log — action: "expired-breakglass-revoked", the exception name, incident ID, owner, expiry, and deletion time. It never prints the ServiceAccount token or the exception’s full spec.

Audit sink: the CronJob’s pod logs, retained the same way as every other workload’s logs in this cluster (the platform’s log-retention policy, not a Signal Forge–specific pipeline — this reaper’s own audit trail deliberately does not depend on the observability stack it’s revoking exceptions from being healthy).

Manual revoke (if the reaper hasn’t run yet, or is itself down)

kubectl -n otel-lab get policyexceptions
kubectl -n otel-lab delete policyexception <name>

Confirm the deletion by re-running kubectl -n otel-lab get policyexceptions — a manual revoke is not itself logged by the reaper (it only logs deletions it performed), so note the manual action and the time in the post-incident review instead.

Post-incident review

Due within one business day of the exception’s expires-at. At minimum, record: why the exception was needed, what it actually bypassed, whether it was auto-revoked by the reaper or manually revoked, and whether the underlying condition that required it has been fixed (an exception that recurs for the same workload is a sign the policy or the workload needs to change, not that break-glass needs to become routine).

What this does not cover

  • Live reaper-loop / expiry-under-incident proof. This runbook and the reaper’s RBAC/CronJob are rendered and unit-tested (scripts/ci/tests/test_policy_chart.py), proving the manifests are correct. They have not yet been proven against a live cluster — creating a genuinely expired exception, waiting one reaper interval, and confirming both the NotFound result and the sanitized log line. Until that live proof exists and is retained as evidence, OAP-013 is not complete even though the chart renders — see OBSERVABILITY_AS_POLICY_10_10_TASKS.md.