PlaybooksMonitoring and alerts go blind
Infrastructureserious30 minutes to prepare

Monitoring and alerts go blind

Logs, metrics, traces, uptime checks, or alerts stop reporting while production continues to run.

01DetectConfirm the signal
02ContainStop more damage
03RecoverRestore control
04VerifyProve it works

Your preparation

0 of 0 safeguards ready
0%
Incident worksheet

Make the next decision with evidence

Restore trustworthy signals, establish manual safety checks, and determine what incidents or customer harm occurred while telemetry was missing.

EvidenceDecisionActionProof

Capture before evidence disappears

  • Record the first missing event by source, last confirmed healthy signal, collector errors, dropped volume, sampling, retention, quota, billing, and configuration changes.
  • Compare independent sources: synthetic checks, cloud metrics, application logs, customer reports, queues, database health, and business transactions.
  • List alerts that could not fire, dashboards that show false normality, and deployments or incidents during the blind window.

Decisions that change the response

QuestionAct whenAction
Pause changes?Critical service health or security cannot be observed through an independent signal.Freeze risky deploys and use scheduled manual checks until minimum telemetry returns.
Backfill data?Source records remain and importing them will not duplicate alerts or exceed retention and cost.Backfill into a separate range, label the blind window, and preserve original timestamps.

Proof that recovery worked

  • A generated test error, latency event, security event, and business transaction reach the correct alerts and dashboards.
  • Independent heartbeat alerts detect failure of the monitoring pipeline itself.
  • The blind interval has an explicit impact review and cannot appear as a healthy period in reports.

Controls to put in place

  • Monitor telemetry freshness, volume, drop rate, quota, and billing from outside the main monitoring stack.
  • Keep at least one independent customer-path synthetic check and status channel.
  • Test alerts end to end and assign owners to every critical signal.
Tabletop drill

Block a staging telemetry collector. Detect missing freshness externally, freeze a test deploy, run manual checks, restore ingestion, and reconcile the blind interval.

Escalate when

Use vendor or platform support when data loss, quota, account access, or regional ingestion is involved; involve security if blindness coincides with suspicious changes.

What this means

No alert does not mean no incident. Failures may continue unnoticed, and the evidence needed to understand them may be missing.

Warning signs

  • Dashboards become flat, empty, or perfectly quiet.
  • Heartbeats, log volume, or synthetic checks stop together.
  • Alert tests fail or notifications do not reach responders.
  • The monitoring vendor or ingestion credential reports errors.

Recover now

First 15 minutes

  1. Check production directly through customer-critical journeys.
  2. Use independent provider, infrastructure, and application signals.
  3. Preserve local logs and extend retention before buffers overwrite them.
  4. Establish a manual watch and update interval until visibility returns.

Today

  1. Restore collection, credentials, agents, quotas, or vendor access.
  2. Determine what happened during the blind window from independent logs.
  3. Replay buffered telemetry carefully without overwhelming ingestion.
  4. Add a dead-man alert that detects missing monitoring itself.

Verify recovery

  • Known test events appear end to end.
  • Alerts reach at least two independent destinations.
  • Customer journeys and background work have current signals.
  • The blind window has an explicit incident review.

Prepare now

Access

  • Two people can access monitoring and notification settings.

Backups and evidence

  • Critical logs also exist outside the primary monitoring vendor.
  • Missing-heartbeat and zero-volume alerts are configured.

Contacts and ownership

  • Every customer-critical journey has an owner and signal.

Practice

  • A test alert and a missing-telemetry alert are exercised regularly.

Common mistakes

  • Trusting a green dashboard whose inputs are dead.
  • Repairing telemetry before checking production.
  • Using the same provider for every monitoring path.

Sources

Last reviewed July 19, 2026Guidance changes. Confirm provider-specific actions in the linked official sources.