Confirm the symptom and scope
Test from outside your own network and distinguish DNS, TLS, connection, application, database, and dependency failures. Record the first known time and a representative error. Avoid repeatedly restarting components before collecting basic evidence.
Check recent deployments, DNS changes, certificate events, provider notices, resource pressure, and dependency health.
Stabilize with the smallest reversible action
Rollback a clearly correlated deployment, disable a failing optional integration, or route to a known healthy origin. Broad changes made under pressure can create a second incident and erase the signal from the first.
Preserve logs, metrics, and configuration snapshots. If compromise is suspected, isolate affected access and follow an incident-response process rather than cleaning evidence in place.
- Name one incident owner.
- Keep a timestamped action log.
- State what is known, unknown, and being tested.
- Set the next communication time.
Verify recovery and follow through
A green homepage is not full recovery. Test the failed path, writes, background jobs, mail, and monitoring from outside. Watch for queued work and retry storms after dependencies return.
After stabilization, document the causal chain and the detection or recovery improvement. Focus on system conditions rather than blame.
Test the original failing action and its downstream side effects, confirm external monitoring sees recovery, and keep heightened observation through at least one normal workload cycle.