Describe impact before infrastructure

Users need to know which actions are impaired, whether data is delayed, and what they should avoid doing. Internal service names rarely answer those questions. Define public components around user-visible capabilities such as authentication, market data, order submission, account state, deposits, or withdrawals.

State whether the issue affects all users, a region, one market, one interface, or a subset of requests. Use absolute UTC timestamps and distinguish detected time from incident start when they differ.

Set an update contract

Publish the next update time even when the root cause is unknown. A fixed cadence reduces repeated support requests and prevents long silent periods. Each update should identify what changed, what remains affected, and what evidence is being monitored.

Avoid declaring recovery from one successful request. Confirm an appropriate observation window, queue drainage, data freshness, and dependent-service behavior.

  • Initial acknowledgement
  • Scope and user impact
  • Mitigation in progress
  • Recovery verification
  • Resolved notice
  • Post-incident reference

Keep the status channel independent

Host the status interface outside the primary application failure boundary where practical. Protect the publishing account with strong authentication and maintain a second authorized publisher. A status page that depends on the failed login system cannot perform its core job.

Link security and support routes from the status page so users can distinguish an availability incident from an account-specific or security concern.

Referenced resources

Verification checkpoint

Run a tabletop incident in which the primary application and login are unavailable, then confirm an authorized operator can still publish and update the public status record.