Prepare before an incident

  • Define service owners and on-call rotations.
  • Document runbooks for common failure modes.
  • Subscribe to the status page notifications.

During an incident

  1. Assess impact and declare an incident internally.
  2. Check the XenonCloud status page for active incidents.
  3. Open a support ticket with relevant IDs and logs.
  4. Implement mitigations such as failing over to another region or scaling out.

After the incident

Once normal operation is restored, perform a post-incident review and update your runbooks and monitoring alerts based on what you learned.