Prepare before an incident
- Define service owners and on-call rotations.
- Document runbooks for common failure modes.
- Subscribe to the status page notifications.
During an incident
- Assess impact and declare an incident internally.
- Check the XenonCloud status page for active incidents.
- Open a support ticket with relevant IDs and logs.
- Implement mitigations such as failing over to another region or scaling out.
After the incident
Once normal operation is restored, perform a post-incident review and update your runbooks and monitoring alerts based on what you learned.