Monitoring & backups
Observability, alerting and backup practice.
Operating a system is part of building it.
Monitoring
- Health checks for every service, used by deployments and load balancers.
- Uptime checks from outside the hosting provider.
- Application metrics — latency, error rate, saturation — with dashboards per service.
- Structured logs with request identifiers, retained according to the system's requirements.
Alerting
Alerts are reserved for conditions that need a person to act, and each alert links to a runbook entry describing what to do. Noise is treated as a defect.
Backups
- Automated, encrypted backups of every stateful store.
- Retention defined per system and documented.
- Restores are tested. A backup strategy is incomplete until a restore has been performed and timed.
High availability
Where downtime is costly we design for redundancy: multiple instances behind health-checked load balancing, managed database failover and stateless application tiers. We describe availability in terms of architecture and tested failover, and do not publish uptime figures we have not measured.
Incidents
Incidents are followed by a short written review: what happened, the impact, how it was detected and what changes prevent it from recurring.