Choosing the target against cost
The availability target sits where business need meets budget. As the target tightens, cost rises exponentially rather than linearly, and that curve belongs on the table during the decision.
- Targets are set per service
- The cost curve is shown before the decision
- Accepted downtime is put in writing
- Components that cannot meet the target are listed
Removing single points of failure
Redundancy is bounded by the weakest link. Servers may be paired while a single switch, circuit or power feed breaks the chain. The map is drawn and each single point is either closed or consciously accepted.
- Power, cooling and circuits belong on the map
- Accepted single points are recorded with their reasons
- Software licensing can block redundancy and is checked
- Dependent external services are assessed separately
Proving it with failure tests
Redundancy is only known to work by deliberately failing a component. Planned failure tests run in stages: first a redundant component, then a whole node, then a link.
- Tests run in stages inside a planned window
- Takeover time is measured and recorded
- Findings from tests change the design
- Automatic failover is tested against false positives
Surviving maintenance and change
A large share of outages comes from change rather than failure. High availability includes performing planned maintenance without interruption; otherwise redundancy exists only on paper.
- Patching and upgrades can run without interruption
- Changes are trialed on the standby node first
- A rollback path is ready for every change
- Maintenance windows are planned against redundancy headroom