SCNET · Enterprise IT · Ankara, Türkiye

Sanal Çekirdek

Not whether something fails — but how long the service stays down.

High availability is a design decision rather than a product: which component can fail without stopping the service. Every additional nine raises cost and complexity together.

'No outages ever' is neither measurable nor achievable. The measurable target is this: in which failure scenario does the service recover, and how quickly.

Choosing the target against cost

The availability target sits where business need meets budget. As the target tightens, cost rises exponentially rather than linearly, and that curve belongs on the table during the decision.

  • Targets are set per service
  • The cost curve is shown before the decision
  • Accepted downtime is put in writing
  • Components that cannot meet the target are listed

Removing single points of failure

Redundancy is bounded by the weakest link. Servers may be paired while a single switch, circuit or power feed breaks the chain. The map is drawn and each single point is either closed or consciously accepted.

  • Power, cooling and circuits belong on the map
  • Accepted single points are recorded with their reasons
  • Software licensing can block redundancy and is checked
  • Dependent external services are assessed separately

Proving it with failure tests

Redundancy is only known to work by deliberately failing a component. Planned failure tests run in stages: first a redundant component, then a whole node, then a link.

  • Tests run in stages inside a planned window
  • Takeover time is measured and recorded
  • Findings from tests change the design
  • Automatic failover is tested against false positives

Surviving maintenance and change

A large share of outages comes from change rather than failure. High availability includes performing planned maintenance without interruption; otherwise redundancy exists only on paper.

  • Patching and upgrades can run without interruption
  • Changes are trialed on the standby node first
  • A rollback path is ready for every change
  • Maintenance windows are planned against redundancy headroom

How we work

  1. Choose per-service targets alongside cost
  2. Map the single points of failure
  3. Accept what cannot be closed, with reasons
  4. Test failure scenarios in a planned window
  5. Make maintenance possible without interruption

How success is measured

  • Every critical service has a written availability target
  • Accepted single points are documented and known
  • Takeover time is measured and inside target
  • Planned maintenance produces no outage

Frequently asked questions

Is redundancy the same as disaster recovery?

No. Redundancy absorbs a component failure in the same place; disaster recovery absorbs the loss of an entire facility. They answer different scenarios and do not substitute for each other.

How many nines are enough?

That is answered per service. Within one organization a few hours of downtime may be acceptable for one service while another is limited to minutes; a single target is never applied to everything.

Does moving to cloud make us highly available automatically?

No. Cloud offers the means for redundancy but the design remains your decision; an application built in one region depends on one region in cloud too.

Let's discuss how much downtime each service can accept and look at the cost curve together.

Agree the downtime you can accept