Disaster Recovery

Disaster RecoveryFuture Computing Solutions, Inc.

Every disaster recovery conversation eventually comes down to two numbers, and most organizations have never actually agreed on either one before an outage forces the question. Both describe the same event from a different angle — one is about data, the other is about time.

Recovery Point Objective (RPO)

How much data can you afford to lose, measured in time. If your systems back up every six hours and something fails right before the next backup, you lose up to six hours of work — that six-hour window is your RPO. A tighter RPO means more frequent replication, which costs more. A looser one is cheaper but means more gets lost when something goes wrong. There is no universally correct number; there is only the number your organization has consciously chosen, versus the number nobody chose that you discover during an actual incident.

Recovery Time Objective (RTO)

How long systems can stay down before the outage becomes a crisis, measured in time. A four-hour RTO means the plan is built to have things running again within four hours of the failure — not four hours to start trying, four hours to actually be back. RTO is a design target for your infrastructure and your runbook, not a hope.

The diagram below shows how the two relate on an actual timeline.

RTO and RPO on a recovery timeline A timeline with three points: the last backup, the moment an incident occurs, and when systems are fully restored. The span between the last backup and the incident is the Recovery Point Objective — the data that could be lost. The span between the incident and full restoration is the Recovery Time Objective — how long systems stay down. Last backup Incident occurs Fully restored Normal operations RPO Data that could be lost RTO Time before you are back
Normal operations
RPO — the gap since your last good backup
RTO — the time it takes to be fully back

Where the numbers actually come from

RTO and RPO are not technical settings a vendor configures for you. They are business decisions — what an hour of the order system being down actually costs, what a day of lost patient records actually means, what a compliance framework requires as a floor. Our job is to help you set numbers that reflect the real cost of downtime for each system, rather than a single target applied to everything from the accounting server to the guest Wi-Fi. Then we build the replication, the failover and the runbook to actually hit them, and test that they hold.

What the engagement covers

  • An RTO and RPO set per system, not one number applied to everything you run
  • Replication and backup architecture built to hit the targets you agreed to
  • A runbook your team can follow during an incident without waiting on us
  • Scheduled failover testing, so the numbers are proven rather than assumed

What a tested plan is worth

An untested recovery plan is a document. A tested one is the difference between an outage that costs a bad afternoon and one that costs a bad quarter.

A number you chose

Most organizations discover their real RTO and RPO during an outage, by finding out what actually happened. Setting them in advance means the number you planned for is the number you get.

Costed against the risk

A tighter target costs more to build. We help you spend that money on the systems where an hour of downtime is expensive, and not on the ones where it is not.

Proven, not assumed

A failover test on the calendar means the first time your plan runs for real is not during an actual incident. We build the test into the engagement.

Do you know your RTO and RPO today?

Most organizations do not, until an outage forces the question. Tell us what you are running and we can help you find out.