What's included
An untested DR plan is a hypothesis. The schedule below is shaped by the escalating test sequence, because each test costs a change window and each one finds things the previous one could not:
- Business impact analysis, Application and service inventory, dependency mapping, the impact workshops, and the RTO and RPO targets each service is held to. Everything downstream is priced off these numbers. Milestone: RTO/RPO baseline approved.
- DR strategy & architecture, Recovery tier assignment, DR site or region selection, replication and backup architecture, network and DNS failover design, and the cost review against the targets. Milestone: architecture approved.
- Build & replication, DR infrastructure, storage and database replication, backup policy changes, identity and security controls at the DR site, and replication lag monitoring. Milestone: replication steady state verified.
- Runbooks & documentation, A failover runbook per recovery tier, rollback and failback procedures, the crisis communication tree, and the declaration criteria that say who is allowed to call it. Milestone: runbooks published.
- Test sequence, Tabletop first, then a partial failover of tier 1 applications with a rollback, then a full failover with business validation and a failback, with remediation time budgeted after each. Milestone: full failover test passed.
- Approval & maintenance, Test report and residual risk, executive approval, responder training, the annual test calendar, and the change control hook that keeps new applications from silently landing outside the plan. Milestone: plan approved.
Who uses this template
IT resilience leads, infrastructure teams, and business continuity managers use this timeline to build and prove recovery capability. It moves from business impact analysis and RTO/RPO targets through replication build and runbooks to tabletop, partial, and full failover tests, ensuring the organization can actually recover within its targets rather than only on paper.
How to customize it
- Set the RTO and RPO per service, not per organisation, a payments service and an internal wiki should not share a tier.
- Book both change windows early; the full failover window usually needs executive sign-off and a quiet trading period, which are calendar constraints, not technical ones.
- Keep the rollback row next to every test, a test with no rehearsed way back is an outage waiting for a bad day.
- Add rows per application tier if you are failing over in groups rather than all at once.
- Extend the remediation window after the partial test; that is where most of the real findings appear.
- Add the annual retest as dated rows so the plan does not quietly expire twelve months after approval.
Scheduling tips
- Test the failback, not just the failover. Running at the DR site is half the exercise; most organisations discover the expensive problems on the way home.
- Do the tabletop before anything technical. It is cheap, it needs no change window, and it reliably finds missing contact details, unclear declaration authority and runbook steps that assume knowledge nobody wrote down.
- Validate with the business, not with a ping. A service that responds is not a service that works; have real users complete real transactions during the full failover test.
- Watch replication lag as a live metric. An RPO you are not measuring continuously is an RPO you will only verify during an incident.
- Hook DR into change control. Every new application added without a recovery tier widens the gap between the plan and reality, and the gap is only visible at test time.
Related templates
- Data Center Build Schedule Template
- Cloud Migration Project Plan Template
- Internal Audit Plan Template
- Browse all Gantt chart templates
Frequently asked questions
How long does it take to build and test a DR plan?
Commonly nine to fifteen months from business impact analysis to an approved, tested plan. The template uses about fifteen months. The build is predictable; the test sequence at the end is what stretches, because each test needs a change window and a remediation cycle after it.
What is the difference between RTO and RPO?
RTO is how long you can be down, the time to restore service. RPO is how much data you can afford to lose, the age of the last usable copy. RTO drives standby infrastructure, RPO drives replication frequency, and together they determine most of the cost of the plan.
Why three tests rather than one?
Because they find different things. A tabletop finds gaps in the runbook and the decision chain for the price of a meeting room. A partial failover finds technical faults on a limited blast radius. A full failover with business validation is the only thing that proves the RTO. Each needs the previous one to have been fixed first.
Do we need a change window for the tests?
For the partial and full failover, yes, they move production traffic and carry real risk. Book them with a rehearsed rollback and a defined abort criterion. The tabletop needs no window, which is exactly why it should be exhausted first.
How does this relate to business continuity?
Disaster recovery is the technology subset: restoring systems and data. Business continuity is wider and covers people, premises and process. This template covers the DR side, though the crisis communication and declaration rows are shared with any continuity plan you run.
Is the disaster recovery template free?
Yes. Free Excel, PowerPoint and CSV downloads, and free online editing with no account.