Ravindra BagaleCourses & study guides Track your progress

Guides

Backup and Disaster Recovery: RPO, RTO and Restore Testing

Backups copy data so you can restore it. Disaster recovery (DR) is the plan to bring services back after ransomware, hardware loss, cloud region trouble or human error. RPO (Recovery Point Objective) is how much data loss you can tolerate; RTO (Recovery Time Objective) is how long recovery may take. Untested backups are hopes — restore drills make them real.

Friends! Saying "we have a backup" and staying quiet — but did you try a restore? Ransomware / laptop die / accidental delete — that is when RPO and RTO make sense at the kitchen table. Today: resilient backups, DR basics, and restore testing — no attack recipes, recovery discipline.

Quick answer

Resilient backup / DR vertical checklist:

  1. Write RPO and RTO in plain numbers (example: RPO 24h, RTO 8h for a small site).
  2. Keep 3 copies on 2 kinds of media with 1 offline or immutable copy (3-2-1 idea).
  3. Protect backup admin with unique passwords + MFA; never browse the web as backup admin.
  4. Version / retain enough history to survive delayed ransomware discovery.
  5. Segment backup network or accounts from everyday laptops.
  6. Test restore of one folder or one VM on a schedule; write the minutes it took.
  7. Document who declares disaster and which runbook to open.

Tiny mental model:

RPO = "how far back is an acceptable restore point?"
RTO = "how fast must we be working again?"
Backup without restore test = unknown
Offline / immutable copy = ransomware shock absorber

What do I need before this guide?

  • Something worth protecting (laptop documents, a lab VM disk, or a small business NAS).
  • A second place to store copies (external disk you can unplug, or reputable cloud backup with versioning).
  • Optional: Ransomware protection, IR first 24 hours.

Educational + own lab only

Practice restores only on your machines, lab VMs or systems you administer with permission. Do not run destructive recovery tests on production without a change window and owner approval. This guide teaches resilience — not how to attack backups.

What are RPO and RTO in kitchen language?

Backups, RPO and RTO Primary systems copy to backups. RPO is how much data you can afford to lose; RTO is how fast you must restore. A restore test proves both. Primary live systems Backups RPO = data loss window you accept offline / versioned Restore RTO = time to recover copy

Primary systems copy to backups. RPO is how much data loss you accept; RTO is how fast you must restore — a restore test proves both.

  1. RPO — If we restore, how many hours (or minutes) of new work may be missing? Nightly backup ⇒ roughly up to ~24h RPO unless you add more frequent copies.
  2. RTO — After we say “recover now”, how long until users work again? Includes finding media, restore time, DNS cutover, smoke tests.
  3. Stricter numbers cost more money and discipline — write what the business truly needs.
  4. Different systems can have different RPO/RTO (payments tighter than a marketing wiki).

How does a resilient backup design look?

Vertical building blocks:

  1. Primary data — laptops, servers, SaaS exports, databases.
  2. Frequent copy — snapshots or sync to another disk/region.
  3. Offline / immutable — weekly disk in a drawer, object-lock style cloud retention, or vendor immutable vault (concept).
  4. Access control — separate backup credentials; MFA; least privilege.
  5. Monitoring — failed backup jobs alert a human.
  6. Restore runbook — steps, owners, estimated timing.
  7. DR site / alternate laptop — where you actually work during recovery.

Real incident: WannaCry (2017) — patch speed and recovery pressure

WannaCry (2017) spread widely by abusing unpatched systems and caused global operational pain. Public lessons mixed fast patching with the ugly truth that organisations without clean, tested restores faced longer outages. Ransomware families since then often steal data too — so backups alone are not a full answer, but without them, negotiation and downtime get worse.

Takeaways (vertical):

  1. What happened (theme) — worm-like ransomware at global scale; unpatched Windows estates suffered.
  2. Care-take — patch critical remote services quickly (pair with malware guide).
  3. Care-take — offline backups survive when online shares get encrypted.
  4. Care-take — measure restore time before the press-release day.
  5. Care-take — document who owns DR decisions at 2 a.m. IST.
  6. Bonus — later pipeline and hospital outages also showed that paid downtime hurts even when data later returns.

Red Team vs Blue Team (awareness only)

Red Team — what attackers try (high-level)

  • Encrypt or delete online backups they can reach with the same admin account.
  • Wait days so short retention windows roll off clean copies.
  • Target backup consoles after stealing identity.

Blue Team — defend, detect, respond

  • Immutable / offline copies attackers cannot touch with one stolen laptop admin.
  • Separate identity for backup systems; alert on backup deletion jobs.
  • Tabletop + actual restore drills.
  • After ransomware: restore from known-clean points; rotate credentials.

How do I implement backups and DR step by step?

Step 1 — Write RPO and RTO on one page

  1. List crown-jewel data (finance files, customer DB, identity config exports).
  2. Assign RPO/RTO numbers the owner agrees with.
  3. Note dependencies (DNS, TLS certs, cloud accounts).

Step 2 — Pick the 3-2-1 shape for your size

  1. Laptop: cloud backup + monthly offline disk clone of critical folders.
  2. Homelab / SMB server: nightly image or file backup + weekly offline.
  3. Cloud: enable provider snapshots and confirm retention; export critical config.

Step 3 — Harden the backup path

  1. Unique vault password + MFA.
  2. Do not map backup shares as writable to every PC.
  3. Keep one copy unreachable from the everyday admin session when you can.

Step 4 — Schedule the restore test

  1. Quarterly (or monthly for tiny sets): restore one folder or one lab VM to a scratch location.
  2. Time it; compare to RTO.
  3. Open a sample file; confirm it is not corrupt.
  4. Write results in the same notebook as incidents.

Step 5 — Wire DR communication

  1. Who declares “we restore now”?
  2. Status page / customer message template (honest, no blame theatre).
  3. Order of restore: identity → DNS/web → database → apps.

Step 6 — Lab idea (authorised only)

  1. On a lab VM, create three text files; back them up.
  2. Delete them on purpose.
  3. Restore; record minutes.
  4. Keep malware samples out of shared networks — this is a restore drill, not a malware zoo.

Ravindra Bagale's Tip

💡 Saying "3-2-1" in interviews is fine — a stronger line is: "Last restore test 12 Aug, 37 minutes, RTO target 2 hours." Numbers beat slogans. Do not sleep on a green backup tick — restore one folder. Got it?

What does a personal vs small-business checklist look like?

Personal laptop (vertical)

  1. Cloud backup of Documents with versioning on.
  2. Monthly offline copy of the same folder tree.
  3. MFA on the cloud account.
  4. Restore test: one school/project folder per month.
  5. Full-disk encryption + strong login (pairs with crypto guide).

Small web shop / NAS (vertical)

  1. Nightly file or image backup to a second disk.
  2. Weekly offline disk swapped and labelled with the date (IST).
  3. Database dump before app upgrades.
  4. Written RTO for “site down” vs “laptop lost”.
  5. After ransomware scare drills: who pulls the offline disk first?

Care-take — make recovery boring (in a good way)

  1. Name an owner for backup failures.
  2. Alert on missed jobs within hours, not weeks.
  3. Keep restore media and cables labelled.
  4. After staff exit: revoke backup console access same day.
  5. Pair with MFA and patching — backups are one layer.
  6. Review RPO/RTO yearly when the business grows.

How do I fix common backup / DR mistakes?

Ghabru naka 😅 — these are the usual ones:

Symptom Likely cause Fix
Backup green, restore fails Never tested; wrong credentials Scheduled restore drill to scratch disk
Ransomware hit backups too Same admin reachable shares Offline/immutable copy; separate vault identity
Restore too slow vs promise RTO fantasy; slow media Re-measure; stage warmer copies; rewrite RTO
Cloud only, account locked Single provider identity failure Second copy / export; break-glass owner
“Full disk” backup skipped weeks No monitoring Ticket on failed job; capacity alerts
Chaos during outage No owner / no order One-page DR runbook with named roles

Try it at home

On your PC only:

  1. Write your personal RPO/RTO for Documents (example: RPO 1 day, RTO 4 hours).
  2. Confirm you have at least two copies of something important.
  3. Restore one small folder to a temp path; note the time.
  4. Unplug or disconnect one offline copy after the test succeeds.

Got it? Backup = copy; DR = bring service back. RPO = data loss window; RTO = downtime window. 3-2-1 + MFA on vault + restore test. WannaCry/patch + offline copy lessons. Next: crypto/TLS guide — what is behind the padlock.

Frequently asked questions

What is RPO?

Recovery Point Objective: how much recent data you can afford to lose if you restore.

What is RTO?

Recovery Time Objective: how long recovery may take before users must work again.

Why test restores?

Green backup jobs can still fail at restore time — drills prove timing and integrity.

What is the 3-2-1 idea?

Keep three copies on two kinds of media with one offline or immutable copy.

How does WannaCry relate?

Unpatched systems and weak recovery planning made outages worse — patch fast and keep offline copies.

Where are related lessons on this site?

RDS backup/restore lessons, IR lifecycle and malware defence chapters.