Every Magento store has backups. Far fewer have restores - tested, timed, trusted restores. The gap only becomes visible on the worst day of the year. Disaster recovery planning is the discipline of closing that gap before you need it: knowing what is backed up, how fast you can bring it back, and having practiced the bringing-back.
What Actually Needs Backing Up
- The database: the crown jewels - products, orders, customers. Daily minimum; for busy stores, point-in-time recovery via binlog archiving or managed-database PITR
- Media:
pub/media- product images are irreplaceable merchant work. Object storage (S3) with versioning makes this nearly free - Code: your git repository is the backup for code; the deployed artifact is reproducible from it. If production has drifted from git (hotfixes applied live), that is a process bug to fix, not a backup to take
- Configuration:
app/etc/env.php,.htaccessrules, Nginx/Apache vhosts, cron definitions - small files whose absence turns a restore into archaeology
Magento’s built-in bin/magento setup:backup exists but is too coarse for production - it does not give you scheduling, offsite storage or point-in-time. Use infrastructure-level backups: managed database snapshots, filesystem/object-store sync, VM snapshots.
RPO and RTO: The Two Numbers
- RPO (recovery point objective): how much data you can afford to lose. Daily backup = up to 24 hours of orders gone. For most stores, that argues for hourly or continuous database replication
- RTO (recovery time objective): how long you can be down. Restoring a 60GB database takes real time - measure it. “We restore in 4 hours” is a claim; a timed drill makes it a fact
Write both numbers down and get merchant sign-off. They drive architecture: RPO in minutes means replication, RTO in minutes means standby infrastructure.
The Runbook
A disaster at 2am is not the time to improvise. The runbook is a one-page document covering:
- Detection: who gets alerted, by what (uptime monitor, New Relic)
- Decision: who can declare disaster and trigger restore
- Steps: exact commands/credentials locations to restore database, media, config; DNS failover if infrastructure-level
- Verification: the checklist proving the restored store works - place a test order, check cron, check payment
- Communication: who tells customers, and what the maintenance page says
Test the Restore
The universal rule: a backup you have never restored is a hypothesis. Twice a year, restore the full backup to a clean environment and run the verification checklist. Time it. Fix what broke. The first drill always finds something - a missing config file, a credentials gap, a media sync that silently stopped in March.
Common Gaps We Find in Audits
- Backups on the same server as the store (server loss = backup loss)
- Media never backed up because “it’s big”
- No offsite copy - ransomware and provider failure both punish this
- Nobody knows the restore procedure except one developer who left in 2024
Backups are insurance; the runbook and the drill are what turn insurance into a plan. Spend the day. The alternative is spending the worst night of your career discovering which parts of your backup were real.