Proxmox disaster recovery: the complete guide
Everything you need to plan and run disaster recovery on Proxmox VE: what Proxmox gives you, where the gaps are, and how to achieve tested, orchestrated recovery across sites.
Disaster recovery vs backup
Backup and DR are not the same thing. Backup protects against data loss over time: long retention, point-in-time restores (this is what Proxmox Backup Server does well). Disaster recovery is about getting workloads running again at another site quickly after a failure: low RPO, fast RTO, and a tested failover process. You need both.
What Proxmox VE gives you natively
- ZFS replication (
pvesr): asynchronous, down to ~1-minute intervals, ZFS-only, and between nodes of the same cluster. - Ceph RBD mirroring: snapshot or journal-based replication between Ceph clusters, a Ceph feature you configure and operate yourself outside Proxmox's own tooling.
- Proxmox Backup Server: efficient incremental backups with deduplication.
What's missing is the orchestration: recovery plans, non-disruptive DR testing, one-click failover and failback, and network re-mapping, plus replication for storage other than ZFS and Ceph.
RPO and RTO, briefly
RPO (Recovery Point Objective) is how much data you can afford to lose, set by your replication interval. RTO (Recovery Time Objective) is how long recovery may take, set by how automated and tested your failover is. Good DR drives both down: frequent replication for RPO, orchestrated recovery plans for RTO.
Non-disruptive DR testing
A DR plan you've never tested is a guess. The gold standard is non-disruptive test failover: boot your replicated VMs in an isolated network, verify they come up and applications work, then tear it down, all without touching production or pausing replication. This is exactly what tools like VMware SRM and Zerto provide, and what Asternodis brings to Proxmox.
Failover and failback
Planned failover gracefully moves workloads to the recovery site (final sync, quiesce, power on). Unplanned failover recovers from the last replicated point when the primary is gone. Failback returns workloads once the primary is healthy, the step most DIY setups get wrong. Asternodis makes it near-instant by reverse-replicating changes back to the primary the whole time your recovered guests run at the recovery site, so failback applies only a small delta instead of re-copying whole disks. (A copy you deliberately leave powered off has nothing tracking its changes, so its failback reads the whole disk instead: safe, just not instant.) And because every operation runs on the controller with a live, resumable progress log, an interrupted failover or failback recovers itself rather than leaving a guest stuck half-way. A guest captured crash-consistent has its filesystem checked and repaired (and a missing bootloader reinstalled) on a throwaway clone before it boots, so it comes up instead of dropping to a rescue shell; and if a recovery copy ever diverges, one click rebuilds it.
Getting SRM/Zerto-class DR on Proxmox
Asternodis is purpose-built for this. It installs on a Linux VM, connects to your Proxmox clusters, and adds near-continuous replication (on any storage), point-in-time recovery with points that are continuously integrity-scrubbed against silent bit-rot, ransomware/corruption detection, guest-liveness monitoring that catches a frozen VM hiding behind a green RPO, boot-time guest filesystem self-heal, non-disruptive DR testing, recovery plans with go/no-go preflight, and one-click failover with near-instant failback, plus role-based access, two-person failover approval, and a tamper-evident audit trail. Managed from one web UI at every site, with no single point of failure. See how it compares to VMware SRM and Zerto.
Frequently asked questions
Does Proxmox have built-in disaster recovery?
Proxmox VE offers built-in ZFS replication (pvesr, between nodes of one cluster), Ceph's own RBD mirroring if you set it up yourself, and backups via Proxmox Backup Server, but no DR orchestration: no recovery plans, non-disruptive testing, or one-click failover. Asternodis adds that layer.
What is a good RPO for Proxmox VMs?
Asternodis lets you set a target per guest from about 1 minute up, on any storage (ZFS, Ceph, LVM-thin or a directory) and replicates continuously toward it. The right RPO depends on how much data you can afford to lose per workload, and what you achieve depends on the guest's write rate and your link between sites, not on which storage it lives on.
Can I do disaster recovery on LVM-thin storage in Proxmox?
Not natively: Proxmox's built-in replication is ZFS-only, and Ceph has its own mirroring. Asternodis's storage-agnostic engine, built on QEMU dirty bitmaps, replicates VMs on LVM-thin and other backends too.
Can Asternodis protect LXC containers, or only VMs?
Both. LXC containers are protected at full parity with VMs: replication, non-disruptive DR testing, failover with re-IP, and failback (including to a different node at the home site). Containers replicate over ZFS send/receive, or an incremental rsync of the root filesystem on LVM-thin and directory storage, and get application-consistent snapshots by briefly freezing the container on the host, with no in-guest agent required. Proxmox itself can replicate containers on ZFS but offers no DR orchestration for them; tools like VMware SRM and Zerto can't protect LXC at all.