Replication & recovery operations
The Replication page is the heart of Asternodis: enroll guests for protection, watch their recovery objective, and run every recovery operation.
Protect a guest
Under Protect VMs, pick guests from a pair’s source side and enroll them. For each you can set a target RPO (recovery point objective): how far behind the recovery copy is allowed to fall, from ~1 minute up. Asternodis replicates continuously toward that target on any storage backend.
Replication is storage-agnostic: whether your guests live on ZFS, Ceph, LVM (thin or thick), or a directory, the same protection model applies. Open a guest’s detail page for its per-VM replication history, health, and settings. Select several guests at once to sync or unprotect them in bulk rather than one at a time.
On a license with a protected-node allowance, protecting a guest on a node that has no protected guests yet is refused once the allowance is full; see Node allowance.
Where the replica lives
Each protected guest’s recovery copy (its shadow) can be placed on any
storage at the recovery site: a directory, NFS or CIFS share, or a native ZFS,
LVM-thin, or Ceph/RBD volume. Set a default for the whole pair, or override it
per guest, so replicas land where you have capacity instead of on the node’s root
filesystem. The picker checks what a candidate storage actually sits on rather than
just its name, so a thin pool that shares the node’s boot disk (the case on a default
Proxmox install, where local-lvm and local are the same volume group) is flagged
too, not only local itself. On block pools (ZFS, LVM-thin, Ceph/RBD), recovery points
are captured as native snapshots. Ceph/RBD also works as a source storage: a
guest whose disks live on Ceph is protected the same way as one on any other backend.
Thick LVM and iSCSI storages can’t hold a replica (the storage picker marks them cannot
hold a VM’s replica), though a guest whose disks live on them can still be protected.
See every replica across your sites, and reclaim orphaned ones, on the Shadows page.
Replication mode. By default Asternodis uses its universal dirty-bitmap engine on every backend. For ZFS-to-ZFS pairs, an optional native ZFS send/receive mode is available as an experimental feature (enable it in Settings).
One caveat: CIFS/SMB source disks need cache=writeback
This is about the disk your protected guest runs on, not where the replica lands.
Proxmox opens a CIFS/SMB-backed disk with its cache mode set to none (direct I/O) by default. Copying a running guest’s disk on that path is not coherent with what the guest is writing: the copy finishes and every layer reports success, but the replica’s filesystem journal comes back invalid and the recovered guest will not boot. Nothing downstream can detect it, because nothing downstream was told anything went wrong.
So Asternodis refuses to replicate that combination rather than hand you a replica that
looks fine until the day you need it. The refusal names the guest and the exact disks, and
the guest’s detail page offers a one-click Fix disk cache button that sets
cache=writeback and leaves every other disk option alone.
Two things worth knowing:
- The guest is not rebooted for you. Proxmox stores a disk-option change on a running guest as pending, so it takes effect at the guest’s next stop/start. Choosing when to bounce a production VM is your call. Replication stays refused until then.
- Stopping the guest is also a fix. A stopped guest is seeded offline, which is safe on any cache mode.
This affects VMs only. Containers on CIFS replicate by file sync rather than a block copy, so they are unaffected. Other file storage (a directory or NFS) is unaffected too: Proxmox does not default those to direct I/O.
Watch the initial sync
A guest’s first sync is a full seed and can take a while over the wire. The Replication page shows it live as a progress bar with percent, transfer rate, and ETA, sized to the guest’s actual allocated data, not its (often much larger) provisioned size. Every sync after that ships only changed blocks and is near-instant.
The guest keeps running normally throughout: its console and Proxmox tooling stay reachable even during that first full seed, so replication never locks you out of a VM.
Containers (LXC)
Asternodis protects LXC containers at full parity with VMs: replication, non-disruptive DR testing, failover with re-IP, failback, and retained recovery points. A few details are container-specific, and Asternodis handles them for you:
- How they replicate. A container has no QEMU block layer, so it replicates over ZFS send/receive when its root filesystem is on ZFS, or an incremental rsync of that filesystem on every other backing: LVM-thin, Ceph/RBD, or a directory, NFS or CIFS share. Either way the recovery copy lands on ZFS at the recovery site, so testing, failover, and failback behave exactly as they do for a VM. This is automatic: there’s nothing to choose, and the CIFS cache caveat above does not apply to containers.
- Application-consistent without an agent. Containers don’t run the QEMU guest agent. Instead Asternodis briefly pauses the container on its Proxmox node with the kernel’s cgroup freezer (the mechanism Proxmox’s own snapshots use), flushes its filesystems, takes the snapshot of every volume inside the pause, and releases it, reading the release back before it calls the point application-consistent. No in-guest software is involved. It falls back to a crash-consistent snapshot, with the reason on the point, when the container can’t be paused: it is stopped, a backup or another Proxmox operation holds it, it is already paused, or the host runs cgroup v1.
- Re-IP on recovery. A recovered container is re-addressed by rewriting its network configuration directly, so re-IP and network mapping work without cloud-init or a guest agent.
- Failback, including to another node. With no dirty-bitmap reverse lane, a container returns home by shipping the ZFS delta accumulated at the recovery site back to the source (a full reseed only if no common base exists). As with a VM, a container on ZFS can also be failed back to a different healthy node at the home site, where the failback screen offers it.
The boot-time fsck and bootloader self-heal described below is VM-specific; a container’s
recovery point is application-consistent by default, and crash-consistent only where the
container could not be paused.
Reset config baseline and Re-seed replica
Two actions on a protected guest’s page look alike and do different things:
- Reset config baseline records the guest’s current Proxmox configuration as the reference that drift detection compares against. It copies no data. Use it after you changed the guest on purpose (more memory, a new disk) and the drift badge should stop pointing at it.
- Re-seed replica releases the guest’s replica on the recovery site and copies the whole disk again from scratch, right away. Every retained recovery point of that guest is destroyed with the old replica. Use it when the replica is known to be diverged or corrupt. The guest keeps running; the copy shows in the seed indicator like a first seed.
Changing a guest’s Shadow storage does the same release and re-seed, onto the new storage, and starts it immediately from whichever site you changed it on.
Recovery points & point-in-time recovery
Asternodis keeps multiple retained recovery points, not just the latest, so you can roll back to a known-good moment. Points are integrity-checked at capture and application-consistent by default: the guest is quiesced around each capture (a VM through its QEMU guest agent, a container by the host’s cgroup freezer), and any guest that can’t be falls back to crash-consistent on its own. It is a site-wide switch under Settings with a per-guest opt-out.
Older points are pruned automatically on a grandfather-father-son (GFS) schedule (keep the last N hourly, daily, and weekly points), so history stays bounded without manual housekeeping. Manually captured points are always kept. Tune the schedule, or turn auto-deletion off entirely, in Settings; the policy is instance-wide and synced across the pair.
Recovering to a point
There are two ways to put a guest back onto an earlier point, and they differ in where the guest ends up running:
- Restore in place rolls the guest back on the primary, where it already runs. No failover, no site change, and no second operation afterwards to undo one.
- Recovering to a point is a failover with that point chosen: the guest comes up at the recovery site from that moment rather than from the latest replicated state, and comes home later via failback.
Open the guest’s Points list. Each row carries the three actions the decision needs:
- DR test boots that point in isolation, on a throwaway clone, with the guest and its replication untouched. Use it to confirm the point is the one you want (that the data predates the incident, and that it boots) before you commit to it.
- Restore in place rewrites the guest’s own disks at the primary from that point. See Restore in place below.
- Recover to this starts the real recovery. It opens the failover page with that point already selected, so every guard the failover path has still applies: the capacity preflight, the split-brain check, force plus a recorded reason where one is needed, and a live step log. Nothing is promoted until you confirm there.
You can also pick a point from the Recovery point dropdown on the failover form itself: the same list, chosen at the moment of failover instead of from the points table.
Points that failed their integrity check are quarantined and cannot be chosen by either route; the refusal names the newest verified point instead. A point whose underlying snapshot has since been pruned or deleted is refused too, rather than quietly falling back to the latest state.
Two consequences worth deciding on before you start, because both are about the source site’s disk:
- Data written after the point is not recovered. That is the whole purpose, but it means the gap between the point and the incident is gone.
- Once the guest is running from that point at the recovery site, Failback flushes that state home, overwriting the home disk. For ransomware or corruption recovery that is the intended ending. If you decide against it, Abort failover returns the guest to its untouched primary disk instead.
From the CLI, the same two operations take the point’s snapshot ID (read it from the Points list):
asternodis dr test --pair <PAIR_UUID> --vmid <N> --point <SNAP_ID>
asternodis dr failover --pair <PAIR_UUID> --vmid <N> --point <SNAP_ID> --reason "pre-incident restore"
Restore in place
Restore in place rolls a guest back to an earlier recovery point on the primary, the site it already runs at. Nothing fails over, nothing moves, and the guest’s identity, node, addresses and storage are all unchanged. It is one operation with one short outage (a stop and a boot), rather than a failover followed later by a failback.
Use it when the guest itself is the problem and the site is fine: ransomware or mass corruption, a bad upgrade, a destructive change you want to undo. Use a failover instead when the site or the host is the problem: restoring in place needs the primary to be healthy, because that is where the guest is going to run.
Start it from the guest’s Points list, on the row you want to land on. Asternodis runs a preflight first and refuses with a reason rather than starting something it cannot finish: there must be room for the restore, the guest must be in its normal at-home state (not failed over or mid-operation), and it is VM guests only for now, not containers. Every supported home storage is covered, including directory, NFS and CIFS.
What happens, in order: the guest is stopped, its disks are rewritten from the chosen point, and it is started again on the same node with its original boot order restored. The operation gets its own live page with ordered steps and a step log, the same as a failover, and it is recorded in the audit log. You can stop a running restore up until the point where stopping would leave the disks worse than finishing. Past that the button refuses and says so.
Your recovery points are kept. After the swap the replica re-syncs in place (the
whole-disk copy is written over the existing replica instead of recreating it), so the
point you restored from and the points before it stay available to fail over to. The
confirm page says so before you start: how many points are kept, and how much extra
recovery-storage space the re-sync needs until retention prunes the points that share the
old blocks. If the recovery storage can’t take the re-sync right now it pauses, with a
critical resync_blocked alert, rather than destroy
anything; and a replica that can’t be re-synced in place is flagged on the confirm page
with the loss of its points named, so you can copy out what you need first. The guest’s
other recovery sites, if it has any, are asked to keep their points too, and the
operation’s log reports each answer. A guest replicated by the experimental ZFS
send/receive mode can’t be restored in place, because that would destroy the history it is
restoring from.
Two things to decide before you start, because both concern the guest’s current disk:
- Data written after the point is gone. That is the purpose, but the window between the point and the incident goes with it.
- The rollback happens on the production disks. If you want to check the point first, run a DR test from the same row: it boots that point in isolation and changes nothing.
Deleting points yourself
Retention handles the routine pruning, but you can also clear points by hand from a guest’s Points list: one at a time, or by selecting many and deleting them in one go. It is the quickest way to reclaim space after a burst of manual snapshots, or to clear out points a re-seed has already invalidated.
Two things make a bulk delete safe to click. The request is re-checked against the server: every selected point must still belong to the guest and pair the list was showing, so a page left open while something else changed can never delete another guest’s history. Mismatches are reported back per point instead of being actioned. And the delete runs to completion even if you navigate away, so you are never left with half the snapshots gone and their entries still listed. Points already deleted count as success, and every batch is recorded in the audit log.
Continuous integrity scrubbing
A recovery point that passes at capture can still rot later (a flipped bit on disk, a bad block) and normally you’d only find out when you tried to recover to it. Asternodis guards against that by re-scrubbing points in the background over their whole retention life, and each kind of recovery storage is checked in the way it can be proven:
- ZFS, Ceph/RBD and container points are chain-verified and capture-checked against the storage’s own checksums: a point’s new blocks are read back within minutes of capture, and every point’s full data is re-read on a rolling cycle (target: 24 hours).
- Directory, NFS, CIFS and LVM-thin points have no storage checksum to read against, so they are measured by digest: Asternodis records a per-chunk digest when it first measures the point, then re-reads it and compares against what changed since. That is a weaker claim, worded as one: it detects change since the measurement, not damage that landed before it.
If the contents have drifted, the point is automatically quarantined (flagged in the points list, with an alert) and can no longer be chosen for a failover or DR test. The points list shows what backs each point’s check and when it last ran, and the Overview’s Copies held here panel shows whether each lane is keeping up (see Fleet coverage). The result: you fail over to a point whose data is known-good now, not one that was merely good once.
The scrub reads from its own isolated copy of the point, and its reads are rate-capped so they can’t crowd out failed-over guests on the same storage: replication keeps running at its normal RPO, and a DR test can start while a check is in progress. A digest comparison that can’t finish between two captures is held back rather than started. An admin sets the read rate under Settings → Replication limits → Recovery-point integrity reads.
Repair now, on a point’s row, runs the same repairing filesystem check the recovery path runs before boot (on a disposable clone of the point, never the point or the live replica), so you can prove a standing filesystem warning is fixable, or learn a point genuinely can’t be repaired, without running an actual DR test or failover.
Every row also carries a consistency badge: app-consistent when the guest was quiesced while the point was taken (through the guest agent for a VM, the cgroup freezer for a container), or crash-consistent when it was not (no responsive agent, a container that could not be paused, or app-consistent capture switched off for the site or for that guest), so you can tell at a glance which points a database or mail server would recover from cleanly. The advice shown for a crash-consistent point names the actual reason: an unresponsive agent, a freeze that lapsed, or the site’s Application-consistent points switch or the guest’s own opt-out being off.
Getting told when something is wrong
Asternodis raises alerts for the conditions you would not otherwise notice until you needed the DR copy, among them an RPO breach, a guest that has never completed a first sync (so it has no replica at all), a guest whose agent has gone silent, a wedged QMP monitor, a recovery point that failed its integrity check or its filesystem scrub, a recovered guest that did not boot cleanly, a stalled reverse sync, and split brain. Alerts lists every kind, what it means and what clears it.
They always appear on the Alerts page, which holds the complete list and your silences. The notification bell and this page’s banner show what is firing and not silenced: the banner lists the critical alerts and counts the advisories. Every alert names the guest it is about, says in plain language what the condition means, and links to the page where you act on it, so an alert is the start of a fix rather than a line to go and interpret.
Three actions keep that list meaningful:
- Silence an alert you have judged and do not want to keep seeing. Silences are listed separately so a muted condition stays visible as a decision you made, not as a gap.
- Dismiss one that needs no action.
- Dismiss all, on this page’s alert banner, clears every firing alert it shows at once, which is worth having when one root cause lights up a dozen guests and you have already dealt with it. It leaves anomaly alerts alone, because those need a review.
Alerts that resolve themselves need none of this; the actions are for the ones that want a
human. To have alerts leave the UI, turn on delivery: Alerting in this page’s
header, then set a webhook URL. It is off by default, so nothing is delivered until you
configure it. Send test in that dialog makes one real POST of a sample alert
(status: "test") to the URL as entered. Prove the channel before the first breach
does; the outcome is shown in the dialog and recorded in the audit log. The test is a
single attempt by design, so an endpoint that only answers on a retry shows up as a
failed test; real deliveries retry, as below.
- One webhook, any receiver. The payload is a fixed JSON document (
event,status,kind,vmid,name,detail,timestamp), so ntfy or a script of your own can take it directly; Slack, Teams and PagerDuty each expect their own format, so put a small relay in front of those. There is no built-in email. Point the webhook at whatever already pages you. - Real deliveries retry: up to three attempts, with a short backoff between them. If
all three fail, the failure is written to the audit log as
alert_delivery_failed, which is worth a periodic look because a webhook that started 404ing after an endpoint change otherwise looks exactly like a quiet week. - Alerts resolve themselves when the condition ends, and any that need a human judgement can be dismissed from the banner.
Check it end to end. Point the webhook at your receiver, then breach an RPO on a test guest (pause its replication and wait until its last sync is more than twice its RPO old: that is the breach threshold) and confirm the message arrives. An alerting path nobody has ever seen fire is not an alerting path.
Ransomware & corruption guard
A change-rate anomaly detector watches replication. If a guest suddenly looks mass-encrypted or corrupted, Asternodis pauses replication to preserve a clean pre-attack recovery point instead of overwriting it with the damaged data. Tune the policy with Anomaly detection in this page’s header: the Detect change-rate anomalies on/off checkbox, Mass threshold (% of disk), Trip multiplier, Min baseline runs, Floor (MiB), which is the size a sync must reach before it can trip, and pause vs alert-only. It is configured there rather than in Settings, whose General → Other policies section points to it.
Pausing the source is only half of it, because the recovery site keeps its own schedule. While an anomaly is under review, that site holds its retention work too:
- No new recovery point is captured for the guest. The replica is frozen at whatever the tripping sync left behind, so a point taken now would record the damaged state as though it were a recovery option.
- Nothing older than the detection is pruned. Your last clean point cannot age out of the grandfather-father-son schedule while you are still deciding what to do.
Both resume automatically the moment you acknowledge the anomaly.
Reviewing one from the Replication page shows the signature (the tripping sync’s volume against the guest’s baseline, the ratio, and the detection time) and offers the three ways out. Restore in place is usually the one you want: pick a point from before the detection time and the guest rolls back where it already runs, with no site change. Fail over to a pre-detection point instead if you also need to move off the primary. Or acknowledge and resume if you have decided the burst was legitimate: a large import or a full-disk rewrite will trip the same signature.
The detector weighs how much a guest shipped against both what it stores and its free space: shipping more than the guest holds is decisive proof of a rewrite, while a match against free space is evidence of a routine trim or disk optimize, and the review says plainly when both readings fit and the numbers alone can’t tell them apart. A guest that trips the detector on a known, recurring benign pattern can be exempted individually, with a required reason, instead of only being able to loosen detection for the whole appliance. The review shows the actual evidence used to flag the guest (bytes stored versus shipped, free space, and the disk’s discard/SSD flags), so you don’t have to go find those numbers by hand.
Dismiss and Dismiss all refuse an anomaly alert and point you at Review anomaly: resuming a paused guest’s replication is a decision, never a side effect of clearing a notification. To resume protection, review the anomaly and choose restore in place, fail over, or acknowledge and resume above.
Catch a frozen guest
Replication health is not the same as guest health. If a guest’s OS freezes or crashes, its QEMU process keeps running and its disk stops changing, so replication keeps succeeding on a static image and the guest’s RPO stays green. None of the usual signals go red, and a dead guest can hide behind a healthy-looking dashboard for days.
Asternodis closes that blind spot. It watches each running guest’s QEMU guest agent and, when the agent has been unreachable for more than a few minutes (with a debounce so an ordinary reboot doesn’t trip it), raises an agent-unreachable alert, independent of RPO. The alert clears on its own the moment the agent answers again.
When it fires, a one-click Fix agent button appears on the guest (and in the Tasks bell) that reboots the wedged guest to bring it back into service. It tries a QMP reset first, which keeps replication incremental, and falls back to a full power-cycle only if the monitor can’t be reached, then waits for the agent to respond. Like failover and failback, it stays available even under an expired license.
Recover from backups too
Beyond live replication, Asternodis can recover from Proxmox Backup Server and other backup-capable storages (directory, NFS, CIFS): a fallback path when you need to restore from a backup rather than a replicated point.
Boot-time filesystem self-heal
A guest without the QEMU guest agent is captured crash-consistent (the equivalent of
pulling the power), so its filesystem can need a journal replay or minor repair before it
will boot cleanly. Asternodis can do that repair for you: on failover (the Auto-repair
filesystem & bootloader before boot option, on by default) and automatically for every
DR test, it runs an offline fsck
and, if the bootloader is missing, an offline GRUB reinstall before the guest
starts.
This is safe because it always runs on a throwaway clone of the recovery point, never
the point or the live replica: a bad repair is discarded with the clone and your protected
data is never touched. It turns a crash-consistent capture that would otherwise drop to a
repair or GRUB-rescue shell into a guest that just comes up. (Encrypted volumes, ReFS and ZFS
members are skipped with the reason shown, never failed. Beyond ext4 and XFS, a recovery also
checks FAT, NTFS, btrfs and the filesystems inside a guest’s own LVM volumes, and repairs what
has a safe offline repair; a Windows NTFS volume is replayed and marked so Windows runs its
own chkdsk at first boot, and btrfs is checked but never repaired offline. The extended
repairs are on by default, with a per-site switch under
Settings.) The health scrub also flags a guest
whose filesystem looks dirty ahead of time, so you know before you need to recover. Pre-boot
self-heal runs for a guest whose recovery replica is on directory, NFS or CIFS storage as
well as on a native volume.
Automatic guest-agent install
A missing or unresponsive guest agent is the most common reason a point comes back crash-consistent instead of application-consistent (the others are the site’s Application-consistent points switch and the guest’s own opt-out, above). Asternodis closes that gap itself: a recovered Linux or Windows copy with no working agent gets one installed before boot, on the same throwaway clone the filesystem self-heal uses (on by default, for every DR test and failover, with nothing for you to do). Failback carries the installed agent home with the rest of the disk. There’s an off-switch in Settings for anyone who wants recovered copies left exactly as captured.
Separately, Install guest agent on a guest’s own page installs the agent into the live guest at its home site, not a recovered copy: either a brief restart (the guest’s disks are edited offline the same way, then a full re-copy runs since the restart resets change tracking) or a live install over SSH with no downtime.
Recovery operations
From Recovery operations you drive, per guest or pair:
- Failover: bring guests up at the recovery site. Before it starts, a go/no-go
preflight checks, per guest, that the recovery node can host it (memory, vCPU,
storage, machine type) and that every NIC lands on a network bridge that exists on the
node that will run it. A target bridge that doesn’t exist there, Linux or Open vSwitch, is
a definite failure, not an unverified maybe, and a check that couldn’t get an answer is
reported as unverified, never as a pass. A preflight that fails, or couldn’t run, never
blocks a recovery, but you acknowledge it first (Fail over anyway?) on every route:
one guest, a recovery plan, a whole site, a planned move; warnings alone need no
acknowledgement.
asternodis dr failoverprints the same verdict. If a VM’s EFI firmware or TPM state can’t be restored, the step log names which step failed and why. The same is true of a DR test and of failback to a chosen node. - DR test: boot replicated guests in an isolated network to prove recovery without touching production or pausing replication. Clean up test VMs when done.
- Recovery points / health: inspect points and run readiness checks.
- Planned migration (Move): a graceful, near-zero-downtime cutover to the other site.
- Failback: after a failover, return the guest to its home site. While it runs at the recovery site, Asternodis reverse-replicates changes back to the primary continuously, so failback applies only a small delta and is near-instant. Forward protection then resumes on its own once the guest is home. That speed depends on the copy running: a live guest is what lets Asternodis track its changed blocks. Fail over with Power on unchecked, or stop the DR copy afterwards, and there is nothing tracking changes: the copy is safe and nothing is lost, but proving the two disks match means reading and comparing the whole disk, so that failback takes as long as a whole-disk read rather than seconds. If the original home node is overloaded, full, or gone, you can fail a guest back to a different healthy node in the same source cluster instead, where the failback screen offers it. Which route a guest gets depends on where its disks are, and Asternodis reads that from the source before anything is touched, so the node picker says up front what is open and why not. A guest whose disks all sit on storage only its home node can read (a non-shared ZFS pool, node-local LVM-thin, or a local directory) gets fresh disks on the node you pick. That is not the small-delta route described above: for a VM it copies the whole disk and destroys that guest’s retained recovery points when the recovery shadow is rebuilt against the new disks, however caught up the reverse sync was. The failback screen says so before you commit; export or recover anything you need first. A VM whose disks all sit on storage both nodes can see (shared storage such as Ceph/RBD or NFS) is moved instead: nothing is copied, the reverse sync brings the disk up to date and the guest’s configuration moves to the node you pick, so it costs no more than an in-place failback and keeps its points. Containers can use the route only on ZFS, and keep their points: their DR dataset is promoted into the shadow, so the history survives and forward protection resumes incrementally. Some guests can’t use it at all: one managed by Proxmox HA, a cluster that has lost quorum or has a single node, disks spread across several storages (or across shared and local storage), thick LVM or iSCSI. For those the screen says so: fail back to the original home node, or use Abort failover. For a VM failing back to its original node, the points survive only when forward protection can resume incrementally, which needs a reverse baseline that was proven fast at failover (the planned-failover quiesce described under Settings). The recovery copy’s form does not matter: a copy stored as a qcow2 file (directory, NFS or CephFS recovery storage) gets the same differential resync as a native volume (ZFS, LVM-thin or Ceph/RBD): the failback ships only the failover delta into the file, with the file’s receiver proven stopped first, and verifies it. A failover whose baseline was not proven (an unplanned, down-primary failover, a point-in-time failover, or a check that was declined), or a fast baseline that was withdrawn afterwards (a failback verify that found the primary’s disk diverged and armed a full rebuild, a forced reverse re-seed, a rollback-home), fails back normally but the first forward cycle afterwards copies the whole disk, which rebuilds the shadow and destroys that guest’s retained recovery points. The failback screen shows a red banner naming which of these applies (and, for a withdrawn baseline, what withdrew it), and the operation log says so again at failback time (“Forward replication will re-seed this guest’s whole disk after the failback (…)”), so there is never a silent loss. A failback that is interrupted after the guest has booted at home (a controller restart, a stalled database) is completed by clicking Fail back again: Asternodis asks the primary whether the guest is already running there and, if so, only finishes the bookkeeping: nothing is copied. If the interruption came earlier, while changes were still being merged into the home disk, the retry says so and asks you to stop the guest first: that merge cannot be applied to a running guest. A guest that was forced live from a replica that had not finished seeding carries that fact until it comes home, and its in-place failback is refused until you acknowledge it separately on the failback screen: that failback would copy a possibly partial image over the primary’s home disk, which is still the last intact copy of the guest. Abort failover returns it to that untouched disk; where the failback screen offers a fresh-disk copy to another node, that leaves it alone too.
- Which guests fail back in place: a VM comes home by shipping only what changed at the
recovery site when its home disks are on ZFS, LVM-thin, Ceph/RBD, a
directory, NFS or CIFS file storage, or thick LVM in a volume group that belongs to
one node, including LVM on an iSCSI LUN. The last four sit behind In-place failback for
non-ZFS home disks in Settings, which is on by
default. Two shapes are refused, and the failover confirm page says so before you fail
over rather than after: a thick volume group shared across the cluster (
shared 1: a SAN or FC/iSCSI LUN every node sees), because a snapshot there rewrites volume-group metadata every node shares, and Ceph/RBD configured withkrbdor an RBD namespace. Those guests come home via Abort failover. - Abort failover: undo a failover that shouldn’t have happened. As long as no reverse changes have been written to the primary’s disk, Asternodis stops the recovery copy and brings the original guest back up on its untouched disk: no flush, no reseed. Once the DR copy has changes worth keeping, use Failback instead. If the guest turns out to be already running at home (an interrupted failback that returned it without recording it), the abort says so and only records it home; it never claims the disk was untouched.
- Rebuild DR copy: the recovery path when a reverse lane stalls or diverges. One
click (or
asternodis dr reseed) re-seeds the DR copy from scratch so the next failback has a clean baseline. It replaces the old have-to-edit-the-database workaround. It appears in the reverse-sync panel of a guest that is failed over and has a reverse lane, so a VM on any home storage the reverse lane supports, not a container: containers have no continuous reverse lane (their failback ships a dataset differential on demand), and their pages say so instead.
Higher-risk operations can require dual-control (two-person approval) if you’ve enabled it in Settings.
Every operation runs on the server and gets its own live progress page (ordered steps, bytes transferred, timing, and any error), reachable from the Tasks bell or directly by URL. You can leave the page or close the browser; the operation keeps going, and it recovers on its own if the controller restarts mid-run.
You can start any of these from either site: Asternodis routes each to the controller that owns it, runs it exactly once, and the progress log names which controller ran it. Even with the primary completely down, the recovery site drives failover on its own. See Driving DR from either site.
Next: group guests into runbooks with Recovery Plans.