📚 Browse docs

Site Pairs

The Site Pairs page manages the links between this controller and its peer recovery sites. Each pair is one production site ↔ one recovery site.

Creating and joining

  • Create a pair on one side to generate a pairing code and reveal this site’s fingerprint.
  • Join on the other side with that code and the peer’s address.

The two controllers pin each other’s fingerprints and bring up a mutually authenticated link. See Create a DR Site Pair for the full handshake and internet/tunnel guidance.

What a pair shows

For each pair you can see its role (primary/recovery), the peer’s name and reachability, the guests protected over it, and its current DR state (protected, testing, failed over, etc.).

If the peer advertised its web console URL, Asternodis shows an open peer console link so you can jump to the other site’s UI in a click.

Driving DR from either site

Both sites run the identical web UI: same pages, same buttons. You can drive every pair operation (failover, DR test, recovery points, re-IP, planned migration, protect/unprotect, failback) from whichever controller you can reach, so you’re never locked out of recovery.

Behind the scenes each action is carried out by the one controller that owns it, routed automatically over the pair’s authenticated link: recovery-side actions (failover, DR test, points) run on the recovery; source-side actions (protect, sync) run on the primary. Because every request funnels to that single owner, two operators (or both sites) can’t drive the same guest into a split-brain: a second attempt simply joins the operation already running. Each operation’s log names which controller actually ran it (e.g. “running on asternodis-sfo (recovery)”), so it’s always clear where work happened.

A guest’s detail screen is the same on both sites: the same health, configuration, live metrics, run history, and operations, with each section talking to the correct controller.

Disaster mode: when the primary is gone

The whole point of DR is that the primary can be down. When it is, the recovery site stays fully operable on its own: pages render immediately instead of hanging on the unreachable peer, and failover, site-wide failover, DR tests, and recovery points all work locally. Source-owned actions (protect, sync) show a clear “source unreachable” state instead of spinning. Declaring a disaster offers a force failover that promotes the recovery copy without waiting on the dead primary. When the primary comes back, Asternodis reconciles cleanly, without double-fencing the guest.

Removing a pair

Deleting a pair unpairs both sides and stops replication over it. Guests protected only on that pair are no longer replicated. Asternodis also removes the pair’s key material (the TLS certificate that encrypts the disk-transfer connection between the sites, and any ZFS encryption key) while keeping the operation history of the guests that were protected across it.

A peer-link-degraded alert fires when the connection to the paired site is slow or unreachable, and clears itself once the link answers normally again, so a degraded link is visible before it costs a stopped guest a full reseed. A slow or momentarily unresponsive peer is never mistaken for a definite answer: aborting a failover, failing back, and the checks that decide whether a stopped guest needs re-seeding retry a flaky connection first, and refuse safely when the peer still doesn’t answer.

Network check

Replication opens TCP connections between the nodes of the two sites (one port per disk, in both directions) and SSH between nodes for containers and ZFS-send guests. A firewall that drops them shows up as a seed that fails after a long wait, or a failback that fails when it is needed. Each pair’s page carries a Network check card, on both consoles, that finds out now: Run network check (operators and admins) tests both directions and says exactly what to open, including the firewall rule to paste where Asternodis knows the node runs pve-firewall. It changes nothing: it holds a few ports open on each site’s nodes for a minute or two and closes them itself. A daily check runs on its own and a new pair starts one; the card attaches to a check that is already running rather than starting a second. The report gives a verdict, the fix for each closed link, and an Addresses used table showing the address each node is dialled at and where it came from; an admin can Set data address for a node of this site when the other site can’t use the derived one. A path it proves closed raises the critical data_path_blocked alert; a check that could not finish says so and proves nothing either way. The port range itself is set under Settings.

Next: Replication to protect guests and drive recovery.