📚 Browse docs

Alerts

The Alerts page is the complete list of what Asternodis is warning you about, across both sites of every pair, together with the conditions you have chosen to silence. The notification bell and the banner on the Replication page show a subset; this page shows everything, and it is the one place a silence can be seen or lifted.

Any signed-in user can view it. Dismissing, silencing and un-silencing need at least the Operator role; configuring delivery needs Admin. The page refreshes every 15 seconds.

Firing

Each row is one alert, and carries:

  • What it concerns: the guest’s VMID and name, or, for an alert about the controller itself or the link to the paired site, what it is about.
  • Its kind, such as rpo_breach. Every kind is described below.
  • The site that raised it, when that is the other site of the pair.
  • When it fired, and the alert’s own explanation of the condition.
  • A link to where you act on it: Open this guest, View recovery points, Open Shadows, Open Site Pairs, Open Networking, Open the network check, Open the audit log, Open Retention settings, Open Settings or Review the anomaly, depending on the kind.

Alerts raised at the other site are listed first. If the other site’s alerts cannot be read, the page says so above the list, which then covers this site only, rather than implying that nothing is firing there.

Critical and advisory

Every kind is either critical (a recovery would not work, protection has stopped, or the controller itself is at risk) or advisory (protection is working, and something is degraded, historical or waiting on housekeeping). The Replication page’s banner lists only critical alerts and shows advisory ones as a count; this page, the bell and the webhook carry both.

Dismiss or silence

Dismiss means “I have seen this”. The alert is marked resolved for everyone, on whichever site raised it; it stays in the alert history with its resolved time, and the dismissal is recorded in the audit log. Dismissing does not make a live problem go away: an RPO breach, a silent guest agent, a degraded peer link or low controller disk, for example, is raised again by its next check if it is still true.

Silence means “I know, and I do not want to be told again”, with an optional reason for the next person. A silence applies to the condition (the guest, the pair and the kind of alert together) rather than to one alert row, so it holds while the alert clears and fires again underneath it. Asternodis keeps checking the condition; a silenced alert is left out of the notification bell, the Replication banner and the badge count, and is listed under Silenced on this page instead. A silence governs what the UI shows, not what is delivered: silencing resolves the alert that is firing, and alerts raised afterwards for the silenced condition still reach your webhook.

A change-rate anomaly alert can be neither dismissed nor silenced. Under the default anomaly policy the guest’s replication is paused, and clearing the notification would hide that: its row shows cannot be silenced: use Review anomaly where the Silence button would be, and Dismiss is refused with an explanation. Follow Review the anomaly to the guest’s page and use Review anomaly there: acknowledging it records your review, resumes replication and resolves the alert. See Ransomware & corruption guard.

The audit-log alerts (audit_chain_broken, peer_audit_diverged, peer_audit_anchoring_lost) are the same shape: they can’t be dismissed or silenced, by you or by the paired site, because that would hide a finding that has not been resolved. Their rows say cannot be silenced: see the Audit page, which is where a finding is acknowledged or a paired site’s log accepted.

Clearing many at once

When one root cause lights up a dozen guests, clear them together. Dismiss all, on the Replication page’s alert banner (shown while a critical alert is firing), resolves every firing alert on both sites, including the advisory ones the banner does not list. The notification bell’s Clear all does the same for the alerts it shows, and also hides the operations it lists (in that browser only). Neither touches a silenced condition, and neither clears a change-rate anomaly: those stay firing, and the Dismiss all confirmation says how many it is leaving.

Silenced

Each silence shows the guest’s VMID (or this site, for a condition that is not about a guest), the kind, the site it lives on when that is the other one, whether the condition is firing right now or not currently firing, when and by whom it was silenced, and the reason. Un-silence lifts it; a condition marked firing right now reappears straight away, because silencing never stopped it being checked. If the other site’s silences cannot be read, the page says so rather than showing an empty list as though nothing were silenced.

How alerts clear

Most alerts describe a condition, and Asternodis resolves them itself when the condition ends: the RPO is back within bounds, the agent answers, the link recovers. Two describe something that happened (recovery points destroyed by a re-seed, and a recovered guest that failed to boot), and each occurrence raises its own alert. The tables below say what clears each kind. Alerts for a guest you stop protecting are resolved too, and resolved alerts stay in the history.

Alert kinds

Replication and protection

KindSeverityWhat it means
rpo_breachCriticalThe guest’s last successful sync is more than twice its RPO target old; when its failed syncs point to the source guest being gone, the alert says so. Not raised while the guest is failed over. Clears once it is back within twice its target, or when it fails over.
never_seededCriticalThe guest has never completed a sync, so it has no replica at all, and its latest attempt failed. A first seed that is still running never raises it. Clears on the first successful sync, or when the guest fails over or its protection is disabled.
drift_criticalCriticalA disk was added to the guest, or grown, since its replicated baseline, so its recovery point is incomplete. Clears once a re-seed takes in the new shape, or Reset config baseline accepts it.
monitor_wedgedCriticalQEMU has stopped servicing the guest’s QMP monitor, so replication cannot run; it fires once three consecutive checks have seen it. It does not pass on its own: restarting the guest (qm stop then qm start on its node) restores the monitor, and the alert clears once the monitor answers again.
anomalyCriticalThe change-rate guard judged a sync to look like ransomware or mass corruption; under the default policy, the guest’s replication is paused. Cannot be dismissed or silenced. Clears when you acknowledge it in the anomaly review.
peer_link_degradedCriticalThe paired site failed to answer a health probe, or took five seconds or more, on two consecutive probes (they run every five minutes). Clears when it answers normally again.
agent_unreachableAdvisoryA running VM’s QEMU guest agent, which had been answering, has stopped. The alert says whether the guest’s disk is still changing (only the agent has stopped) or not (the guest is likely frozen); Fix agent on the guest’s page is the remedy. Clears when the agent answers.
ct_left_frozenCriticalA container Asternodis paused to take an application-consistent point may still be paused: the release could not be read back from the kernel, so it may be serving nothing. The alert names the container’s cgroup and the exact command to thaw it; the node also releases it on a timer of its own. Clears when a later capture releases it and reads the release back.
resync_blockedCriticalA guest restored in place is not replicating forward: its replica is stale and the in-place re-sync is paused (usually no room on the recovery storage for the rewrite, or a replica a seed can’t be written over). Its recovery points are kept and the replica is untouched. Clears when the re-sync resumes or completes, the restore mark is dropped on purpose, or the guest stops being protected.

Recovery points

KindSeverityWhat it means
point_integrityCriticalA recovery point failed its integrity check and is quarantined: it can no longer be chosen for a failover or DR test. Clears once no quarantined point remains for the guest.
point_snapshot_lostCriticalOne or more recovery points no longer have the snapshot they were made of, so they cannot be recovered from; they are marked bad in the points list. Clears once no point is missing its snapshot, for example after you remove the affected points.
point_fs_unrepairableCriticalA flagged point was repair-tested on a disposable clone and its filesystem was still inconsistent afterwards, so a recovery from it may not boot. Recover from a different point. Clears once no such point remains.
point_fs_checkAdvisoryA recovery point’s guest filesystem needs repair, so it may not boot cleanly as captured. Clears once no such point remains, or each one has been repair-tested clean.
recovery_points_clearedAdvisoryA re-seed recreated the guest’s replica and destroyed every retained recovery point with it, because each point is a snapshot of that replica. The replica itself is current; what is lost is the history. Each re-seed raises its own alert, which clears when the retention schedule captures the next point.
replica_damagedCriticalThe live replica holds a block the storage can’t read, which a failover to the latest state and a failback would read. Forward replication only rewrites what the source changes, so it can’t repair this; re-seed the replica. Clears when the damaged ranges read clean, the pool stops listing the volume, or the replica is re-seeded.
pool_read_errorsCriticalReads of a guest’s recovery points fail with I/O errors that reproduce while the storage pool reports itself unhealthy, so nothing was quarantined: a pool-wide fault says nothing about one point, but recovery can’t come from a pool that can’t be read. Clears when a point of the guest reads back clean.
retention_prune_heldCriticalRetention pruning has been held for over a day for guests whose replication keeps landing data, so points are piling up past policy and can fill the recovery pool: a full thin pool fails every replica on it. Pair-wide. Clears when no guest of the pair is held: a confirmation landed, or pruning caught up.
anomaly_state_unconfirmedAdvisoryFor over an hour this site could not confirm a guest’s change-rate anomaly state with its source, so retention pruning is held (and automatic capture may be) and a re-seed of the affected guests is refused until it can be. The copy and every captured point are intact. Clears when the state is confirmed.
anomaly_mirror_undeliveredAdvisoryA change-rate anomaly state the source owes the recovery site has gone undelivered for over an hour. The source keeps retrying, and the recovery site pulls the same state itself as soon as it can reach this one, so this names a degraded path, not a stopped one. Clears when it is delivered.
app_consistent_default_pendingAdvisoryApplication-consistent capture is becoming this site’s default. The site has been capturing crash-consistent points and keeps doing so until the stated time; after that each guest is briefly quiesced at capture, unless you choose Keep off under Settings → Retention. Clears when a choice is recorded or the grace period ends.

Recovery-site readiness

KindSeverityWhat it means
firmware_state_unreplicatedCriticalThe guest boots with Secure Boot enforcing and its UEFI variable store is not held at the recovery site, so a recovery would boot it against the recovery node’s own Secure Boot keys. If those keys are older than the ones the guest was installed under, its bootloader is refused. Clears once a replication cycle carries the guest’s variables.
machine_unsupportedAdvisoryThe recovery node’s QEMU cannot provide the machine type the guest is pinned to. Where a supported substitute exists, recovery uses it so the guest still boots, with changed virtual hardware; upgrading the recovery node recovers it exactly as captured. Clears when the node can provide the type.
recovery_bridge_missingCriticalA NIC of a guest this site recovers would land on a bridge that does not exist on the node that would run it (Open vSwitch bridges included), so a failover would import the disks and leave the guest stopped: Proxmox refuses to start a guest whose NIC names a bridge it can’t attach to. Checked every five minutes; clears by itself once the bridge exists there.
data_path_blockedCriticalThe network check proved a path replication needs is closed (a node can’t open a TCP connection to a node of the other site on the data ports), so seeds can’t land and a failback over it will fail. Open the pair’s Network check for the fix for each closed link. A node that is switched off, or a port nothing on the pair dials, doesn’t raise it. A later check that still finds the path closed raises it again; it clears when a check that looked at least as widely finds nothing closed, or the pair is removed.

Failover and failback

KindSeverityWhat it means
recovery_boot_failedCriticalA VM recovered by a failover or DR test appears not to have booted: its serial console showed a known failure (GRUB rescue, kernel panic, emergency mode) or its disks showed no activity, confirmed on a second look with no guest agent answering. Each failed boot raises its own alert, and a DR test also fails with the same reason. A failover’s alert clears once the guest is no longer failed over.
split_brainCriticalThe guest is failed over, so its home copy is fenced (stopped and locked) while the recovery copy runs. The alert stands for as long as the guest is failed over and clears when it is home again. If the home copy was found running and could not be stopped, or could not be locked, the alert says so and what to do by hand.
stranded_recovery_conflictCriticalWhile restarting a guest at its primary after an interrupted failover, Asternodis found that the guest’s recovery state had changed at the same moment, most likely because the failover completed, so both copies may be running. Check which one is live before writing to either, and dismiss the alert once you have: it does not clear itself when the conflict passes.
reverse_stalledCriticalA failed-over guest’s changes have stopped reaching its home site, so everything written at the recovery site since exists only there. Clears when the reverse sync runs again, or the guest is no longer failed over.
reip_revert_failedAdvisoryA guest came home from a failover, but its addressing could not be reverted, so it is very likely still on its DR address. Fix home addressing on the guest retries. Clears once the guest no longer holds a DR address.
failback_leftoversAdvisoryA failback brought the guest home but could not finish cleaning up on its home node: a rollback snapshot or a reverse-sync marker was left behind. The guest is running at home and replicating. The cleanup is retried every 10 minutes, and the alert clears once it succeeds.

Cleanup

KindSeverityWhat it means
dr_test_teardown_failedAdvisoryA DR test’s isolated copy could not be removed; the alert says what is left of it on the node. Retry Stop test, or remove it by hand as the alert describes. Clears when a Stop test succeeds.
pending_release_stuckAdvisoryA guest was unprotected, and six hours later the recovery site has still not accepted the recovery-side teardown for it. The teardown is retried every 10 minutes, and each retry that fails raises or refreshes this alert; until it lands, the replica is likely stranded at the recovery site (see Shadows). It stops once the teardown lands, the guest is protected again, or the pair is removed.

The controller

KindSeverityWhat it means
controller_disk_lowCriticalThe volume holding the controller’s data directory has less than 2 GiB free, or less than 10%. Checked every six hours; clears once free space is back above both thresholds.
audit_chain_brokenCriticalThe audit log’s hash chain no longer verifies: an entry was edited, inserted or deleted. Entries before the break still verify. Checked hourly; clears if the chain verifies again.
peer_audit_divergedCriticalA paired site’s audit log no longer agrees with what that site signed for this one: it ends before an entry this site anchored, holds a different hash there, answered with something that isn’t a valid statement from it, or the site now paired here isn’t the one that was anchored. It never clears by itself (not even if the log matches again) and can’t be dismissed or silenced here or by the paired site: an admin Accepts the peer’s current log on the Audit page as a recorded decision.
peer_audit_anchoring_lostCriticalThis site can no longer keep anchoring a paired site it had anchored: the peer stopped answering the anchor request after having answered it (a downgrade to a build that doesn’t), has produced no usable statement for ten minutes, or can no longer be asked at all. Not raised for a peer that is simply down (that is peer_link_degraded). Clears on the first successful pull, when an admin acknowledges the peer no longer anchors, or when the pair is removed.

Delivery outside the UI

Alerts always appear in the UI. To have them leave it, an admin turns on webhook delivery with Alerting in the Replication page header, enters the webhook URL and, optionally, a per-attempt timeout (10 seconds by default), and saves. Delivery is off by default, and turning it on does not replay alerts raised while it was off.

  • Every kind is delivered, both when it fires and when it resolves, as a JSON POST with event ("alert"), status (firing or resolved), kind, vmid, name, detail and timestamp (UTC, RFC 3339). Any 2xx response counts as delivered.
  • A real delivery makes up to three attempts, with a short backoff between them. If all three fail, the failure is written to the audit log as alert_delivery_failed.
  • Send test posts one sample alert (status: "test", kind: "test") to the URL and timeout as entered, so you can prove the channel before saving. It is a single attempt by design: an endpoint that only answers on a retry is exactly what a test should expose. The result is shown in the dialog and recorded in the audit log.
  • Each site delivers the alerts it raises. Recovery-point checks and boot checks run at the recovery site, so both controllers need the policy. Alerting is one of the settings settings sync mirrors across a pair (on by default), so while sync is on, configuring it at one site covers both.