Atlas Knowledge Base
Dashboard
High Availability

High Availability


SBN Monitor runs as a pair of identical core servers in an active/passive arrangement, hosted on separate cloud instances in different availability zones. Both cores run the same build and the same sensors configuration. The active core owns alerting; the passive core is a warm standby that probes the full estate and routes every alert to its log sink until it takes over. Core names are role-neutral (monitor-a and monitor-b in the examples), because the roles can swap: which box is active is set in each box's ha: block in monitor-config, described in Configuration.

What each core does

  1. Active core - alert routing is armed. It health-checks the standby and raises a warning through normal routing if the standby goes dark, so losing the standby is itself an alert. It pushes the sensors file to the standby on every successful reload and on a periodic reconcile.
  2. Passive core - probes everything the active core probes, with every alert gated to the log sink. It runs the promotion engine and receives configuration from the active core. It also hosts the Prometheus/Grafana history stack, which scrapes both cores' /metrics from outside the active core's failure domain.

Promotion

Promotion is automatic. The standby promotes only when all three conditions hold at once:

  1. The peer is unreachable on both its public URL and its direct private-network URL for a configured number of consecutive checks.
  2. The standby's own internet egress verifies against an independent third-party anchor URL.
  3. An agent quorum - a configured fraction of the expected agents - is currently connected to the standby.

Together these distinguish "the peer died" from "I am the isolated one" without a consensus protocol: a standby that has lost its own network sees its egress check fail and its agents disconnect, and stays passive.

On promotion the standby arms its alert routing, resets every remote sensor's staleness clock so that only a genuinely dark agent alarms, journals a PROMOTED event, and raises exactly one Priority-1 ("monitor core down - standby promoted") through the SRM sink.

Demotion and handback

Handback is an operator action: an authenticated POST /api/demote, or the banner on the promoted core's dashboard. Automatic handback would invite flap loops between the cores.

Demote refuses with a 409 while the peer is unreachable, since dropping to passive with the peer still down would only lead to re-promotion a few checks later. The core stays active until the primary answers. A recovered primary boots passive under the boot rule below, so the sequence after an outage is: wait for the primary to come back, demote the standby, and the primary takes the active role back automatically.

Boot rule

A configured-active box that reboots while the standby is promoted comes up passive: at start it asks the peer its current role over the config-sync channel and, finding the peer active, defers until the operator demotes the standby. This prevents a dual-active pair after a reboot. The rule depends on config sync being configured between the boxes; a box with no config-sync pairing uses its configured role.

Agent dual-dial

Agents enroll against one URL and learn the full core set centrally: the cores deliver the endpoint list (ha.agentEndpoints) over the agent transport, and each agent persists it and holds one connection per core, with independent reconnect and backoff per core. Every agent executes a single manifest and duplicates every reading to every connected core - the manifests are identical because configuration is synced, and readings are tiny.

Dual-reporting keeps the standby's trigger-timing state warm, so a promotion starts from live data, and the connected fleet doubles as the promotion quorum. An agent that only supports a single endpoint keeps working unchanged against its enrollment URL.

Configuration sync

On every successful reload, and again on a periodic reconcile, the active core pushes the central sensors file to the standby over a mutually authenticated TLS channel on the private network. The channel authenticates with client certificates minted from the transport CA - the same CA that backs agent identity - and both boxes hold that CA, so the channel stays valid when the roles swap. The standby applies the received file through its normal hot-reload path.

The per-box monitor-config file is never synced: it carries credentials, the box's role, and peer addresses, which differ per box by design.

Agent enrollment and token minting stay on the active core. The standby accepts existing agent certificates and issues no new identities.

One stable dashboard address

A single neutral hostname (health.example.com) always reaches the active core's dashboard. DNS points the name at both origins; a request that lands on the passive box is redirected with a 302 to the active peer's dashboard, path and query preserved, while the active box serves normally. If the passive box cannot reach its peer, it serves a terminal "this core is passive" page.

Three exceptions are served locally on whichever box the request lands on:

  1. GET /healthz/role - unauthenticated role probe (200 active / 503 passive), used by monitoring.
  2. The public status page (/status) - customer-facing status must stay reachable regardless of which core is up.
  3. Read-only state endpoints (GET /api/state, GET /api/events) - the desktop tray reads these, and both cores hold warm state through dual-reporting.

The passive core's own dashboard shows a prominent "you are on the passive server" banner linking to the active core.

DNS records


Record

Type

Points at

monitor-a.example.com

A

core A's public address

monitor-b.example.com

A

core B's public address

health.example.com

two A records, or a CNAME to each origin behind a proxy

both cores - the app-level redirect above does the steering

Each box serves TLS for its own name and for the neutral name, since a neutral-host request can terminate at either origin.

Pick one canonical hostname per core and use it everywhere: the URLs in ha.agentEndpoints and the URL each agent enrolled against must match string-for-string. Agents identify a core by its URL string, so two spellings of the same core (say, a legacy alias and the current name) are treated as two different cores, and every agent runs duplicate sessions against that box. When a core's public name changes, update agentEndpoints on both cores and the agents' configured URL together.

Network ports


Port

Where it is open

Carries

443/tcp

both cores, public

dashboard and API, agent websocket at /agent/v1, the public status page, /metrics scrapes - terminated by the core itself or by a reverse proxy in front of it; on core B also Grafana at /grafana/

80/tcp

both cores, public

redirect to HTTPS and ACME certificate challenges

22/tcp

both cores

SSH administration; the nightly backup push crosses it on the private network

8099/tcp

private network only

the monitor process when a reverse proxy fronts it - the proxy forwards to it locally, and the peer's direct health check reaches it across the private network; the cloud firewall blocks it publicly. A core that terminates HTTPS itself listens on 443 directly and has no listener here

8098/tcp

private network only

the config-sync mTLS listener on the passive core; also answers the role query behind the boot rule

9090/tcp, 3000/tcp

loopback on core B

Prometheus and Grafana; reachable from outside only through the reverse proxy at /grafana/

Monitored hosts open no inbound ports: agents connect outbound to 443/tcp on both cores, and enrollment, readings, manifests, and fleet updates all travel over that connection.

Journals and history

Alert journals are append-only and per core, with every entry tagged by the core that wrote it. They are merged only when comparing the two cores' records, never replicated live. Trigger-timing state stays per-core too - dual-reporting keeps both warm.

Long-term history lives in the Prometheus/Grafana stack on the standby, which scrapes both cores. Three gauges cover the pair itself: sbnmonitor_ha_role (1 = active, 0 = passive), sbnmonitor_ha_promotion_pending_checks, and sbnmonitor_ha_peer_reachable.

Backups

Each core archives its durable state nightly - sensors file, monitor-config, journals, enrollment-token ledger, and the transport CA state - and pushes the archive to its peer over the private network at staggered times, through a restricted receive-only account. Each side keeps the newest seven daily snapshots of its peer, so losing either box leaves a current copy of its state on the survivor. A copy of the transport CA is also held in credential custody outside the pair, for the case where both boxes are lost.



Was this helpful?