Configuration
SBN-Monitor keeps its configuration in two files. A central sensors file on the core describes everything being monitored. A per-box monitor-config file holds the credentials and deployment details belonging to that machine. Monitored hosts hold neither.
The central sensors file
One file on the active core carries the whole monitoring picture: sensor definitions, role templates, the host list with each host's roles and variables, tags, the named alert routes, the components shown on the public status page, and scheduled maintenance windows. Anything describing what should be watched and how it should behave lives here.
A trimmed example — one host with its roles, one directly-assigned sensor, one status-page component:
The latency fields set how long a bad state must persist, the channel limits set the warning and error values, and the staleness bound covers agent silence.
The file hot-reloads: saving an edit applies it to the running core within seconds.
In a two-core deployment the active core pushes this file to the standby on every successful reload and again on a periodic reconcile, so both cores always probe an identical set. The standby applies it through its own hot-reload path.
Route definitions
A sensor's route: names where its alerts go, and it takes either a single name or a list of them. The names are the estate's own vocabulary, declared in a routes: section: each entry names the delivery kind it uses - the Priority-1 lane, an ordinary service request, a Jira issue, an email, a signal into an SBN installation, or log-only - and the values that kind should file under, such as the SR type and priority, the Jira project and issue type, the email recipients, or the SBN account prefix. Six names ship predefined, one per kind, and an entry using one of those names redefines what it means here, so what a route is called and what it does are both site decisions. What each kind does with an alert is covered in Alerting.
A route definition never carries a credential, because this file is version-controlled and shared. Base URLs, logins and tokens stay in the per-box monitor-config, and a definition supplies only the filing values. Where a definition leaves one of those unset, the leg's own monitor-config value applies.
Combinations are checked at load. A set may pair the Priority-1 lane or an ordinary service request with a Jira issue, and both together are how a customer-facing record and an engineering follow-up are raised from one condition. Refused: the Priority-1 lane together with an ordinary service request, since the Priority-1 lane already sends its warnings and restores to the SR leg; log-only together with anything else, since log-only means no external system is touched; and two routes of the same kind, whatever they are called. An undefined name and an empty set are refused with a message naming the problem.
A sensor that names no route is routed by the routing: policy: one block mapping each priority level - critical, high, normal, low - to the route names its Down, Warning and escalation take. A level or column left out takes the shipped default, and each cell obeys the same combination rules a sensor's own route does. The policy in force, and how many sensors each cell serves, is read and edited on the dashboard's Routing view.
A configuration that routes to a leg whose credentials are missing is refused at load and at reload, naming the route and what it lacks, so a route that cannot deliver is found when the file is saved rather than at the first alarm.
Standard roles
A library of standard roles ships built into the monitor. A host lists the roles it fills and gets every sensor those roles define, so a new machine of a familiar kind is covered by naming what it is. A site-defined role with the same name replaces the library one wholesale.
Role | What it monitors | Parameters |
|---|---|---|
| Reachability of a network device — switches, firewalls, gateways, receiver hardware. | none |
| Reachability plus a name-resolution check against the server. | none |
| Reachability plus an HTTP health check on a receiver automation host. | none |
| Public HTTP availability and TLS certificate expiry. | none |
| The SBN tunnel services and the tunnel dashboard service. | none |
| The core and probe services of a PRTG installation. | none |
| The SCS-VR receiver service and its SQL Server instance. | none |
| The SBN web front-end application pools in IIS. | none |
| The database and backup services, blocked processes, deadlocks, long-running queries, user-connection and kernel-memory usage, report and email queue depth, and the server error log. |
|
| Replication Server partition usage and the replication error log. |
|
| One service check per SBN Windows service the host runs. |
|
| Back-compat alias covering |
|
| A Microsoft SQL Server host — blocked processes, long-running queries and user-connection usage. |
|
| The SBN Media service on a host, its local API and message-broker endpoints, and its application log. |
|
| One process check per SBN Media product the host runs. |
|
| A whole-mesh health roll-up, reported from one designated host per mesh. |
|
| The secure SIP listeners on a host running the SIP server. | none |
| The relay control endpoint on a host running browser media. | none |
| The analytics database endpoint on an insights host. | none |
| The tunnel service on a farm member, its public connection port and its peer-coordination endpoint. |
|
| The tunnel dashboard service, plus HTTP availability and TLS certificate expiry on its endpoint. |
|
Every host with an agent also gets processor, memory and per-volume disk sensors automatically, without naming a role for them.
Role parameters
A role's sensors carry placeholders rather than fixed names, and the host fills them in with a vars: block. One role therefore fits many machines, each supplying its own service names, addresses and paths.
A var holding a list expands to one sensor per element: a host's list of SBN Windows services becomes one service check each.
On an agentless target the host's address is available as a builtin placeholder, so a role written for a device needs no vars of its own.
Credentials are named rather than written — a var carries the environment-variable name, with the value in the environment of the machine running the probe.
To change what a role itself contains, whether that is its sensors, their timing or their thresholds, define a role of the same name in the site file's roles: block.
Var | What it is | Values / example |
|---|---|---|
| Database server suffix, used to form the server and backup service names. |
|
| Which database the health checks talk to. |
|
| Address and port of the database server. |
|
| The monitoring database holding the health procedures. |
|
| Name of the environment variable holding the monitoring login, never the login itself. |
|
| Name of the environment variable holding its password. |
|
| Full path to the database server error log, forward slashes even on Windows. |
|
| Full path to the Replication Server error log, same path convention. |
|
| List of SBN Windows services on the host, one check each. |
|
| The SBN Media service exactly as installed on the box. |
|
| Full path to the active SBN Media log, forward slashes even on Windows. |
|
| List of media product processes, named exactly as they run including any |
|
| Full path to the SBN Media command-line tool on the box. |
|
| Name of the environment variable holding the mesh administration token. |
|
| The tunnel server service exactly as installed on the box. |
|
| The tunnel dashboard service exactly as installed on the box. |
|
| Hostname of the tunnel dashboard's public endpoint. |
|
The per-box monitor-config
Each core box has its own monitor-config carrying secrets and machine-specific state:
- Alerting credentials — the account the core raises urgent pages with and publishes into the monitored SBN instance under.
- Jira leg — the Jira base URL, whether the monitor authenticates with an account and API token or with a token alone, the credential itself, the default project and issue type, and the transition name a clear applies. Leave the block out and no issue is ever opened.
- Email leg — the relay servers walked in order, the From address and name, a default subject prefix and the send timeout. Who receives a message is on the route, never here. Leave the block out and no mail is sent.
- SBN signal leg — the API Engine address and credential of the SBN installation that receives monitor signals, the source id the installation knows the monitor by, and the account family prefix (
ZRECMONunless set). Leave the block out and no signal is sent. - Dashboard — the operator sign-in credential (the shipped default is changed from the dashboard on first sign-in), whether the core terminates HTTPS itself or sits behind a reverse proxy, the release-tree directory it serves at
/dist/, and an optional per-sensor Grafana link template. - Install-page gate — a separate, deliberately weaker credential pair for the agent install page, its artifact directory, and the address agents dial.
- Alert journal — path, size cap and backup count for the durable alert log on that box.
- HA block — this box's role, its peer's name and addresses, the promotion thresholds, the endpoint set agents are told to dial, and the config-sync pairing.
- Agent auto-update — the target agent release, the signed-artifact directory, the address agents download from, the canary list and the bake time.
- Management API — which of the API surfaces this box exposes, and the keys that authenticate automation against them.
- Core self-update — the release signing keys this box will accept a replacement binary from.
A trimmed example of the same file:
The dashboard.tls block selects how the core terminates HTTPS. acme obtains and renews a public certificate itself for the names listed; file serves an operator-supplied certificate and key from certDir and reloads them when the files change; self-signed generates one certificate and keeps it across restarts. With any of the three the core listens on 443 directly, binds 80 to redirect plain HTTP and to answer the certificate challenge, and serves the agent release tree at /dist/ from distDir. A mode that cannot produce a certificate refuses to start. off, the default, leaves the listener plain HTTP for a core behind a reverse proxy, which then sets behindTlsProxy: true so the session cookie is issued for HTTPS only. List only the box's own names under hostnames, never the shared neutral name; where a proxy or CDN in front of the pair presents the neutral name to the core, use file mode, which serves the one certificate for every name asked for.
The ha: block is present only on a two-core pair; a single-core deployment leaves it out.
This file is deployed out of band, owned by the service user and readable only by that user. Each box is edited independently and the file is never copied between cores, because role, peer and paths differ per box by design.
Secrets a sensor needs — a database monitoring login, an API key for an HTTP check — are referenced in the sensors file by environment-variable name. The value lives in the environment of the machine running the probe.
Monitored hosts
Probe agents carry no configuration of their own. An agent enrolls once and then receives its manifest from the core over the connection it already holds, listing exactly the checks it should execute. Adding, removing or retuning a check on a monitored host is done on the core.
The one host-local setting is an optional environment file bounding the agent's resource footprint: how many probes it runs at once, and the CPU and memory levels at which it sheds its own work to protect the host. The shipped defaults suit most machines.
Making a change live
A sensors-file edit on the active core goes live on hot reload, with nothing restarting anywhere. The standby core receives the same file automatically, and every affected agent receives an updated manifest within seconds.
A monitor-config edit applies to the box it was made on and takes effect when that box's service restarts. On an HA pair, restart one box at a time so the estate stays covered.
Overrides made from the dashboard
A field changed from the dashboard is stored as an override on top of the sensors file rather than written into it. The hand-maintained file stays exactly as it was; the stored overrides are laid over it on every load, and for the fields they name the override is what runs.
An override can be written on a single sensor or at the shared level the value came from, and every change is recorded with the operator, the time, the reason given and the values either side. The editing surface itself is covered in Dashboard.
An override is checked against the whole prospective configuration before it is stored, so a change that would not load is rejected and nothing moves. Only the active core owns the override set; the standby refuses the change and points at the active peer. If the file is later edited by hand in a way the stored overrides no longer fit, the load falls back to the unoverridden file and the dashboard says the overrides are not in force. An override whose sensor, role or host no longer exists after a file edit is dropped at the next load and the drop is recorded as the system's own.
To settle an override permanently, write the value into the sensors file and delete the override.
The management API
Each core can expose an HTTP management API for automation. Every surface is off by default and turned on one at a time, and a surface that is off is not mounted at all - it answers 404 like any address the core does not serve. Turning one on never implies another, and listing API keys on its own turns on nothing.
Surface | What it allows |
|---|---|
Configuration read | Fetch the live sensors file and list the backups beside it. |
Configuration write | Replace the live sensors file. |
Service log | Read this core's own service log. |
Self-update | Replace this core's binary, roll it back, and report what is installed. |
Automation authenticates with a key presented as a bearer credential; the operator's dashboard login keeps working on the same addresses, so a person and a script each use what suits them. Keys are per box and never copied between cores, several can be listed at once so one is rotated by adding the new key and then removing the old, and a key is never written to a log, to the permanent record, or into a diagnostic bundle. Replacing the configuration and updating the binary accept a key only.
Reading and replacing the configuration
A read of https://<core>/api/config returns the live sensors file and an entity tag for it. A replacement must quote that tag: a request without one is refused, and so is one quoting a tag that has since moved, so two people editing at the same time cannot silently overwrite each other. A replacement also carries a short tag naming the change, and it is accepted on the active core only.
Nothing is written until the whole submitted file has been trial-loaded the way a hot reload would load it. A file that fails is rejected with the reason it failed and the box is left byte for byte as it was. A file that passes is first backed up beside the live one under the date and the change tag, then replaced in one move and hot-reloaded, and the reply says how many sensors were added, changed and removed. A backup is never overwritten: a second replacement on the same day under the same tag is refused and asks for a fresh tag, before the live file is touched. The change is recorded either way, including the case where the new file was stored but the reload that followed refused it. There is no second step to push the file to the standby - the reload the write triggers carries it there.
Core self-update
A core can replace its own binary from a release archive signed with the same signing chain the agent fleet uses. The signature is checked against keys pinned in that box's own configuration before anything is installed, because an authenticated upload proves who sent it and never what was sent. Verification covers the signature over the release manifest, the checksum of the bytes that arrived, the signature over the artifact itself, and which product the release is for - so a validly signed agent release is refused by a core. Only then is the new binary unpacked beside the live one, the running binary kept as the previous one, and the new binary moved into place; one call restores the previous binary. A refusal leaves the box unchanged, and every attempt is recorded with the versions either side and what came of it.
The surface cannot be enabled without pinned keys - a core told to accept updates with nothing to verify them against refuses to start. Installing a new binary ends with the core shutting down cleanly for the service manager to start again, so the service unit is set up to treat that exit as a restart rather than a stop.
Where a new setting belongs
Kind of setting | File |
|---|---|
A check, threshold, timing or routing choice | Sensors file |
A role template or a host's role assignment | Sensors file |
A tag, maintenance window or status-page component | Sensors file |
A route definition: its name, its kind and what it files under | Sensors file |
A password, key or account | monitor-config |
This box's role, peer or listening address | monitor-config |
A filesystem path or artifact directory | monitor-config |
Shared monitoring intent belongs in the sensors file, where both cores see it and the agents inherit it. Anything a second box would answer differently belongs in that box's monitor-config.