Alerting
Thresholds and tiers
Every sensor takes a reading on its own schedule and compares it against its own thresholds. Most sensors carry two tiers: a Warning level for a condition worth attention but not urgent (a disk filling, a certificate nearing expiry) and a Down level for something broken (a service stopped, a query failing, a host gone silent). Capacity sensors can trip on either an absolute amount or a percentage, whichever matches the concern - a 2 TB volume and a 20 GB volume at the same percent free are not the same situation.
Trigger timing
A single bad reading rarely means a real problem. Each sensor has a hold time - anywhere from one second to several minutes - that a Down condition must persist before it counts as a real Down. A service that flickers and recovers inside its hold window never raises an alert. The hold time is set per sensor, so a latency-sensitive check can react in a second while a noisy one waits out the noise.
Where an alert goes
Routine conditions - the Warning tier, and Down conditions on sensors that are not on the Priority-1 lane - raise an ordinary service request in SRM, which the monitor closes itself on the restore that clears the sensor. A sensor can be routed elsewhere instead, or as well: to a Jira issue, to an email, or into an SBN installation as a signal on a receiver account, each described below.
A Priority-1 condition takes a different path. A Down alarm on a sensor on the Priority-1 lane raises a Priority-1 service request in SRM automatically, carrying the current P1 behaviour - the request, the on-call SMS, the email. That sensor's warnings and all of its restores still take the ordinary service-request leg, and the Priority-1 request itself is closed by a person. If the same Priority-1 condition recurs inside the re-open window - a configurable period, an hour by default, that SRM itself enforces - it is folded into the request already open rather than opening a fresh one each cycle, so a flapping fault does not bury on-call in duplicate tickets. The two paths are independent by design: the Priority-1 route to SRM does not depend on SBN being healthy, so "SBN is unreachable" is itself a Priority-1 that still gets out.
Route names and combinations
Where a sensor's alerts go is set by its route, and a sensor can name more than one. Six delivery kinds exist: the Priority-1 lane, an ordinary monitor-raised service request, a Jira issue, an email, a signal into an SBN installation, and log-only. The names sensors use for them are the estate's own. Six ship predefined, one per kind, and the sensors file can redefine any of those names or declare new ones, each naming the kind it uses and the values it files under - so a site can run two Jira lanes into different projects, or two service-request lanes at different priorities, or two email lanes to different teams. The permanent record, the alert log and the dashboard all read back the name that was written.
Naming two routes on one sensor is how a single condition reaches two places: a Priority-1 page and a Jira issue for the engineering follow-up, or a customer-facing service request and an issue. Pairs that would only double up are refused when the file loads rather than at the first alarm - the Priority-1 lane together with an ordinary service request, because the Priority-1 lane already sends its warnings and restores down the service-request leg; log-only with anything at all, because log-only means no external system is touched; and two routes of the same kind, whatever they are called.
Jira issues
A sensor routed to Jira opens one issue for itself when it alarms. The monitor keeps no local list of what it has open: it searches Jira before it creates anything, matching its own label and the summary it files that sensor under. So the same sensor alarming again comments on the issue already open instead of opening a second one, and an issue opened by one core is closable by the other after a failover. Every later transition while the condition persists is another comment on that issue. When the sensor clears, the monitor comments the recovery and then applies the configured close transition; a workflow that no longer offers that transition gives an error naming both sides, and the recovery comment has already landed, so a human has the issue either way.
The project, the issue type, the priority and any extra labels belong to the route, so two Jira lanes can file into different projects from the same estate. The base address, the credentials and the name of the close transition are per-box operator configuration and are never written in the sensors file. Leave the Jira configuration out and no issue is ever opened.
A sensor routed to email sends one plain-text message per event, the alarm and the clear alike; the restore message is the close, and there is nothing to reconcile after a restart or a failover. The subject reads the priority, the tier, the sensor and its host and the state, with RESTORED on the clear; the body carries the sensor, host, priority, tier, status, value, time, routes and message, and a link to the dashboard when the estate has a single dashboard address. Who receives the message is on the route: a to: list, an optional cc: list, and an optional subject prefix. How mail leaves the box - the relay servers, the From address, a default prefix and the timeout - is per-box operator configuration. There is deliberately no box-level default recipient: a route that names nobody is refused when the file loads. A route can be set to build and log every message without sending, for proving the wording before anyone is mailed.
Signals into SBN
A sensor routed to SBN arrives in an SBN installation the way a receiver-side technical fault does: as a signal on a receiver-fault account, delivered through the API Engine over HTTPS. Each monitored host owns its own account in a family with a shared prefix - ZRECMON01, ZRECMON02 and so on, ninety-nine per family - so an operator reads which server an alarm concerns off the account alone. A number is allocated the first time a host routes to SBN, recorded in the permanent record, and kept across restarts, reloads and failovers; a decommissioned host's number is released from the management API and handed out next. With the family full, an allocation is refused and recorded rather than doubled onto another server's account.
The alarm is an event and the restore is a restore on the same account and zone, so the installation's own handling clears it; the zone is a stable code derived from the sensor name, and the event code follows the sensor's priority - P1 for critical, P2 for high, P3 for normal, P4 for low - so the installation can rank a monitor signal without reading its text. The installation must hold the account family and the translation for those four codes. The API Engine address and credential, the source id and the account prefix are per-box operator configuration; a route can name a different prefix.
Routing by priority
A sensor states how much it matters - critical, high, normal or low, set on the sensor, on its host, or on a role, resolved in that order - and a routing policy states where each level goes. The policy is a table of priority against tier: for each level, the routes a Down takes, the routes a Warning takes, and the routes an escalation takes. A sensor that names its own route keeps it whatever the tier; one that names none is routed by its priority's row. The shipped table sends a critical Down to the Priority-1 lane and its warnings to a service request, a high Down and warning to a service request, a normal Down to a service request and its warnings to the log, and everything on low to the log; escalations go to the log. The Routing view on the dashboard shows the table in force, with the number of sensors each cell currently serves, and a change made there is recorded and pushed to the peer core.
Log-only while bringing a sensor up
A brand-new sensor can be routed to log-only. It probes and evaluates normally and writes everything to the permanent record, but pages no one - so you can watch a new check behave for a day or a week, confirm its thresholds and timing are right, and only then let it alert for real.
Escalation, repeats, and restores
Repeat reminders and escalation run on the service request the monitor raised, on the same rules the team already uses for any other request. When a sensor clears, the restore follows the same route the alert took, so a request the monitor opened is closed by the monitor, a Jira issue is commented and closed, an email announces the recovery, an SBN signal restores the alarm, and an on-call engineer sees the all-clear.
Latched log states
Log-reading sensors - the log file tail and the OS event log - hold on to what they found. A warning or a down raised by a matched line does not clear because a later check happened to read no matching lines: an error written to a log is an event someone has to see, and a quiet minute after it is not a recovery. The state stays until an operator explicitly clears it.
Clearing is its own action, separate from acknowledging. An acknowledgement annotates the state and changes nothing about alerting, and it behaves identically on every sensor type; clearing is what ends a latched log state. While a state is latched, further matches update the message and can raise its severity. Problems with the file itself - a log that has gone missing, or a pattern matching nothing - are not latched and clear on their own once the file is back.
Maintenance windows and pausing
Planned work should not page anyone. A maintenance window suppresses alerting for a chosen set of sensors for a bounded time; the sensors keep probing and their state stays visible, but no alert leaves while the window is open. For an ad-hoc quiet, any sensor - or a whole tagged group - can be snoozed for a set duration from the dashboard and resumes automatically when the timer runs out, with no reminder needed to switch it back on.
A permanent record
Every alert and every restore is written to a permanent record the moment it happens: which sensor, which tier, where it was routed, and whether delivery succeeded. The record is durable, and it is what lets the team reconstruct exactly what happened and when, long after the condition has cleared.