Atlas Knowledge Base
Dashboard
Sensors and Roles

Sensors and Roles


What a sensor is

A sensor is a single check with its own settings: the thresholds that decide warning from down, the trigger timing that says how long a bad reading must persist before it counts, where its alerts route, and whether it is currently paused. Each sensor is independent. A capacity sensor reports both a percentage and the absolute amount, and its thresholds can be set on either, because a nearly-full small volume and a nearly-full large one are different situations.

Direct assignment versus roles

A sensor can be attached to a host directly, one at a time, which suits a one-off check. But most hosts do the same job as other hosts, and typing the same set of checks onto each of them by hand does not scale. That is what roles are for.

A role is a named bundle of related sensors kept in a shared role library. Assigning a role to a host instantiates every sensor in it for that host at once. A database-server role, for example, brings the service checks, the disk checks, and the database health checks that any database server should have, all together. A host can carry more than one role and composes what it actually runs from them.

Per-host variables

Roles are templates, not fixed copies. A role reads per-host variables when it expands, so one library role fits many hosts. A database list or a set of service names supplied on the host lets the same role produce one sensor per database or per service on that particular machine - the list drives the expansion.

Per-site overrides

Library defaults are a starting point. A per-site override on a host beats the library default, so a single host can raise a threshold, change a timing, or opt out of one sensor from a role without editing the shared template and without disturbing every other host that uses it.

Tags and grouping

Every sensor, host, and role template carries free-form tags, and a sensor's effective tags are the union of its own, its host's, and its role's. Tags are how the estate is organized: what used to be a fixed stack or service-type field is now just a tag, and a sensor can carry several. The dashboard groups and filters by tag - a sensor with several tags appears under each of its groups

  1. and rolls status up per node so the worst thing under a branch is visible without expanding it.

Sensor type and owning host stay available as built-in grouping dimensions.

Log-only while bringing a sensor up

A new sensor can be routed to log-only: it probes and records exactly what it would have alerted on, but pages no one. This is how a sensor is brought up safely - watch it behave against real conditions, confirm its thresholds and timing are right, then switch its routing to live. It is also the standby core's normal mode for everything, which is why a takeover is warm.

A short role template looks like this, with placeholder values only:

role: database-server
sensors:
- type: service
name: db-service
route: log # raise to live after parallel-run validation
- type: disk
name: data-volume
warnPercent: 85
downPercent: 95

Central, live configuration

All of this - sensors, roles, host assignments, overrides, tags, routing - lives in one place on the core. Nothing is deployed to the monitored hosts but the agent itself. Changes hot-reload without restarting the cores or the agents: the core hands each agent its updated manifest over the existing connection, and the change appears live within seconds.



Was this helpful?