Atlas Knowledge Base
Dashboard
Agent Updates

Agent Updates


Central distribution

Agent updates are distributed centrally, so nobody logs into a monitored host to upgrade it. The active core publishes a new agent release over the same secure connection the agent already holds, names the target version, and lets each agent pull and apply it on its own. The estate can hold hundreds of hosts; none of them is touched by hand.

Release verification

An agent never runs a release just because it arrived over a trusted connection. Before it applies anything it verifies the release is ours and that the files are intact and unchanged, and it refuses any version older than the one it is already running. Even a compromised core cannot push arbitrary code to the fleet. It can at most choose which of our genuine, not-older builds an agent receives.

Update sequence

Once a release is verified the agent applies it carefully and reversibly:

  1. It stages the new binary alongside the running one in its protected state directory.
  2. It swaps to the new binary and restarts under its own service supervisor.
  3. It checks its own health - it must reconnect and re-authenticate to a core within a few minutes.
  4. If that self-check fails, it rolls back automatically to the previous binary and reports the rollback.

The previous binary is kept until the new one has proved itself, so recovery never needs a person on the host.

Staged rollout

Rollouts are staged. A named canary set of agents is advised of the new version first and given time to apply it and prove healthy. Only after that bake time does the rest of the fleet receive the same advisory. A build that fails on the canaries surfaces its trouble early, while almost the whole estate is still on the known-good version.

Update records

Every step leaves a permanent record: the advisory, each apply, each failure, each rollback. Each agent's current version is visible centrally, so at any moment an operator can see which hosts are on the target version and which are still catching up.

What an operator sees, and what to do

During a rollout the version column moves through the fleet in waves - canaries first, then the rest after the stage delay - and the recorded update events accumulate as hosts report in. This is the normal picture and needs no action.

If an agent reports a failed update, the right response is to do nothing to the host. The agent has already rolled itself back to the version it was running and is reporting normally on that version. The failed-update record is there so the release can be looked at centrally.



Was this helpful?