What the charts show
CPU is a fraction of total capacity across every core, fully loaded reads as 100%, never a raw per-core number. Memory used is total minus what the kernel reports as available, not minus what it reports as free, which is the reading that actually predicts an out-of-memory kill. Disk is tracked per mounted filesystem, up to 32 per machine. Load average is the kernel's raw 1, 5, and 15 minute numbers, not normalized by core count, since a load of 4 means something different on a 2-core box than on a 32-core one and normalizing would hide that.
A sample lands about once a minute. The very first sample after the agent starts establishes a baseline for the CPU reading and is not charted, so the first real data point appears roughly a minute after the agent comes up.
Network, processes, containers, services
Agents from version 0.2 also report, once a minute: bytes per second in and out per network interface (loopback excluded); the top ten processes by CPU and by memory, by process id and executable name only, never a command line; every container on a Linux host from cgroup v2 (Docker, containerd, CRI-O, Podman, LXC, and Kubernetes pods), with its CPU share and memory against its limit; and the status of any services you name on the host's page (active, inactive, failed). The host's page shows the latest reading of each under Right now, and a network throughput chart alongside the four core ones. A family the platform cannot measure is simply absent: containers on macOS and Windows, load average on Windows.
Threshold alerts
Every host has three rules, CPU, memory and disk, each reading as "at or above N% for M minutes, clears under N% for M minutes", editable on the host's server health page. A rule fires only once the reading has held past its line for the whole sustained window, not on a single sample, since one high reading is often routine, a backup job, a build, a garbage-collection pause, and the point is to catch a real state, not a moment. Disk is judged per filesystem.
| Rule (defaults) | Fires | Clears |
|---|---|---|
| Disk, on | at or above 90% used for 2 minutes | under 88% used for 5 minutes |
| Memory, on | at or above 90% used for 5 minutes | under 88% used for 5 minutes |
| CPU, off until you turn it on | at or above 90% busy for 5 minutes | under 80% busy for 5 minutes |
The firing and clearing lines are deliberately different (control-theory hysteresis), so a reading sitting right on the edge does not open and close on ordinary measurement noise, and the form will not save a recovery line at or above the alert line. Windows run from 1 to 60 minutes. A gap in the series, an agent that was offline for part of the window, reads as undecided rather than as a verdict. A condition that keeps flapping open and closed gets suppressed after repeated cycles in a day, replaced by one notice that it is flapping, rather than repeating the same alert over and over.
Alerts flow through the same email, Slack, webhook, and PagerDuty channels every other alert on the account uses, and quote the host's own rule. CPU alerting is off by default because CPU is the flap-happiest reading; the sustained window is what makes it safe to turn on. A host nobody has tuned follows the defaults above, including any future change to them; Reset to default on a custom rule returns it there.