Skip to content

What do you need to do?

Docs

Server health

Every machine running the Monitor agent gets CPU, memory, disk, and load charts with history, at no separate cost and with nothing extra to install.

On this page

What the charts show

CPU is a fraction of total capacity across every core, fully loaded reads as 100%, never a raw per-core number. Memory used is total minus what the kernel reports as available, not minus what it reports as free, which is the reading that actually predicts an out-of-memory kill. Disk is tracked per mounted filesystem, up to 32 per machine. Load average is the kernel's raw 1, 5, and 15 minute numbers, not normalized by core count, since a load of 4 means something different on a 2-core box than on a 32-core one and normalizing would hide that.

A sample lands about once a minute. The very first sample after the agent starts establishes a baseline for the CPU reading and is not charted, so the first real data point appears roughly a minute after the agent comes up.

Network, processes, containers, services

Agents from version 0.2 also report, once a minute: bytes per second in and out per network interface (loopback excluded); the top ten processes by CPU and by memory, by process id and executable name only, never a command line; every container on a Linux host from cgroup v2 (Docker, containerd, CRI-O, Podman, LXC, and Kubernetes pods), with its CPU share and memory against its limit; and the status of any services you name on the host's page (active, inactive, failed). The host's page shows the latest reading of each under Right now, and a network throughput chart alongside the four core ones. A family the platform cannot measure is simply absent: containers on macOS and Windows, load average on Windows.

Threshold alerts

Every host has three rules, CPU, memory and disk, each reading as "at or above N% for M minutes, clears under N% for M minutes", editable on the host's server health page. A rule fires only once the reading has held past its line for the whole sustained window, not on a single sample, since one high reading is often routine, a backup job, a build, a garbage-collection pause, and the point is to catch a real state, not a moment. Disk is judged per filesystem.

Rule (defaults)FiresClears
Disk, onat or above 90% used for 2 minutesunder 88% used for 5 minutes
Memory, onat or above 90% used for 5 minutesunder 88% used for 5 minutes
CPU, off until you turn it onat or above 90% busy for 5 minutesunder 80% busy for 5 minutes

The firing and clearing lines are deliberately different (control-theory hysteresis), so a reading sitting right on the edge does not open and close on ordinary measurement noise, and the form will not save a recovery line at or above the alert line. Windows run from 1 to 60 minutes. A gap in the series, an agent that was offline for part of the window, reads as undecided rather than as a verdict. A condition that keeps flapping open and closed gets suppressed after repeated cycles in a day, replaced by one notice that it is flapping, rather than repeating the same alert over and over.

Alerts flow through the same email, Slack, webhook, and PagerDuty channels every other alert on the account uses, and quote the host's own rule. CPU alerting is off by default because CPU is the flap-happiest reading; the sustained window is what makes it safe to turn on. A host nobody has tuned follows the defaults above, including any future change to them; Reset to default on a custom rule returns it there.