Skip to content

What do you need to do?

Docs

API reference

The REST API and MCP server, documented in full: authentication, permission scopes, rate limits, every endpoint, webhooks, PagerDuty, and custom domains.

On this page

Both the REST API and the MCP server expose the same operations against the same data, but they are not gated identically (roadmap D-3). The REST API stays paid plans only (Growth or Scale: the free tier has no programmatic access to it at all). The MCP server serves twenty-three tools and is split by operation: its read tools (checks, status pages, incidents, maintenance, error projects and issues) work on every plan, free included; its write tools require Growth or Scale, the same rule as every REST route; and its seven outage tools (is_service_down, get_regional_readings, get_report_volume, list_recent_incidents, get_stack, internet_weather, is_the_internet_down) read RealUptime's own public measurements and are also served keyless at POST /public. Full write access on either surface, and REST access of any kind, remain the actual monetization hook, enforced in code (apps/web/lib/api-auth.ts, apps/mcp/index.ts), not just on the pricing page.

Authentication

Both surfaces use the same key, generated from the dashboard's API & MCP access section on any plan, free included. The key is shown once, at creation. It's stored hashed (SHA-256) server-side, so if you lose it, generate a new one and revoke the old.

Authorization: Bearer ru_live_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
FailureStatusSurface
Missing Authorization header401Both
Key doesn't exist, was revoked, or has expired401Both. An expired key is refused exactly as a revoked one is, with the same message and code, see Key expiry below.
Key belongs to a free-tier account403REST only. MCP instead allows the four read tools and refuses only the five write tools, see "Free-tier access" under MCP server below.
Key's permission scope doesn't allow the action (see Permission scopes below)403 / isError: trueBoth
Key carries an IP allowlist and the request came from outside it (see IP allowlists below)403, RU-1007Both
The account is suspended403, RU-1009Both. The key itself is otherwise valid; the response names the reason plainly ("This account is suspended. Contact support@realuptime.io for help.") rather than the generic invalid-key message, since the caller already holds a real credential. Checked before the free-tier gate, so a suspended free-tier account gets this message, not the free-tier one. Takes effect on the very next request; nothing about a key or session is cached across requests.

Permission scopes

Every key has a scope, chosen at creation from the dashboard and shown next to the key's name:

ScopeCan do
readList and get: GET /checks, GET /checks/:id, GET /status-pages, GET /status-pages/:id, GET /status-pages/:id/components, GET /notification-channels, GET /incidents, and the equivalent MCP read tools
read_writeEverything read can, plus create/update/delete: POST /checks, PATCH /checks/:id, PATCH /checks/:id/assertions, PATCH /checks/:id/auth, DELETE /checks/:id, POST/PATCH/DELETE on /status-pages, /status-pages/:id/components and /notification-channels, POST /incidents, POST /incidents/:id/updates, and the equivalent MCP write tools

The dashboard defaults new keys to read (the safer default). Pick read_write explicitly if the key needs to create or modify monitors, and if you're on Growth or Scale: a free-tier account can only create or hold read keys, since neither surface's write path works for a free account regardless of the key's own scope (apps/web/app/account/api-keys-actions.ts). Enforcement is server-side on every mutating REST route (requireWritePermission in apps/web/lib/api-auth.ts) and every mutating MCP tool (checkWritePermission in apps/mcp/index.ts), not just hidden in the dashboard form: a read key hitting a write route gets a clean 403 ({ "error": "This API key is read-only. Generate a read-write key to perform this action." }); the MCP equivalent sets isError: true on the tool result with the same message.

Keys created before scopes existed keep full read_write access (grandfathered, migration 018_api_key_permissions.sql). They were minted under an all-powers regime with no scope concept at all, so narrowing them retroactively would silently break whatever they're already wired into. Those keys are now identified as such, and can be narrowed in place: see "Legacy keys from before scopes existed" below.

Legacy keys from before scopes existed

Until REA-666, a key holding full read_write because a customer chose it and a key holding full read_write because migration 018_api_key_permissions.sql handed it over were the same row. api_keys now carries scope_origin (migration 202609061627_api_key_scope_origin.sql), which is chosen for every key created since scopes existed and legacy_default for a key backfilled as pre-scope. Changing a key's permission or its resource scopes settles it at chosen, whichever direction the change goes.

In the dashboard. A key that is still on legacy_default AND still holds read-write on every resource group carries a note in the API & MCP access list saying its access dates from before scopes existed, with two one-click paths out: narrow it to read, or give it access to specific resource groups instead. Both take effect on the very next request the key makes, both write their usual audit row (api_key.permission_changed, api_key.resource_scopes_updated), and both clear the note. Narrowing is reversible from the same row menu that has always been there.

The 30-day notice window. An account holding one or more of these keys is emailed once, ever, naming the keys by name and visible prefix and stating that we will ask again in 30 days. The notice is recorded as an api_key.legacy_notice row in the account's audit log, carrying the key ids, the window length and the instant it ends. Nothing about the key changes when the notice is sent: the keys keep working exactly as before, and RealUptime does not narrow, expire, or revoke a customer's key on its own. The window exists so a future forced-narrowing pass can point at the date the account was told, rather than changing a live credential with no warning behind it. That pass is not built.

Renaming a key (also from the API & MCP access section) never changes its scope or its plaintext value, only its display name. Renaming a key, changing its permission scope, and changing its expiry (see Key expiry below) after creation are all recorded in the account audit log, as api_key.renamed, api_key.permission_changed and api_key.expiry_changed respectively, each with the old and new value.

IP allowlists

Any key can optionally be restricted to a set of source addresses, set when the key is generated or later from the key's row menu (Edit IP allowlist) in the dashboard's API & MCP access section. Leave it blank (the default) and the key is valid from anywhere, exactly as before the feature existed.

  • Entries are IPv4 or IPv6, as a bare address (203.0.113.9, 2001:db8::1,

meaning /32 and /128) or a CIDR range (198.51.100.0/24, 2001:db8::/32). Up to 20 entries per key; separate them with newlines, commas, or spaces. Host bits are zeroed on save, so 198.51.100.77/24 is stored as 198.51.100.0/24.

  • The catch-all ranges 0.0.0.0/0 and ::/0 (any /0) are refused with a

message saying to clear the list instead: an allowlist that admits everyone is not a restriction and the dashboard should never show one as such.

  • Enforcement is server-side on BOTH surfaces, in the authentication step

before any route handler or MCP tool runs: the REST middleware (authenticateApiRequest in apps/web/lib/api-auth.ts) and the MCP server's authenticate (apps/mcp/index.ts). The address checked is the one Fly's edge proxy sets (Fly-Client-IP, the same header the failed-auth rate limiter trusts); X-Forwarded-For is never consulted, so a client cannot claim an allowed address by sending a header.

  • A request from outside the list answers 403 with code RU-1007

({ "error": "This API key is restricted to an IP allowlist and this request did not come from an allowed address.", "code": "RU-1007", ... }). The response never echoes the list. The refusal is counted by the per-IP failed-authentication limiter, so hammering a restricted key from the wrong network also earns a 429.

  • A restricted key whose request carries no determinable client address is

refused, not allowed: once a key is restricted, "could not tell where this came from" fails closed.

  • Changes apply on the next request: both surfaces read the list from the

key row on every call, with no cache. Every change is recorded in the account audit log as api_key.allowlist_updated with the before and after lists.

Typical use: a CI runner or integration host with a fixed egress address gets its own key restricted to that address, while a developer's laptop key stays unrestricted. Find the real public egress address from the host that makes the requests (for example curl https://api.ipify.org), not from the machine you are reading this on.

New-network notice for unrestricted keys (REA-669). A key with no allowlist is valid from anywhere, which means it produces no signal at all when it starts being used from a network it has never been seen on before. For a key with no allowlist, the first authenticated request from a network it has not used before (truncated to /24 for IPv4, /48 for IPv6) sends the account owner one notice, through whichever alert channels the account's default alert rule has configured, naming the key, the network, and the time. Later requests from the same network never repeat the notice, and the notice itself is capped at one per key per day even if several genuinely new networks appear in the same window; recorded in the account audit log as api_key.new_network_notice. A key that already carries an allowlist never gets this notice: it has already opted into a stricter control that refuses the request outright, so there is nothing left for this notice to add.

Resource scopes

Any key can additionally be scoped to a subset of resources (REA-475), set when the key is generated or later from the key's row menu (Edit resource scopes) in the dashboard's API & MCP access section. Leave it blank (the default) and the key stays a legacy global-scope key: its read or read_write permission (see Permission scopes above) applies to every resource, exactly as before this feature existed.

There are four resource groups:

ResourceCovers
monitors/checks and everything under it (including results, exports and heartbeat runs), /agents/:id/maintenance, /agents/:id/metrics/export, /capacity, /metrics, and the MCP check tools
status_pages/status-pages and its components, /incidents and /incidents/:id/updates, /maintenance, and the MCP status-page, incident and maintenance tools
errors/errors/projects, /errors/issues, the error-event export, POST /errors/releases, and the MCP error tools
account/notification-channels and other account-level configuration that isn't itself a monitor, status page, or error project

Reading incidents (GET /incidents, GET /incidents/export, MCP list_incidents) needs read on either status_pages or monitors: many incidents belong to a check with no status page.

A resource-scoped key is written as one resource:level pair per resource it needs, where level is read or read_write:

monitors:read_write
errors:read

This grants full access to monitors and read-only access to errors, and nothing at all to status pages or account configuration: once a key lists any resource, it loses access to every resource NOT listed, regardless of the permission chosen when the key was created. There is no partial-credit fallback to that global permission for an unlisted resource.

Enforcement composes with the Permission scopes check above rather than replacing it: requireWritePermission/requireResourcePermission (apps/web/lib/api-auth.ts) and checkWritePermission (apps/mcp/index.ts) both call the single hasResourceAccess rule in packages/db/api-key-scopes.ts, so REST and MCP can't drift on what a scope allows. Every route and tool is gated, reads included: a GET needs read on its resource and a mutation needs read_write. A resource-scoped key that lacks the resource gets a 403 with RU-1008, naming it: { "error": "This API key has no read access to errors. Grant it access from account settings.", "code": "RU-1008", ... } (MCP returns the same body as a tool error). A legacy global-scope key that is only read still gets RU-1004 on a mutation, as before. A key whose resource-scope list is empty is a legacy key, not a key scoped to nothing. apps/web/app/api/v1/read-scope-gate.test.ts and apps/mcp/tool-scopes.test.ts fail for a v1 GET route or an MCP tool added without its resource.

Every key created before this feature, and every key created without an explicit resource-scope entry, stores resource_scopes as null and is completely unaffected: this is additive on top of the existing global permission column, never a replacement for it. Changes apply on the next request (both surfaces read the column on every call, no cache) and are recorded in the account audit log as api_key.resource_scopes_updated with the before and after grants.

Key expiry

Any key can optionally be given an expiry, chosen when the key is generated or later from the key's row menu (Change expiry) in the dashboard's API & MCP access section. The choices are 30 days, 90 days, 1 year, and no expiry.

No expiry is the default and always has been. Every key created before this feature never expires, and a new key never expires unless you pick a length. Nothing about an existing key changed.

  • The expiry is a hard stop. Past it, both surfaces refuse the key with the

same 401 and the same RU-1002 message a revoked key gets ({ "error": "Invalid or revoked API key." }). That is deliberate: an expired key and a revoked key are indistinguishable from outside, so a caller holding a dead string learns it is dead and nothing further.

  • The comparison happens in the authentication query itself, against the

database's clock. There is no sweep and no grace period: the key works up to the instant, and not after it. Nothing sets revoked_at, so an expired key stays visible on your key list rather than vanishing, and you can revoke it, replace it, or give it a new date.

  • The account's contact address gets an email **7 days before, 1 day before,

and once on expiry**. The mail names the key and its visible prefix, never its value. Each of the three is sent at most once per key per date.

  • Changing the date resets those notices, so a key you extend earns its

warnings again against the new date. Clearing the date stops them entirely, and the key never expires again.

  • Changes apply on the next request (both surfaces read the key row on every

call, no cache) and are recorded in the account audit log as api_key.expiry_changed with the before and after instants.

Typical use: a key handed to a contractor or a short-lived pipeline gets 30 or 90 days, so nobody has to remember to clean it up; a key wired into a long-running production integration keeps no expiry, and the IP allowlist above is the tighter control for that case.

Who a key belongs to

A key belongs to the ACCOUNT, not to the person who generated it. It grants the same company-wide scope whoever holds it, every colleague with API-key permission can see and revoke it, and it keeps working if its creator's role changes.

The key list names whoever created each key, so you can tell which is which when several people share an account.

Removing someone from the team revokes the keys they created. This is immediate and cannot be undone: the next request on such a key gets a 401, and the fix is to generate a replacement and update whatever was using the old one. Offboarding a colleague is exactly when a credential they created should stop working, so if a shared integration was built on one person's key, move it to a key someone still on the team generated before removing them. Changing their role does not revoke anything.

Key governance

What exists today for managing a key over its lifetime, and what REA-621 identified as the gaps still open. This section is honest about which is which: the "Today" list is built and shipped, the "Target model" list is not built yet and carries no date.

Today

  • Permission scope (read / read_write), set at creation and

changeable after (see Permission scopes above).

  • Resource scopes, an optional per-resource-group narrowing on top of

the permission scope (see Resource scopes above).

  • A per-key IP allowlist, optional, enforced on every request on both

surfaces (see IP allowlists above).

  • `last_used_at`, shown on the key's row as "Last used" or "Never

used", updated only on a request the key was actually allowed to make (packages/db/api-keys.ts markApiKeyUsed).

  • Revoke, immediate on both surfaces, no cache in between.
  • Attribution: every key names who created it (created_by_user_id,

packages/db/migrations/057_api_key_created_by.sql), shown on the row, and removing that person from the team revokes every key they created (see Who a key belongs to above).

  • Audit coverage is partial, not complete. api_key.created,

api_key.revoked, api_key.allowlist_updated, and api_key.resource_scopes_updated are all written to the account audit log with before/after state where relevant (apps/web/app/account/api-keys-actions.ts, packages/db/audit-log.ts). Two mutations are not: renaming a key (renameApiKeyAction) and changing its permission scope between read and read_write (changeApiKeyPermissionAction) both write to the database with no matching audit row. A permission-scope change is the more material of the two gaps: it changes what a live credential can do, the same class of event api_key.allowlist_updated and api_key.resource_scopes_updated already log.

  • No account-level key inventory export. The account's own data export

(GET /api/account/export, backed by exportAccountData in packages/db/accounts.ts) returns checks, status pages, incidents, subscribers, and maintenance windows. It does not query api_keys at all, so a key's name, scope, allowlist, or last-used timestamp is not in that bundle. The customer audit log does have its own CSV export (apps/web/app/account/audit/export/route.ts) and that export does surface the four audited api_key.* events, but a log of past events is not a current inventory of live keys and their present settings.

  • Uplink (staff) can read a customer's key inventory, metadata only

(REA-667, shipped 2026-09-08). The account detail page renders ApiKeyInventorySection (apps/web/app/uplink/(console)/accounts/[id]/api-key-inventory-section.tsx) from listAllApiKeysForAccount (packages/db/api-keys.ts), which never selects key_hash, so neither the plaintext nor the hash can reach the page. It is read-only, lists revoked keys alongside live ones with their revoked_at, and is gated by the same capability check as the rest of that page. Do not confuse it with /uplink/keys (apps/web/app/uplink/(console)/keys/page.tsx), which is the signed-in staff member's own passkey list and has nothing to do with customer API keys.

  • Unused-key notices (REA-668). Once a day, a fleet job checks every live

key against two rules and emails the account owner about any that match: used before, but not in the last 90 days (last_used_at older than the threshold); or never used, and created more than 90 days ago (a brand-new key gets time to be wired up before it counts as unused). 90 days is the default threshold; an operator can change it with the API_KEY_UNUSED_THRESHOLD_DAYS environment variable. The same key does not get a second notice for 30 days after the first (API_KEY_UNUSED_SUPPRESS_DAYS), so an ignored notice repeats on a monthly cadence rather than daily, and a key that gets used again, or is revoked, simply stops qualifying on the next sweep. The notice links to the dashboard's key list (https://dashboard.realuptime.io/account/api-keys), never a one-click revoke: revoking a live credential from an email link needs a signed, single-use token this codebase does not have yet, so the email only ever points the owner at the page where they can look at the key and decide. Every notice is also recorded in the account audit log as api_key.unused_notice, alongside the key's id, its prefix, its last-used date (or that it was never used), and how many days it had gone unused -- never the key's hash or plaintext. Implementation: packages/db/api-key-unused-alerts.ts, wired into apps/probe on the same 24h cadence as the certificate/domain expiry sweep.

Target model (not built; no date)

  • Optional expiry on keys, with an automated nudge to the account

owner as the expiry date nears, on the pattern packages/checker/expiry-alerts.ts already established for TLS certificate and domain expiry (a threshold-crossing sweep enqueuing an email through email_outbox). No column, sweep, or nudge exists for API keys today; every key is valid until revoked.

  • A path to identify and narrow legacy pre-scope keys. Keys created

before packages/db/migrations/018_api_key_permissions.sql shipped keep full read_write access, described in Permission scopes above as "grandfathered." There is no column or flag distinguishing a legacy grandfathered key from a key someone deliberately chose read_write for after scopes existed; both look identical in api_keys. The only narrowing path today is fully manual: generate a new read key and revoke the old one, one key and one customer at a time, with nothing in-app identifying which live keys are candidates. A target build adds a way to identify them (by comparing created_at against the migration's deploy date) and a narrower or force-rotate path with advance notice, rather than only the manual regenerate-and-revoke available now.

  • Alerts on a key used from a new network when no allowlist is set.

The IP allowlist is opt-in and enforced only when a customer sets one; an unrestricted key (the default, and every pre-feature key) generates no signal at all when it starts being used from a network it has never been used from before.

  • Rename and permission-scope-change audit rows, closing the two gaps

named under Today above.

Rate limits

Per-API-key token-bucket budgets (packages/db/rate-limit.ts), shared by the REST API, the MCP server, and the dashboard's own check-creation form (the same limiter instances/keying scheme, not separate per-surface quotas):

BudgetLimitApplies to
Read120/minGET /checks, GET /checks/:id, GET /metrics, GET /capacity, GET /status-pages, GET /status-pages/:id, GET /status-pages/:id/components, GET /notification-channels, GET /incidents, GET /maintenance, MCP list_checks, MCP get_check_status, MCP list_status_pages, MCP list_incidents, MCP get_active_maintenance_window
Write30/minPOST /checks, PATCH /checks/:id, PATCH /checks/:id/assertions, PATCH /checks/:id/auth, DELETE /checks/:id, every POST/PATCH/DELETE under /status-pages, /status-pages/:id/components and /notification-channels, POST /incidents, POST /incidents/:id/updates, POST /maintenance, DELETE /maintenance/:id, MCP create_check, MCP update_check_regions, MCP update_check_assertions, MCP delete_check, MCP create_incident, MCP add_incident_update, MCP create_maintenance_window, MCP end_maintenance_window

A REST request over budget gets:

json
{ "error": "Rate limit exceeded. Try again shortly." }

with status 429 and a Retry-After header (whole seconds until the bucket has a token again). An MCP call over budget gets the same error message and a retryAfterSeconds field, but delivered as a tool result with isError: true. MCP tool errors aren't surfaced as HTTP status codes, so there's no 429 or Retry-After header on that surface (see MCP server below).

This is a single-process, in-memory limiter: budgets are per web/MCP instance, not cluster-wide. That is acceptable precisely because these two budgets are keyed by API key, an attributable and revocable identity that a single instance already sees enough of to self-throttle.

The two keyless public surfaces cannot rely on that, having no identity to key on, so both use the Postgres-backed limiter (packages/db/rate-limit-db.ts) and genuinely hold across the fleet:

BudgetLimitApplies to
Keyless outage JSON120/min per IPGET /api/v1/outages/:slug (RU-6003)
Keyless MCP60/min per IPPOST /public on the MCP host (RU-6004)

Both are per-IP rather than per-key, and the keyless MCP budget is deliberately the tightest of the four: see "Keyless access" under MCP server below for the sizing argument.

Unusual usage alerts

Rate limits cap traffic; they do not tell you when a key starts doing something it has never done. For that, each key's usage is compared with its own last 14 days, every 15 minutes, and a notice goes to the account's alert channels (the account's email address by default, plus Slack, PagerDuty, webhooks or any other channel the account routes alerts to) when the key's requests so far today (UTC) reach ten times its daily average and at least 2,000.

A key needs at least 7 days of use before it can trigger an alert. Each key alerts at most once a day. The email states today's number and the average it was judged against. Nothing is blocked or throttled by an alert: it is a notice, and revoking the key is your decision.

Usage

Every authenticated REST and MCP request is now recorded against the key that made it, bucketed by day and by the region that served it, with whether it ended in an error (a status of 400 or above, or an MCP tool error) and how many items it returned (a list's length, 1 for a single object, the rows of an export). It is a monitoring and support record, not a meter: nothing is billed from these numbers, and they change nothing about the budgets above or about what any plan includes. The recording happens off the request path, so it adds no latency and no failure mode to a call.

You can see it on the API keys page in your account settings, under "Usage, last 30 days". For each key that is live, or that made calls in the window:

  • requests per UTC day
  • the total
  • the split between the REST API and the MCP server

Counts land within about a minute of a call. A key that made no calls says so rather than showing zero.

REST API

Base URL: https://realuptime.io/api/v1 (or http://localhost:3000/api/v1 locally).

GET /api/v1

Unauthenticated endpoint index: no Authorization header required, since this returns a static list of routes, never account data. Safe to hit for discovery (including automated/LLM tool discovery) before you have a key.

bash
curl https://realuptime.io/api/v1
json
{
  "name": "RealUptime API v1",
  "documentation": "https://realuptime.io/docs/api",
  "authentication": "Authorization: Bearer <api key>. Growth/Scale plans only: generate a key from the dashboard's API & MCP access section.",
  "endpoints": [
    { "method": "GET", "path": "/api/v1/checks" },
    { "method": "POST", "path": "/api/v1/checks" },
    { "method": "GET", "path": "/api/v1/checks/:id" },
    { "method": "PATCH", "path": "/api/v1/checks/:id" },
    { "method": "PATCH", "path": "/api/v1/checks/:id/assertions" },
    { "method": "DELETE", "path": "/api/v1/checks/:id" },
    { "method": "GET", "path": "/api/v1/status-pages" },
    { "method": "GET", "path": "/api/v1/incidents" },
    { "method": "POST", "path": "/api/v1/incidents" },
    { "method": "POST", "path": "/api/v1/incidents/:id/updates" },
    {
      "method": "GET",
      "path": "/api/v1/outages/:slug",
      "auth": "none",
      "description": "Public outage readings for one tracked service, mirroring https://realuptime.io/outages/:slug. No API key. Optional ?region=iad|sjc|fra|nrt."
    },
    {
      "method": "GET",
      "path": "/api/v1/outages/stacks/:key",
      "auth": "none",
      "description": "Public reading for one stack page, mirroring https://realuptime.io/outages/stack/:key. No API key. :key is the stack slug or id."
    },
    {
      "method": "GET",
      "path": "/api/v1/outages/internet-weather",
      "auth": "none",
      "description": "Public internet weather, mirroring https://realuptime.io/outages/internet-weather: per-region p50/p95/p99 probe latency over the outage catalog, 7-day baseline and slower-than-usual state. No API key."
    }
  ],
  "mcp": "https://mcp.realuptime.io/mcp"
}

This index does not currently list the /api/v1/errors/* routes or the keyless /api/errors/v1/ingest|sourcemaps endpoints; they are documented in their own sections below regardless.

GET /checks

Every check row also carries paused_at (null while live; set when the operator has paused it, in which case no region probes it and it renders as Paused everywhere) and tags (a normalized lowercase string array, at most 10 of at most 32 characters). Both are readable here and writable from the dashboard only for now.

List every monitor on the account: every type, not only the ones REST/MCP can create (see "Monitor types beyond http" below).

bash
curl https://realuptime.io/api/v1/checks \
  -H "Authorization: Bearer ru_live_..."
json
{
  "checks": [
    { "id": "...", "account_id": "...", "name": "API", "url": "https://api.example.com/health", "type": "http", "tcp_host": null, "tcp_port": null, "tcp_tls": false, "interval_seconds": 60, "selected_regions": ["iad", "sjc", "fra", "nrt"] }
  ]
}

Monitor types beyond http

The dashboard can also create heartbeat, tcp, dns, smtp, and ping monitors, and this endpoint (and MCP's list_checks/get_check_status) returns ALL of an account's checks regardless of type: a customer with a TCP or heartbeat monitor sees it here too, not just their http ones. type is one of "http" | "heartbeat" | "tcp" | "dns" | "smtp" | "multistep" | "ping"; tcp_host/tcp_port/tcp_tls are non-null only on a tcp row (null/null/false otherwise), smtp_host/smtp_port/smtp_starttls only on an smtp row, and ping_host only on a ping row, matching the Check interface in packages/db/index.ts. url is null for heartbeat, tcp, smtp, ping, and multistep checks (a multistep check's targets live per-step, in check_steps; see GET /checks/:id below). Neither surface here exposes dns_hostname / dns_record_type / dns_expected_value for a dns check, or the heartbeat-specific grace period/ping-token fields on the check itself: listChecksForAccount and getCheckForAccount (packages/db/checks.ts) don't select those columns (a heartbeat's own fields are the heartbeat object on GET /checks/:id, and its runs are GET /checks/:id/runs), so a dns check's type field reads "dns" here with no way to see what it resolves. REST and MCP create_check can still only create http checks (see POST /checks below): the other six types are dashboard-only in v1.

The eleven assertion_* fields (Monitor Phase 4 plus REA-176's JSON-path group, below) ARE exposed by both listChecksForAccount and getCheckForAccount, unlike the dns fields above -- every field is null on a non-http row or an http row with no assertions configured.

What smtp checks, and why it isn't a tcp check on port 25

A tcp check proves a socket opens. It does not prove a mail server is answering: a listener that accepts the connection and then never speaks (an overloaded MTA, a load balancer in front of a dead backend, a greylisting tarpit) looks perfectly healthy to a connect-only probe. An smtp check reads the 220 greeting and completes an EHLO, and if smtp_starttls is set it also negotiates STARTTLS and requires the TLS handshake to finish. Recorded latency is the time to complete that conversation.

It never authenticates and never sends mail. There is no AUTH, MAIL FROM, RCPT TO, or DATA anywhere in the prober (packages/db/guarded-smtp.ts), so there is no code path that could be made to relay a message, and no credential of yours is ever presented to your server. Do not configure one; there is no field for it.

smtp_port is restricted to 25, 465, 587, or 2525, enforced in the input schema, in the SSRF guard, and as a database constraint. A tcp check may target an arbitrary port because it writes zero application bytes; an smtp check writes SMTP verbs, so it only writes them where a mail server is what is supposed to be listening. Port 465 is implicit TLS from the first byte, so smtp_starttls is ignored there rather than rejected.

What ping checks measure, and why it isn't raw ICMP

A ping check reports round-trip latency and packet loss for a bare host, with no port and no application protocol to configure -- the check for a router, a printer, or anything else that speaks neither HTTP nor a known TCP service. ping_host is non-null only on a ping row.

It is not raw ICMP. Our cloud probes run as an unprivileged process on Fly Machines (no CAP_NET_RAW, and no host-level net.ipv4.ping_group_range grant reachable from inside a Machine either), and Node has no built-in ICMP socket support regardless. Rather than silently downgrade without saying so, the transport is TCP-connect-time against whichever of a host's commonly-open ports (443, then 80, then 22) accepts a connection first, reusing the exact guarded, address-pinned socket path the tcp prober uses (packages/checker/ping-probe.ts; see that file's header for the full reasoning). A host that answers real ICMP but has all three of those ports closed reports down here -- for that case, a tcp check against the specific open port you know about is the more precise tool.

POST /checks

Create an http monitor: the only type this endpoint (and MCP's create_check) can create; heartbeat, tcp, dns, smtp, ping, and multistep monitors are created from the dashboard. Enforces the account's tier allowance (TIER_LIMITS in packages/db/checks.ts, an alias for INCLUDED_MONITORS in packages/db/allowances.ts: 10 on free (REA-186), 25 on Growth, 150 on Scale as of the 2026-08-15 pricing pass): every type counts as one monitor in that same allowance, multistep included.

What happens past that allowance depends on overage pricing (packages/db/allowances.ts), which is environment-gated and off by default:

  • Free is always a hard cap: overage requires a payment method, which a

free account has none of. Free always gets 403 once it hits its allowance.

  • Growth/Scale, with no overage rate configured

(MONITOR_OVERAGE_CENTS_PER_UNIT unset), behave exactly like free: 403 once the allowance is hit. This is the safe default and, unless the owner has priced overage, the only behavior in production today.

  • Growth/Scale, with an overage rate configured, can keep creating

monitors past the allowance and get billed per extra unit per month ($MONITOR_OVERAGE_CENTS_PER_UNIT/100 each), up to an optional ceiling: MONITOR_OVERAGE_MAX_UNITS (an absolute unit count) and/or MONITOR_OVERAGE_CEILING_MULTIPLIER (a per-tier multiplier of that tier's own INCLUDED_MONITORS (REA-148) -- e.g. 3 caps Growth's extras at 75 and Scale's at 450). Whichever of the two ceilings is lower wins; past it it's 403 again.

Either way the error body is the same: { "error": "Your <tier> plan allows up to <limit> monitors. Upgrade to add more." }: on the overage-ceiling path "upgrade" is not literally true (the ceiling is a flat environment setting, not tier-scoped), but no caller has ever hit it in production.

The limit check and the insert run inside one transaction guarded by a per-account Postgres advisory lock (pg_advisory_xact_lock, createCheck in packages/db/checks.ts), so concurrent create calls for the same account queue up one at a time instead of racing past the cap. A single-statement insert ... where (select count(*) ...) < limit was tried first but isn't safe under Postgres's default READ COMMITTED isolation, so it was replaced with the lock-then-count-then-insert sequence.

Billing sync gap (REST `DELETE /checks/:id`, MCP `create_check`, MCP `delete_check`): after a successful create, this REST route calls syncMonitorOverage (apps/web/lib/stripe-overage.ts) to keep the account's Stripe "additional monitor" line item in step with its real monitor count: same as the dashboard's create AND delete actions. REST's DELETE /checks/:id and MCP's create_check/delete_check do not call it, so a monitor created or deleted only through those three paths leaves Stripe's overage line item stale until the hourly reconciler (apps/web/lib/overage-drift-worker.ts) corrects it. That reconciler exists specifically to catch drift like this (it reports every correction as an operational error first), so this is a bounded, monitored gap rather than a silent one, but it means an MCP-only customer's bill can lag their real usage by up to an hour rather than updating inline like the dashboard's does.

bash
curl -X POST https://realuptime.io/api/v1/checks \
  -H "Authorization: Bearer ru_live_..." \
  -H "Content-Type: application/json" \
  -d '{"name":"API","url":"https://api.example.com/health"}'
FieldTypeRequired
namestringyes
urlstring (must include scheme, e.g. https://)yes
intervalSecondsnumberno, defaults to 60
regionsarray of "iad" | "sjc" | "fra" | "nrt" | "ord" | "yyz" | "lhr" | "sin" | "syd" | "gru"no, defaults to the four core regions (iad, sjc, fra, nrt); the other six are Growth and Scale only
latencyThresholdMsnumber, 1-10000no, defaults to null (no threshold)
assertionBodyOp"contains" | "not_contains"no
assertionBodyValuestring, the literal text to look forrequired if assertionBodyOp is set
assertionBodyCaseSensitivebooleanno, defaults to true
assertionHeaderNamestringno
assertionHeaderOp"equals" | "contains"required if assertionHeaderName is set
assertionHeaderValuestringrequired if assertionHeaderName is set
assertionStatusMinnumber, 100-599no
assertionStatusMaxnumber, 100-599, >= assertionStatusMinrequired if assertionStatusMin is set
assertionJsonPathstring, a dot/bracket path (e.g. data.items[0].status)no
assertionJsonOp"equals" | "contains" | "exists"required if assertionJsonPath is set
assertionJsonValuestringrequired for equals/contains, and rejected for exists

Returns 201 with { "check": { ... } } (the returned object also carries an internal isFirstCheck boolean, true only when this was the account's very first monitor, safe to ignore), or 400 if name/url are missing or the wrong type, name exceeds 200 characters, name/url contain a control character, url doesn't parse, regions is present but empty or contains a duplicate or unknown region, intervalSeconds is present but not a positive finite number, an assertion field is present without its required pair (see "Response assertions" below), the request body isn't valid JSON, or the target fails the safety check below.

Regions

Every check probes from the four core regions (iad/sjc/fra/nrt) by default (DEFAULT_CHECK_REGIONS in packages/db/checks.ts). Ten regions are live (LIVE_REGIONS in packages/db/region-catalog.ts); the six beyond the core four (ord, yyz, lhr, sin, syd, gru) are Growth and Scale only and are never added by default, so a Free check cannot silently gain regions its plan does not include. Pass regions at creation to choose which live regions probe this check; omitting it (or every check created before this field existed) keeps the core-four behavior exactly. At least one region is required: an empty array is rejected with 400. Each region may only appear once: a duplicate (e.g. ["iad","iad"]) is also rejected with 400. Both rules come from the same shared schema (regionsShape in packages/db/api-schemas.ts) enforced identically by the dashboard form and the MCP create_check/update_check_regions tools.

bash
curl -X POST https://realuptime.io/api/v1/checks \
  -H "Authorization: Bearer ru_live_..." \
  -H "Content-Type: application/json" \
  -d '{"name":"API","url":"https://api.example.com/health","regions":["iad","fra"]}'

To change an existing check's regions later, use PATCH /checks/:id (below). The scheduler (getDueChecks in packages/db/index.ts) only ever considers a check due in a region it selected; the public status page and its history bars render only a check's selected regions, never a phantom "no data" row for a region it was never checked from.

Latency threshold (REA-359)

latencyThresholdMs is an optional per-check "degraded above X ms" knob. null/omitted (the default) means no threshold, and behaves exactly like every check created before this field existed. When set, a region's probe that SUCCEEDS but takes longer than the threshold reads degraded for that region instead of operational -- the same degraded state a region already reaches when it goes down (computeRegionTransition in packages/db/index.ts), so this rides the existing status rollup, history and alerting with no new state.

Bounds are [1, 10000] milliseconds, matching the probe's own default timeout ceiling (packages/checker/index.ts): a threshold at or above the timeout could never fire, since a probe that slow has already failed rather than succeeded. An out-of-range value is rejected with 400 (RU-3010), not clamped -- silently moving a customer's threshold would quietly change which regions read degraded.

Storing a threshold is accepted on every tier; whether it's actually honoured at probe time is gated to Growth and up, the same paid-tier line response snapshots sit on (allowances.ts's tierAllowsLatencyThreshold). A Free-tier check may set the field, but no region will read degraded for it until the account is on a plan that enforces it -- re-checked on every probe, so a downgrade silently stops enforcement with no separate cleanup step.

bash
curl -X POST https://realuptime.io/api/v1/checks \
  -H "Authorization: Bearer ru_live_..." \
  -H "Content-Type: application/json" \
  -d '{"name":"API","url":"https://api.example.com/health","latencyThresholdMs":2000}'

To change an existing check's threshold later (or clear it back to null), use PATCH /checks/:id (below).

Response assertions (http checks only)

By default an http check is up when the status is 2xx. Response assertions (Monitor Phase 4, plus the JSON group added by REA-176) let it judge itself on more than the status code, in four independent groups -- set any, all, or none:

  • Body: assertionBodyOp ("contains" or "not_contains") plus

assertionBodyValue, the literal substring to look for. Never a regex -- the value is matched literally, never compiled as a pattern (a customer-authored regex evaluated inside our own fleet process on every tick is a ReDoS surface this API refuses to open). assertionBodyCaseSensitive defaults to true.

  • Header: assertionHeaderName, assertionHeaderOp ("equals" or

"contains"), and assertionHeaderValue, checked against the named response header (case-insensitive name, case-sensitive value).

  • Status override: assertionStatusMin/assertionStatusMax (both

100-599, max >= min) REPLACE the default 2xx-is-up rule with an inclusive range -- useful for a check that expects, say, a 401 from an auth-required endpoint. Set both to the same value for an exact status.

  • JSON field: assertionJsonPath addresses one field in a JSON response

body, and assertionJsonOp says what to require of it: equals and contains compare it against assertionJsonValue, and exists only requires the path to resolve (so it takes no value, and passing one is a 400 rather than being ignored).

Each group's fields must be set together or not at all: assertionBodyOp with no assertionBodyValue (or vice versa) is a 400, and likewise for the header trio, the status pair, and the JSON path/op pair.

##### JSON path syntax

Deliberately a small grammar, and everything outside it is an error rather than something quietly ignored:

FormMeans
status, data.statusobject keys
items[0], [0]array indices (non-negative integers)
["odd.key"], ['odd key']quoted key, for a key containing a dot or bracket
a leading $accepted and ignored, so $.data.id and data.id are the same path

At most 32 segments. Not supported, and refused with a named error: wildcards ([*], .*), filter expressions ([?(@.x > 1)]), and recursive descent ($..name). A parser that silently dropped a [*] would evaluate a different assertion than the one you wrote and then report its result as yours, which is worse than refusing it.

Why a path is allowed where a regex is not. The body assertion above is never a regex, on purpose (a customer-authored pattern evaluated inside our probe process every 60 seconds is a ReDoS surface). A path is not the same bet: a regex is a program whose cost is super-linear in the input, while a path is an address -- a finite list of literal segments walked once against an already-parsed object, costing O(segments) and nothing in the size of the body. The grammar above has no alternation, no repetition, and no branching, which is why it has no worst case to trigger.

Evaluation details worth knowing. The body is parsed as JSON, so a non-JSON body fails the assertion with "the response body is not valid JSON" (and, past the 256KB cap, says so specifically rather than blaming your JSON). equals compares against the value's text form and fails with a named reason if the path resolves to an object or array, which cannot equal a string. contains is a substring test against that text form, which for an object or array is its compact JSON -- so contains on a tags array matches an element inside it. For exists, a field explicitly set to null COUNTS as existing: the API said null on purpose, which is different from the key being absent.

bash
curl -X POST https://realuptime.io/api/v1/checks \
  -H "Authorization: Bearer ru_live_..." \
  -H "Content-Type: application/json" \
  -d '{"name":"Shop","url":"https://shop.example.com","assertionBodyOp":"contains","assertionBodyValue":"Add to cart"}'
bash
curl -X POST https://realuptime.io/api/v1/checks \
  -H "Authorization: Bearer ru_live_..." \
  -H "Content-Type: application/json" \
  -d '{"name":"API health","url":"https://api.example.com/health","assertionJsonPath":"data.status","assertionJsonOp":"equals","assertionJsonValue":"ok"}'

An assertion failure is a check DOWN with a distinct, honest reason (e.g. Assertion failed: body is missing "Add to cart") in check_results.error, flowing through the same hysteresis/incident/alert pipeline as any other down reason -- latency is still recorded. The body is read only when a body assertion is configured and the status check has already passed, capped at 256KB (truncated past that, evaluated against what was read). Evaluated identically for a fleet-run check and an agent-bound one.

Omitting every assertion field (or updating a check back to none) is exact current behavior: no body read, no assertion evaluated, status-only as before this feature existed.

To change an existing check's assertions later, use PATCH /checks/:id/assertions (below).

Monitor target validation

Every check-creation path (this endpoint, the dashboard "add monitor" form, and the MCP create_check tool) runs the target URL through the same guard (validateMonitorTarget in packages/db/target-guard.ts) before it's ever saved:

  • Scheme must be http:// or https://.
  • Port must be the scheme default or one of 80, 443, 8080, 8443.
  • The hostname (or every IP it resolves to, if it's not a literal IP) is

rejected if it's loopback, private, link-local, carrier-grade NAT, multicast, reserved/documentation, or the link-local metadata range (which covers 169.254.169.254). This is a summary of checks against the standard blocked IPv4/IPv6 ranges, not an exhaustive list here.

  • localhost, bare/no-dot hostnames, and hostnames ending in .internal,

.flycast, or .local are rejected outright, without a DNS lookup.

On a validation failure the dashboard returns a fixed generic message ("That URL can't be monitored. Use a public http:// or https:// address."); the REST API returns the validator's specific reason string verbatim (e.g. "That target resolves to a private or reserved address and can't be used."); the MCP tool call returns the same specific reason string and fails with isError: true.

This same check runs again at probe time, on every redirect hop (packages/checker/index.ts): a target that passed validation at creation can still redirect to a private address later (DNS rebinding, or an operator changing what the target redirects to), so each hop through up to 3 redirects (4 requests total) is independently re-validated. A redirect to an unsafe target, or exceeding the hop limit, fails the probe rather than following it.

intervalSeconds: the per-tier floor

intervalSeconds is partially validated, then either refused or clamped, depending on where it falls. The shared zod schema (checkCreateShape in packages/db/api-schemas.ts) rejects with 400 any value that isn't a finite, positive number: wrong type, NaN/Infinity, zero, or negative (e.g. intervalSeconds: -100 or "fast" both return 400 and never reach createCheck).

Every plan has its own minimum check interval (TIER_MIN_INTERVAL_SECONDS in packages/db/checks.ts):

PlanFastest interval
Free60s
Growth60s
Scale30s

A schema-valid request between the absolute technical floor (30s) and your plan's own floor is refused, not silently slowed down: POST /checks returns 400 with { "error": "Checks run no faster than every 60s on this plan. Upgrade to Scale for 30s checks.", "code": "RU-3006" } (see https://realuptime.io/kb/errors/ru-3006 for the full entry). This is deliberate: silently running the check slower than you asked would hide the reason from you, so the API tells you instead of guessing for you. validateIntervalForTier in packages/db/checks.ts is the single shared validator behind this refusal; the dashboard create form, REST POST /checks, and MCP create_check all call it before creating anything, so the refusal and its wording are identical on every surface.

An out-of-range value below the absolute 30s floor or above MAX_INTERVAL_SECONDS (86400, 24 hours) is still treated as obviously-wrong input and clamped rather than refused, same as before this ladder existed: rounded to the nearest integer, then bounded into [30, 86400]. Below 30 becomes 30; above 86400 becomes 86400; omitted becomes your plan's own floor (clampIntervalSeconds in packages/db/checks.ts). The REST API never returns 400 solely because a value lands in this clamp range; the saved check simply reflects the clamped value, which may differ from what you sent. The MCP create_check tool uses the same shared input schema and validator, so it refuses and clamps identically to the REST API.

Existing checks are unaffected by a plan change either way: intervalSeconds is set only at creation and there is no edit surface for it today, so a check created on Scale and then downgraded keeps running at its original interval.

GET /checks/:id

One monitor's current status, aggregated across the check's selected regions (up to 4) plus the raw per-region breakdown.

bash
curl https://realuptime.io/api/v1/checks/<id> \
  -H "Authorization: Bearer ru_live_..."
json
{
  "check": { "id": "...", "account_id": "...", "name": "API", "url": "...", "type": "http", "tcp_host": null, "tcp_port": null, "tcp_tls": false, "interval_seconds": 60, "selected_regions": ["iad", "sjc", "fra", "nrt"] },
  "status": "operational",
  "regions": [
    {
      "check_id": "...",
      "region": "iad",
      "current_state": "operational",
      "consecutive_fail_count": 0,
      "consecutive_ok_count": 1,
      "last_changed_at": "2026-08-04T01:29:02.228Z",
      "last_checked_at": "2026-08-04T01:29:02.228Z"
    }
  ],
  "locus_mix": { "public_regions": 1, "private_locations": 0 }
}

Note the two different "regions" concepts in this response: check.selected_regions is which regions this check is configured to probe from (see Regions above); the top-level regions array is per-region live status, and only ever contains entries for regions that have reported at least once. It can never contain an entry for a region outside selected_regions.

Where a reading came from

region is a fleet region code, or a private location's locus: agent: followed by the first eight hex characters of the id of the agent that produced it. A monitor bound to a private location has exactly one such entry and no fleet entries, because it runs from one machine the customer owns rather than from RealUptime's regions. A mixed monitor (Scale and MSP, private locations phase 2) runs from its selected regions and one private location, so it has entries of both kinds, and status counts both.

Two additional fields make that distinction explicit rather than leaving every consumer to re-derive it from a string prefix:

FieldAlways presentWhat it says
locus_mixyes{ "public_regions": N, "private_locations": M }, the mix the status above was computed over
private_locationsonly when M > 0[{ "locus": "agent:a1b2c3d4", "name": "prod-vpc-1" }], the customer's own name for each private locus
split_reasonyesFor a mixed monitor, which side is failing when exactly one is: "external_only" (every region down, the private location up: reachable from inside the network, failing from the internet) or "internal_only" (the reverse). null otherwise, including for every monitor that is not mixed and whenever a side has no fresh reading

This is a provenance rule, not a convenience (docs/private-probe-locations.md section 2.1). A 99.99% computed partly from a machine in the customer's own rack is a different claim than 99.99% from ten independent vantages, and an API that hands both back under one array without saying which is which invites the first claim to be published as the second. The stored locus keeps its agent: prefix permanently, so the name is a join this response performs for you and never a rewrite of the history.

Both fields are additive. A consumer reading check, status and regions sees exactly what it saw before.

status is operational (no fresh region down), degraded (some fresh regions down, or full down coverage isn't confirmed yet), down (every one of the check's selected regions is fresh and down), stale (at least one region has ever reported, but every report that currently exists is 3 minutes old or more. A selected region that's never been probed at all doesn't block this; it's simply absent from consideration, same as it is from the regions array), or unknown (no region has ever reported for this check). This is the same staleness-aware aggregation the public status page uses (aggregateStatusForDisplay in packages/db/index.ts). A region whose last report is more than 3 minutes old is excluded from the aggregation entirely, so a check that's actually down but hasn't been probed recently reports stale, not a stale operational; likewise down only fires when every one of the check's selected regions is confirmed fresh and down. For example, a check with all 4 live regions selected, 3-of-4 fresh-down with the 4th unprobed reports degraded, not down; a check configured via PATCH /checks/:id with fewer than 4 selected regions can reach down with correspondingly fewer fresh-down regions. regions may have fewer than 4 entries for a monitor that hasn't been probed from every region yet.

Returns 404 if the check doesn't exist or belongs to a different account. Ownership is always scoped to the authenticated key's account, there's no way to fetch another account's check by guessing its id.

Heartbeat monitors: the heartbeat object (REA-1052)

A heartbeat check's response also carries heartbeat, the fields the shared check shape does not (it is absent on every other type):

FieldTypeMeaning
state"operational" | "down"The monitor's own state. status above reads unknown for a heartbeat, which has no regions.
last_ping_atISO string or nullLast success ping. null means it has never been pinged (unarmed, not down).
grace_secondsnumberGrace after the interval before a missing ping counts as down.
max_runtime_secondsnumber or nullA started run that has not finished within this is stuck and takes the monitor down once. null is off.
latchedbooleanTrue while the monitor is held down after repeated failures (see "Runs" below).
latched_sinceISO string or nullWhen that hold began.
last_runrun or nullThe most recent run, in the shape GET /checks/:id/runs returns.

GET /checks/:id/runs

A heartbeat monitor's runs, newest first (REA-1052). A run is one execution of the job: /api/ping/<token>/start opens it, the plain ping (success) or /api/ping/<token>/fail closes it, paired by ?rid= when the job sends one, else with the most recent open run. A monitor that never sends /start and never fails has no runs, and behaves exactly as heartbeats always have.

bash
curl "https://realuptime.io/api/v1/checks/<id>/runs?limit=20" \
  -H "Authorization: Bearer ru_live_..."
QueryMeaning
limit1 to 500, default 50.
beforeISO 8601 timestamp: runs that started before it. Pass the previous page's next_before.
json
{
  "runs": [
    {
      "id": "6f1c...",
      "run_id": null,
      "started_at": "2026-10-02T03:00:00.000Z",
      "finished_at": "2026-10-02T03:01:30.000Z",
      "outcome": "success",
      "duration_ms": 90000,
      "stuck": false,
      "exit_code": null,
      "exit_detail": null
    }
  ],
  "visible_since": "2026-09-02T03:05:00.000Z",
  "next_before": null
}

outcome is running, success, fail, or superseded (a later /start without a run id replaced it: a job that starts again without finishing is taken to have abandoned the earlier run, which is never reported stuck). duration_ms is null until a run has both a start and a finish; a /fail with no open run is a finish-only row. stuck is true when the run outlived max_runtime_seconds. exit_code / exit_detail are what a failed run reported in its body (a bare integer is an exit code, anything else a message, cut to 500 characters). Reads are clamped to the plan's raw-history window (visible_since: 30 days on Free, 90 on paid plans), the same window raw check results get; runs are kept for 1095 days. 404 (RU-3003) if the check is not one of this account's heartbeat monitors; 400 (RU-3001) for a malformed limit or before.

When a failure takes the monitor down. A /fail, and a run that outlives max_runtime_seconds, take an operational monitor down at once and alert with their own wording ("run failed", "run stuck"), distinct from a missed ping; the next success recovers it. A monitor that has failed and recovered three times in 24 hours is held down on the next failure instead of flapping (latched): successes are still recorded, and it recovers once one full interval plus grace passes with a success and no failed or stuck run. Pausing the monitor clears the hold. A missed ping never counts toward it.

PATCH /checks/:id

A sparse patch: send only the fields you're changing, any field left out keeps its current value. Covers name, url, interval and which regions probe this check. To edit response assertions instead, use PATCH /checks/:id/assertions below -- a separate route on purpose, so updating one concern never requires resending the other.

bash
curl -X PATCH https://realuptime.io/api/v1/checks/<id> \
  -H "Authorization: Bearer ru_live_..." \
  -H "Content-Type: application/json" \
  -d '{"name":"API (renamed)","interval_seconds":30}'
FieldTypeRequired
namestring, 1-200 charsno
urlstringno
interval_secondspositive numberno
regionsarray of "iad" | "sjc" | "fra" | "nrt" | "ord" | "yyz" | "lhr" | "sin" | "syd" | "gru"no
latencyThresholdMsnumber, 1-10000, or null to clearno
maxRuntimeSecondsnumber, 1-86400, or null to turn off (heartbeat monitors only)no

At least one field must be present. name, url and interval_seconds run through the exact same checks POST /checks does: url is re-validated against the SSRF target guard (a public http:///https:// address), and interval_seconds is refused with RU-3006 if it's below the account's tier floor (rather than silently clamped). name/url/interval_seconds are scoped to http checks, same as the dashboard's edit form -- patching them on a tcp/dns/heartbeat monitor returns 404. regions behaves as before: missing/empty/containing an unknown or duplicate region is a 400. latencyThresholdMs follows POST /checks's rules above (bounds [1, 10000], refused with RU-3010 rather than clamped) with one addition: pass null explicitly to clear an existing threshold back to "off" -- omitting the field entirely leaves it unchanged, the same present-vs-absent distinction every other sparse-patch field on this route has. maxRuntimeSeconds (REA-1052) applies to heartbeat monitors only, a 400 (RU-3001) on any other type: a run that started and has not finished within it alerts once as stuck. null turns stuck-run alerting off. A heartbeat's response also carries the heartbeat object GET /checks/:id returns.

Returns 200 with { "check": { ... } } on success, 400 if the request body isn't valid JSON, has no recognized field, or fails one of the checks above, or 404 if the check doesn't exist or isn't owned by this account. Every accepted name/url/interval change writes a monitor.updated audit row. A regions change is picked up by the scheduler on its very next tick per region (no restart or propagation delay); the public status page and history bars stop rendering a dropped region on their next render (revalidate = 30, matching the page's existing ISR window), and reach every visitor within 60 seconds: expireTime in apps/web/next.config.mjs caps how long a shared cache may go on serving the render made before the change.

PATCH /checks/:id/assertions

Replaces an http check's response assertions (Monitor Phase 4) -- a FULL replace, not a sparse patch. An empty body ({}) clears every assertion; a body naming only a body assertion clears whatever header, status, or JSON assertion existed before. See "Response assertions" under POST /checks above for the field meanings.

bash
curl -X PATCH https://realuptime.io/api/v1/checks/<id>/assertions \
  -H "Authorization: Bearer ru_live_..." \
  -H "Content-Type: application/json" \
  -d '{"assertionBodyOp":"contains","assertionBodyValue":"Add to cart"}'

Returns 200 with { "check": { ... } } on success, 400 if the request body isn't valid JSON, an assertion field is present without its required pair, a status range has assertionStatusMax below assertionStatusMin, or a JSON assertion is malformed (unparseable path, exists with a value, or equals/contains without one), 404 if the check doesn't exist or isn't owned by this account, or 400 (not 404) if the check exists and is owned but isn't an http check -- assertions have nothing to attach to on a tcp/dns/smtp/heartbeat check. An assertion-consistency 400 (from this route or from POST /checks) carries "code": "RU-3004" alongside the human message; a malformed-JSON or schema-level 400 carries "code": "RU-3001".

PATCH /checks/:id/auth

Replaces the authentication a private-location http check sends (REA-1014, docs/private-probe-locations.md section 3.6). A FULL replace, like assertions: {} removes all of it, and removing it is allowed on any plan.

json
{
  "headers": [{ "name": "Authorization", "value": "Bearer ${SECRET:BILLING_API_TOKEN}" }],
  "userinfo": null
}

Every value is a reference template. ${SECRET:NAME} names a secret the private location reads on its own machine, from REALUPTIME_SECRET_NAME or the file named by REALUPTIME_SECRETS_FILE. The credential's value is never sent to RealUptime and never stored. userinfo is user:${SECRET:PASSWORD}, sent by the agent as HTTP Basic; the password part must be references only.

FieldRule
headers[].nameAuthorization, Proxy-Authorization, Cookie, X-Api-Key, or a name the location allows with REALUPTIME_AUTH_HEADERS. At most 8 headers, each once.
headers[].valueUp to 1024 characters, at least one ${SECRET:NAME} reference, no control characters. NAME is capital letters, digits and underscores, starting with a letter. Literal text beside a reference that looks like a credential (a long run mixing letters and digits) is refused.
userinfouser:${SECRET:PASSWORD}, or null. Not together with an Authorization header.

Returns { "check_id": "...", "auth": { "headers": [...], "userinfo": null } }. 400 RU-3001 for a malformed body or an unknown field, 400 RU-3013 for a configuration that breaks a rule above, 403 RU-3014 for a check that cannot carry authentication right now (not an http check that runs from a private location, the account below Growth on both Status and Monitor, or the location's agent older than 0.4.0), 404 RU-3003 if the check is not on this account. GET /checks/:id includes the same auth object on a check that has any. Each change writes a monitor.updated audit row with field: "auth", whose before and after states are the templates.

DELETE /checks/:id

Removes the monitor and its history. Returns 204 on success, 404 if not found/not owned by this account. Does not sync Stripe overage billing inline: see the "Billing sync gap" note under POST /checks above.

Exports: GET /checks/export and GET /checks/:id/results/export

Both download a file rather than returning a JSON envelope, and both are the one place in v1 that accepts a dashboard session as well as a bearer key: the Monitor board and each monitor's detail page link straight to them, so a signed-in customer on any plan (including Free, for whom the rest of v1 is gated) can export their own data. A request that sends an Authorization: Bearer header goes through the ordinary key gate (paid plan, read scope) and gets the API's error codes; a request with no bearer header is answered from the session cookie instead (401 with no session, 403 for a role without view).

Rate limit: 20 exports per minute per account, whichever way you authenticated, through the durable limiter (RU-2002, with Retry-After). This is separate from the per-key read budget.

?format=csv|json (default json) on both. Anything else is 400 RU-3001. Bodies are streamed; the JSON envelope is

json
{ "checks": [ ... ], "meta": { "exported_at": "...", "tag_filter": [], "rows": 12 } }

(results instead of checks for the history export), and the CSV has a header line, one row per record (arrays such as selected_regions and tags are space-joined, timestamps are ISO-8601 UTC), and a final # {...} comment line carrying the same meta object.

`GET /checks/export[&tag=a&tag=b]` exports every check's configuration: id, name, type, target fields, interval, regions, agent binding, the response assertions, paused_at, tags, and created_at. No secrets are ever included (a heartbeat's ping token is not exportable). The repeatable tag parameter narrows to checks carrying any of those tags, the same filter the board uses, so "export what I am looking at" is one link.

`GET /checks/:id/results/export` exports one check's raw result history (checked_at, region, ok, status_code, latency_ms, error), newest first. The window is what your plan can see (meta.since/meta.window_days, from the Monitor history ladder, capped at the plan's raw-result visibility window: 30 days on Free, 90 on Growth and Scale). Rows older than that are not exported.

They are, however, still there. RealUptime keeps its own measurements for up to three years on every plan, free included, and the plan window governs how far back you can read them rather than how long they exist. Moving to a plan with a longer window widens this export retroactively, with no backfill and nothing to request: the same call simply returns more.

404 RU-3003 for an unknown or foreign check id.

Paused monitors and tags themselves are dashboard features today (pause/resume, tag editing and bulk actions live on the Monitor board); the REST surface reads them back in GET /checks (paused_at, tags) and in these exports, and does not yet write them.

GET /checks/results: time-ranged, multi-check results query (REA-564)

The query-shaped twin of GET /checks/:id/results/export above, built for a polling consumer rather than a file download -- the motivating case is a Grafana panel calling $__timeFrom()/$__timeTo() on every repaint, which the export endpoint can't serve without pulling the account's entire retention window on every refresh. Same session-or-bearer auth as the exports, but its own rate-limit scope: 120 requests/minute per account (RU-2002, Retry-After), well above the exports' 20/minute since a dashboard with several panels legitimately polls every few seconds.

GET /checks/results?from=2026-08-01T00:00:00Z&to=2026-08-08T00:00:00Z&check_ids=<id1>,<id2>&region=iad&step=5m&limit=1000
ParamRequiredMeaning
fromyesISO-8601 start of the range (inclusive).
toyesISO-8601 end of the range (exclusive), must be after from.
check_idsnoComma-separated check ids. Omitted means every check the account owns -- a 40-monitor dashboard asks once, not 40 times. Up to 200 ids per call. An id belonging to another account, or that never existed, is silently excluded rather than erroring.
regionnoOne of the live probe regions. Omitted means every region.
stepnoDownsample interval: a number plus s/m/h/d (30s, 5m, 1h, 1d), between 10s and 7 days. Omitted returns raw rows.
cursornoOpaque keyset cursor from a previous response's next_cursor.
limitnoRow cap for this response: 1-5000, default 1000.

from/to are clamped to the account's Monitor retention window (resolveMonitorRetention/resultExportWindowForAccount, the exact function the export above uses) rather than rejected: a request reaching further back than the plan allows gets the tier's actual floor instead, and meta.clamped says whether that happened.

Without step, the response is raw rows (checked_at, region, ok, status_code, latency_ms, error, plus check_id so a multi-check page can be attributed):

json
{
  "results": [
    { "check_id": "...", "region": "iad", "ok": true, "status_code": 200, "latency_ms": 118, "error": null, "checked_at": "2026-08-01T00:00:30Z" }
  ],
  "next_cursor": null,
  "meta": {
    "from": "2026-08-01T00:00:00.000Z", "to": "2026-08-08T00:00:00.000Z",
    "requested_from": "2026-08-01T00:00:00.000Z", "requested_to": "2026-08-08T00:00:00.000Z",
    "clamped": false, "tier": "free", "history_days": 14,
    "check_ids": ["..."], "region": null, "step": null, "row_cap": 1000, "rows": 1
  }
}

With step, the response is downsampled buckets instead (check_id/region/bucket_start, sample_count, ok_count, p50_latency_ms, p95_latency_ms -- sample_count counts every result, ok_count only the passing ones, and both latency percentiles exclude results with no timing), under a "buckets" key instead of "results". Buckets align to the Unix epoch, not to from, so the same step always produces the same boundaries across repeated polls.

Pagination is keyset: when a range holds more than row_cap rows (or buckets), next_cursor names the next page; pass it back as cursor on the following request. next_cursor is null once a page reaches to with nothing left.

400 RU-3012 for a missing/invalid from/to, from not before to, an invalid check_ids entry or too many of them, an out-of-range step, an unrecognized region, a malformed cursor, or a limit outside [1, 5000].

Exports, the rest of the product line (REA-451)

The Monitor exports above were the first of these; every product now has a complete export path so nothing you build on RealUptime is stuck here. Same shape throughout: ?format=csv|json (default json), session or bearer key, streamed bodies, a meta object naming the export window, RU-2002/Retry-After on the shared data-export rate-limit scope (20 a minute per account, separate from the Monitor exports' own scope so the two cannot starve each other).

`GET /incidents/export` -- every incident across every check on the account, newest first. No retention window: incidents are kept indefinitely, so this is the account's complete incident history. CSV carries the summary row per incident (status, timing, update count); JSON additionally carries each incident's full update timeline, since a CSV row has no honest way to hold a nested list.

`GET /errors/projects/:id/export` -- one Errors project's issues and events. format=csv streams the issue summary table only (title, status, first/last seen, event count, release markers) -- a CSV row cannot hold an issue's nested events. format=json streams both issues (the durable record, never pruned) and events (bounded by the account's resolved Errors retention window -- meta.events_since/meta.events_retention_days match the same window the retention job enforces).

`GET|POST|DELETE /agents/:id/maintenance` -- planned maintenance on one monitor agent's server. POST with { "minutes": 30, "reason": "kernel update", "hosts": ["www.example.com"] } (minutes 1-1440, required; reason and hosts optional) opens a window, or extends the one already active; DELETE ends it now; GET returns it. Each answers { "window": { ... } }, or { "window": null } when nothing is active. While a window is active the server raises no offline or server-health alert, and every check that runs from the agent, or targets its reported host name or one of hosts, opens no incident, sends nothing and does not count against uptime. When it ends, anything still down is reported from that moment. Accepts a read_write API key, or the agent's own rua_ token with :id set to the agent's id or to self, so a server can flag itself; the agent does this with realuptime-agent maintenance. The rua_ token half answers identically on ingest.realuptime.io (and, on the internal stack, telemetry.realuptime.io): it is shared code, not a second implementation.

`GET /agents/:id/metrics/export` -- one monitor agent's raw server-metrics history (CPU, memory, load), newest first, bounded by the 7-day raw-sample retention window every tier currently holds (meta.since/meta.window_days).

Export everything. The account settings page's "Export everything" section is a manifest, not a single zip: the existing account-data bundle (GET /account/export from the dashboard) already covers accounts, checks, status pages, incidents, subscribers, API keys (metadata only, never a hash or plaintext), and an uptime summary in one JSON file, and the manifest adds direct links to the per-project (Errors) and per-agent (server metrics) exports above plus the standalone incident- history export, since those three are inherently scoped to a resource a flat account-wide file cannot represent cleanly. A streaming zip of every export in one download is a real feature for a later ticket, not a this-afternoon addition on top of four keyset-paginated generators.

Status subscriber lists export separately, at GET /dashboard/status-page/subscribers/export?page=<id> (CSV, dashboard session only): it is PII, gated on manage_status_pages, and every export writes a subscribers.exported audit-log row with the filters and row count, same as any other bulk read of personal data in this codebase.

GET /metrics

A Prometheus/OpenMetrics-compatible scrape endpoint (REA-565): the current state of every check the account owns, one sample per check per region. This is the only v1 route whose success body is not JSON. Deliberately not the same route as apps/web/app/api/agent/v1/metrics, which is an unrelated internal ingest endpoint for a different consumer and is untouched by this feature.

  • Auth: a bearer key on the same paid-tier gate as the rest of v1 (see

Authentication above), with read on monitors (see Resource scopes above).

  • Content-Type: text/plain; version=0.0.4; charset=utf-8 on success --

the classic Prometheus text exposition format promtool check metrics validates, and what Prometheus, VictoriaMetrics, Grafana Alloy, and Datadog's OpenMetrics check all scrape. An auth failure or rate limit still answers with the standard JSON error envelope; only a 200 is text.

  • Current state only. No history and no time-range parameter -- a

sibling issue (REA-564) covers historical queries. This endpoint always describes "right now", from the same region_state table (not raw check_results.ok) the dashboard's own status reads use, so a single transient failure that RealUptime's two-consecutive-failure hysteresis absorbs everywhere else does not flap realuptime_check_up here either.

  • Caching: the snapshot behind a scrape is cached per account for up to

10 seconds server-side. Prometheus's own documentation recommends a scrape interval of 15 seconds or slower for a target like this one; the cache exists to absorb overlapping requests within that window (an HA Prometheus pair scraping independently, a retried scrape, manual curl testing next to a live scraper), not to make scrapes slower than that interval stale.

Series:

MetricLabelsMeaning
realuptime_check_upcheck_id, check_name, region1 unless the region's current hysteresis state is down, else 0
realuptime_check_latency_millisecondscheck_id, check_name, regionLatency of the most recent completed probe from that region. Absent (not zero) when the last probe produced no reading
realuptime_check_status_codecheck_id, check_name, regionHTTP status of the most recent completed probe from that region. Absent for check types with no status code
realuptime_check_last_checked_timestamp_secondscheck_id, check_name, regionUnix timestamp (seconds) of the most recent completed probe from that region
realuptime_check_consecutive_failurescheck_id, check_name, regionConsecutive failed probes from that region since the last success
realuptime_incident_opencheck_id, check_name1 while a non-resolved incident (page-less or published) is attached to the check, across every region, else 0

region is a fleet region code (iad, sjc, fra, nrt, and the rest of the live catalog -- see "Regions" in docs/business.md) or an agent locus (agent:<id>) for an agent-bound check. A check with no completed probe yet, including a heartbeat monitor (the fleet never probes one at all), contributes no per-region sample, only realuptime_incident_open -- every series is present with a stable label SET on every scrape regardless of current data, so a PromQL query written today keeps working the day the account's first incident actually opens.

Cardinality. Labels are check_id, check_name, and region only -- never a URL, error string, or anything else unbounded. A scrape produces at most 5 * C * R + C samples, where C is the account's check count and R its average regions per check (five per-region gauges, plus one realuptime_incident_open per check regardless of region). Worked examples: a Scale account at its 150-monitor allowance, each check reporting from all ten live regions, tops out at 7,650 samples; a Growth account at its 25-monitor allowance, each check at its 8 entitled regions, tops out at 1,025.

Example scrape config (prometheus.yml):

yaml
scrape_configs:
  - job_name: realuptime
    scheme: https
    metrics_path: /api/v1/metrics
    static_configs:
      - targets: ["api.realuptime.io"]
    authorization:
      credentials: "<your read-scoped API key>"
    scrape_interval: 30s

VictoriaMetrics and Grafana Alloy accept the identical config shape. Datadog's OpenMetrics check instead takes openmetrics_endpoint plus a Bearer <key> header in its headers map -- consult its own prometheus.d/conf.yaml documentation for the exact field names.

Starter Grafana dashboard. A ready-to-import dashboard JSON -- checks-down and open-incident counts, an up/down table by check and region, and latency/consecutive-failure panels -- is served at /grafana/realuptime-monitor-dashboard.json. In Grafana, Dashboards -> New -> Import, paste that URL (or upload the downloaded file), and pick the Prometheus data source scraping this endpoint when prompted.

GET /capacity

Latest CPU, memory, and disk for every installed Monitor agent on the account, plus a 7-day daily trend (REA-681, SOC 2 A1.1). Same idea as GET /agents/:id/metrics/export above, reshaped for "every agent, current state and trended" instead of "one agent's raw history": production's own Uplink Platform page and its deep-health capacity signal read the internal instance's own three-host fleet this way, with a read-scoped key (UPLINK_HOST_CAPACITY_API_KEY, docs/runbooks/self-monitoring.md), but the route itself is a plain account-scoped read available to any account.

  • Auth: a bearer key on the same paid-tier gate as the rest of v1, read

or read_write.

  • No parameters. Every non-revoked agent on the account is returned.
json
{
  "hosts": [
    {
      "agent_id": "5f3d...",
      "name": "vps1",
      "vantage": "host",
      "reporting": true,
      "last_seen_at": "2026-09-06T12:03:00.000Z",
      "latest": {
        "sampled_at": "2026-09-06T12:03:00.000Z",
        "cpu_used_ratio": 0.18,
        "cpu_cores": 2,
        "memory_used_ratio": 0.61,
        "disk_used_ratio": 0.42,
        "disk_mount_point": "/",
        "load1": 0.32,
        "load5": 0.29,
        "load15": 0.25
      },
      "trend_daily": [
        {
          "bucket_start": "2026-08-31T00:00:00.000Z",
          "cpu_used_ratio_avg": 0.15,
          "cpu_used_ratio_max": 0.31,
          "memory_used_ratio_avg": 0.58,
          "memory_used_ratio_max": 0.63,
          "disk_used_ratio_avg": 0.40,
          "disk_used_ratio_max": 0.42
        }
      ]
    }
  ],
  "meta": { "generated_at": "2026-09-06T12:03:05.000Z", "window_days": 7 }
}

reporting is false, and latest is null, for an agent that has never sampled or whose last sample is older than the reporting-staleness window (5 minutes; the agent samples every 60 seconds) -- never a fabricated "healthy" reading for a host that has gone quiet. trend_daily carries whatever daily buckets exist inside the 7-day window (meta.window_days), oldest first; a young agent or one with gaps in its history returns a shorter (or empty) list rather than backfilling zeros. disk_used_ratio* is null when the host has no disk sample yet; otherwise it describes the single mount with the largest total capacity at the latest reading (in practice the root filesystem), the same mount used for the trend -- a secondary volume filling up would not surface here.

GET /status-pages

json
{
  "statusPages": [
    {
      "id": "...", "account_id": "...", "slug": "acme", "name": "Status",
      "custom_domain": null, "page_title": null, "page_description": null,
      "logo_url": null, "domain_status": "none", "domain_error": null
    }
  ]
}

Each page also carries an additive vendor_components array (REA-251), empty when the page has none:

json
"vendor_components": [
  {
    "id": "...", "display_name": "Payments (Stripe)", "source": "vendor:stripe",
    "vendor_slug": "stripe", "notify_subscribers": false,
    "outages_url": "https://realuptime.io/outages/stripe"
  }
]

A vendor component mirrors a tracked third party from the Outages catalog on the page, as RealUptime's own probe reading of that vendor (labelled third-party on the public page). This field is configuration, not the live reading: read the live verdict from GET /api/v1/outages/<slug> or the public page. Vendor components are managed in the dashboard; this API does not create them.

Returns every status page the account holds, oldest first. Free and Growth accounts include one page, so their response is the same 1-item array it has always been; Scale includes up to 10 (lead decision 2026-08-17), so a Scale account that added pages in the dashboard sees them all here. The array shape predates multiple pages and did not change, so this is not a breaking change for any existing client. The first element is the account's default page: the one signup auto-creates on the first monitor, and the one every surface that does not name a page explicitly still means. page_title/page_description/logo_url are null until a customer sets branding. custom_domain reflects the raw hostname column regardless of verification state; whether it's actually live is domain_status (one of "none"/"pending_dns"/"issuing"/"issued"/"error", see "Custom domains" below), not the mere presence of a value. domain_error is populated only when domain_status is "error".

The object deliberately says nothing about page password protection (see "Password-protected pages" below). The password hash is never selectable through this API, and the protection flag is not exposed either: this endpoint is authenticated as the account that owns the page, so it could safely report the flag, but adding it would create a second place for "is this page private" to be answered from, and the value of having exactly one is worth more than the field. Read protection state in the dashboard.

POST /status-pages

Creates an ADDITIONAL status page (REA-226; the Terraform provider's realuptime_status_page). Body: { "name": "...", "slug": "..." }; slug is optional (omitted, a neutral status-xxxxxx one is generated). Slugs are lowercase letters, digits and single hyphens, at most 63 characters, one global namespace, reserved words refused. Returns 201 { "statusPage": {...} } with the same entity as GET /status-pages plus the branding columns (hide_branding, accent_color, homepage_url, support_url, theme_preset, font_choice, locale). Refusals: 400 RU-5008 (malformed body or slug), 403 RU-2006 (the plan's page allowance is used: Free includes one page), 409 RU-5010 (slug taken, including by a rename redirect). Writes the same status_page.created audit row the dashboard does, with via: "api" and the key id.

GET /status-pages/:id, PATCH /status-pages/:id, DELETE /status-pages/:id

GET returns { "statusPage": {...} } or 404 RU-5009. PATCH is a partial update of name, slug and brand chrome through the same updateStatusPageBranding the dashboard uses; every field optional, an omitted field is untouched, an explicit null clears a nullable one:

json
{
  "name": "Acme Status", "slug": "acme",
  "pageTitle": "Acme platform status", "pageDescription": null,
  "logoUrl": "https://.../logo.png", "hideBranding": true,
  "accentColor": "#1f6feb", "homepageUrl": "https://...", "supportUrl": "https://...",
  "themePreset": "slate", "fontChoice": "inter", "locale": "en-GB"
}

Allowed values: themePreset in signal|midnight|slate|ocean|plum|graphite, fontChoice in inter|system|space-grotesk, locale in the twelve tags STATUS_PAGE_LOCALES lists (packages/db/status-page-branding.ts). Links must be https and public; the accent must be #rrggbb and readable on white. A slug rename keeps the old slug redirecting. Refusals: 400 RU-5008, 404 RU-5009, 409 RU-5010. DELETE removes the page and everything published on it (components, groups, incidents, maintenance windows, subscribers, custom domain row cascade); 204, or 409 RU-5011 for the account's last page, 404 RU-5009 otherwise. Custom domains and passwords are not exposed here, for the reasons above and under "Custom domains".

Status page components: /status-pages/:id/components[/:componentId]

A component is a check published on a page. GET .../components lists them in display order as { "components": [ { "id", "status_page_id", "check_id", "display_name", "description", "sort_order", "group_id" } ] }. POST takes { "checkId": "...", "displayName": "...", "description": "..." } (displayName defaults to the check's name) and returns 201 { "component" }; refusals 400 RU-5008 (malformed, or the check is agent-bound: those are private by design, see addComponentToPage in packages/db/checks.ts), 404 RU-5009 (page) / RU-3003 (check), 409 RU-5012 (the check is already on this page). GET/PATCH/DELETE .../components/:componentId read, partially update (displayName, description; null clears the description) and unpublish a component. Unpublishing leaves the check running; a maintenance window whose whole scope was that component is ended (active) or deleted (future) rather than silently widened to the page, the same rule deleting a check applies. Ordering and grouping are dashboard-only.

Notification channels: /notification-channels[/:id]

The channel registry (packages/db/notification-channels.ts): Discord, Microsoft Teams, Opsgenie, incident.io, FireHydrant and generic JSON endpoints. Email, Slack, plain webhooks and PagerDuty stay on alert rules and are not here. GET lists { "channels": [...] }; every entity is REDACTED: config_summary (a URL origin, "EU region, API key stored"), never config. POST takes { "type": "discord|msteams|opsgenie|incidentio|firehydrant|generic_json", "name": "...", "config": {...} } where config is { "webhookUrl" } for discord/msteams/firehydrant, { "url" } for generic_json, { "apiKey", "region": "us|eu" } for opsgenie, { "apiKey", "alertSourceConfigId" } for incidentio (its HTTP alert source's own API key and config ID, from incident.io's Alert Sources > HTTP settings); the type's own validator and the SSRF guard run, and the channel is created UNVERIFIED (verified_at null until a test send succeeds from the dashboard). Refusals 400 RU-3007, 403 RU-2007 (plan ceiling), 409 RU-3009 (duplicate name). PATCH /:id takes any of name, config (full replace; clears verification and the last test result) and enabled (false soft-disables, keeping alert-rule attachments); type is immutable. DELETE /:id hard-deletes the channel and its stored credential (204). 404 RU-3008 for an unknown or foreign id. Create and delete write the same notification_channel.added / .removed audit rows the dashboard does.

Opsgenie and incident.io are the two registry channel types with PagerDuty's exact alias/dedup-key resolve semantics: Opsgenie closes the alert its down alert opened, by alias, and incident.io posts a resolved status keyed on the same deduplication_key, rather than either one raising a second alert titled "recovered." Like the PagerDuty resolve event (see "PagerDuty notifications" below), their recovery event is not subject to scheduled-maintenance suppression: a maintenance window active at recovery time silences a NEW alert on these two channels, never the close of one they already opened. Discord, Microsoft Teams, FireHydrant and generic JSON have no such lifecycle and keep the same maintenance suppression Slack, operator email and webhooks do (REA-783, same class as REA-768's PagerDuty fix).

Where a status page lives

A status page is served at its own subdomain, at the root of that host:

SurfaceAddress
The pagehttps://<slug>.realuptime.io/
Uptime badge (SVG)https://<slug>.realuptime.io/badge.svg (?days=30 for a 30-day window)
Embeddable widgethttps://<slug>.realuptime.io/embed (?days=30)
Logo proxyhttps://<slug>.realuptime.io/logo
Atom feed (incidents + maintenance)https://<slug>.realuptime.io/feed.atom
RSS feed (incidents + maintenance)https://<slug>.realuptime.io/feed.xml
iCalendar feed (scheduled maintenance)https://<slug>.realuptime.io/maintenance.ics

The older realuptime.io/status/<slug> form (and its /badge.svg, /embed and /logo sub-paths) permanently redirects to the address above, so links, README badges, and embedded iframes pasted before this change keep working without an edit. The redirect is a 308, which preserves the request method, so a form POST to an old address completes rather than being downgraded to a GET.

Renaming a page's slug moves its address to the new subdomain, and the old subdomain redirects to the new one. That works across a rename chain the same way the old path form did.

Nothing else in RealUptime is reachable on a status subdomain: the dashboard, /login, the REST API and the marketing site all return 404 there. This is the same guarantee custom domains have, for the same reason: a hostname carrying a customer's brand should not also expose ours.

Password-protected pages

Growth and Scale accounts can put a shared password in front of their public status page. A visitor without it gets a password form instead of any status, and the page's uptime badge (https://<slug>.realuptime.io/badge.svg), embeddable widget (https://<slug>.realuptime.io/embed), Atom/RSS feeds (/feed.atom, /feed.xml) and iCalendar feed (/maintenance.ics) all return 404 for as long as the page is protected, so a private page never reports its status from a third-party site or app.

Like custom domains and webhook notifications, this is dashboard-only: there is no REST or MCP endpoint to set, change, or clear a page password, and free-tier accounts see an upgrade prompt instead of the form (enforced server-side in the dashboard action, not just hidden in the UI).

One asymmetry worth knowing, because it differs from every other paid status-page benefit: the tier gates SETTING a password and never gates honouring one. A page that is already protected stays protected after a downgrade to Free, and its owner can still turn protection off on any tier. A billing event must never publish a page a customer made private, and must never trap them on one either.

Page types: public, private, internal

Every status page has a type, set in the dashboard when the page is created (and changeable afterward): public, private, or internal. public is the default and today's behaviour, unchanged. Growth and Scale accounts can set a page to private or internal; Free cannot (same asymmetry as password protection above: the tier gates SETTING a restricted type, never honouring one already set).

  • `private` is for customers. A visitor gets in with the page's shared

password (see above) or, if the account has single sign-on configured with a verified domain, by signing in with an account on that domain. The page is withdrawn from search engines (noindex) and from robots.txt on its own hostname.

  • `internal` is for your own team. Only a signed-in member of the owning

account can read it -- there is no password bypass. It is withdrawn from search engines and robots.txt like private, and additionally renders no subscriber sign-up form and no "Powered by RealUptime" attribution, since neither makes sense on a page meant only for employees.

Like password protection and custom domains, this is dashboard-only: there is no REST or MCP endpoint to read or set a page's type. GET /status-pages says nothing about it for the same reason it says nothing about password protection -- read it in the dashboard.

Custom domains

Growth and Scale accounts can serve their status page at their own hostname (e.g. status.yourcompany.com) instead of <slug>.realuptime.io. Like webhook notifications, this is dashboard-only: there is no REST endpoint to attach or detach a domain, and free-tier accounts see an upgrade prompt instead of the form (enforced server-side in the dashboard action, not just hidden in the UI).

Setup flow

  1. From the dashboard's "Custom domain" section, enter a hostname. It's

validated as a plain public DNS name: no scheme or path, not an IP address, not realuptime.io or any subdomain of it, not an internal or reserved name (.internal, .flycast, .local, .localhost, .arpa, plus the RFC 2606 reserved TLDs .test/.example/.invalid, and known cloud metadata hostnames). Any hostname containing a non-ASCII character is rejected outright. There is no IDN/punycode-acceptance path in v1, so a visually-similar homograph domain simply never validates. Uniqueness is enforced across all accounts by a database constraint, not just an application-level check, so two accounts racing to attach the exact same hostname can't both win.

  1. Create a CNAME record for your hostname pointing to

realuptime-web.fly.dev.

  1. The dashboard polls Fly's certificate API on every page load and shows

one of five honest states: None (none, no domain attached, the default), Waiting for DNS (pending_dns, the CNAME hasn't been observed yet), Issuing certificate (issuing, DNS looks correct and Let's Encrypt issuance is in progress), Live (issued, serving your status page over HTTPS), or Error (error, something went wrong registering the domain; remove it and try again).

  1. Once issued, requests to your hostname serve exactly your status page

(/) and nothing else: the rest of RealUptime (dashboard, login, REST API, marketing pages) is not reachable through a customer's own domain. Canonical URLs and Open Graph metadata on that render use your domain, not <slug>.realuptime.io.

  1. Detaching a domain from the dashboard removes the Fly certificate

registration and clears the hostname, freeing it up to be claimed by any account (including a different one).

Why TLS issuance is the ownership proof

There is no separate domain-ownership challenge (a TXT record, an email link). Let's Encrypt only issues a certificate after confirming, via the CNAME, that the requester controls the hostname's DNS; reaching issued already proves that control. This is the same trust model most "bring your own domain" SaaS products use. A hostname that never gets pointed at realuptime-web.fly.dev simply stays at pending_dns forever: it never routes traffic and never gets a certificate.

GET /incidents

bash
curl "https://realuptime.io/api/v1/incidents?limit=50" \
  -H "Authorization: Bearer ru_live_..."

Optional ?limit= query param, 1–500, defaults to 100. An out-of-range or non-numeric limit returns 400 rather than being silently clamped or falling back to the default. This is shared with the MCP list_incidents tool via the incidentsListShape/incidentsListSchema definitions in packages/db/api-schemas.ts. Returns every incident across every monitor on the account, newest first. "Every monitor" is literal: since migration 065 this includes page-less incidents, opened for a monitor that is on no status page at all (the RealUptime Monitor shape, the internal-docs repo's monitor-plan.md). Those carry "status_page_id": null; everything else about them is identical, including acknowledgement and escalation. Before 065 a private monitor could go down without producing an incident at all, so this list simply had nothing to say about it. A client that assumed status_page_id was always a string should treat it as nullable.

json
{
  "incidents": [
    {
      "id": "...", "status_page_id": "...", "check_id": "...", "region": "nrt",
      "title": "Asia-Pacific is down",
      "body": "Our Asia-Pacific probe is reporting API as down. Other regions are unaffected.",
      "status": "investigating", "opened_at": "...", "resolved_at": null,
      "acknowledged_at": null, "acknowledged_by": null,
      "regions": [
        { "region": "nrt", "label": "Asia-Pacific", "failed_at": "...", "recovered_at": null },
        { "region": "fra", "label": "Europe", "failed_at": "...", "recovered_at": "..." }
      ]
    },
    {
      "id": "...", "status_page_id": null, "check_id": "...", "region": "iad",
      "title": "US-East is down",
      "body": "Our US-East probe is reporting Internal API as down. Other regions are unaffected.",
      "status": "investigating", "opened_at": "...", "resolved_at": null,
      "acknowledged_at": null, "acknowledged_by": null,
      "regions": [
        { "region": "iad", "label": "US-East", "failed_at": "...", "recovered_at": null }
      ]
    }
  ]
}

`regions` (REA-924) is every region failing under this one incident, and
each one's recovery. An incident is scoped to (status page, monitor): a
monitor probed from four regions that goes down everywhere is ONE incident
carrying four entries here, not four incidents, and its subscribers are
mailed once. `region` is unchanged and still a single region code: the
FIRST region that failed, which is the one the incident's own title and body
name. A region that has come back carries a `recovered_at`; the incident
resolves once every entry has one. A heartbeat monitor has no regional
dimension, so its incident carries `"region": null` and an empty
`regions`.

acknowledged_at/acknowledged_by (on-call, migration 054) record which PERSON on the account's team took responsibility for the incident and when; acknowledged_by is a users.id, or null if nobody has yet, or if the person who did has since left the team. Acknowledging and releasing an incident are dashboard-only: like page passwords, custom domains, and webhook endpoints, there is no REST or MCP way to set or clear them, so every incident returned here and by MCP's list_incidents / create_incident / add_incident_update carries the field but neither surface can change it.

POST /incidents

Opens a new incident against one of the account's own status pages and monitors (components), and writes the opening timeline update in the same transaction.

statusPageId is required and stays required, deliberately, even though migration 065 made incidents.status_page_id nullable in the database. An operator-authored incident is an announcement, and an announcement with no page has no audience; page-less incidents exist only so a private monitor's automatic outage can be acknowledged and escalated, and the prober is their only writer. The same rule applies to MCP's create_incident. If you want to open one against a private monitor, put the monitor on a status page first.

Requires the read_write scope; subject to the 30/min write rate limit.

bash
curl -X POST https://realuptime.io/api/v1/incidents \
  -H "Authorization: Bearer ru_live_..." \
  -H "Content-Type: application/json" \
  -d '{"statusPageId":"...","checkId":"...","region":"iad","title":"API is down","body":"Investigating elevated error rates."}'
FieldTypeRequired
statusPageIdUUIDyes
checkIdUUIDyes
region"iad" | "sjc" | "fra" | "nrt" | "ord" | "yyz" | "lhr" | "sin" | "syd" | "gru"yes
titlestring, 1-200 charactersyes, unless templateId is given
bodystring, 1-5000 charactersyes, unless templateId is given
templateIdUUIDno (REA-188)

title and body reject control characters (the same screen every other free-text field in this API applies). Validated through the shared incidentCreateWithTemplateShape/incidentCreateWithTemplateSchema definitions in packages/db/api-schemas.ts (an additive superset of the older incidentCreateShape/incidentCreateSchema, which MCP's create_incident tool below still uses directly), so REST and MCP agree on every bound both share.

`templateId` (additive, REA-188). Names one of the account's incident templates (managed under the status page's incident settings in the dashboard). When given, title and body become optional: the template's title pattern and initial-update body are expanded against the tiny {{component}}/{{time}} placeholder set ({{component}} is the check's own name, {{time}} is the current time in UTC) and used as the incident's title/body. An explicit title and/or body in the same request still wins over the template's for that field. A request with neither templateId nor both title and body returns 400. A templateId that does not belong to the account returns 404, the same as a foreign statusPageId/checkId.

checkId must be a component of `statusPageId`, and region must be one of the regions that check is actually probed from (regions on the monitor, see PATCH /checks/:id). Both are enforced, and both refuse with 404:

  • A check the account owns but which is not on that page would publish an

incident on a public page for a component that page does not list, while notifying nobody (the subscriber fan-out keys off the page's component list). Private monitor-product checks with no public page are exactly this case.

  • A region the check is not probed from is a claim with no observation

behind it, and the automatic recovery path can never clear it, because a region we do not probe never transitions.

Returns 201 with { "incident": { ... } } (same shape as the entries in GET /incidents, acknowledgement fields included) on success. Returns 400 if the request body isn't valid JSON or fails validation, or 404 if statusPageId/checkId don't both belong to the authenticated account, the check isn't a component of the page, or the region isn't probed: all four are enforced by the same insert that creates the row, not a separate lookup, so a foreign, mismatched or unpublished id is indistinguishable from "not found" and never reaches a 500.

Sends email. On success, every confirmed subscriber of that status page gets an incident notification email. This is the first REST write with that side effect: every other write route only touches monitor rows. There is no way to suppress it per-request; if you don't want subscribers notified for a given incident, don't call this endpoint until you're ready for that notification to go out.

POST /incidents/:id/updates

Posts a staged update against an existing incident, advancing its lifecycle status (investigating -> identified -> monitoring -> resolved) in the same call. Requires the read_write scope; subject to the 30/min write rate limit.

bash
curl -X POST https://realuptime.io/api/v1/incidents/<id>/updates \
  -H "Authorization: Bearer ru_live_..." \
  -H "Content-Type: application/json" \
  -d '{"status":"identified","body":"Root cause found, deploying a fix."}'
FieldTypeRequired
status"investigating" | "identified" | "monitoring" | "resolved"yes
bodystring, 1-5000 charactersyes
forcebooleanno, see below

Validated through the shared incidentIdShape/incidentUpdateCreateShape definitions in packages/db/api-schemas.ts, shared with the MCP add_incident_update tool below. Posting resolved stamps the incident's resolved_at; posting any other status after a resolved incident clears it (re-opening the incident).

Resolving while the monitor is still failing

Behavior change, 2026-08-17. No API version bump. Resolving over a live outage used to succeed silently here; it now needs force.

status: "resolved" returns 409 when our own probes currently read the incident's component down or degraded, unless the request body carries force: true. With force: true the call proceeds exactly as it did before. Every other status ignores force entirely, and no other endpoint is affected, so a caller that was not resolving over a live outage needs no change.

The 409 body names what we see, in the same words a person would read:

json
{
  "error": "Our last probe still reports Checkout API from iad down. Marking this incident resolved tells every subscriber and every status page visitor that the outage is over. Send force: true to resolve anyway.",
  "check": "Checkout API",
  "region": "iad",
  "state": "down"
}

state uses the same staleness-aware vocabulary as GET /checks/:id, and the same aggregation produces it. Only down and degraded trigger the conflict: operational, stale (the region stopped reporting) and unknown (it never reported) are not observations we can hold against the caller, so they resolve without force.

This exists because the dashboard has always asked the same question before publishing Resolved, and a status page saying Resolved over a live outage is the one claim this product cannot afford. It is a prompt, not a policy: an operator can legitimately know the fix landed before the next probe round confirms it, which is what force is for. A degraded component counts as live for exactly the same reason a down one does.

The automatic path cannot clean this up afterwards, which is why the check happens before the write: the prober only opens an incident on a transition INTO down, and a check that is already down never transitions.

Returns 201 with { "update": { ... } } on success. Returns 400 for a malformed body, 404 if the incident id isn't a valid UUID or doesn't belong to the authenticated account, 409 for the resolve conflict above. Concurrent updates to the same incident are serialized server-side (a transaction-scoped advisory lock, see addIncidentUpdate in packages/db/incidents.ts); an identical update (same status, same body) submitted twice within 10 seconds is treated as a double-submit and returns the first update's row again rather than creating a duplicate timeline entry.

Sends email. On success (and not deduped as a double-submit), every confirmed subscriber of the incident's status page gets a notification email with the new status and update text. Same no-suppression caveat as POST /incidents above.

Inbound incident webhook

POST /api/v1/incidents/inbound/<token> opens, updates and resolves incidents on one of your status pages from an alerting tool: Prometheus Alertmanager, Grafana, Datadog, New Relic, or anything that can POST JSON. The token in the URL is the credential. It is minted under Account settings > Integrations > Inbound incident webhooks, shown once, and can be rotated or revoked there; each webhook is bound to one status page and carries a default component and region for payloads that name neither. No API key is involved and the plan tier is not consulted. A suspended account's token reads exactly like an unknown or revoked one (404), matching how a revoked token is already handled: nothing here distinguishes any of the three.

bash
curl -X POST https://realuptime.io/api/v1/incidents/inbound/<token> \
  -H "Content-Type: application/json" \
  -d '{"key":"checkout-5xx","status":"firing","title":"Checkout API returning 5xx","body":"5xx ratio is 7.2% over 5m.","component":"Checkout API","region":"fra","severity":"critical"}'

The generic shape

One event, or { "events": [ ... ] } for up to 50. Only key and status are required; extra fields are ignored, so a vendor template can carry its own alongside.

FieldTypeRequiredMeaning
keystring, 1-200yesYour stable identity for the alert. A later resolved with the same key closes what this opened; a repeated firing with the same key updates instead of duplicating. Aliases accepted: id, alert_id, alertId, fingerprint.
status"firing" | "resolved"yesAliases accepted in status, state or event_type: triggered/open/alerting/critical/warning/active/problem read as firing; ok/closed/recovered/normal read as resolved.
titlestring, up to 200noIncident title. Falls back to key. Aliases: summary, name.
bodystring, up to 5000noFirst update text. Falls back to the title. Aliases: description, message, text.
componentstringnoA check NAME or id that is a component of the webhook's status page. Unmatched hints fall back to the default and are recorded in the audit row. Aliases: check, service.
regioniad | sjc | fra | nrtnoMust be a region the component is probed from; otherwise the default (or the component's first region) is used.
severitystringnoPrefixed onto the update text, e.g. [critical]. Alias: priority.

What each event does

PayloadExisting link for this keyResult
firingnoneopens an incident (status investigating), remembers the link
firingopen incidentposts an update if the body changed, otherwise no-op
firingresolved incidentopens a NEW incident and repoints the link
resolvedopen incidentposts the resolved update
resolvedresolved / noneno-op

Every accepted event is audit-logged (inbound_incident.received, with the action taken and the incident id); every refusal is too (inbound_incident.rejected, with the reason). Subscribers are emailed exactly as they are for a dashboard-posted incident.

Adapters: Alertmanager and Grafana

Point Alertmanager's webhook receiver, or a Grafana webhook contact point, at the URL unchanged. The body is recognized by its alerts array (Grafana is told apart by its orgId / title + message extras) and every alert becomes one event: the fingerprint is the key (older Alertmanagers without one get a hash of the sorted labels), annotations.summary is the title (else alertname on instance), annotations.description plus the generatorURL is the body, the severity label is the severity. Two optional labels steer where the incident lands:

yaml
# alert rule
labels:
  realuptime_component: "Checkout API"   # a check name or id on the webhook's page
  realuptime_region: "fra"               # iad | sjc | fra | nrt

Alertmanager's group_by and repeat_interval are respected by design: one request can carry many alerts, and a repeated notification of the same firing is a no-op rather than a second incident. Force an adapter with ?source=alertmanager or ?source=grafana if a proxy rewrites the body; ?source=generic forces the generic parser.

Datadog and New Relic

Both let you write the webhook body yourself, so the generic shape is the adapter. Datadog webhook integration, payload template:

json
{"key":"$ALERT_ID","status":"$ALERT_TRANSITION","title":"$EVENT_TITLE","body":"$EVENT_MSG","severity":"$ALERT_PRIORITY"}

($ALERT_TRANSITION is Triggered / Recovered, both understood.) New Relic workflow, webhook destination, custom payload:

json
{"key":"{{issueId}}","status":"{{state}}","title":"{{issueTitle}}","body":"{{annotations.description.[0]}}","severity":"{{priority}}"}

({{state}} is ACTIVATED / CLOSED; send firing / resolved literally if your template language can branch, otherwise map CLOSED to resolved with a conditional.) Add "component" and "region" to either template to target a specific component.

Signing (optional)

Create the webhook with Also mint a signing secret and every request must carry X-RealUptime-Signature: sha256=<hex> where <hex> is the HMAC-SHA256 of the raw request body under that secret. Sign the exact bytes you send; re-serialized JSON fails. Rotating the token rotates the secret.

Responses

Returns 200 with one result per event, in order:

json
{"source":"alertmanager","results":[{"key":"a1b2c3","action":"opened","incidentId":"…"},{"key":"ffff","action":"ignored","incidentId":null,"reason":"resolved for an alert we never saw fire"}]}

action is opened, updated, resolved, ignored or refused. A whole payload of ignored is still 200: nothing to do is a successful delivery from the sender's point of view, and a 4xx would make it retry. Errors: 400 RU-5005 (not JSON, or no recognized shape; error says which field), 401 RU-5006 (signing secret set, signature missing or wrong), 404 RU-5004 (unknown or revoked token, indistinguishable from outside), 429 RU-5007 (more than 60 requests per minute for one token; Retry-After says when). Responses are Cache-Control: no-store.

Issue-tracker integrations (outbound)

The reverse direction has no API of its own: Create ticket on an Errors issue page or on an incident in the dashboard files a ticket into a GitHub repository, GitLab project, Jira project, Linear team, Bitbucket repository or Azure DevOps project connected under Account settings > Integrations, stores the link, and shows it on the issue or incident. One ticket per issue per connection: pressing it twice returns the same link. The KB how-to per provider lists the token scope each needs.

Maintenance windows

A maintenance window suppresses alerts for one of your status pages (or a subset of its components) without pausing the underlying checks: history stays continuous, and the public page says "under maintenance" instead of "down". This is the right primitive for planned work; pausing a check (there is no API for that today) would instead lose the readings taken during the window.

Every route below reuses the exact same createMaintenanceWindow/ updateMaintenanceWindow code the dashboard's own maintenance form calls (packages/db/maintenance.ts), so subscriber notification, alert suppression and the status page's render behave identically whether the window was scheduled here or by hand.

The one difference from the dashboard: this API never creates an open-ended window. The dashboard lets an operator explicitly choose "no end time"; a caller through the API has no UI to notice a forgotten window gone stale, so POST /maintenance requires an end, one way or another (see below). A window left open forever is worse than no maintenance window at all, because it silently keeps suppressing every alert for the page.

GET /maintenance

Returns the window currently active for one of your status pages, or null if nothing is. Requires the read scope.

bash
curl "https://realuptime.io/api/v1/maintenance?statusPageId=..." \
  -H "Authorization: Bearer ru_live_..."
Query paramTypeRequired
statusPageIdUUIDyes

Returns 200 with { "window": { ... } | null }. Returns 400 if statusPageId is missing or not a UUID.

POST /maintenance

Schedules a new maintenance window. Requires the read_write scope; subject to the 30/min write rate limit.

bash
curl -X POST https://realuptime.io/api/v1/maintenance \
  -H "Authorization: Bearer ru_live_..." \
  -H "Content-Type: application/json" \
  -d '{"statusPageId":"...","title":"Database maintenance","body":"Applying a security patch.","durationMinutes":30}'
FieldTypeRequired
statusPageIdUUIDyes
titlestring, 1-200 charactersyes
bodystring, 1-5000 charactersyes
startsAtISO 8601 timestampno, defaults to now
endsAtISO 8601 timestampone of endsAt/durationMinutes, not both
durationMinutesinteger, 1-43200 (30 days)one of endsAt/durationMinutes, not both
componentIdsarray of UUIDno, omitted or empty means the whole page
noDowntimebooleanno, defaults to false

noDowntime: true makes the window planned work: while it is active, a component it covers going down opens no incident, sends no alert or subscriber email, and the time is neutral in public uptime and the day bars; the component reads "Maintenance" on the page and its mirror, and the page's /status.json reports its status as unknown, the value a manual component in maintenance also reports there. A component still down when the window ends is reported from that moment, never backdated. Omitted or false, the window behaves as every window always has: operator alerts muted, incidents and downtime still recorded. The field is returned on every window as no_downtime.

Exactly one of endsAt/durationMinutes must be given; when durationMinutes is given, endsAt is startsAt + durationMinutes. Validated through the shared maintenanceCreateShape/maintenanceCreateSchema definitions in packages/db/api-schemas.ts.

Returns 201 with { "window": { ... } } on success. Returns 400 if the body isn't valid JSON, fails validation, or gives neither/both of endsAt/durationMinutes. Returns 404 if statusPageId, or any id in componentIds, doesn't belong to the authenticated account.

Sends email. On success, every confirmed subscriber of that status page gets a maintenance notification email, the same as scheduling one in the dashboard.

DELETE /maintenance/:id

Ends an active window early, by moving its ends_at to now. Requires the read_write scope; subject to the 30/min write rate limit.

bash
curl -X DELETE https://realuptime.io/api/v1/maintenance/<id> \
  -H "Authorization: Bearer ru_live_..."

This is a plain edit through updateMaintenanceWindow, not a delete of the row: the window and its history stay, only its end time moves. Alert suppression ends immediately. Sends exactly one subscriber email: the "maintenance complete" notice, the same one an operator gets from editing a window's end time to "now" in the dashboard -- not the "scheduled" re-announce, which is skipped for an edit that lands a window in the past.

A window that hasn't started yet, or is already over, is returned unchanged: there is nothing to end early about either one.

Returns 200 with { "window": { ... } }. Returns 404 if the window id isn't a valid UUID or doesn't belong to the authenticated account.

Platform health (public, keyless)

GET /api/health/db

Whether the machine that answered can reach the database right now: one select 1 through the shared pool, bounded at 5 seconds. It is an operational endpoint for the platform's own self-heal and for operators, not part of the versioned /api/v1 surface, and its shape may change without notice.

json
{ "ok": true, "db": true, "latency_ms": 4, "machine": { "id": "3d8d9e2b651e89", "region": "iad" } }

Returns 200 when the query succeeds. Returns 503 with "ok": false, "db": false and an error naming the failure class when it does not. machine is the id and region of the Fly machine that served the request, or null outside Fly, so repeated calls show which machine is unhealthy where an aggregate health answer would blur it. Responses are never cached.

Outages (public, keyless)

GET /api/v1/outages/:slug

The keyless JSON mirror of one https://realuptime.io/outages/:slug page (the internal-docs repo's outages-plan.md). No Authorization header, no account, no rate-limit budget shared with the bearer-authed routes above: bound instead by a per-IP limiter plus a hard cache (Cache-Control: public, max-age=120, s-maxage=120, stale-while-revalidate=600), with Access-Control-Allow-Origin: * since this is public measurement meant to be quoted from anywhere.

Deliberately per-service only: there is no index route and no ?all= parameter, so a caller names the one service it wants. Optional ?region=iad|sjc|fra|nrt narrows regions_read to that region.

bash
curl https://realuptime.io/api/v1/outages/github

Response fields: schema_version (currently 1, so a consumer can pin), service, verdict, a one-sentence summary, measured_at, freshness, regions_read (only the regions asked about), facets (the per-function breakdown), plate (the page's headline level and clause), last_outage, reliability (12 calendar months), coverage (regions_monitored, regions_not_monitored, measurement_depth, first_observed_at), recent_events, report_volume (the user-report second signal, available: false with a reason when no report data applies, never blended into verdict), and source (url, methodology_url, publisher, attribution, method).

recent_events (REA-219) is the probe-detected outage event history: { available: true, window_days: 90, limit: 10, tracking_since, events, method }, where events lists the events that started in the last 90 days, most recent first, at most 10, plus the open event if there is one. Each event carries id, started_at, ended_at (null while open), open, duration_ms / duration_minutes (so far, while open), peak_level (down when the primary surface failed from every monitored region, otherwise partially_down), worst_status, the peak regions, the distinct facets, the per-facet per-region legs (each with its own started_at / ended_at), scope, intermittent, reopen_count, and detected_by, which is always "probe" (the database admits nothing else). An empty events array means no event in that window since tracking_since, not a clean history before we watched the service; last_outage still names a closed event older than the window. available: false with a reason is returned only when the answer was built without the event read, and is never an empty list.

A service RealUptime does not track returns 404 with a real body (verdict: "not_covered", error code RU-6001) rather than an empty one: "we do not measure this" is the honest answer, not silence that could be mistaken for "healthy". A malformed slug or region returns 400 (RU-6002); exceeding the per-IP rate limit returns 429 (RU-6003) with a Retry-After header and Cache-Control: no-store.

This endpoint and the MCP is_service_down tool read the same buildOutageAnswer function, so the two surfaces can never answer the same question differently.

GET /api/v1/outages/stacks/:key

The keyless JSON mirror of one "my stack" page, https://realuptime.io/outages/stack/:key (REA-13, Outages Phase 6). A stack is a named set of tracked services, built anonymously at /outages/stack or from a Monitor dashboard; :key is its slug (a slugified name plus a random suffix, e.g. acme-prod-k7f3q2m9) or its id. Same posture as the per-service endpoint: no Authorization header, the same per-IP limiter and cache headers, Access-Control-Allow-Origin: *. There is no list endpoint: a caller holds the key or it does not.

bash
curl https://realuptime.io/api/v1/outages/stacks/acme-prod-k7f3q2m9

Response fields: schema_version (1), stack (id, slug, name, url, published, member_count, created_at), summary (level, the worst member's plate level; label; tone; headline, e.g. "2 of 6 services down or degraded"; detail; and counts of down, degraded, awaiting_data, operational, total), services (one entry per member: slug, name, category, url, short_url (the rlup.io/down/<slug> form), measured_at, freshness, plate with the same level / label / tone / clause vocabulary as the per-service answer, open_event, last_outage, coverage), and source.

Every member's level is computed by the same @realuptime/db/outage-plate pipeline the member's own page runs, over the same rows; the roll-up counts members by tone and is never an average. The signal layer (user reports, public chatter, the vendor's claim) is not read for a stack and is published as null under plate.signals, so reachable_reports_elevated cannot appear here; a member's own page still shows it. published only says whether the HTML page is indexable: a stack is public by URL either way.

An unknown key returns 404 with a real body (summary.level: "not_found", error code RU-6005), which says nothing about any service's health; a malformed key returns 400 (RU-6002); over budget returns 429 (RU-6003). The MCP get_stack tool reads the same buildOutageStackAnswer.

GET /api/v1/outages/internet-weather

The keyless JSON mirror of https://realuptime.io/outages/internet-weather (REA-52, the public internet-latency observatory). Per live probe region (iad US-East, sjc US-West, fra Europe, nrt Asia-Pacific): the latest hour's p50/p95/p99 response time across the ~195 public services in the outage catalog, the region's 7-day baseline, and a hysteresis-gated state. Same posture as the other keyless endpoints: no Authorization header, the same per-IP limiter, Access-Control-Allow-Origin: *, cached five minutes to match the page's ISR. Takes no parameters.

bash
curl https://realuptime.io/api/v1/outages/internet-weather

Response fields: schema_version (1), measured_at (the newest hourly bucket across regions, or null), regions (one entry per live region: region, label, state of "normal" / "slower" / "no_baseline" / "no_data", a plain sentence, latest_hour with bucket_start, p50_ms, p95_ms, p99_ms, sample_count, service_count or null, baseline_7d with p50_ms, p95_ms, hours or null, ratio_to_baseline, slower_for_hours, and sparkline_24h of 24 hourly {bucket_start, p50_ms} with null where an hour was under the sample floor), daily_30d (one row per region per UTC day with a daily bucket), coverage (tracking_since, days, service_count, region_count, typical_claim_allowed, typical_claim_unlocks_after_days), methodology (one paragraph) and source (url, methodology_url, publisher, data_is_not).

What the numbers are: percentiles of response times RealUptime's own probes measured from a named fleet region against the public, unauthenticated endpoints of the catalog's services, over successful (2xx) readings with a recorded latency only. What they are not: any single service's latency, a crowd report, or customer monitoring traffic, which never enters this aggregate. state is "slower" only after three consecutive hours with a median above 1.5x the 7-day baseline, and clears on the first hour at or under 1.25x, so a region hovering at the line does not flap. Do not describe a figure as "typical for the region" unless coverage.typical_claim_allowed is true (the record is at least 30 days old and at least three services stand behind the latest day); say "compared with the last 7 days" instead. The MCP internet_weather tool returns the same buildInternetWeatherAnswer.

Errors

The RealUptime Errors surface (the internal-docs repo's errors-plan.md, Phase 3; release metadata REA-185). Same bearer authentication and paid-tier gate as every route above, plus one extra requirement: the account must hold the Errors product (a 403 with "This account does not have RealUptime Errors." otherwise). The three GET routes spend the read rate limit; the one write route, POST /errors/releases, spends the write rate limit like every other mutation on this API. Issue status, grouping rules, and project management stay in the dashboard: this write route exists only for release announcements, the one thing a CI or deploy script needs.

An honesty note that applies to every response here: total counts are EXACT (counted in the same query shape as the returned page, never estimated), and event counts always travel with the period's drop counters (droppedOverQuota / droppedRateLimited / droppedMalformed / droppedClient) so a consumer can never quote a count without its asterisk being in the same response.

GET /errors/projects

The account's Errors projects (id, name, key hint, revoked flag), the quota line (tier, eventsPerMonth, retentionDays), and usageThisPeriod with the full drop counters.

GET /errors/projects/:id/issues

Issue search, the same filters as the dashboard, as query parameters: q (substring over title and culprit, case-insensitive, max 200 chars), release (issues that saw at least one event tagged with it), status (open | resolved | ignored), window over last-seen (24h | 7d | 30d | all), tag, limit (1–100, default 100). Answers { project, filters, total, truncated, issues, retentionDays, usageThisPeriod }; truncated: true means issues is a page of total. Out-of-vocabulary filter values are rejected with a 400, never silently coerced.

tag is key:value for an exact match, or a bare key for "the tag is set at all"; the value is split at the FIRST colon, so a value containing colons (a URL, a timestamp) survives intact. It matches issues with at least one RETAINED event carrying the tag, and that boundary is real: tags live on events, and raw events are pruned at the plan's retention window while issue records outlive them. An issue whose only tagged occurrences have aged out will not appear under a tag filter even though an unfiltered search still lists it. retentionDays in the same response is what tells you how far back the filter could have looked.

GET /errors/issues/:id

One issue in full: the issue row (status, counts, first/last seen, release marks, regression mark), its per-release breakdown, and recent raw occurrences with their scrubbed context and breadcrumb trails (each occurrence carries breadcrumbs_dropped, the count of older entries its bounded trail is not showing). An empty occurrences past the plan's retention window is stated via retentionDays, not silence.

POST /errors/releases

Announces a release's deploy/commit metadata, so a release created implicitly (from the release tag on an ingested event) or explicitly by this call gains: an optional git commitSha, a repoUrl, a deployedAt timestamp, and a link to the release it supersedes. This is the one write route on this surface, meant for a CI or deploy script.

Body: { "projectId": "<uuid>", "release": "v2.4.1", "commitSha": "a1b2c3d", "repoUrl": "https://github.com/acme/widgets", "deployedAt": "2026-08-22T00:00:00Z", "previousRelease": "v2.4.0" }. Only projectId and release are required.

  • Idempotent on `(projectId, release)`. Calling this twice for the same

pair updates the same release (created: false the second time) rather than creating a duplicate -- safe to retry from a flaky CI step.

  • Fields you omit are left alone; a field you explicitly send as null

clears it. This lets a first call announce just commitSha and a later call add repoUrl without erasing the first.

  • `previousRelease` names an ALREADY-registered release by its version

string (not an id) within the same project; naming one that doesn't exist yet is refused with RU-4008. Omit it and realuptime infers the most recently first-seen OTHER release for the project -- correct for a pipeline that announces every release in order, and the reason the field is optional.

  • Response: `{ "release": { "id", "projectId", "release", "firstSeen",

"lastSeen", "commitSha", "repoUrl", "deployedAt", "previousReleaseId" }, "created": true }, 201 on first creation, 200` on an update.

  • Honesty scope: realuptime never clones, fetches, or scans the

repository at repoUrl. It is stored exactly as given and used only to build a compare-range LINK (/compare/shaA...shaB on GitHub, /-/compare/shaA...shaB on GitLab) on the dashboard's release and issue pages -- "first seen in v2.4.1; view the diff range," never "caused by." An unrecognized host renders no link at all rather than a guessed one.

See `kb/announce-releases-from-ci` for copy-pasteable curl and CLI examples.

Inbound release webhook: POST /errors/releases/inbound/<token> (Vercel deploy-hook connector)

A second door into the same write path, for a deploy platform that cannot be scripted into calling POST /errors/releases directly. The token in the URL IS the credential -- minted once per project under Errors > Settings for that project, shown once, and rotatable. No bearer key or product tier is checked: the token was minted by someone who could already call POST /errors/releases for that project, and this path does nothing that person could not.

Point Vercel's project Settings > Webhooks (or an Integration Console webhook) at the URL, event Deployment Succeeded (deployment.succeeded; deployment.promoted and deployment.ready are also accepted, since they carry the same "this is now live" fact from a different trigger). Every event body Vercel documents (https://vercel.com/docs/webhooks/webhooks-api) is accepted as-is; nothing is templated. The release string registered is the deployment's own URL (falling back to its dpl_... id when there is none); commitSha and repoUrl are read from payload.deployment.meta's GitHub/GitLab/Bitbucket fields when Vercel sent them, and left null for a deploy with no linked repo -- nothing here is guessed. Idempotent on (project, release), the same as the direct route: a redelivered webhook re-registers the same row.

An optional signing secret, minted alongside the token, requires every POST to carry X-RealUptime-Signature: sha256=<hex HMAC-SHA256 of the raw body> -- this is realuptime's OWN secret, not Vercel's native x-vercel-signature header, which this connector does not check; Vercel's own webhooks cannot be configured to sign with an arbitrary secret, so this is only useful behind a relay that can. Response: { "source": "vercel", "release": "my-app-abc123.vercel.app", "releaseId": "<uuid>", "created": true }, 201/200 like the direct route. Refusals: 404 RU-4013 (unknown or revoked token), 400 RU-4014 (an event type that isn't a live deployment, or a body that isn't Vercel's documented shape), 401 RU-4015 (signature missing or wrong, only when a secret is set), 429 RU-4016 (rate limited, per token). Each occurrence also carries its wire-v2 context (see "Event context" just below): user_context, tags, context, device, and context_dropped. All four maps are null on an event that did not carry them, including every event stored before the v2 fields existed. context_dropped is the number of identity/tag/context entries an ingest size cap dropped from that event: non-zero means the maps you are reading are a SUBSET, and the number says how much of one.

SDKs

Five first-party SDKs speak the ingest wire contract. All five are MIT, dependency-free (standard library only beyond PHP's ext-json/ext-curl), scrub PII client-side by default against the same scrub-vectors.json fixture, and are pinned to the same limits and the same event/batch field names by their own wire-contract tests. JS, Python, and Ruby publish to their real package registries (npm/PyPI/RubyGems) via scripts/publish-sdk-registries.mjs (REA-272); each also keeps a public GitHub mirror, published by scripts/publish-sdk-mirrors.mjs (REA-209), as a secondary install path for pinning an exact commit. PHP is registered on Packagist but has no tagged release yet, so it still installs from its mirror only. The Go module needs no such stopgap: a Go module installs straight from a public git repo with plain go get, no registry publish step ever.

SDKSourceInstallIntegrations
JS/Nodepackages/errors-jsnpm install @realuptime/errors (or npm install github:RealUptimeHQ/realuptime-errors-js)Next.js, Express, generic wrap() for serverless handlers
Pythonpackages/errors-pypip install realuptime-errors (or pip install "git+https://github.com/RealUptimeHQ/realuptime-errors-py")Django, Flask, FastAPI/Starlette, Celery
Rubypackages/errors-rubygem "realuptime-errors" (or git: "https://github.com/RealUptimeHQ/realuptime-errors-ruby")Rails (Railtie: middleware, ActiveJob, request context), Rack, Sidekiq
Gopackages/errors-gogo get github.com/RealUptimeHQ/realuptime-errors-gonet/http middleware, log/slog handler
PHP/Laravelpackages/errors-phpcomposer require realuptime/errors:dev-main against the mirror (Packagist has no tagged release yet)Laravel service provider, exception handler, failed-queue-job capture

The batch's sdk field names which one sent it (realuptime-errors-js/<v>, realuptime-errors-py/<v>, realuptime-errors-ruby/<v>, realuptime-errors-go/<v>, realuptime-errors-php/<v>), and device.runtime names the runtime (node, python, ruby, go, php).

Event context: identity, tags, custom context, device

The SDKs can attach four additive things to an event (setUser / setTag / setContext in the JS SDK, set_user / set_tag / set_context in the Python and Ruby ones, SetUser / SetTag / SetContext in the Go one). They are optional everywhere: an event that carries none of them is byte-for-byte the event the ingest endpoint has always accepted, and an SDK that predates them keeps working untouched.

  • `user` is { id, email, username }. id is meant to be your own

opaque identifier and rides as-is. `email` and `username` are `[scrubbed]` by default, replaced inside your process before the event is serialized. Opt one back in by exact name in SDK config (allowFields: ["user.email"]), per field, in code. There is no server-side switch for this, deliberately: a default that a dashboard toggle can disable is a default that leaks before anyone notices. An opted-in value still passes the pattern scrub, so opting a field back in can never smuggle a raw token through.

  • `tags` is a flat string map of short labels (tenant, plan,

framework). Values pass the pattern scrub; keys do not, because a key is a label you chose and scrubbing it would make a tag filter depend on whether the value beside it looked like a card number.

  • `context` is a flat string map of free-form application state. Same

scrub treatment as tags. Read with the occurrence, never searched.

  • `device` is { runtime, runtimeVersion, platform, arch }, read from

your own process. It never contains a hostname, a local IP, or a container id: those identify a machine rather than a runtime, and add nothing release and environment do not already say.

Size caps, and what happens at them. At most 20 tags and 20 context entries per event, keys up to 64 characters, values up to 512, and a combined 4KB budget across identity, tags, context and device. Going over does NOT cost you the error report: the entries that fit are kept, the rest are dropped, and the number dropped is stored on the event and shown on the issue. Only an absurd payload (more than 200 entries in one map) is refused outright, with RU-4009 and reason: "oversize_context".

Billing: an event carrying all of this is still ONE event. Quota and metered overage count events, not bytes. The context makes an event bigger, and the caps bound how much bigger, but it can never make an event cost more than an event without it.

Affected-user counts (REA-526)

"1,000 events" says nothing about how many people were actually hurt; "40 users hit this" is a decision. When events on an issue carry user.id (above), the issue tracks how many DISTINCT identifiers it has seen and shows it as affected_user_count in the API and "affects N users" in the dashboard -- an issue list can also sort by it (sort=impact on GET /errors/projects/:id/issues, alongside the default sort=recent).

The count is exact up to affected_user_count_capped: once one issue has seen more than 5,000 distinct user identifiers, tracking new ones stops for that issue (the number this product shows never silently drifts stale without saying so), affected_user_count_capped becomes true, and the dashboard renders the number as a floor ("5,000+ users") rather than an exact count. This is a Sentry-basic, not-APM feature: it is not an exact distinct count computed live over an issue's full retained history (that does not scale, and it would shrink as retention prunes old events), it is a bounded per-issue record of user identifiers maintained at ingest time.

The identifier is stored exactly as sent (user.id already rides the wire unhashed and unscrubbed by design -- see above), not hashed again for this feature. There is no separate opt-in: any account already sending user.id gets impact counts for free, retroactively from the first event after this feature shipped (existing issues start at 0, since their pre-existing raw events do not get replayed into this new counter).

Source-map upload (ingest-key authed, not bearer-authed)

POST /api/errors/v1/ingest/<project_key>'s sibling, POST /api/errors/v1/sourcemaps/<project_key>: uploads JS source maps for one release using the project's DSN ingest key (the same single secret a CI pipeline already has -- write-only, one project). Body: { "release": "v2.4.1", "files": [{ "name": "assets/app.js.map", "content": "<map JSON as a string>" }] }. Limits: 10 files per request, 2MB per file, 50 files per release (a 409 with reason: "release_cap" past it); re-uploading a name replaces it. Malformed uploads are refused whole with a 400. Frames that cannot be mapped at display time render raw with a stated reason -- see the Errors product docs.

User feedback (ingest-key authed, opt-in, REA-257)

POST /api/errors/v1/feedback/<project_key>: the JS SDK's opt-in "what happened?" dialog (init({ feedback: {...} }), packages/errors-js/ feedback.ts) posts here, attaching a user's own words to the exact occurrence they describe. Same DSN-style ingest key as /api/errors/v1/ingest/<project_key>, rate-limited and size-capped separately (a report is a human at a keyboard, not a buffered event batch).

Body: { "eventId": "<client-minted uuid>", "comments": "...", "name": "...", "email": "...", "sdk": "realuptime-errors-js/0.4.0" }. eventId and comments are required; name, email, and sdk are optional. Body capped at 16KB, refused whole over that with a 413.

Two independent switches gate acceptance, both required:

  • The project's feedback toggle, off by default, turned on per project

under Errors → Settings.

  • The SDK's own init({ feedback: {...} }) opt-in -- an event only ever

carries the client-minted id this endpoint needs when the integration turned feedback on client-side, so a project that flips its toggle on later still can't attach reports to older events.

Refusals: 404 for a junk-shaped or unknown key; 403 with reason: "key_revoked" for a revoked key (the same loud, machine-readable refusal ingest uses); 403 with reason: "feedback_disabled" when the project's toggle is off; 404 with reason: "event_not_found" when the event id does not resolve (usually the event batch still in flight ahead of the report -- the SDK's transport retries this one specifically); 429 with reason: "rate_limited" over budget.

comments, name, and email are scrubbed client-side by the SDK before they leave the page (the same pattern rules as everything else) and scrubbed again server-side as the second net. Reports show up on the issue detail page under "User feedback," with a count on the issue's row in the issues list.

Sentry SDK compatibility (REA-280)

GlitchTip's move: any of Sentry's ~100 official and community SDKs can point at realuptime unmodified, no re-instrumentation required. A Sentry DSN has the shape https://<key>@host/<project_id>; the SDK derives two POST endpoints from it and carries the key in X-Sentry-Auth or the sentry_key query param.

  • POST /api/<project_id>/envelope/ -- current SDKs, Sentry's

newline-delimited envelope format (spec).

  • POST /api/<project_id>/store/ -- older SDK versions and some

captureMessage fallbacks, a bare event JSON object.

Both authenticate the same way our native ingest route does: our own ingest key (rue_<hex>, ERROR_KEY_SHAPE) already is a project-scoped, safe-to-expose public credential, so the DSN we hand out uses it in BOTH slots -- https://<key>@ingest.realuptime.io/<key> for a newly created or rotated project (an existing project's already-issued DSN keeps working against realuptime.io, since both hosts serve the identical handler; see `docs/infrastructure.md`) -- and {project_id} alone is sufficient to authenticate; a sentry_key present in the header or query must agree with it. Auth failures answer 401 (Sentry SDKs treat 401 as "stop trying," same as our native route's 404/403); a rate-limited request answers 429 with Retry-After, reusing the exact ingest rate limiter and drop counters the native route uses.

Translation happens once, at the edge (apps/web/app/api/sentry-translate.ts): a Sentry event's exception, message/logentry, breadcrumbs, tags, contexts, user, release, and environment become the identical ErrorEventInput our own SDKs produce, then flow through the exact same ingestErrorEvents transaction (apps/web/app/api/sentry-common.ts). Grouping, PII scrub rules, quota, the per-project ceiling, and alerting all apply identically -- the same error reported through Sentry's SDK or ours groups into ONE issue, never two.

What does not translate: transaction envelope items (Sentry's performance/tracing product, which realuptime does not ship) are accepted and discarded, answering success without being stored or billed. attachment items are refused with a clear 415 rather than silently dropped, since attachments are not stored today. gzip- and deflate-compressed bodies (both common from Sentry SDKs) are decompressed before parsing, under the same size ceiling as a native ingest batch.

Find a project's Sentry-compatible DSN under Errors → Settings, revealed alongside the native ingest URL at project creation and key rotation (see also the Use your existing Sentry SDK how-to).

Slack notifications

The first alert channel: paste an incoming webhook URL from the dashboard's "Slack" section under Notifications, and RealUptime posts a message on the same down/recovery transitions that drive every other channel below. Like webhook notifications and PagerDuty, this feature has no tier gate and is managed from the dashboard, not via the REST API: there is no endpoint to set, read, or clear the webhook URL programmatically.

Webhook notifications

A third alert channel alongside Slack and operator email (roadmap F-3). Configure one or more HTTPS endpoint URLs from the dashboard's "Webhook notifications" section; RealUptime POSTs a signed JSON payload to each one on the same down/recovery transitions that already drive Slack messages and operator email, respecting the same scheduled-maintenance suppression as those two channels. That suppression is per component since migration 060: a maintenance window silences only the components it covers, and only a window that names no components covers the whole status page.

This feature is separate from the REST API/MCP surface above: it has no tier gate (Slack alerts and operator email don't either), and endpoints are managed from the dashboard, not via the REST API.

Endpoint requirements

  • URL must use https://. Plain http:// is rejected.
  • There is a ceiling on how many enabled endpoints one account may hold at

once, matching that plan's included monitor count (see MAX_NOTIFICATION_ENDPOINTS in packages/db/allowances.ts for the current numbers). This is a safety bound on outbound fan-out, not a plan feature: every transition sends one POST per endpoint, so the endpoint count multiplies outbound traffic and something has to bound it. It sits far above normal use. Removing an endpoint frees its slot immediately.

  • The target goes through the same SSRF safety check as a monitor URL

(see "Monitor target validation" above): no loopback, private, link-local, carrier-grade NAT, multicast, reserved, or metadata-range destinations, checked at the time you save the endpoint AND again immediately before every delivery attempt (a URL that resolves safely today can resolve somewhere unsafe later).

  • A signing secret is generated when you add the endpoint and shown

once. Copy it down immediately. It's stored server-side in plaintext (required so RealUptime can re-sign every future delivery) but is never re-exposed through the dashboard or any API response after that first display; remove the endpoint and add a new one to rotate it.

Payload

json
{
  "event": "down",
  "check": { "id": "...", "name": "API", "url": "https://api.example.com/health" },
  "region": "iad",
  "state": "down",
  "timestamp": "2026-08-06T01:29:02.228Z",
  "incident": { "id": "...", "title": "US-East is down" }
}
FieldTypeNotes
eventstringdown or recovery
check.id / check.name / check.urlstringThe monitor that transitioned
regionstringiad, sjc, fra, or nrt (the region that transitioned, not necessarily every region)
statestringdown or operational (the region's new state after this transition)
timestampstringISO 8601, when the delivery was enqueued
incidentobject or nullThe incident this transition opened or resolved, if the check belongs to a status page; null for a standalone monitor with no public page

This shape is stable: existing fields will not be renamed or removed, but new fields may be added, so parse it tolerantly (ignore unknown keys).

Vendor-outage context (REA-532)

A down payload gains one more field when the account is linked to a catalog vendor (its own watchlist, or an inferred dependency link) with a probe-measured outage open right now:

json
"vendorOutageContext": {
  "vendorSlug": "cloudflare",
  "vendorName": "Cloudflare",
  "startedAt": "2026-08-31T12:00:00.000Z",
  "outageUrl": "https://realuptime.io/outages/cloudflare"
}

Absent entirely (not null) on every payload where no such vendor outage is open, and never present on a storm-collapse digest. This states coincidence, never causation: the same wording the operator email, Slack, and PagerDuty channels carry for this signal is "coincides with," never "because of" or "caused by," and only ever names an outage our own probes measured, never a vendor's own status-page claim.

Unusual API key usage (api_usage_anomaly)

Sent to every enabled endpoint when the account's alert routing includes webhooks and one of its API keys crosses the unusual usage threshold (see "Unusual usage alerts" under Rate limits). At most once a day per key. The key is named by its name and prefix only.

json
{
  "event": "api_usage_anomaly",
  "key": { "id": "...", "name": "ci", "prefix": "ru_live_abc123" },
  "kind": "volume_spike",
  "today": 25400,
  "baselinePerDay": 1200,
  "baselineDays": 14,
  "timestamp": "2026-09-27T12:00:00.000Z"
}

On-call escalation pages (escalation_page)

An escalation policy step whose target is "Webhook endpoints (all enabled)" (Growth and Scale; see the dashboard's On-call section) POSTs a different event to every enabled endpoint when an unacknowledged incident reaches that step. Same signature headers, same retries; X-Realuptime-Event is escalation_page.

json
{
  "event": "escalation_page",
  "check": { "id": "...", "name": "API" },
  "incident": { "id": "...", "title": "US-East is down" },
  "step": { "position": 2, "waitMinutes": 15 },
  "dashboardUrl": "https://dashboard.realuptime.io/status",
  "timestamp": "2026-08-23T09:00:00.000Z"
}

step.position is 1-based, as shown in the policy editor. The page is sent once per step reached; acknowledging the incident on the dashboard stops the ladder, and nothing is sent on acknowledgement or resolution through this event. Escalation steps that target a notification channel (Discord, Teams, Opsgenie, Zapier) or the built-in Slack or PagerDuty channel are delivered through those channels' own formats and are not webhook events.

Verifying the signature

Every request carries an event header and, during the migration described below, two signature headers:

X-Realuptime-Event: down
X-Realuptime-Signature: 5f4e5c1b9a3d...
X-Realuptime-Signature-V2: t=1717000000,v1=88df1c2b4a9e...

`X-Realuptime-Signature-V2` (recommended, REA-771). The value is t=<unix seconds>,v1=<hex>, where <hex> is an HMAC-SHA256 of ` ${t}.${rawBody} (the timestamp, a literal ., then the exact raw request body bytes), hex-encoded, keyed with your endpoint's signing secret. Reject anything whose t is more than **5 minutes** from your own clock, in either direction: that tolerance, not the HMAC alone, is what stops a captured delivery from being replayed later, since a replayed request's t` never advances and eventually falls outside the window.

js
import { createHmac, timingSafeEqual } from "node:crypto";

const TOLERANCE_SECONDS = 5 * 60;

function isValidSignatureV2(secret, rawBody, headerValue, nowSeconds = Math.floor(Date.now() / 1000)) {
  const parts = Object.fromEntries(headerValue.split(",").map((p) => p.split("=")));
  const t = Number(parts.t);
  if (!Number.isFinite(t) || Math.abs(nowSeconds - t) > TOLERANCE_SECONDS) return false;
  const expected = createHmac("sha256", secret).update(`${parts.t}.${rawBody}`, "utf8").digest("hex");
  const expectedBuf = Buffer.from(expected, "hex");
  const providedBuf = Buffer.from(parts.v1 ?? "", "hex");
  if (expectedBuf.length !== providedBuf.length) return false;
  return timingSafeEqual(expectedBuf, providedBuf);
}

Worked example: secret whsec_abc123, raw body {"event":"down"}, t=1717000000 signs the string 1717000000.{"event":"down"}; the v1 value is that string's HMAC-SHA256 under whsec_abc123, hex-encoded. Change the secret, the timestamp, or one byte of the body, and the digest no longer matches, that's the whole check.

The JS and Python SDKs ship a verifier that already does this (both schemes, tolerance included), so most integrations never need to hand-roll the recipe above:

js
import { verifyWebhookSignature } from "@realuptime/sdk";
const ok = await verifyWebhookSignature({ headers: request.headers, rawBody, secret });
python
from realuptime_sdk import verify_webhook_signature
ok = verify_webhook_signature(request.headers, raw_body, secret)

`X-Realuptime-Signature` (deprecated, removed December 6, 2026). The original scheme: an HMAC-SHA256 of the exact raw request body bytes alone, hex-encoded, no timestamp, keyed with the same signing secret.

js
import { createHmac, timingSafeEqual } from "node:crypto";

function isValidSignatureV1(secret, rawBody, providedSignatureHex) {
  const expected = createHmac("sha256", secret).update(rawBody, "utf8").digest("hex");
  const expectedBuf = Buffer.from(expected, "hex");
  const providedBuf = Buffer.from(providedSignatureHex, "hex");
  if (expectedBuf.length !== providedBuf.length) return false;
  return timingSafeEqual(expectedBuf, providedBuf);
}

It has no replay protection at all: a captured delivery's v1 signature stays valid forever, since there is nothing in it that ages. Both headers are sent on every delivery today so an existing integration keeps working unchanged; `X-Realuptime-Signature` stops being sent on December 6, 2026 (90 days from this feature's release), after which X-Realuptime-Signature-V2 is the only signature header. Migrate before that date.

In both schemes, use the raw, unparsed request body for the HMAC, not a re-serialized copy of the parsed JSON: re-serializing can change key order or whitespace and would compute a different signature than the one that was sent. Compare with a constant-time comparison, never === or ==.

Delivery guarantees

  • Retried with exponential backoff (60s, doubling: 60s, 120s, 240s, 480s) up

to 5 attempts, then dead-lettered. Same outbox pattern as operator email (see "Notification delivery retries" in the roadmap): a claim-based PostgreSQL queue, not a separate queue service.

  • A 5-second timeout per attempt. A slow or hanging endpoint counts as a

failed attempt and is retried, not held open.

  • Any non-2xx response, a timeout, a connection error, or the endpoint

failing a re-check of the safety rules above all count as a failed attempt.

  • Redirects are not followed. Respond 2xx directly at the URL you

registered.

  • Response bodies are never read or trusted for anything beyond a short

snippet in delivery logs. Delivery success is judged purely by HTTP status code.

PagerDuty notifications

A fourth alert channel alongside Slack, operator email, and generic webhooks. Paste your PagerDuty Events API v2 integration (routing) key from the dashboard's "PagerDuty" section under Notifications; RealUptime sends a trigger event when a region goes down and a matching resolve event with the same dedup_key on recovery, so the incident PagerDuty opened closes automatically on your side.

Unlike Slack, operator email, and generic webhooks, the resolve event is not subject to scheduled-maintenance suppression: a maintenance window active at recovery time silences a NEW trigger, never the resolve of an incident PagerDuty is already holding open. If a check went down before a window started, PagerDuty opened a real incident, and that incident closes the moment the check recovers, window or no window; otherwise it would stay open in PagerDuty until someone noticed and closed it by hand (REA-768). The Opsgenie and incident.io registry channels (see "Notification channels" above) share this exact alias/dedup-key resolve behavior and the same maintenance-bypass-on-recovery rule, for the identical reason (REA-783).

Like webhook notifications, this feature has no tier gate and is managed from the dashboard, not via the REST API. One integration per account; a new key you save replaces the previous one. The pasted key must be 8-256 characters after trimming. PagerDuty doesn't publish a strict format for this key, so this is a loose sanity check, not real key validation.

Delivery guarantees

Same outbox pattern, retry schedule (60s, doubling: 60s/120s/240s/480s, 5 attempts then dead-lettered), and 5-second per-attempt timeout as webhook notifications above. Delivery success is judged purely by HTTP status code (PagerDuty's Events API v2 returns 202 on a successfully queued event).

Slack app

An installable Slack app (REA-232), distinct from the webhook-based Slack notification channel: this one runs /realuptime slash commands and, once wired as a notification channel target, posts messages via chat.postMessage rather than a fixed incoming-webhook URL.

  • Install: GET /api/slack/install (linked from Account settings >

Integrations, gated on manage_notifications) redirects to Slack's OAuth v2 authorize screen requesting the commands and chat:write bot scopes. GET /api/slack/oauth/callback exchanges the code and stores the bot token for the installing account, scoped one installation per (account, Slack workspace). CSRF is enforced the same way the sign-in OAuth flows enforce it: a random state value round-tripped through an httpOnly cookie, with the installing account's id folded into the state itself.

  • Slash commands: POST /api/slack/commands verifies Slack's request

signature (HMAC-SHA256 over v0:{timestamp}:{raw body} with the app's signing secret, hex-encoded, v0= prefixed) and refuses anything more than 5 minutes old. A second guard rejects an exact replay of an already-answered request inside that window. /realuptime status, /realuptime status <name>, /realuptime incidents and /realuptime help all answer with an ephemeral message, visible only to the person who ran the command.

  • Events: POST /api/slack/events handles the one-time

url_verification handshake (echoes challenge) and app_uninstalled (marks the installation for that workspace uninstalled, the same effect as removing the app from Account settings > Integrations).

  • Secrets: the bot token is stored like a tracker connection's

credential or a webhook channel's config -- never rendered back to a browser, replaced by reinstalling rather than edited in place.

  • Dark until the owner creates the Slack app and vaults

SLACK_CLIENT_ID/SLACK_CLIENT_SECRET/SLACK_SIGNING_SECRET; see deploy/slack/manifest.json and deploy/slack/RUNBOOK.md.

MCP server

https://mcp.realuptime.io/mcp: a stateless `StreamableHTTPServerTransport` endpoint (POST, JSON-RPC 2.0, no session state, same Authorization: Bearer header as the REST API on every request). Health check at /health.

Request bodies over 100KB are rejected before parsing. This is checked first against the Content-Length header (fast-path, before auth even runs), then enforced again against the actual bytes read as a fallback. Every input this server takes is a handful of short strings, so 100KB has no legitimate use here; an oversized request gets a plain 413 with { "error": "Request body too large." }, not a JSON-RPC error envelope.

There is a second, keyless endpoint at POST /public on the same host, serving seven of the tools below with no API key at all. See "Keyless access" below for what it serves and why.

Point any MCP client at /mcp with the bearer key set. The twenty-three tools below mirror the REST API exactly, except the seven outage tools, which read RealUptime's own public measurements rather than the caller's account:

ToolEquivalent toArguments
list_checksGET /checksnone
get_check_statusGET /checks/:idcheckId (uuid)
create_checkPOST /checksname, url, intervalSeconds? (number ≥ 1; not integer- or range-checked at the schema layer, rounded and clamped to 60–86400 downstream, see below), regions? (array, defaults to the four core regions), plus the optional assertionBody*/assertionHeader*/assertionStatus* fields (see "Response assertions" under POST /checks in the REST section above)
update_check_regionsPATCH /checks/:idcheckId (uuid), regions (array, at least one required)
update_check_assertionsPATCH /checks/:id/assertionscheckId (uuid), plus the same optional assertion fields as create_check -- a full replace, http checks only
delete_checkDELETE /checks/:idcheckId (uuid)
list_status_pagesGET /status-pagesnone
list_incidentsGET /incidentslimit? (1–500)
create_incidentPOST /incidentsstatusPageId (uuid), checkId (uuid), region ("iad" | "sjc" | "fra" | "nrt" | "ord" | "yyz" | "lhr" | "sin" | "syd" | "gru"), title (1-200 characters), body (1-5000 characters)
add_incident_updatePOST /incidents/:id/updatesincidentId (uuid), status ("investigating" | "identified" | "monitoring" | "resolved"), body (1-5000 characters), force? (boolean, see below)
get_active_maintenance_windowGET /maintenancestatusPageId (uuid)
create_maintenance_windowPOST /maintenancestatusPageId (uuid), title (1-200 characters), body (1-5000 characters), startsAt? (ISO 8601), exactly one of endsAt? (ISO 8601) or durationMinutes? (1-43200), componentIds? (array of uuid), noDowntime? (boolean)
end_maintenance_windowDELETE /maintenance/:idwindowId (uuid)
list_error_projectsGET /errors/projectsnone
list_error_issuesGET /errors/projects/:id/issuesprojectId (uuid), q?, release?, status? ("open" | "resolved" | "ignored"), window? ("24h" | "7d" | "30d" | "all"), limit? (1–100)
get_error_issueGET /errors/issues/:idissueId (uuid)
is_service_downGET /api/v1/outages/:slugservice (catalog slug), region? ("iad" | "sjc" | "fra" | "nrt" | "ord" | "yyz" | "lhr" | "sin" | "syd" | "gru")
get_regional_readingsGET /api/v1/outages/:slugservice (catalog slug)
get_report_volumeGET /api/v1/outages/:slugservice (catalog slug)
list_recent_incidentsGET /api/v1/outages/:slug (recent_events, last_outage, reliability)service (catalog slug)
get_stackGET /api/v1/outages/stacks/:keystack (stack slug or id)
internet_weatherGET /api/v1/outages/internet-weathernone
is_the_internet_downGET /api/v1/outages/internet-weather (the provider-correlation verdict: ok, provider_incident, or several_unrelated, with the measured dependency map)none

The seven outage tools are the odd ones out on this surface: they read RealUptime's own probe measurements of third-party public services, not the caller's account, so nothing in them is scoped by account id and nothing in them is tier-gated. Their answer body is built by the same buildOutageAnswer (packages/db/outage-answer.ts) the public JSON endpoint uses, so the two surfaces cannot answer the same question differently. They are the seven tools the keyless endpoint serves.

Keyless access (POST /public)

https://mcp.realuptime.io/public serves the seven outage tools with no API key, so an MCP client can be added and asked "is Stripe down" without its user signing up for anything.

This is the citation pillar applied to MCP (the internal-docs repo's outages-plan.md, "Strategy pillar: be the source AI assistants cite"): a key is a wall between an assistant and a citation, and every byte this endpoint returns is public measurement we are actively trying to have quoted. The keyless per-service JSON endpoint (GET /api/v1/outages/:slug) already made the same trade for the same data; this is its MCP twin.

A separate path, not an anonymous mode

The rejected design was "on /mcp, if there is no Authorization header, build a reduced server". It is fewer lines and worse, for three reasons:

  1. It turns an auth failure into a partial success. A customer whose key

expired, or whose proxy stripped the header, would stop getting RU-1001/RU-1002 and start getting a working four-tool session. A credential fault that presents as "some of my tools disappeared" is much harder to diagnose than one that presents as 401.

  1. It would blunt an existing control. The failed-auth per-IP limiter

counts a missing header as a failure deliberately, so that the cheapest possible junk request is not the one thing left unbounded. If "no header" became a legitimate mode, that count could no longer be taken.

  1. **It would make the allowlist a runtime condition instead of a structural

fact.** buildPublicServer() takes no arguments: there is no Account, no key id and no permission anywhere in its scope, so no handler it registers can read account data even through a bug. "Unreachable" is a property of the object graph, not of a branch someone has to remember to write.

No key is ever read on `/public`. The endpoint does not look at the Authorization header at all. A keyless endpoint that verified a key it then ignored would be a free, unauthenticated oracle for testing whether a stolen key is still live; ignoring the header entirely means there is nothing to probe. A client that sends one anyway simply gets the keyless tool set.

No cookies, no state. Same stateless transport as /mcp (sessionIdGenerator: undefined), so there is no session to fixate and nothing to correlate two calls with. No Set-Cookie and no mcp-session-id is ever emitted. Nothing is written anywhere except the shared rate-limit bucket row, which holds an address and a count.

What it serves

Exactly seven tools, and tools/list returns exactly these:

ToolWhy it is safe keyless
is_service_downRead-only, public measurement, no account scope, no outbound request
get_regional_readingsSame
get_report_volumeSame
list_recent_incidentsSame
get_stackSame; the stack key is a shareable handle, not an account scope
internet_weatherSame; a public aggregate over the catalog's own probe readings, no arguments
is_the_internet_downSame; the whole-catalog provider-correlation read over those same probe readings, no arguments

Every other tool stays keyed: all five mutating tools, and every read tool scoped to the caller's own account (monitors, status pages, incidents, and the three Errors reads). Calling one on /public answers HTTP 403 with RU-1006, naming the keyed endpoint as the fix rather than the bare "tool not found" an unregistered name would produce. A JSON-RPC batch carrying one keyed call among allowed ones is refused whole, for the obvious reason that checking only the top-level object would miss it. A name in neither list falls through to the SDK's own "tool not found", so a typo is never told that a tool exists.

apps/mcp/keyless-tools.test.ts asserts the served set equals the allowlist in both directions, and that the keyless and keyed lists together account for every tool buildServer registers, so a tool added later cannot land in neither list by omission.

Rate limit

60 requests per minute per IP, Postgres-backed (keylessMcpDbRateLimiter) so it holds across every Fly machine rather than being multiplied by the machine count. Over budget answers HTTP 429 with Retry-After, RU-6004, and a retryAfterSeconds field.

The budget is deliberately half the keyed read budget (readRateLimiter, 120/min per key) and half the keyless outage JSON endpoint's (120/min per IP). Both comparisons are load-bearing:

  • Stricter than keyed, because a keyless call cannot be attributed to an

account, cannot be revoked, and buys no relationship. A public on-ramp must never be the cheapest way to consume us, or the rational move for a heavy consumer would be to stay anonymous.

  • Stricter than the keyless JSON endpoint, because that one is hard-cached

(Cache-Control: public, max-age=120) and a spike against one popular service collapses to a handful of database reads. A JSON-RPC POST is cacheable by nothing, so every call here is a real read.

Being honest about the size: one assistant question costs roughly four requests (initialize, notifications/initialized, tools/list, tools/call), so 60/min is about fifteen questions a minute from one address. That is far above a single assistant session and far below what a shared egress address (a corporate NAT, or an assistant vendor's own connector infrastructure) can produce. A shared address genuinely can exhaust this, which is why RU-6004 names the keyed endpoint as the fix instead of pretending the limit is generous, and why the keyed endpoint remains the documented fallback exactly as the plan intended.

The bucket is keyed off Fly's fly-client-ip, never X-Forwarded-For, whose leftmost hop the client picks: an XFF-keyed budget on an anonymous endpoint would be no budget at all. No usable client IP means no limiting rather than one shared bucket, matching every other limiter in this codebase, since the alternative is that one header-less request throttles every keyless consumer at once.

Abuse posture

  • SSRF reach is unchanged, and it is zero. None of the seven tools takes a

URL or drives any outbound request; they read rows our own probes already wrote. This is what separates them from the public web tools (/tools/is-it-down, /tools/ssl-checker, /tools/status-page-grader), which do make outbound requests and are bounded six times harder for exactly that reason.

  • Same 100KB request-body cap as /mcp, checked against Content-Length

before the rate limiter is even charged, then again against the bytes actually read.

  • Response size is bounded by the catalog shape, not by caller input: at

most ten regions (every live probe region) times two endpoint roles per service, with bounded recent lists. There is no parameter that widens a response.

  • No bulk export. One service per call, exactly like the JSON endpoint,

and no catalog-listing tool. Per-service reads serve citation; a bulk dump serves a competitor rebuilding the dataset without running any probes. The catalog listing stays on the human hub page at /outages.

  • Nothing is persisted about a keyless caller beyond the rate-limit bucket

row.

Free-tier access

Unlike the REST API (paid plans only, see Authentication above), MCP read access is free: list_checks, get_check_status, list_status_pages, list_incidents, get_active_maintenance_window, the three Errors read tools (list_error_projects, list_error_issues, get_error_issue -- accounts holding the Errors product only), and all seven outage tools (no product or tier requirement at all) work with a free-tier key. create_check, update_check_regions, update_check_assertions, delete_check, create_incident, add_incident_update, create_maintenance_window, and end_maintenance_window all require a Growth or Scale plan (no gating change from response assertions -- the new tool is a write tool, gated exactly like every other one); a free-tier key calling one of them gets isError: true with an upgrade message instead of running (checkFreeTierWrite in apps/mcp/index.ts), the same shape as a read-scope key hitting a write tool. The two checks are independent: a paid account's read-scope key is still refused by the permission-scope check, and a free account's key is refused by the tier check regardless of its own stored permission scope (a grandfathered read_write key from before an account downgraded to free, for example).

create_incident and add_incident_update require a read_write key (a read key gets isError: true with the same "read-only" message the REST routes return) and share the write rate limit. Both send email to the status page's confirmed subscribers on success, same as their REST equivalents above: see POST /incidents and POST /incidents/:id/updates for the exact conditions and caveats.

add_incident_update carries the same resolve guard as its REST twin (behavior change, 2026-08-17): status: "resolved" returns isError: true with a message telling you to pass force: true when our probes currently read the component down or degraded. Both surfaces read one shared predicate (isLiveOutageState, packages/db/incident-live-state.ts), so neither can quietly let something through the other would stop. Passing force: true proceeds exactly as before, and every other status ignores it. See POST /incidents/:id/updates above for which states count as live and why.

get_check_status's status field uses the exact same staleness-aware aggregateStatusForDisplay aggregation as GET /checks/:id (see above): operational, degraded, down, stale, or unknown. This was not always true: before this fix, get_check_status aggregated with the raw aggregateStatus, which has no staleness concept, so a check with stale per-region data could read operational via MCP while the public status page correctly read stale for the same check.

create_check's intervalSeconds argument uses the same shared schema as the REST endpoint (checkCreateShape in packages/db/api-schemas.ts, see intervalSeconds clamping above): any finite value ≥ 1 passes schema validation and is clamped to [60, 86400] downstream by clampIntervalSeconds, identically on both surfaces; it is not separately range-checked at the schema layer. create_check's regions argument and update_check_regions's regions argument both use the same regionsShape (see Regions above): at least one region required, unknown region values rejected. checkId (get_check_status, delete_check, update_check_regions) and limit (list_incidents) are likewise validated via shared schemas (checkIdShape, incidentsListShape) rather than each tool declaring its own inline shape. create_check, update_check_regions, update_check_assertions, and delete_check all apply the write rate limit; list_checks, get_check_status, list_status_pages, and list_incidents all apply the read rate limit: the same budgets as the REST routes (see Rate limits above), not a separate MCP allowance. create_check also applies the same target-validation guard as the REST endpoint (see above). create_check, update_check_regions, update_check_assertions, and delete_check all also require a read_write key (see Permission scopes above): a read key gets isError: true with a "read-only" message instead of reaching the database, the same rule the REST API's POST/PATCH/DELETE routes enforce.

create_check and delete_check do not sync Stripe overage billing inline the way the REST POST /checks route and the dashboard do: see the "Billing sync gap" note under POST /checks above.

Every tool returns the same JSON shape as its REST equivalent, serialized as a text content block. The exception is delete_check, whose REST equivalent (DELETE /checks/:id) returns an empty 204 No Content body; the MCP tool instead returns { "deleted": true }. A "not found", over-limit, unsafe-target, or rate-limited condition sets isError: true on the tool result rather than throwing or returning an HTTP error status; the JSON-RPC call itself still succeeds at the transport level (a rate-limited call gets a retryAfterSeconds field inside the JSON text content, not an HTTP 429 or Retry-After header; those only exist on the REST surface).

Manual protocol check

bash
curl -X POST https://mcp.realuptime.io/mcp \
  -H "Authorization: Bearer ru_live_..." \
  -H "Content-Type: application/json" \
  -H "Accept: application/json, text/event-stream" \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"test","version":"1"}}}'

Each request is independent (stateless transport): a client library normally handles the initialize → tools/list → tools/call sequence for you; this is only useful for a manual sanity check.

The keyless endpoint takes the same request with no Authorization header:

bash
curl -X POST https://mcp.realuptime.io/public \
  -H "Content-Type: application/json" \
  -H "Accept: application/json, text/event-stream" \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'

Errors

Status codes in use across the REST API:

StatusMeaning
400Malformed request body, missing/invalid required field, or a monitor target that failed the safety check
401Missing, invalid, revoked, or expired API key (see Key expiry above)
403Key belongs to a free-tier account, the account is at its monitor limit for its tier, the key's scope is read and the route requires read_write (see Permission scopes above), or the key carries an IP allowlist and the request came from outside it (RU-1007, see IP allowlists above)
404Check doesn't exist, or exists but isn't owned by this account
405HTTP method not implemented on that route (e.g. DELETE /checks, PUT /incidents), returned automatically by the framework, not by application code
409Resolving an incident whose monitor we still read down or degraded, without force: true (see POST /incidents/:id/updates)
429Rate limit exceeded (see Rate limits above)
500Unhandled server error

The MCP server surfaces the equivalent failures as isError: true tool results instead of HTTP status codes (the transport call itself still returns 200/JSON-RPC success), except request-body-too-large (413) and malformed JSON (400), which are rejected before a tool ever runs; a handler exception, which the server catches and turns into a plain 500; and some argument-validation failures on a tool's input schema, which the underlying MCP SDK can reject as a JSON-RPC-level error before a normal tool result is ever produced, rather than as isError: true.

The RU-XXXX error code registry

Some error responses (the public GET /api/v1/outages/:slug endpoint above is the main example today; authentication failures in apps/mcp/index.ts are another) carry a machine-checkable code field alongside the human error message, drawn from the registry in packages/db/error-codes.ts. Each code is banded by the surface it belongs to: RU-1xxx auth, RU-2xxx billing, RU-3xxx monitors, RU-4xxx errors product, RU-5xxx status, RU-6xxx outages. Every code has a public, linkable explanation at https://realuptime.io/kb/errors/ru-1003-style URLs (lowercased code; errorCodeKbUrl in packages/db/error-codes.ts builds it), so a consumer, human or assistant, that surfaces a bare code to an end user can point them at a real page instead of an opaque string. Not every error response carries a code yet; where one is absent, the error message and HTTP status above are the only signal.

Page-view beacon (internal, not the public API)

POST /api/ga/pv takes { path, search, title, sid } from the site's own page-view component (app/_components/ga-pageview.tsx) and always answers 204 with no body. This is not part of the public REST API: it exists so server-side web analytics (REA-833, lib/ga4.ts) can count client-side navigations. It forwards one page_view over the GA4 Measurement Protocol only when measurement is allowed for the visitor (lib/consent-region.ts: consent first in the EU/EEA/UK/Switzerland or when the country is unknown), and on first use sets the ru_cid first-party client id cookie. It sends nothing for crawlers, for hosts that are not realuptime's own, or when GA4_MEASUREMENT_ID and GA4_API_SECRET are unset.

Page-leave event beacon (internal, not the public API)

POST /api/ga/event is the Google Analytics half of the one dashboard funnel event that fires while a page unloads, resource_create_abandoned, which the browser can only deliver with navigator.sendBeacon (lib/analytics-beacon.ts). This is not part of the public REST API. It accepts only the events in PAGE_LEAVE_EVENTS, with only the fields the dashboard event list allows (lib/measurement-scope.ts), after the same checks as the page-view beacon, and only for a visitor that already has a ru_cid cookie: it never sets one. It always answers 204 with no body.

Two first-party routes serve the site's own consent banner and the /privacy control (lib/measurement-consent.ts). Neither is part of the public REST API.

  • GET /api/measurement/region answers { "required": boolean }: whether this

visitor's location needs consent before any measurement (lib/consent-region.ts). Statically rendered pages cannot know the visitor's country, so the banner asks here once per page load, and only for a visitor with no stored answer. The country comes from Cloudflare and is trusted only with the zone's edge secret; an unknown country answers required: true. Sent with cache-control: private, no-store; nothing is stored or logged.

  • POST /api/measurement/forget expires the ru_cid client id cookie when a

visitor declines measurement. The cookie is httpOnly, so the browser cannot delete it itself; this route answers 204 with the matching expiry, on the same Path and Domain POST /api/ga/pv writes it with. Idempotent.

Agent protocol (internal wire contract, not the public API)

POST /api/agent/v1/poll, POST /api/agent/v1/results and POST /api/agent/v1/metrics are the wire contract between the RealUptime Monitor agent and this server (the internal-docs repo's monitor-plan.md phases 1 and 2). This is not part of the public REST API. It is spoken by one program we ship, versioned with that program, and it is documented here so the two halves cannot drift, not as a surface a customer integrates against. Nothing on it is covered by the REST API's tier gate, its key format, its rate limits, or its error vocabulary.

Credential

A per-agent bearer token, rua_ followed by 32 random bytes as hex, minted once when the agent is registered on /monitoring and shown exactly once. The server stores only its SHA-256 hash, so a lost token cannot be recovered and is replaced by registering a new agent.

Authorization: Bearer rua_<64 hex characters>

Every auth failure on all three routes answers 404 with a tiny JSON body: a missing header, a malformed token, an unknown token and a revoked token are deliberately indistinguishable, so the endpoints cannot confirm whether a guessed token exists or whether a stolen one has been revoked yet. Every response is no-store. The wrong HTTP verb against a route that exists is a routing fact, not an auth decision, and is not folded into that 404: it answers 405 with an Allow header naming the verb that works.

Rate limit: 120 requests per minute per agent, counted on the token's hash and consumed before the lookup, so a flood of unknown tokens is throttled too. Over the limit answers 429 with Retry-After: 60. One budget covers all three routes: the token is the unit being protected, and a separate bucket per endpoint would let a compromised token spend three times as much.

POST /api/agent/v1/poll

The body is {}, or, from agents that include REA-1013, the location's egress policy mode and the verdicts since its previous poll (counts only, never a target):

json
{ "egress": { "mode": "report", "refused": 0, "wouldRefuse": 2, "byRule": { "public": 2 } } }

The server sums these per agent per day for the dashboard. A missing, malformed or oversized body is ignored, never refused: the poll's answer does not depend on it. Returns the checks bound to this agent, and stamps the agent's last-seen time as a side effect of the same call.

json
{
  "checks": [
    {
      "id": "b2c3d4e5-...",
      "type": "http",
      "url": "https://10.0.0.5:8080/health",
      "tcpHost": null,
      "tcpPort": null,
      "tcpTls": false,
      "dnsHostname": null,
      "dnsRecordType": null,
      "dnsExpectedValue": null,
      "intervalSeconds": 60
    }
  ]
}

type is http, tcp, dns, or ping. Exactly the fields that type needs are non-null. A revoked agent is served nothing, because its token no longer resolves at all.

The vocabulary is closed, and the agent enforces it too

This response is a data contract with a finite vocabulary, never a command channel (docs/private-probe-locations.md section 3.2). The vocabulary is enforced in three places that must agree, and a value absent from any one of them is refused rather than passed through:

  1. checks_agent_id_probed_types in the database, which admits a binding only

for http, tcp, dns and ping. multistep, smtp, browser and heartbeat are fleet-only.

  1. This route's serializer, which builds the body from typed columns and

cannot emit a field that has no column.

  1. parseChecks in apps/agent/api.ts, whose parseRules holds the same

four types, the six dns record types (A, AAAA, CNAME, MX, TXT, NS) and the tcp port rule (an integer in 1 to 65535, never one of the RFC 862-865 amplification ports 7, 9, 13, 17, 19). A check carrying anything else is skipped, or has that field nulled and reports as a failed result naming the missing configuration.

The third is the one that matters most: an old agent facing a server that learned a new verb refuses the verb. A compromised server can reuse the capabilities a deployed agent was compiled with, and cannot teach it new ones. Adding a verb is therefore a decision that needs its own design document, not a ticket.

What the agent does with the response, whatever it says

Four bounds are applied on the customer's machine, from that machine's own environment, and no field in this response can move any of them (docs/private-probe-locations.md section 3.4):

BoundDefaultBehaviour past it
Cadence floor60sA shorter intervalSeconds is clamped up, and counted on the agent's log line.
Concurrent probes8Checks queue instead of bursting.
Probes per minute600The excess is skipped, and each skipped check reports the reason as a failed result.
Assigned checks250A longer list is truncated in the order it arrived, with a log line.

The agent also decides which targets it will dial: cloud metadata endpoints, link-local, multicast and reserved addresses are refused unconditionally, public addresses are refused under its egress policy, and an optional allowlist held on the customer's machine narrows it further. A refused target arrives back on /results as a failed result reading Blocked by this location's local policy. None of this is server-configurable, deliberately: see apps/agent/README.md, "Where this agent will dial".

Since protocol v2 (REA-181) the response also carries services: the opt-in list of service/unit names the customer asked the dashboard to report status for on this host, as an array of strings (absent or empty means none). It is the one server-pushed setting that changes what the agent reads locally, and it is constrained on both sides to a conservative name shape so it can only ever name a unit, never a path or a command.

json
{ "checks": [ ... ], "services": ["nginx", "postgresql@16"] }

Authenticated checks: secret references, never values (REA-1014)

From agent 0.4.0 the poll body also says what the agent is:

json
{ "agent": { "version": "0.4.0", "capabilities": ["secret_refs"], "secretHeaderNames": ["X-Internal-Auth"] } }

secretHeaderNames is the list the customer set with REALUPTIME_AUTH_HEADERS on that machine, beyond the four every agent allows. The server stores the block when it changes (packages/db/monitor-agent-self-report.ts) for the dashboard, and a body with no agent block is recorded as an agent that reports nothing. A malformed block is ignored.

An http check that carries authentication is then served with an auth block of reference templates:

json
{
  "id": "b2c3d4e5-...",
  "type": "http",
  "url": "https://10.0.0.5:8443/admin",
  "auth": {
    "headers": [{ "name": "Authorization", "value": "Bearer ${SECRET:BILLING_API_TOKEN}" }],
    "userinfo": null
  }
}

The server never holds a credential's value, so the poll cannot carry one. The agent resolves each ${SECRET:NAME} at dial time from REALUPTIME_SECRET_<NAME> or the file named by REALUPTIME_SECRETS_FILE, substitutes only in header values and in userinfo (sent as HTTP Basic), never in the host, path, port or query, sends only header names its own machine allows, fails the check with This location has no secret named NAME when a name does not resolve, and redacts every resolved value from its logs and from every error string it posts to /results (docs/private-probe-locations.md section 3.6). parseChecks accepts the block only on an http check and only in exactly this shape; an unknown key anywhere inside it refuses the whole check.

The block is served only to an agent whose THIS poll declared `secret_refs` at version 0.4.0 or later. Any other agent, including 0.3.0, which shipped with no secret support, gets the same check with "url": null and no auth, which every agent release reports as the failed result Check is missing configuration: no url without dialling. An older agent therefore never sends the literal ${SECRET:...} text and never probes the target without the credential it was configured with.

POST /api/agent/v1/results

json
{
  "results": [
    {
      "checkId": "b2c3d4e5-...",
      "ok": false,
      "statusCode": 503,
      "latencyMs": 122,
      "error": "connection refused",
      "checkedAt": "2026-08-17T14:04:05.000Z"
    }
  ]
}

checkId, ok and checkedAt are required; statusCode, latencyMs and error are optional and may be null. At most 100 results per call: a larger batch is refused with 400 so a buffering agent slices its backlog rather than sending one request that times out and is retried forever.

json
{ "accepted": 1, "rejected": 0, "clamped": 0 }

Three rules govern what happens to each result:

  • Accepted results enter the same pipeline a regional probe's result does:

the raw row, the hysteresis state machine (two consecutive failures to go down, one success to recover), incidents, escalation, and the full alert fan-out. The locus recorded is agent: followed by the first eight characters of the agent id, in the same region column the four fleet regions use.

  • **A result for a check this agent is not bound to is dropped and counted in

rejected.** It never fails the batch: erroring would confirm which ids exist, and one stale id would block every genuine result behind it.

  • `checkedAt` is clamped and counted in `clamped` when it is in the future

or more than six hours old. Buffering through a connectivity loss is expected, so timestamps inside that window are honoured exactly and the history shows an outage where it happened; outside it, server time is used instead, so a wrong clock cannot invent history a customer has already read, or make a check look freshly checked when nothing checked it.

  • A result more than ten minutes old is history only. It is written where

it happened but does not move the check's live state, so it opens no incident and sends no alert: by the time it arrives, fresher results say what is true now. The response does not count these separately.

POST /api/agent/v1/metrics

Server health from the machine the agent runs on (the internal-docs repo's monitor-plan.md phase 2). Separate from /results because it answers a different question: not whether a target is up, but what this machine is doing.

json
{
  "vantage": "host",
  "vantageDetail": null,
  "collectorVersion": "0.2.0",
  "samples": [
    {
      "sampledAt": "2026-08-18T14:04:05.000Z",
      "cpuUsedRatio": 0.42,
      "cpuCores": 8,
      "memoryTotalBytes": 17179869184,
      "memoryUsedBytes": 9663676416,
      "load1": 1.2,
      "load5": 0.9,
      "load15": 0.7,
      "filesystems": [
        { "mountPoint": "/", "totalBytes": 500107862016, "usedBytes": 462742192128 }
      ]
    }
  ]
}

At most 100 samples per call, each with at most 32 filesystems. A larger batch is refused with 400, so a buffering agent slices its backlog rather than sending one request that times out and is retried forever.

Protocol v2 (REA-181)

A v2 batch adds "protocolVersion": 2, an optional host header, and up to four optional families on each sample. Everything is additive: a batch with no protocolVersion is v1 and is handled exactly as before, and a v2 batch against a server that predates this section is accepted with the additions ignored. An unknown protocolVersion is refused with 400.

json
{
  "vantage": "host",
  "protocolVersion": 2,
  "host": { "hostname": "vps-1", "os": "linux", "osVersion": "Ubuntu 24.04", "arch": "x64", "cluster": "prod-eu", "node": "vps-1" },
  "samples": [
    {
      "sampledAt": "...", "cpuUsedRatio": 0.42, "cpuCores": 8, "memoryTotalBytes": 1, "memoryUsedBytes": 1, "filesystems": [],
      "network":    [{ "name": "eth0", "rxBytesPerSec": 104857, "txBytesPerSec": 52428, "rxErrors": 0, "txErrors": 0, "rxDropped": 0, "txDropped": 0 }],
      "processes":  [{ "pid": 4242, "name": "postgres", "cpuRatio": 0.12, "memoryBytes": 536870912 }],
      "containers": [{ "id": "3f9c...", "name": null, "runtime": "docker", "cpuRatio": 0.05, "memoryUsedBytes": 104857600, "memoryLimitBytes": null }],
      "services":   [{ "name": "nginx", "status": "active" }]
    }
  ]
}
FamilyCapNotes
network32bytes per second in/out computed by the collector from two counter readings; rxErrors/txErrors/rxDropped/txDropped are deltas over that interval, not lifetime counters; loopback excluded
processes20top 10 by CPU plus top 10 by memory, deduped; name is the executable name only, never a command line; cpuRatio is a fraction of TOTAL capacity like cpuUsedRatio
containers64cgroup v2 on Linux; runtime is one of docker, containerd, cri-o, podman, kubernetes, lxc; cpuRatio null on the warm-up sample; memoryLimitBytes null when unlimited
services64status is active, inactive, failed or unknown, for each name in the poll's services list
hostoneos is linux, darwin, windows or other; cluster and node are labels the dashboard groups by; every field is at most 128 characters

A family that is ABSENT means the collector does not measure it on this platform (containers on macOS, load on Windows); an EMPTY array means it measured and found none. The distinction is preserved in storage. A malformed element refuses the batch whole, same as a malformed filesystem entry, and the 400 names the field (processes.cpuRatio must be a fraction ...).

The four families are stored as jsonb beside the core sample (server_metric_details, cascading from it) and follow raw-sample retention (7 days on every tier, the internal-docs repo's retention.md). host is stored on the agent, last write wins.

json
{ "accepted": 1, "rejectedStale": 0, "rejectedFuture": 0, "rejectedDuplicate": 0 }

Units

Every number has exactly one legal form. These are enforced by CHECK constraints in the database as well as by the route, so a collector that gets one wrong fails on its first batch instead of drawing a plausible chart that is wrong by a factor of a hundred.

FieldUnit
cpuUsedRatiofraction of TOTAL capacity across all cores, 0 to 1
cpuCoreswhole number of cores that total is across
memoryTotalBytes, memoryUsedBytesbytes; used is total minus AVAILABLE
totalBytes, usedBytes (filesystems)bytes
load1, load5, load15raw kernel load averages, unnormalised
  • CPU is a ratio, not a percentage, and not per-core. 0.87, never 87.

A percentage invites the 0-100 versus 0-1 ambiguity at every boundary the number crosses, and a per-core figure cannot be compared between two machines without also knowing the core count. "How full is this machine" is the question both a chart and an alert threshold ask, and it is capacity relative. cpuCores travels alongside so a per-core view stays derivable.

  • Memory `used` is total minus available, not total minus free. On Linux,

cache and buffers are reclaimable, so total-minus-free reports a healthy machine as permanently full.

  • Percentages are never reported. The server derives them from the totals,

so a chart and an alert cannot disagree about what "90% full" meant.

  • Load is all three or none. A platform without load averages omits all

three and must never send zeroes, which read as an idle machine.

Vantage

vantage is required and is either host or container. vantageDetail is optional free text for humans ("docker", "lxc", "kubernetes").

A collector running inside a container reads the CONTAINER's cgroup limits for CPU and memory and its overlay filesystem for disk. Those are real numbers, and reported as host metrics they are a lie nothing downstream can detect: the chart looks plausible and the thresholds are wrong by an unknowable factor.

So the server pins whichever vantage an agent first reports and refuses any later batch that disagrees with 409, naming both vantages. An operator genuinely moving a collector from the host into a container registers a new agent, which is already the answer for rotating a token. The vantage is also stored on every sample, so history read a month later still says what the number was a measurement of.

What happens to each sample

  • `sampledAt` is dropped, not clamped. Accepted between 24 hours old and

one minute into the future; anything outside that is counted in rejectedStale or rejectedFuture and never written. This is deliberately unlike /results, which rewrites an out-of-window checkedAt to server time: clamping a time series would invent a data point at an instant nothing was measured, and would collapse a whole buffered batch onto one row.

  • **A repeated instant is counted in rejectedDuplicate and the first writer

keeps it.** A redelivered batch is therefore idempotent. A count that stays non-zero across FRESH batches means two collectors are running on one token; neither can overwrite the other's history, and the number is how that shows up.

  • A malformed batch is refused whole, with `400` naming the field. Unlike a

result naming a stale check id, there is no benign reason for a metric sample to be malformed, and half a batch on a chart is worse than none: nobody would think to distrust it.

Why agent-bound checks are never probed by us

An agent exists to watch things only the customer's own network can reach, so its targets are private by design and skip the target safety check every fleet monitor goes through. The other half of that trade is absolute: a check bound to an agent is never returned to the regional fleet's scheduler, so no machine of ours ever dials a private address. Neither rule is safe without the other.

Why two auth implementations

apps/web/lib/api-auth.ts and the authenticate() function in apps/mcp/index.ts both verify the key and look up the account, but no longer apply an identical tier rule: the REST implementation still rejects a free-tier account outright, while the MCP implementation lets a free-tier account through and leaves tier enforcement to each write tool (checkFreeTierWrite in apps/mcp/index.ts), so its four read tools stay free. apps/mcp deliberately has no dependency on the Next.js app: it's a standalone service that only depends on packages/db, so this logic is duplicated, and now diverges on purpose, rather than shared across that boundary.

OpenAPI spec and CLI (REA-180)

OpenAPI 3.1 document

bash
curl https://realuptime.io/api/v1/openapi.json

Public and unauthenticated, cached for a day: it's a static description of the API, never account data, generated at request time from the same code every route runs, not hand-written per endpoint. Two things feed the generator (apps/web/lib/openapi/build-spec.ts):

  • Every path, parameter, and request body comes straight from

apps/web/app/api/v1/openapi-manifest.ts, which references the exact zod schemas the routes validate with (packages/db/api-schemas.ts) -- never a retyped copy. A drift test (apps/web/app/api/v1/openapi-drift.test.ts) walks every real route.ts file on disk and fails the suite if the manifest's path set, or the HTTP methods it declares for a file, disagree with what the file actually exports -- the same allowlist pattern apps/mcp/keyless-tools.test.ts uses for the keyless tool set.

  • Every RU-XXXX error code on every documented response is looked up

in packages/db/error-codes.ts at generation time (an unknown or retired code fails the build, not a customer's request) and carries its public KB link (https://realuptime.io/kb/errors/<code>). The same test file separately checks every code that actually appears in a route's own source is named in the manifest for that route, and vice versa, so a new refusal added to a route without documenting it here fails the suite too.

Response body schemas (packages/db/openapi-entities.ts) are best-effort documentation, not a validation boundary: they name the fields a response actually returns but mark the object additionalProperties: true, so they never claim to be an exhaustive contract the way the request schemas are. The worked JSON examples throughout this document remain the more exact reference for a response shape.

CLI

@realuptime/cli (workspace package packages/cli, bin name realuptime) is a zero-dependency Node 22 command-line client over this REST API and the public outage endpoint. Customer-facing page: https://realuptime.io/docs/cli. Published: CLI 0.2.0 is on npm (@realuptime/cli) and the Homebrew tap realuptimehq/realuptime, with signed GitHub releases on RealUptimeHQ/realuptime-cli. scripts/publish-cli.mjs runs the full pipeline (GitHub release + npm publish); the registry token is vaulted.

Install (REA-223):

bash
# npm (primary)
npm install -g @realuptime/cli

# Homebrew (tap mirrored from homebrew-realuptime/ in this repo)
brew install realuptimehq/realuptime/realuptime

# Signed release tarball, if you want to verify the binary directly
npm install -g https://github.com/RealUptimeHQ/realuptime-cli/releases/download/cli-0.2.0/realuptime-cli-0.2.0.tgz

The monorepo is private, so releases and the source mirror live on the public RealUptimeHQ/realuptime-cli repo; every release carries SHA256SUMS and a cosign SHA256SUMS.sig verifiable against https://realuptime.io/.well-known/cosign.pub. The package ships compiled JavaScript in dist/ (Node refuses to strip types under node_modules), built by packages/cli/tsconfig.build.json.

Errors SDKs. Five languages share one wire contract (the internal-docs repo's errors-plan.md, "The SDKs: three languages, one contract"):

SDKpackageinstall today
JS/Node@realuptime/errors (packages/errors-js)npm install @realuptime/errors (mirror: npm install github:RealUptimeHQ/realuptime-errors-js)
Pythonrealuptime-errors (packages/errors-py)pip install realuptime-errors (mirror: pip install "git+https://github.com/RealUptimeHQ/realuptime-errors-py")
Rubyrealuptime-errors (packages/errors-ruby)gem "realuptime-errors" (mirror: git: "https://github.com/RealUptimeHQ/realuptime-errors-ruby")
Gorealuptime-errors-go (packages/errors-go)go get github.com/RealUptimeHQ/realuptime-errors-go
PHP/Laravelrealuptime/errors (packages/errors-php)composer require realuptime/errors:dev-main against the mirror (Packagist is registered but has no tagged release yet)

JS, Python and Ruby are live on their registries (npm, PyPI, RubyGems) and that is the primary install path now; their public GitHub mirrors scripts/publish-sdk-mirrors.mjs keeps in sync remain available as a fallback for pinning an exact commit (docs/ci.md, "Publishing the SDK mirrors"). Go resolves straight from the public git repo, no registry step ever. PHP is still mirror-only: Packagist has the realuptime/errors name registered but zero tagged versions, so composer require realuptime/errors alone finds nothing yet.

bash
export REALUPTIME_API_KEY=ru_live_...

realuptime checks list
realuptime checks create --name API --url https://api.example.com/health
realuptime checks delete <id>

realuptime incidents list --limit 20
realuptime incidents create --status-page-id <id> --check-id <id> --region iad --title "API is down" --body "Investigating."
realuptime incidents update <id> --status resolved --body "Fixed." [--force]

realuptime status-pages list

realuptime maintenance start --status-page-id <id> --title "DB maintenance" --body "Applying a patch." --duration-minutes 30
realuptime maintenance start --status-page-id <id> --title "DB maintenance" --body "Applying a patch." --ends-at 2026-09-02T04:00:00Z --component-ids <id1>,<id2>
realuptime maintenance end <id>
realuptime maintenance status --status-page-id <id>

# Keyless -- no API key needed, hits the public /outages/:slug JSON endpoint.
realuptime outages github --region iad

realuptime errors issues list --project-id <id> --status open
realuptime errors releases announce --project-id <id> --release v2.4.1 --commit-sha a1b2c3d --repo-url https://github.com/acme/widgets

# Source maps (REA-254): ingest-key authed, so no API key -- the DSN is the credential.
# Walks dist for .map files, names them relative to dist, POSTs in batches of 10.
REALUPTIME_ERRORS_DSN=https://ingest.realuptime.io/api/errors/v1/ingest/rue_... \
  realuptime errors sourcemaps upload dist --release v2.4.1 [--url-prefix ~/static] [--dry-run]

# Every command accepts --json for machine-readable output, and surfaces a
# refusal's RU-XXXX code and KB link exactly as the API returned them.
realuptime checks list --json

--api-key/REALUPTIME_API_KEY and --api-url/REALUPTIME_API_URL (default https://realuptime.io/api/v1) configure every command except outages, which never sends a key, and errors sourcemaps upload, which is authenticated by the project's ingest key (--dsn/REALUPTIME_ERRORS_DSN, or --ingest-key/REALUPTIME_ERRORS_INGEST_KEY with the origin taken from --api-url) because the endpoint it wraps, POST /api/errors/v1/sourcemaps/<key>, is. The same upload core backs the SDK's Vite and webpack plugins (@realuptime/errors/vite, /webpack); see packages/cli/README.md and packages/errors-js/README.md. A refusal prints the API's own error message, its RU-XXXX code, and the KB link, then exits non-zero -- the CLI never invents its own error copy for a server-side refusal.

Monitoring as code: apply / export / validate (REA-246)

The CLI also reconciles an account to a declarative file (packages/cli/mac/, full guide at /docs/monitoring-as-code, example at packages/cli/examples/realuptime.yaml):

bash
realuptime export -f realuptime.yaml                 # bootstrap from the account (round-trips)
realuptime validate -f realuptime.yaml               # parse + validate offline
realuptime apply -f realuptime.yaml --dry-run        # plan; exit 2 if anything would change
realuptime apply -f realuptime.yaml                  # create / update
realuptime apply -f realuptime.yaml --prune --yes    # also delete http checks not in the file
realuptime apply -f realuptime.yaml --allow-replace  # accept delete+recreate for url/intervalSeconds
  • Field model: a check in the file is exactly the POST /checks body

above (checkCreateShape), same names, same rules -- the Terraform provider (REA-226) shares it. packages/cli/mac/definition-drift.test.ts pins the CLI's hand-ported validator (zero-dependency, so it can't import zod at runtime) to the zod schemas on a vector table.

  • Key: name, unique within the file. Duplicate names on the account,

or a file name that collides with a non-http check, are plan errors (exit 1, nothing written).

  • What updates in place: regions (PATCH /checks/:id) and the

assertion fields (PATCH /checks/:id/assertions). url / intervalSeconds have no edit endpoint, so a change is a replace (DELETE+POST, history lost) refused without --allow-replace. Omitting regions means all live regions; omitting intervalSeconds means unmanaged; omitting every assertion field means none (cleared if the account has some).

  • Prune: only http checks, only with --prune AND --yes. Other types

are reported and never touched.

  • Order / failure: deletes, replaces, creates, updates; stops at the

first refusal printing the API's own message and code, exit 1; re-runs recompute from live state, so they're idempotent.

  • Exit codes: 0 applied / nothing to do / clean dry run, 1 error,

2 --dry-run with pending changes.

  • YAML: a strict subset the CLI parses itself (mappings, lists, flow

lists of scalars, quoted/plain scalars, comments). Anchors, tags, block scalars, flow mappings, multi-doc are refused with a line number. .json with the same keys is always accepted.

  • Scope: v1 declares checks only -- GET /status-pages is read-only

and components / notification channels / alert rules have no v1 resource. Those top-level keys are refused with a message saying so.

JS/TS and Python SDKs, and Postman collection (REA-514)

Two typed clients over this REST API, both generated from the same openapi/v1.json the CLI and Terraform provider are hand-written against, plus a Postman collection for exploring the surface without writing code. pnpm generate:sdk (repo root) regenerates all three, and openapi-spec-file-drift.test.ts fails the gate if openapi/v1.json falls behind buildOpenApiSpec() (the function GET /api/v1/openapi.json itself serves) without someone re-running it.

SDKpackageinstall
JS/TS@realuptime/sdk (packages/sdk-js)coming soon on npm
Pythonrealuptime-sdk (packages/sdk-python)coming soon on PyPI
ts
import { RealUptimeClient } from "@realuptime/sdk";
const client = new RealUptimeClient({ apiKey: process.env.REALUPTIME_API_KEY });
const { checks } = await client.checks.list();
python
from realuptime_sdk import RealUptimeClient
client = RealUptimeClient(api_key=os.environ["REALUPTIME_API_KEY"])
checks = client.list_checks()["checks"]

Both are zero-third-party-dependency (built on fetch / urllib.request respectively, the same stance packages/cli takes) and both raise a RealUptimeApiError carrying the API's own error/code/docs fields untouched on any non-2xx response. Each covers the same named resources as the CLI (checks, status pages + components, notification channels, incidents, the keyless outage reads) with a generic request() escape hatch for everything else the spec documents (errors-tracking, agent metrics export). See packages/sdk-js/README.md and packages/sdk-python/README.md for full usage.

Generated vs hand-written, same split the OpenAPI document itself makes: request/response types are generated straight from openapi/v1.json (packages/sdk-js/src/types.gen.ts via openapi-typescript; packages/sdk-python/src/realuptime_sdk/models.py via a small JSON-Schema-to-TypedDict walker in scripts/openapi/json-schema-to-python.mjs, since there's no Java runtime dependency this repo otherwise carries and openapi-python-client's generated shape doesn't match the zero-dependency stance the rest of this codebase's client tooling holds). The client methods, error types, and READMEs are hand-written against those generated types, same division the CLI makes between generated route knowledge and hand-written commands.

Postman. postman/realuptime-api.postman_collection.json (Collection v2.1.0, hand-rolled from the spec by scripts/openapi/postman.mjs rather than the openapi-to-postmanv2 package, whose dependency tree is large for what is a straightforward walk of this repo's own spec shape): one request per operation grouped by tag, a bearer-auth entry wired to the collection's {{apiKey}} variable on every route the spec marks as requiring one, and none on the keyless routes. Import the file directly in Postman; not published to the Postman public API network.

Not done by this repo: publishing to npm, PyPI, or the Postman network. Each needs a credential this lane doesn't have (an @realuptime npm org token, a PyPI API token, a Postman API key/team) -- see each package's README "Publishing" section for the exact command to run once those exist.

Terraform provider (REA-226)

packages/terraform-provider-realuptime (module github.com/RealUptimeHQ/terraform-provider-realuptime, built on terraform-plugin-framework) is a third client of this same API, alongside the CLI and the MCP server. Resources realuptime_check, realuptime_status_page, realuptime_status_page_component and realuptime_notification_channel, plus a realuptime_status_page data source, each hand-mapped onto the endpoints above. Attribute names are this API's own names, so a .tf file and a curl command read alike; the one place they differ is the branding write shape, which is camelCase in the request body and snake_case in the response (documented under PATCH /status-pages/:id above), and the provider absorbs that difference rather than exposing both spellings.

The status page, component and notification channel write routes documented above were added FOR this provider: before REA-226 the v1 surface could read status pages and nothing else. There is deliberately still no API route for custom domains, page password protection, component grouping and reordering, or sending a channel test: each is a multi-step or secret-handling flow that stays in the dashboard, and the provider marks the corresponding attributes read-only rather than pretending to manage them.

It is source available, not published to the Terraform Registry. The RealUptimeHQ/realuptime source address does not resolve today; customers build the provider and use a dev_overrides block (see /docs/terraform). scripts/publish-terraform-provider.mjs mirrors the package to its public repo, and the package ships a GoReleaser config and a release workflow so the mirror is release-ready, but the Registry publish itself needs a public repo, a GPG key registered with HashiCorp, and a tagged release. Those are owner steps, listed at the top of the publish script.