Glossary

Monitoring glossary

Alerting, observability and reliability terms — explained in plain language.

Monitoring, alerting and observability come with a lot of vocabulary. This glossary explains the terms you will meet most often — briefly, in plain language, and with a note on where each one matters in practice. For the bigger picture, read What is monitoring?

Alert
A notification that asks a human to look at something, sent when a monitored condition is met — for example an error rate above a threshold. A good alert names the affected system, says what is wrong, and links to the details. See alerting people will not ignore.
Alert fatigue
The state in which a team receives so many irrelevant or duplicate alerts that it starts ignoring them — including the important ones. Reduced by alerting on symptoms, using time windows, deduplicating with a cooldown, and deleting alerts that never lead to action. See how to reduce alert fatigue.
APM (Application Performance Monitoring)
Monitoring of an application from the inside: request latency, throughput, error rates and slow database calls, usually per endpoint and often with traces. APM tools are request-centric and typically see little of what happens in background jobs.
Background job
Work that runs outside of a user request — scheduled tasks, queue consumers, imports, email sending. In .NET, Hangfire is the most widely used framework for this — see Hangfire vs. Quartz.NET vs. TickerQ. Background jobs fail quietly and need their own monitoring.
Black-box monitoring
Checking a system from the outside, the way a user would see it, without knowledge of its internals — for example an HTTP probe against the login page. The counterpart is white-box monitoring.
Cooldown
The minimum time between two notifications for the same ongoing problem. A 30-minute cooldown turns a two-hour outage into a handful of messages instead of hundreds.
Cron expression
A compact string such as 0 3 * * 1-5 that describes a recurring schedule (minute, hour, day of month, month, day of week). Try the cron expression explainer to see when an expression fires.
Dead man's switch
Another name for heartbeat monitoring: an alert that fires when an expected signal stops arriving, rather than when something bad happens.
Error rate
The number or share of failed operations within a time window, such as “5 failed jobs in 10 minutes”. Evaluating errors over a window is far less noisy than alerting on every single exception.
Four golden signals
Latency, traffic, errors and saturation — the four measurements Google’s SRE book recommends for any user-facing service.
Heartbeat
A small, periodic “I am alive” message from a process to a monitoring system. If heartbeats stop arriving within a timeout, the process is considered down — the only reliable way to detect failures that produce no error at all. Pick the timeout with the heartbeat timeout calculator.
Incident
An unplanned interruption or degradation of a service that needs a response. Monitoring exists to detect incidents early; post-mortems exist to prevent them from repeating.
Latency and percentiles (p95, p99)
Latency is how long an operation takes. Because averages hide slow outliers, it is measured as percentiles: a p95 of 800 ms means 95 % of requests are faster than 800 ms and 5 % are slower.
Logs
Timestamped records of individual events. Structured logs (key-value or JSON) can be searched and filtered; free-text logs mostly cannot. Logs are where you investigate after an alert.
Metrics
Numeric measurements sampled over time, such as CPU usage or requests per second. Cheap to store and query — the foundation of dashboards and threshold alerts.
MTTD / MTTR
Mean time to detect and mean time to resolve (or recover). The two numbers monitoring exists to reduce: how long a problem goes unnoticed, and how long it takes to fix once noticed.
Observability
A property of a system: how well its internal state can be understood from the telemetry it emits. Monitoring answers known questions; observability lets you answer new ones.
OpenTelemetry
An open, vendor-neutral standard and set of SDKs for producing traces, metrics and logs. It lets you instrument once and send the data to any compatible backend.
Queue backlog
Work that has been enqueued but is not being processed — because workers are down, saturated, or listening on the wrong queue. See jobs stuck as Enqueued.
RED method
Rate, errors, duration: three metrics that describe any request-driven service.
Retention
How long monitoring data is kept before it is deleted. Short retention saves cost; long retention lets you answer “when did this start?” and compare months or seasons.
Retry
Automatically running a failed operation again. Retries hide transient failures — useful for users, dangerous for monitoring if the failed attempts are never recorded.
SLI, SLO and SLA
A service level indicator (SLI) is a measurement, such as the share of successful jobs. A service level objective (SLO) is the target for it, such as 99.5 %. A service level agreement (SLA) is a contractual promise to a customer, usually looser than the internal SLO.
Synthetic monitoring
Scripted checks that simulate user actions against a live system at regular intervals — from a simple uptime ping to a full login-and-checkout flow.
Traces
A record of one request’s path through multiple services, with the duration of each step (span). Traces answer “where did the time go?” in distributed systems.
Uptime monitoring
Periodically checking whether a service responds, usually from several external locations. The simplest form of black-box monitoring.
USE method
Utilization, saturation, errors: a checklist for every resource (CPU, memory, disks, connection pools, queues) to find bottlenecks.
White-box monitoring
Monitoring based on internal data the system exposes about itself — metrics, logs, traces and state changes. It sees causes that black-box checks cannot, but only from inside.
← Monitoring guide Try QueueHawk free →

Monitoring for the jobs nobody watches

History, error-rate and heartbeat alerts for Hangfire. Free for one application.

Start free