Guide

What is monitoring? A practical guide to metrics, alerts and heartbeats

Monitoring is the practice of continuously collecting signals from a system and turning them into answers to one question: is it working right now, and will I know when it stops? This guide covers what monitoring actually is, which signals matter, how to alert without drowning your team, and the blind spot almost every setup has — work that runs in the background.

What is monitoring?

In software operations, monitoring means collecting data about a running system — measurements, events, logs — evaluating it against expectations, and notifying a human when reality and expectation diverge. The goal is not to collect as much data as possible. The goal is to shorten two intervals:

The worst possible outcome is a problem that a customer detects before you do. Every monitoring decision in this guide can be measured against that: does it make it less likely that the first report comes from outside?

Monitoring vs. observability

The two terms are often used interchangeably, but they describe different things. Monitoring answers questions you knew to ask in advance: is the error rate above 2 %? Did the nightly import run? Observability is a property of a system — how well you can answer questions you did not anticipate, by exploring its telemetry after the fact.

You need both. Observability without monitoring means you can investigate beautifully, but only once someone tells you there is something to investigate. Monitoring without observability means the pager goes off and you have no way to find out why. In practice, most small and mid-sized teams get more value from a small number of well-chosen monitors than from an expensive observability platform nobody has time to explore.

The four kinds of signals

Almost everything a monitoring system consumes falls into one of four categories:

Metrics

Numbers sampled over time: requests per second, CPU usage, queue length, error count. Metrics are cheap to store and fast to query, which makes them the backbone of dashboards and threshold alerts. Their weakness is context — a metric tells you the error rate doubled, not which customer or which code path is affected.

Logs

Timestamped records of discrete things that happened, ideally structured (JSON with named fields) rather than free text. Logs carry the detail metrics lack, but they are expensive at volume and easy to drown in. Treat them as the place you go after an alert, not as the alert itself.

Traces

A trace follows a single request as it crosses services, queues and databases, recording how long each step took. Traces are the fastest way to answer “where did the time go?” in distributed systems. OpenTelemetry has become the standard way to produce them.

Events and state changes

The often-forgotten fourth category: discrete business or lifecycle events such as deploy finished, job enqueued, job failed, payment captured. They are what make monitoring of workflows possible — you cannot tell that a job is stuck from CPU metrics, but you can tell from the fact that it was enqueued twenty minutes ago and never started.

What to monitor: the layers of a system

A useful mental model is to monitor a system from the bottom up and from the outside in:

  1. Infrastructure — hosts, containers, disks, network. Disk-full and out-of-memory are still among the most common causes of outages.
  2. Dependencies — databases, caches, message brokers, third-party APIs. Connection pool exhaustion and slow queries show up here first.
  3. Application (APM) — request latency, error rates, exceptions, per endpoint.
  4. Background work — scheduled jobs, queue consumers, batch imports, email sending. Covered in detail below, because it is where most setups have a gap.
  5. Synthetic / uptime checks — an external probe that calls your public endpoints from outside, exactly like a user would. It catches DNS, TLS and routing problems that no internal metric can see.
  6. Business outcomes — signups per hour, orders per day, invoices sent. When all technical signals are green but this one drops, something is broken in a way you did not think to monitor.

You do not need all six on day one. But you should know which ones you are not covering, and what a failure in that layer would look like from the outside.

Golden signals, RED and USE

Three frameworks help decide which metrics matter, so you do not end up with 400 dashboards and no idea which one to look at.

FrameworkSignalsBest for
Four golden signals (Google SRE)Latency, traffic, errors, saturationAny user-facing service
REDRate, errors, durationRequest-driven services and APIs
USEUtilization, saturation, errorsResources: CPU, memory, disks, pools, queues

The frameworks overlap on purpose. RED describes the work a service does, USE describes the resources it does the work with. A slow endpoint (RED: duration) with a saturated connection pool (USE: saturation) is a diagnosis; either signal alone is just a symptom.

One detail matters more than any framework: measure latency as percentiles, not averages. An average of 200 ms can hide the fact that one request in twenty takes eight seconds. The p95 or p99 is what your unhappiest users experience.

Monitoring the absence of a signal

Most monitoring is built to react to something that happens: an error, a spike, a timeout. The hardest failures to catch are the ones where nothing happens:

None of these produce an error, so no error-based alert fires. The technique that catches them is heartbeat monitoring, also called a dead man's switch: the process sends a small “I am alive” signal at a regular interval, and the monitoring system alerts when the signal stops arriving within an expected window. It inverts the logic — silence becomes the alarm.

Two practical rules make heartbeats reliable: set the timeout to a few multiples of the interval (a 30-second heartbeat with a 90-second timeout tolerates one lost packet without paging anyone), and evaluate it per service, not per host. In containerized environments every deploy replaces hosts; if each old pod counts as “dead”, every release looks like an outage.

Alerting people will not ignore

An alert is a request for a human's attention. Every alert that turns out to be irrelevant makes the next one a little less likely to be read — that is alert fatigue, and it is the most common reason monitoring setups fail in practice even though the data was there. A few rules keep alerts trustworthy:

If you set targets formally, this is where SLOs (service level objectives) come in: you define what “good enough” means — say, 99.5 % of nightly jobs succeed on the first attempt — and alert when you are burning through the allowed failure budget too fast, instead of on every individual failure.

How long to keep monitoring data

Retention is a trade-off between cost and the questions you can still answer. A rough guide:

Be aware of where retention is decided for you. Many tools keep only a short, fixed window by default — Hangfire's built-in storage, for example, expires succeeded jobs after about a day (details here). And keep privacy in mind: monitoring data often contains personal data in exception messages or payloads, so retention should end with actual deletion, not just with hiding old data.

The blind spot: background jobs

Most monitoring stacks grow out of the request path: uptime checks, APM, error trackers that hook into the web framework. They see a request come in and a response go out. What they do not see is the work that happens outside of any request — and in many applications, that is where the business logic that actually makes money lives: invoice runs, data imports, report generation, email campaigns, synchronisation with ERPs and partner systems.

Background jobs have three properties that make them hard to monitor with request-centric tools:

  1. They fail quietly. Most job frameworks retry automatically. A job that fails, retries and eventually succeeds is fine; a job that fails ten times over two days and then gives up may leave no trace that anyone reads.
  2. Their worst failure is silence. A recurring job that stops being scheduled does not throw — it just does not run. Only a heartbeat or a “this should have happened by now” check detects it.
  3. Their history is short-lived. Job frameworks store state for their own operation, not for your analysis. After a deploy or a cleanup, the evidence of what ran last week is often gone.

For background jobs, useful monitoring means: every state change (enqueued, processing, succeeded, failed) recorded outside the application, so it survives restarts; an error-rate rule per application or job type; a heartbeat rule per application that fires when all of its workers go quiet; and a history long enough to answer “when did this start failing?”. For .NET teams using Hangfire, that is exactly what QueueHawk does — see Hangfire monitoring for how it hooks in, or the two alert rules that cover most failures.

A practical monitoring checklist

Use this to find gaps in an existing setup. Every “no” is a failure you would currently learn about from a customer.

Unfamiliar with a term? The monitoring glossary explains all of them in one place. And if you schedule jobs with cron expressions, the free cron expression explainer shows exactly when they will run.

FAQ

What is the difference between monitoring and observability?

Monitoring answers questions you knew to ask in advance: is the error rate above 2 %, did the nightly job run? Observability is the ability to answer questions you did not anticipate, by exploring rich telemetry such as high-cardinality traces and structured logs. Monitoring is a practice built on top of observable systems — you need both.

What are the four golden signals?

Latency, traffic, errors and saturation. They come from Google's Site Reliability Engineering book and describe the health of any user-facing service from the outside: how long requests take, how many there are, how many fail, and how close the service is to its capacity limits.

What is heartbeat monitoring?

Heartbeat monitoring (also called a dead man's switch) expects a process to check in regularly. If the expected signal does not arrive within a time window, an alert fires. It is the only reliable way to detect things that fail silently — a cron job that never started, a worker that crashed, a server that is gone.

How long should monitoring data be kept?

It depends on the question you want to answer. Minutes to days are enough for alerting, weeks cover incident analysis and deploy comparisons, and months to a year are needed for capacity planning, seasonal patterns and customer-facing reporting. Keep raw data short and aggregated data long if volume matters.

Do background jobs need separate monitoring?

Yes. Uptime checks and request metrics only see the synchronous part of an application. Background jobs run outside of any request, often at night, and their most common failure mode is not an error but silence — a job that simply stops running. They need job-level history, error-rate alerts and heartbeat checks.

← All articles Try QueueHawk free →

Close the background-job gap in your monitoring

History, error-rate and heartbeat alerts for Hangfire. Free for one application, no card required.

Start free