Monitoring is the practice of continuously collecting signals from a system and turning them into answers to one question: is it working right now, and will I know when it stops? This guide covers what monitoring actually is, which signals matter, how to alert without drowning your team, and the blind spot almost every setup has — work that runs in the background.
In software operations, monitoring means collecting data about a running system — measurements, events, logs — evaluating it against expectations, and notifying a human when reality and expectation diverge. The goal is not to collect as much data as possible. The goal is to shorten two intervals:
The worst possible outcome is a problem that a customer detects before you do. Every monitoring decision in this guide can be measured against that: does it make it less likely that the first report comes from outside?
The two terms are often used interchangeably, but they describe different things. Monitoring answers questions you knew to ask in advance: is the error rate above 2 %? Did the nightly import run? Observability is a property of a system — how well you can answer questions you did not anticipate, by exploring its telemetry after the fact.
You need both. Observability without monitoring means you can investigate beautifully, but only once someone tells you there is something to investigate. Monitoring without observability means the pager goes off and you have no way to find out why. In practice, most small and mid-sized teams get more value from a small number of well-chosen monitors than from an expensive observability platform nobody has time to explore.
Almost everything a monitoring system consumes falls into one of four categories:
Numbers sampled over time: requests per second, CPU usage, queue length, error count. Metrics are cheap to store and fast to query, which makes them the backbone of dashboards and threshold alerts. Their weakness is context — a metric tells you the error rate doubled, not which customer or which code path is affected.
Timestamped records of discrete things that happened, ideally structured (JSON with named fields) rather than free text. Logs carry the detail metrics lack, but they are expensive at volume and easy to drown in. Treat them as the place you go after an alert, not as the alert itself.
A trace follows a single request as it crosses services, queues and databases, recording how long each step took. Traces are the fastest way to answer “where did the time go?” in distributed systems. OpenTelemetry has become the standard way to produce them.
The often-forgotten fourth category: discrete business or lifecycle events such as deploy finished, job enqueued, job failed, payment captured. They are what make monitoring of workflows possible — you cannot tell that a job is stuck from CPU metrics, but you can tell from the fact that it was enqueued twenty minutes ago and never started.
A useful mental model is to monitor a system from the bottom up and from the outside in:
You do not need all six on day one. But you should know which ones you are not covering, and what a failure in that layer would look like from the outside.
Three frameworks help decide which metrics matter, so you do not end up with 400 dashboards and no idea which one to look at.
| Framework | Signals | Best for |
|---|---|---|
| Four golden signals (Google SRE) | Latency, traffic, errors, saturation | Any user-facing service |
| RED | Rate, errors, duration | Request-driven services and APIs |
| USE | Utilization, saturation, errors | Resources: CPU, memory, disks, pools, queues |
The frameworks overlap on purpose. RED describes the work a service does, USE describes the resources it does the work with. A slow endpoint (RED: duration) with a saturated connection pool (USE: saturation) is a diagnosis; either signal alone is just a symptom.
One detail matters more than any framework: measure latency as percentiles, not averages. An average of 200 ms can hide the fact that one request in twenty takes eight seconds. The p95 or p99 is what your unhappiest users experience.
Most monitoring is built to react to something that happens: an error, a spike, a timeout. The hardest failures to catch are the ones where nothing happens:
None of these produce an error, so no error-based alert fires. The technique that catches them is heartbeat monitoring, also called a dead man's switch: the process sends a small “I am alive” signal at a regular interval, and the monitoring system alerts when the signal stops arriving within an expected window. It inverts the logic — silence becomes the alarm.
Two practical rules make heartbeats reliable: set the timeout to a few multiples of the interval (a 30-second heartbeat with a 90-second timeout tolerates one lost packet without paging anyone), and evaluate it per service, not per host. In containerized environments every deploy replaces hosts; if each old pod counts as “dead”, every release looks like an outage.
An alert is a request for a human's attention. Every alert that turns out to be irrelevant makes the next one a little less likely to be read — that is alert fatigue, and it is the most common reason monitoring setups fail in practice even though the data was there. A few rules keep alerts trustworthy:
If you set targets formally, this is where SLOs (service level objectives) come in: you define what “good enough” means — say, 99.5 % of nightly jobs succeed on the first attempt — and alert when you are burning through the allowed failure budget too fast, instead of on every individual failure.
Retention is a trade-off between cost and the questions you can still answer. A rough guide:
Be aware of where retention is decided for you. Many tools keep only a short, fixed window by default — Hangfire's built-in storage, for example, expires succeeded jobs after about a day (details here). And keep privacy in mind: monitoring data often contains personal data in exception messages or payloads, so retention should end with actual deletion, not just with hiding old data.
Most monitoring stacks grow out of the request path: uptime checks, APM, error trackers that hook into the web framework. They see a request come in and a response go out. What they do not see is the work that happens outside of any request — and in many applications, that is where the business logic that actually makes money lives: invoice runs, data imports, report generation, email campaigns, synchronisation with ERPs and partner systems.
Background jobs have three properties that make them hard to monitor with request-centric tools:
For background jobs, useful monitoring means: every state change (enqueued, processing, succeeded, failed) recorded outside the application, so it survives restarts; an error-rate rule per application or job type; a heartbeat rule per application that fires when all of its workers go quiet; and a history long enough to answer “when did this start failing?”. For .NET teams using Hangfire, that is exactly what QueueHawk does — see Hangfire monitoring for how it hooks in, or the two alert rules that cover most failures.
Use this to find gaps in an existing setup. Every “no” is a failure you would currently learn about from a customer.
Unfamiliar with a term? The monitoring glossary explains all of them in one place. And if you schedule jobs with cron expressions, the free cron expression explainer shows exactly when they will run.
Monitoring answers questions you knew to ask in advance: is the error rate above 2 %, did the nightly job run? Observability is the ability to answer questions you did not anticipate, by exploring rich telemetry such as high-cardinality traces and structured logs. Monitoring is a practice built on top of observable systems — you need both.
Latency, traffic, errors and saturation. They come from Google's Site Reliability Engineering book and describe the health of any user-facing service from the outside: how long requests take, how many there are, how many fail, and how close the service is to its capacity limits.
Heartbeat monitoring (also called a dead man's switch) expects a process to check in regularly. If the expected signal does not arrive within a time window, an alert fires. It is the only reliable way to detect things that fail silently — a cron job that never started, a worker that crashed, a server that is gone.
It depends on the question you want to answer. Minutes to days are enough for alerting, weeks cover incident analysis and deploy comparisons, and months to a year are needed for capacity planning, seasonal patterns and customer-facing reporting. Keep raw data short and aggregated data long if volume matters.
Yes. Uptime checks and request metrics only see the synchronous part of an application. Background jobs run outside of any request, often at night, and their most common failure mode is not an error but silence — a job that simply stops running. They need job-level history, error-rate alerts and heartbeat checks.
History, error-rate and heartbeat alerts for Hangfire. Free for one application, no card required.