Blog

Alert fatigue: 7 causes and how to fix them

Every monitoring setup starts with the fear of missing something, so the first alerts are generous. A few months later the alert channel is muted, the email filter sends everything to a folder, and the incident that finally matters is found by a customer. That is alert fatigue — and it is almost never a people problem. It is a design problem in the alerts.

What alert fatigue actually is

Alert fatigue is the gradual loss of trust in alerts. It happens when too many notifications turn out to be irrelevant, duplicated, or impossible to act on. Each false alarm teaches the team that the next alert can probably wait. The dangerous part is that the learning is correct most of the time — until it is not.

You can usually see it before anyone names it: an alert channel that is muted by half the team, emoji reactions instead of replies, alerts that are acknowledged within seconds (too fast to have looked at anything), and the phrase “oh, that one always fires”.

Seven causes, and the fix for each

1. Alerting on causes instead of symptoms

“CPU above 85 %” or “memory above 90 %” fires often and rarely means users are affected. Alert on what users or customers experience — error rate, failed jobs, work that did not happen — and keep resource metrics on dashboards for diagnosis. A saturated resource that nobody notices is not an incident.

2. Single events instead of time windows

One failed request or one failed job is normal in any real system; timeouts and transient network errors happen. Alert on a pattern: “5 failures within 10 minutes” instead of “any failure”. This one change usually removes most of the noise, and it is how QueueHawk's error-rate rule works for background jobs.

3. No deduplication for ongoing problems

A problem that lasts two hours should not produce a message every minute. Use a cooldown: after an alert fires, stay quiet for a defined period while the condition persists, then send at most a reminder. Thirty minutes is a reasonable starting point for most teams.

4. No “resolved” message

Without a recovery notification, people cannot tell from the channel whether something is still broken. They check manually — or, more often, assume it sorted itself out. Always send the resolution, and send it only once.

5. Alerts without context

“Error rate exceeded” is not actionable. Which application? Which environment — production or staging? Which job or endpoint? What was the last error? An alert that forces the reader to open three tools before they know whether it matters is an alert that gets postponed. Put the application, environment, the affected job types, the last exception and a direct link into the message itself.

6. Every alert goes to every channel

Route by urgency and audience. Production outages go to the team's chat channel (or a paging tool via webhook); staging failures and low-priority warnings go to a separate channel or an email digest. When staging noise lands next to production alerts, production alerts get muted too.

7. Nobody reviews the alerts

Alerts accumulate. Once a month, look at every alert that fired and ask one question: did this lead to an action? If it did not, raise the threshold, widen the window, change the route, or delete it. Deleting an alert that never led to action does not reduce safety; it increases the chance that the remaining ones are read.

The special case: silence

Tuning alerts to be quieter has one trap. The most important failures of scheduled and background work produce no error at all: a scheduler that stopped, a worker that crashed, a server that was removed. Quieting error alerts does nothing for those — they need a separate heartbeat check that fires when an expected signal stops arriving. Choose its timeout carefully: too short and every deploy looks like an outage (a direct source of fatigue), too long and you find out an hour late. The heartbeat timeout calculator helps pick a value.

A quick self-check

For the broader picture, see What is monitoring?; the terms used here are explained in the monitoring glossary.

FAQ

What is alert fatigue?

Alert fatigue is what happens when a team receives so many alerts that are irrelevant, duplicated or not actionable that people stop reacting to them — including to the alerts that matter. It is a trust problem: every false alarm lowers the chance that the next real one gets read.

What causes alert fatigue?

The usual causes are alerts on causes instead of symptoms (CPU, memory), thresholds on single events instead of time windows, no deduplication for ongoing problems, alerts without context, one channel for everything, and alerts nobody ever reviews or deletes.

How many alerts per week is too many?

There is no universal number, but a useful rule from SRE practice is that every page should need a human action. If an on-call person regularly gets more alerts than they can investigate properly — often cited as more than two actionable incidents per shift — the alerting needs tuning, not the person.

Should monitoring send a message when a problem is resolved?

Yes. A resolved message closes the loop, stops people from investigating something that already recovered, and makes it possible to see from the channel history how long an incident lasted.

← All articles Try QueueHawk free →

Alerts that say what broke, where, and when it recovered

Error-rate and heartbeat alerts for Hangfire with cooldowns and resolved messages. Free for one application.

Start free