Free tool

Heartbeat timeout calculator

A heartbeat check fires when a process stops checking in. Set its timeout too short and every late packet or deploy becomes a false alarm; too long and you learn about an outage an hour late. Enter your setup to get a timeout that tolerates normal noise, and see how quickly a real failure would be detected.

How often the process checks in.
How many may go missing without an alert.
Network, GC pauses, batching.
0 for rolling deploys with several replicas.
How often monitoring checks.
Recommended timeout
—
Time to alert after a failure
—

How the calculation works

The model is intentionally simple, so you can check it by hand:

Three rules of thumb

  1. Tolerate at least one missed heartbeat. Two is a sensible default. Zero guarantees false alarms.
  2. Evaluate per service, not per host. In container environments every deploy replaces hosts. If each old instance counts as “down”, every release pages someone. A service is down when none of its instances sends heartbeats.
  3. Shorten the interval before lengthening the timeout. If detection is too slow, a heartbeat every 15 seconds with a 45-second timeout detects faster than a 60-second heartbeat with the same tolerance — at a negligible cost in traffic.

QueueHawk's defaults

For Hangfire, the QueueHawk agent sends a heartbeat every 30 seconds per Hangfire server (HeartbeatIntervalSeconds). Every new application gets a heartbeat rule with a 90-second timeout, evaluated about once a minute, and the rule looks at the application as a whole — it fires only when no server of that application is still sending, so rolling deploys do not trigger it. Both values can be changed per application. More on the idea behind this in monitoring the absence of a signal.

FAQ

What is a good heartbeat timeout?

A common starting point is three times the heartbeat interval plus a small buffer — for a 30-second heartbeat, roughly 90 to 100 seconds. That tolerates two missed heartbeats without a false alarm while still detecting a real outage within a couple of minutes.

Why not set the timeout equal to the interval?

Because heartbeats are never perfectly punctual. Network delays, garbage-collection pauses and batching make some arrive a few seconds late; with a timeout equal to the interval, every late heartbeat becomes an alert, and the team soon learns to ignore them.

How do deploys affect heartbeat alerts?

If a deploy stops the only instance before the new one starts, there is a gap without heartbeats. Either make the timeout longer than that gap, or deploy with several replicas and rolling updates so there is always at least one instance sending heartbeats — and evaluate heartbeats per application, not per server.

Cron expression explainer Try QueueHawk free →

Heartbeat alerts for Hangfire, without the deploy noise

Evaluated per application, not per server. Free for one application.

Start free