A heartbeat check fires when a process stops checking in. Set its timeout too short and every late packet or deploy becomes a false alarm; too long and you learn about an outage an hour late. Enter your setup to get a timeout that tolerates normal noise, and see how quickly a real failure would be detected.
The model is intentionally simple, so you can check it by hand:
For Hangfire, the QueueHawk agent sends a heartbeat every 30 seconds per Hangfire server (HeartbeatIntervalSeconds). Every new application gets a heartbeat rule with a 90-second timeout, evaluated about once a minute, and the rule looks at the application as a whole — it fires only when no server of that application is still sending, so rolling deploys do not trigger it. Both values can be changed per application. More on the idea behind this in monitoring the absence of a signal.
A common starting point is three times the heartbeat interval plus a small buffer — for a 30-second heartbeat, roughly 90 to 100 seconds. That tolerates two missed heartbeats without a false alarm while still detecting a real outage within a couple of minutes.
Because heartbeats are never perfectly punctual. Network delays, garbage-collection pauses and batching make some arrive a few seconds late; with a timeout equal to the interval, every late heartbeat becomes an alert, and the team soon learns to ignore them.
If a deploy stops the only instance before the new one starts, there is a gap without heartbeats. Either make the timeout longer than that gap, or deploy with several replicas and rolling updates so there is always at least one instance sending heartbeats — and evaluate heartbeats per application, not per server.
Evaluated per application, not per server. Free for one application.