I treat heartbeats with suspicion because a missing heartbeat does not prove a node is dead. It may be paused, overloaded, or isolated, so I recommend pairing heartbeat rules with timeout budgets and recovery behavior.
Failure detection is a suspicion, not a fact
I am looking at the messy part here: late messages, duplicates, missing messages, and the decision the system makes.
A coordinator usually waits for several missed heartbeats before taking action. That delay reduces false alarms but increases recovery time. The right values depend on the cost of moving work compared with the cost of waiting.
Worker monitor example
I use the workflow to keep the idea honest: who acts, what can fail, and how the team knows it worked.
A worker sends a heartbeat every ten seconds. After three missed heartbeats, the scheduler moves the job to another worker and marks the first one as suspect.
Heartbeat decision code
When I show code, I want it to reveal the production decision, not just the syntax.
Failure detection should mark a worker as suspect before moving work:
void checkHeartbeat(String workerId, Instant lastSeen, Instant now) {
if (lastSeen.plusSeconds(30).isBefore(now)) {
scheduler.markSuspect(workerId);
}
}
The check marks a worker as suspect only after the heartbeat crosses the timeout. It is a failure signal, not proof, which is why real systems pair it with retries or quorum.
Sequence diagram: suspect after missed heartbeats
The scheduler acts after repeated evidence, not after the first missing signal.
The timeout choice
- Tune heartbeat intervals to the business cost of delay.
- Add jitter so every node does not report at the same instant.
- Alert on patterns, not isolated missed beats.