I do not read partition tolerance as a promise that everything stays perfect. I read it as a reminder to define safe behavior when links break and both sides still want to continue.
A split needs policy, not hope
This is the test I use: what happens when messages are slow, missing, or repeated?
When communication breaks, services still receive user traffic. The design must decide whether to serve old reads, accept writes, queue work, or refuse certain operations.
Cross zone example
The point here is to move from concept to operation: who uses it, what breaks, and what decision changes.
A wallet service cannot reach enough replicas to confirm a withdrawal. It allows balance reads from cache but blocks the withdrawal until quorum is reachable.
Partition aware code
The code should make the decision clear: what the system allows, rejects, retries, or sends for review.
The service keeps low risk reads open and closes high risk writes:
WalletResponse handle(WalletCommand command, boolean quorumReachable) {
if (!quorumReachable && command.isWithdrawal()) {
return WalletResponse.rejected("Quorum unavailable");
}
return wallet.execute(command);
}
The wallet rejects withdrawals when quorum is unavailable, because accepting them during a partition could spend money the system cannot safely verify.
Sequence diagram: tolerate the split with a safe rule
The service tolerates the partition by keeping safe operations available and unsafe operations closed.
The failure assumption
- Separate low risk reads from high risk writes.
- Design user messages for partition behavior.
- Measure how often partition rules are triggered.