Modelled line by line on src/cluster.c, Redis 7.2.14. Every constant, threshold and quorum below is annotated with the line it comes from; defaults come from src/config.c.
A master dies. Its replicas do not simply take over — they wait a delay ranked by
replication offset, then ask the other masters to vote. Two things this normally gets wrong:
the quorum is computed over every master holding at least one slot including the dead
one (cluster.c:5162), and under the default cluster-require-full-coverage
the whole cluster returns CLUSTERDOWN for the entire window, not just the dead master's
slots (cluster.c:5137). Four masters are used here on purpose: with three, the right and
the wrong quorum formula both give 2.
crc16(key) & 0x3FFF (cluster.c:1387). With a {tag} only the
bytes between the first { and the next } are hashed
(cluster.c:1398), which is how related keys are forced onto one master. No closing
}, or an empty {}, falls back to hashing the whole key
(cluster.c:1394). The CRC is XMODEM, polynomial 0x1021, init 0 (crc16.c:32-40).
A node sets PFAIL on a peer locally when the smaller of its ping delay and its
last-data delay exceeds cluster-node-timeout (cluster.c:4861-4865). PFAIL
is only ever a local opinion, which is why each card below shows a different view.
Promotion to FAIL needs (cluster->size / 2) + 1 agreeing
masters (cluster.c:2002). Two conditions usually left out: the node doing the promoting
must itself be holding PFAIL on the target (cluster.c:2004) — the comment at
cluster.c:1486-1490 says outright that it does not just trust other nodes' counts — and
it counts itself when it is a master (cluster.c:2009), so this is a majority of
all masters, not of the other masters. Failure reports expire after
node_timeout × 2 (cluster.c:1496-1497, multiplier at cluster.h:16),
so stale reports cannot accumulate into a quorum.
A replica runs the failover only if it is a replica, its master is FAIL (or this is a
manual failover), cluster-replica-no-failover is not set, and its master
still holds slots (cluster.c:4302-4306).
The delay is 500 + random() % 500 + rank × 1000 milliseconds
(cluster.c:4347-4357): 500ms fixed so the FAIL message can propagate, up to 500ms of
jitter to break ties, and one full second per rank. Rank is the number of sibling
replicas with a strictly greater replication offset, skipping ones flagged NOFAILOVER
(cluster.c:4142-4157), so the freshest replica gets rank 0 and goes first. Rank is
re-checked just before the votes go out and the delay grows if a sibling has since
reported a better offset (cluster.c:4382-4395).
Data age is checked first. It is the time disconnected from the master, reduced by one
node_timeout (cluster.c:4326-4327) because being disconnected at least
that long is exactly what FAIL means. The replica is disqualified when that age exceeds
repl-ping-replica-period × 1000 + node_timeout × factor
(cluster.c:4333-4336) — with stock defaults about 175 seconds, not 150. A factor of
0 disables the check entirely (the guard at cluster.c:4333).
The replica bumps currentEpoch, records it as the election epoch and
broadcasts FAILOVER_AUTH_REQUEST (cluster.c:4411-4415). A replica sends its
master's slot bitmap and configEpoch, not its own (cluster.c:3577-3581,
cluster.c:3604, cluster.c:3621).
A master grants the vote only if: it is a master serving at least one slot
(cluster.c:4039); the request epoch is not older than its currentEpoch
(cluster.c:4045); it has not already voted this epoch (cluster.c:4055); the requester
is a replica whose master is FAIL, unless FORCEACK is set (cluster.c:4066-4067); it
has not voted for another replica of that same master within
node_timeout × 2 (cluster.c:4088); and no slot claimed is currently
served by a master with a greater configEpoch (cluster.c:4102-4119).
Winning needs (cluster->size / 2) + 1 votes (cluster.c:4277), where
size is the count of masters holding at least one slot — and the FAIL flag
does not remove a master from that count (cluster.c:5162; the FAIL test one line
later at cluster.c:5164 feeds a separate reachable_masters counter). The
dead master is still in the denominator while its slots are unclaimed. With the four
masters here that is 3 votes needed from the 3 that can still answer.
An election that does not reach quorum is not retried immediately. It expires after
MAX(node_timeout × 2, 2000) (cluster.c:4291-4292, cluster.c:4404)
and a fresh one at a new epoch is scheduled after twice that again
(cluster.c:4293, cluster.c:4346).
The winner raises its configEpoch to the election epoch (cluster.c:4431-4432), then
clusterFailoverReplaceYourMaster turns it into a master, moves every slot
bit from the old master to itself, updates cluster state and PONGs every node
(cluster.c:4236-4264).
Everyone else converges through clusterUpdateSlotsConfigWith: a slot is
rebound when the claimant's configEpoch beats the current owner's
(cluster.c:2441-2442), and any node whose own master ends up with zero slots
reconfigures as a replica of the claimant (cluster.c:2489-2495). That single rule covers
both the surviving sibling replicas and the old master when it comes back.
CLUSTER FAILOVER is sent to a replica (cluster.c:6491-6493). With no
argument the master must still be reachable, or the command refuses with
"Master is down or failed, please use CLUSTER FAILOVER FORCE"
(cluster.c:6497-6503), and it starts by sending MFSTART to the master
(cluster.c:6524).
FORCE skips the coordination but still needs the votes: the election delay is
zeroed (cluster.c:4359-4363), the data-age check is bypassed (cluster.c:4338),
cluster-replica-no-failover is bypassed (cluster.c:4305), and the request
carries FORCEACK (cluster.c:3999) so masters grant it even though the master is up
(cluster.c:4067).
TAKEOVER implies FORCE (cluster.c:6483) and then skips the election completely:
it calls clusterBumpConfigEpochWithoutConsensus — which just increments
currentEpoch and assigns it to itself (cluster.c:1819-1836) — and claims the slots
(cluster.c:6514-6515). No master votes. It is the only path here that can promote a
replica in a minority partition, and it can produce two masters claiming the same
slots.
int flags; /* cluster.h:123 — int, and it needs the width */
#define CLUSTER_NODE_MASTER 1 /* cluster.h:49 */
#define CLUSTER_NODE_SLAVE 2
#define CLUSTER_NODE_PFAIL 4
#define CLUSTER_NODE_FAIL 8
#define CLUSTER_NODE_MYSELF 16
#define CLUSTER_NODE_HANDSHAKE 32
#define CLUSTER_NODE_NOADDR 64
#define CLUSTER_NODE_MEET 128
#define CLUSTER_NODE_MIGRATE_TO 256
#define CLUSTER_NODE_NOFAILOVER 512 /* cluster.h:58 — not 16; 16 is MYSELF */
#define CLUSTER_NODE_EXTENSIONS_SUPPORTED 1024
/* cluster.c:4142 — verbatim */
int clusterGetSlaveRank(void) {
long long myoffset;
int j, rank = 0;
clusterNode *master;
serverAssert(nodeIsSlave(myself));
master = myself->slaveof;
if (master == NULL) return 0;
myoffset = replicationGetSlaveOffset();
for (j = 0; j < master->numslaves; j++)
if (master->slaves[j] != myself &&
!nodeCantFailover(master->slaves[j]) &&
master->slaves[j]->repl_offset > myoffset) rank++;
return rank;
}
TAKEOVER promotes with no votes at all (cluster.c:6514).