← index

Redis Cluster Failover Simulator

Modelled line by line on src/cluster.c, Redis 7.2.14. Every constant, threshold and quorum below is annotated with the line it comes from; defaults come from src/config.c.

A master dies. Its replicas do not simply take over — they wait a delay ranked by replication offset, then ask the other masters to vote. Two things this normally gets wrong: the quorum is computed over every master holding at least one slot including the dead one (cluster.c:5162), and under the default cluster-require-full-coverage the whole cluster returns CLUSTERDOWN for the entire window, not just the dead master's slots (cluster.c:5137). Four masters are used here on purpose: with three, the right and the wrong quorum formula both give 2.

Cluster

cluster.c:5128 clusterUpdateState
cluster_state: ok
advance
config.c:3223, default 15000
config.c:3174, default 10; 0 disables
config.c:3196, default 10 (seconds)
config.c:3092, default yes

Topology

16384 slots · cluster.h:8

Key → slot

cluster.c:1380 keyHashSlot

crc16(key) & 0x3FFF (cluster.c:1387). With a {tag} only the bytes between the first { and the next } are hashed (cluster.c:1398), which is how related keys are forced onto one master. No closing }, or an empty {}, falls back to hashing the whole key (cluster.c:1394). The CRC is XMODEM, polynomial 0x1021, init 0 (crc16.c:32-40).

What the source does

cluster.c
1  Failure detection: PFAIL, then FAIL

A node sets PFAIL on a peer locally when the smaller of its ping delay and its last-data delay exceeds cluster-node-timeout (cluster.c:4861-4865). PFAIL is only ever a local opinion, which is why each card below shows a different view.

Promotion to FAIL needs (cluster->size / 2) + 1 agreeing masters (cluster.c:2002). Two conditions usually left out: the node doing the promoting must itself be holding PFAIL on the target (cluster.c:2004) — the comment at cluster.c:1486-1490 says outright that it does not just trust other nodes' counts — and it counts itself when it is a master (cluster.c:2009), so this is a majority of all masters, not of the other masters. Failure reports expire after node_timeout × 2 (cluster.c:1496-1497, multiplier at cluster.h:16), so stale reports cannot accumulate into a quorum.

2  Election delay, ranked by replication offset

A replica runs the failover only if it is a replica, its master is FAIL (or this is a manual failover), cluster-replica-no-failover is not set, and its master still holds slots (cluster.c:4302-4306).

The delay is 500 + random() % 500 + rank × 1000 milliseconds (cluster.c:4347-4357): 500ms fixed so the FAIL message can propagate, up to 500ms of jitter to break ties, and one full second per rank. Rank is the number of sibling replicas with a strictly greater replication offset, skipping ones flagged NOFAILOVER (cluster.c:4142-4157), so the freshest replica gets rank 0 and goes first. Rank is re-checked just before the votes go out and the delay grows if a sibling has since reported a better offset (cluster.c:4382-4395).

Data age is checked first. It is the time disconnected from the master, reduced by one node_timeout (cluster.c:4326-4327) because being disconnected at least that long is exactly what FAIL means. The replica is disqualified when that age exceeds repl-ping-replica-period × 1000 + node_timeout × factor (cluster.c:4333-4336) — with stock defaults about 175 seconds, not 150. A factor of 0 disables the check entirely (the guard at cluster.c:4333).

3  Voting and the quorum denominator

The replica bumps currentEpoch, records it as the election epoch and broadcasts FAILOVER_AUTH_REQUEST (cluster.c:4411-4415). A replica sends its master's slot bitmap and configEpoch, not its own (cluster.c:3577-3581, cluster.c:3604, cluster.c:3621).

A master grants the vote only if: it is a master serving at least one slot (cluster.c:4039); the request epoch is not older than its currentEpoch (cluster.c:4045); it has not already voted this epoch (cluster.c:4055); the requester is a replica whose master is FAIL, unless FORCEACK is set (cluster.c:4066-4067); it has not voted for another replica of that same master within node_timeout × 2 (cluster.c:4088); and no slot claimed is currently served by a master with a greater configEpoch (cluster.c:4102-4119).

Winning needs (cluster->size / 2) + 1 votes (cluster.c:4277), where size is the count of masters holding at least one slot — and the FAIL flag does not remove a master from that count (cluster.c:5162; the FAIL test one line later at cluster.c:5164 feeds a separate reachable_masters counter). The dead master is still in the denominator while its slots are unclaimed. With the four masters here that is 3 votes needed from the 3 that can still answer.

An election that does not reach quorum is not retried immediately. It expires after MAX(node_timeout × 2, 2000) (cluster.c:4291-4292, cluster.c:4404) and a fresh one at a new epoch is scheduled after twice that again (cluster.c:4293, cluster.c:4346).

4  Promotion, and what the losers do

The winner raises its configEpoch to the election epoch (cluster.c:4431-4432), then clusterFailoverReplaceYourMaster turns it into a master, moves every slot bit from the old master to itself, updates cluster state and PONGs every node (cluster.c:4236-4264).

Everyone else converges through clusterUpdateSlotsConfigWith: a slot is rebound when the claimant's configEpoch beats the current owner's (cluster.c:2441-2442), and any node whose own master ends up with zero slots reconfigures as a replica of the claimant (cluster.c:2489-2495). That single rule covers both the surviving sibling replicas and the old master when it comes back.

5  Manual failover: FAILOVER, FORCE, TAKEOVER

CLUSTER FAILOVER is sent to a replica (cluster.c:6491-6493). With no argument the master must still be reachable, or the command refuses with "Master is down or failed, please use CLUSTER FAILOVER FORCE" (cluster.c:6497-6503), and it starts by sending MFSTART to the master (cluster.c:6524).

FORCE skips the coordination but still needs the votes: the election delay is zeroed (cluster.c:4359-4363), the data-age check is bypassed (cluster.c:4338), cluster-replica-no-failover is bypassed (cluster.c:4305), and the request carries FORCEACK (cluster.c:3999) so masters grant it even though the master is up (cluster.c:4067).

TAKEOVER implies FORCE (cluster.c:6483) and then skips the election completely: it calls clusterBumpConfigEpochWithoutConsensus — which just increments currentEpoch and assigns it to itself (cluster.c:1819-1836) — and claims the slots (cluster.c:6514-6515). No master votes. It is the only path here that can promote a replica in a minority partition, and it can produce two masters claiming the same slots.

Flags and rank, as declared

cluster.h:49-69 · cluster.c:4142
int flags;              /* cluster.h:123 — int, and it needs the width */

#define CLUSTER_NODE_MASTER               1   /* cluster.h:49 */
#define CLUSTER_NODE_SLAVE                2
#define CLUSTER_NODE_PFAIL                4
#define CLUSTER_NODE_FAIL                 8
#define CLUSTER_NODE_MYSELF              16
#define CLUSTER_NODE_HANDSHAKE           32
#define CLUSTER_NODE_NOADDR              64
#define CLUSTER_NODE_MEET               128
#define CLUSTER_NODE_MIGRATE_TO         256
#define CLUSTER_NODE_NOFAILOVER         512   /* cluster.h:58 — not 16; 16 is MYSELF */
#define CLUSTER_NODE_EXTENSIONS_SUPPORTED 1024

/* cluster.c:4142 — verbatim */
int clusterGetSlaveRank(void) {
    long long myoffset;
    int j, rank = 0;
    clusterNode *master;

    serverAssert(nodeIsSlave(myself));
    master = myself->slaveof;
    if (master == NULL) return 0;

    myoffset = replicationGetSlaveOffset();
    for (j = 0; j < master->numslaves; j++)
        if (master->slaves[j] != myself &&
            !nodeCantFailover(master->slaves[j]) &&
            master->slaves[j]->repl_offset > myoffset) rank++;
    return rank;
}

Node inspector

local view

CLUSTER FAILOVER

cluster.c:6475

TAKEOVER promotes with no votes at all (cluster.c:6514).

Event log