Modelled on src/pubsub.c, Redis 7.2.14. Pattern matching is a line-for-line port of stringmatchlen() from src/util.c; hash slots use the XMODEM CRC-16 of src/crc16.c.
Two facts about Pub/Sub cost people production incidents. The first: a client subscribed to
a channel and to a pattern that matches it receives the message twice —
once as message, once as pmessage — and PUBLISH
counts both (pubsub.c:485-490 then pubsub.c:510-516, with no de-duplication between them).
The second: delivery is attempted exactly once against whoever is connected at that instant.
Nothing is stored, nothing is replayed, nothing is acknowledged. Take a subscriber offline
below and publish; the message is gone.
Each push travels from the publisher to one subscriber and is labelled with its type:
message from the channel dict, pmessage from a matching
pattern, smessage from a shard channel. A client holding both a channel and
a pattern that matches it is flown to twice. A push aimed at a disconnected client drifts
past and fades — that one is gone. The flight order and count are read straight off the
delivery list that produced the integer, so the motion cannot disagree with it. With
prefers-reduced-motion set, every arrival lands at once instead.
Glob, not regular expressions. There is no alternation, no quantifier, no anchor and no
\d-style class: a backslash simply means “the next byte, literally”.
*?[abc][^abc][a-z]\xx literally, inside a set or out (util.c:110, util.c:159)
Ported as-is, minus the case-insensitive path: PUBLISH calls
stringmatchlen(…, 0) at pubsub.c:508, so matching is always
case-sensitive. Assumes ASCII channel names — the C compares bytes, JavaScript
compares UTF-16 units.
Each is a dict from name to a list of clients. Every
PUBLISH walks pubsub_patterns in full and runs the matcher per
entry (pubsub.c:499-517), so patterns cost O(N) in the number of distinct patterns, not
O(1) like a channel lookup.
subscribe 3 elements 1) "subscribe" 2) channel 3) (integer) channels + patterns unsubscribe 3 elements 1) "unsubscribe" 2) channel 3) (integer) channels + patterns psubscribe 3 elements 1) "psubscribe" 2) pattern 3) (integer) channels + patterns ssubscribe 3 elements 1) "ssubscribe" 2) channel 3) (integer) shard channels only message 3 elements 1) "message" 2) channel 3) payload smessage 3 elements 1) "smessage" 2) channel 3) payload pmessage 4 elements 1) "pmessage" 2) pattern 3) channel 4) payload
The integer is not a total. For SUBSCRIBE, PSUBSCRIBE and their
un-subscribes it is clientSubscriptionsCount() — channels plus patterns,
shard channels excluded (pubsub.c:221-223). For SSUBSCRIBE and
SUNSUBSCRIBE it is clientShardSubscriptionsCount() — shard
channels only (pubsub.c:226-228). The two namespaces never mix; only the internal
CLIENT_PUBSUB flag is cleared on the sum of both (pubsub.c:240-242).
A redundant SUBSCRIBE, or an UNSUBSCRIBE for something the client
never held, still gets the full three-element push with the count unchanged. The reply is
emitted outside the dictAdd / dictDelete branch (pubsub.c:267,
pubsub.c:304-306), so there is no “already subscribed” error and no silent
no-op — a client library can count replies against commands sent.
On a RESP2 connection, once CLIENT_PUBSUB is set the server rejects everything
except SUBSCRIBE, SSUBSCRIBE, PSUBSCRIBE, their three
un-subscribes, PING, QUIT and RESET. That is why the
old advice is “one connection to subscribe, another to publish”. RESP3 lifts the
restriction entirely, because pushes arrive on their own frame type
(>3 rather than *3, pubsub.c:110-113) and can no longer be
mistaken for a command reply. Flip the protocol above and try to publish from a subscribed
client.
Shard channels are a separate namespace, added in 7.0. SSUBSCRIBE never sees a
pattern: shard publishing returns before the pattern loop is reached (pubsub.c:493-496), so
PSUBSCRIBE n* will not pick up SPUBLISH news.
In cluster mode the shard channel name is hashed exactly like a key —
crc16(name) & 0x3FFF, honouring a {hashtag} (cluster.c:1380-1399)
— and the command is routed on that slot (key specs at commands.def:4656 and 4681).
Address the wrong node and you get MOVED <slot> <host>:<port>
(cluster.c:7587-7590); name two shard channels in one SSUBSCRIBE that hash
differently and you get CROSSSLOT (cluster.c:7569-7570). Plain
SUBSCRIBE and PUBLISH skip this layer and work on any node
(cluster.c:7397).
The propagation is the whole point. PUBLISH is broadcast to every node in the
cluster whether or not that node has a subscriber — clusterBroadcastMessage
at cluster.c:3960-3964. SPUBLISH goes only to the nodes of its own shard
(cluster.c:3969-3977). Either way the integer you get back is counted before
propagation (pubsub.c:601 then pubsub.c:603): it is the number of receivers on the node that
handled the command, never a cluster-wide total, and it excludes replica subscribers reached
by the replication link (pubsub.c:615-616).
pubsubPublishMessageInternal walks the live subscriber lists and appends to each
client's output buffer. There is no log, no offset, no acknowledgement and no retry anywhere
in that function or its callers. When the last subscriber of a channel leaves, the channel's
entry is deleted outright (pubsub.c:291-295), so there is not even a place a replay could
come from. A subscriber that is disconnected during a PUBLISH has missed the
message permanently; freeClient has already torn down all three of its
subscription sets (networking.c:1634-1636), so on reconnect it must subscribe again and
starts from silence. If you need history, that is what Streams are for.
A subscriber that stays connected but stops reading is a different failure: it is cut off by
client-output-buffer-limit pubsub, whose default is
32mb 8mb 60 (config.c:172, enforced at networking.c:3896-3917). That is the
knob, not client-query-buffer-limit — that one bounds the
inbound command buffer (config.c:3247) and does nothing for a slow reader.