Adding Real-Time Features to a Microservice Architecture Without Breaking It
Microservices assume a request arrives and goes away. A WebSocket arrives and stays for four hours, and every assumption underneath the architecture is quietly wrong for it. How to add a real-time layer without making all five services stateful, and the deploy problem that follows.
Microservices assume a request arrives, gets handled, and goes away. A WebSocket arrives and stays for four hours. Every assumption the architecture is built on — stateless instances, rolling deploys, route anywhere, scale on request rate — is quietly wrong for that connection, and the first time you notice is usually a deploy that disconnects everybody mid-session.
This is how the real-time layer got added to the five-service library system without any service learning what a socket is, and the three decisions that made it survivable.
The mismatch
Why this is not just "add a WebSocket endpoint"
| Microservice assumption | What a WebSocket does | What breaks |
|---|---|---|
| Requests are short and stateless | Holds a connection for hours | Instances can never be drained quickly |
| Any instance can serve any request | State lives on one instance | A message for a user must reach their node |
| Rolling deploy is invisible | Every replaced pod drops its clients | A deploy is a mass disconnect |
| Scale on requests per second | Cost is concurrent connections | Autoscaling watches the wrong number |
| Auth on every request | Auth once, at the handshake | A token expires mid-connection and nothing notices |
The instinct is to add real-time to whichever service owns the data. Lending owns loans, so Lending should push loan updates. That is wrong for one decisive reason: it makes every service stateful, and then every service inherits every row of that table.
The shape
One real-time edge, fed by events
loan.created and is finished. Whether anybody is watching, on which node, is not its problem.The rule that keeps this clean: a domain service publishes a domain event, never a UI message. Lending emits loan.created with ids. The edge decides that this means a toast saying "Renewed until 4 March" for the two people watching that member. Put the wording in Lending and you have coupled a backend service to a frontend copy change.
// Domain event in. Which sockets care is decided HERE, not by Lending.
bus.on('loan.created', async (event) => {
const topics = [
`member:${event.member_id}`, // the borrower's own screens
`branch:${event.branch_id}`, // the desk that checked it out
];
// Publish to Redis, not to sockets. Another node may hold the socket.
for (const topic of topics) {
await redis.publish(`rt:${topic}`, JSON.stringify({
type: 'loan.created',
loan_id: event.loan_id,
due_at: event.due_at,
}));
}
});
// Every node subscribes only to the topics its own clients asked for.
// Subscribing to one firehose is how adding a node makes things slower.
async function subscribe(socket, topic) {
if (!localTopics.has(topic)) {
localTopics.set(topic, new Set());
await redisSub.subscribe(`rt:${topic}`);
}
localTopics.get(topic).add(socket);
}
Identity
Auth across a boundary that only happens once
HTTP re-authenticates every request. A socket authenticates once and then lives for hours, which creates two problems nobody has in request/response.
-
Authenticate at the handshake, before allocating anything
The gateway already verifies the token for HTTP. It does the same for the upgrade request and passes the claims to the edge as headers. The edge never talks to Members — it trusts the gateway, because the gateway is the only route in.
-
Give the connection its own expiry
A token valid for fifteen minutes does not entitle you to a four-hour socket. Store the token's expiry on the connection and close it with a specific code when it passes, so the client knows to refresh and reconnect rather than treating it as a network blip.
-
Authorise the subscription, not just the connection
Being connected is not permission to subscribe to
member:4471. The edge checks the claim against the topic on every subscribe. This is the one authorisation decision the edge is allowed to make, and only because it is structural rather than domain logic. -
Handle revocation as a broadcast
When Members disables a card it emits
member.suspended. The edge closes that member's sockets. Without this, a revoked account keeps receiving updates until it happens to disconnect.
Deploys
The part that actually hurts
A rolling deploy replaces instances. For stateless services nobody notices. For the edge, every replaced instance drops every connection it holds — and they all reconnect at once, which is a thundering herd against the node that is still up.
Standard rolling restart. 40 seconds of the whole estate reconnecting, twice per deploy.
Worse: clients reconnected instantly and in lockstep, so the surviving node took the entire load in one spike and shed some of it, which produced a second wave.
Drain: stop accepting new connections, send a reconnect frame with a random delay between 0 and 30 s, then exit when empty or the deadline passes.
Clients move across gradually. The spike disappears because the herd was told to scatter.
process.on('SIGTERM', async () => {
healthy = false; // load balancer stops sending new ones
for (const socket of connections) {
// A jittered instruction, not a disconnect. The client controls the timing,
// which is what stops every client returning in the same second.
socket.send(JSON.stringify({
type: 'reconnect',
after_ms: Math.floor(Math.random() * 30_000),
}));
}
const deadline = Date.now() + 45_000; // must exceed the jitter window
while (connections.size > 0 && Date.now() < deadline) {
await sleep(500);
}
process.exit(0);
});
Two details are load-bearing. The termination grace period in your orchestrator must be longer than the jitter window, or the platform kills the process mid-drain and you have added latency without removing the problem. And healthy = false has to come first — a node that is draining but still passing health checks keeps receiving the clients it is trying to shed.
Scaling on the right number
Autoscaling on CPU or request rate does nothing useful here: an idle socket costs almost no CPU and generates no requests, right up until the node runs out of file descriptors or memory. The edge scales on concurrent connections per node, with the ceiling set from a real measurement rather than a guess.
Honesty
What this does not solve
- Message ordering across nodes is not guaranteed. Two events published microseconds apart can arrive in either order. Anything order-sensitive carries a sequence number and the client sorts, or you accept it.
- Redis pub/sub is at-most-once. A dropped subscriber misses messages with no replay. For live UI updates that is fine — the next event corrects the view. For anything that must not be missed, the event bus is the durable path and the socket is only a hint to refetch.
- The edge is a single point of failure for real-time, by design. If it is down, the site still works and simply stops updating live. That was an explicit trade: degrade the feature, never the transaction.
- Reconnect storms after a network partition are worse than after a deploy, because nobody told the clients to scatter. The client backoff has to have jitter for the same reason the drain does.
- Local development now needs Redis to see any live behaviour, which is one more thing in the compose file and one more thing new developers hit on day one.
The checklist
- Exactly one service holds sockets. Statefulness must not spread.
- Domain services publish domain events, never UI messages.
- Authenticate at the handshake, through the gateway, before allocating state.
- Give the connection its own expiry and close with a code the client understands.
- Authorise every subscription, not just the connection.
- Revocation closes open sockets — it will not happen by itself.
- Subscribe per topic, never to one global channel.
- Drain with jittered reconnect on SIGTERM, and fail health checks first.
- Grace period longer than the jitter window.
- Scale on concurrent connections, not CPU or request rate.
Where to start on Monday
Draw the line before you write anything: which service will hold a socket, and what event each other service will publish. If the answer is "several services will push updates", stop — that is the version that makes your whole estate stateful, and it is much harder to undo than to avoid.
- realtime edge
- domain events
- Redis pub/sub
- per-topic channels
- handshake auth
- connection drain
- jittered reconnect
- connection-based scaling
The service boundaries this sits on are in the gateway case study. The per-topic subscription rule and the numbers behind it come from the Redis pub/sub benchmarks, and the deploy story continues in Dockerizing real-time services.
Comments (0)
No comments yet
Be the first to share a thought on this article.
Join the conversation