Skip to content
A AhsanLab.Tech
Deployment & CI/CD 15 min read · September 5, 2026

Dockerizing Real-Time Services: CI/CD for WebSocket and SFU Apps

A rolling deploy is invisible for a stateless HTTP service. Do the same to a WebSocket service and you have disconnected everyone; do it to an SFU and you have ended everybody call. The orchestrator did exactly what it was told — what it was told assumed requests are short.

A Ahsan Habib Save

A rolling deploy is invisible for a stateless HTTP service: drain the connections, which take milliseconds, replace the container, move on. Do the same thing to a WebSocket service and you have just disconnected everyone it was holding — and to an SFU, you have ended everybody's call. The orchestrator did exactly what it was told. What it was told assumed requests are short.

This is the capstone of the deployment series and the hardest case in it: containerising and continuously deploying services whose whole purpose is to hold a connection open.

~800calls dropped, deploy onethe orchestrator was right
0after connection drainingclients move gradually
2health endpointsliveness ≠ readiness

The mismatch

Why the defaults are wrong here

Default behaviourFine for HTTP becauseWrong here because
30 s termination graceRequests finish in millisecondsConnections last hours
SIGTERM then SIGKILLNothing is in flightMedia streams and game state are
Replace one instance at a timeLoad rebalances instantlyEvery client on that instance reconnects at once
Health check = process is upUp means readyUp can mean draining, or out of file descriptors
Scale on CPUCPU tracks requestsIdle sockets cost memory and descriptors, not CPU
Any instance serves any clientRequests are independentMedia state lives on one instance

The image

Container basics that matter more for stateful services

Dockerfile — the signal handling is the important part
FROM node:20-slim AS build
WORKDIR /app
COPY package*.json ./
RUN npm ci
COPY . .
RUN npm run build

FROM node:20-slim
WORKDIR /app
COPY --from=build /app/dist ./dist
COPY --from=build /app/node_modules ./node_modules

# EXEC form, so node is PID 1 and receives SIGTERM directly.
# The shell form (CMD node dist/server.js) runs it under /bin/sh, which does
# NOT forward signals — your drain handler never fires and every deploy
# becomes a hard kill. This one line is the whole difference.
CMD ["node", "dist/server.js"]

# Two endpoints, not one. See below.
HEALTHCHECK --interval=10s --timeout=3s CMD node dist/healthcheck.js
The bug that produces "my drain code never runs" CMD npm start or any shell-form CMD puts a shell at PID 1. Shells do not forward SIGTERM to children by default, so your carefully written drain handler is never called and the platform kills the process after the grace period. Use the exec form, or an init like tini. Verify it: docker stop the container and check your drain log line actually appears.

Liveness and readiness are different questions

/healthz — liveness

"Is this process wedged and in need of killing?"

Must stay green while draining. Fail it and the platform kills a container that is politely handing over its clients.

/readyz — readiness

"Should new connections be sent here?"

Goes red the instant draining starts, and also when the instance is near its connection ceiling.

Collapsing these into one endpoint is the most common mistake in this whole area, and it produces a specific, confusing symptom: containers restarting during every deploy, disconnecting the clients they were mid-way through handing over.

The deploy

Drain, do not replace

FOUR PHASES · THE OLD CONTAINER LIVES THROUGH ALL OF THEM 1 · NEW COMES UP /readyz green old still serving everyone 2 · OLD GOES UNREADY /readyz red, /healthz GREEN no new connections routed 3 · TELL CLIENTS TO GO reconnect frame, jittered 0–60 s spread 4 · EXIT WHEN EMPTY or at the deadline whichever comes first terminationGracePeriodSeconds MUST EXCEED the jitter window 60 s of jitter under a 30 s grace period = the platform kills the container mid-drain, and you added latency for nothing maxSurge: 1, maxUnavailable: 0 · capacity for both generations at once a deploy now takes minutes instead of seconds. that is the correct trade.
FIGURE 1 — DRAIN, DO NOT REPLACE · The whole design is that the old container stays alive and healthy while it empties.
server.js — the drain handler
let draining = false;

app.get('/healthz', (_, res) => res.sendStatus(200));            // alive, even draining
app.get('/readyz',  (_, res) =>
  res.sendStatus(draining || sockets.size >= MAX_CONNS ? 503 : 200));

process.on('SIGTERM', async () => {
  draining = true;                       // /readyz goes red; no new routing here

  // Give the load balancer time to notice before we push anyone off.
  await sleep(5_000);

  for (const socket of sockets) {
    socket.send(JSON.stringify({
      type: 'reconnect',
      after_ms: Math.floor(Math.random() * 60_000),   // scatter the herd
    }));
  }

  const deadline = Date.now() + 90_000;  // < terminationGracePeriodSeconds
  while (sockets.size > 0 && Date.now() < deadline) await sleep(1_000);

  log.info('drained', { remaining: sockets.size });
  server.close(() => process.exit(0));
});

The five-second pause before notifying anyone is easy to leave out and it matters: without it, clients start reconnecting while the load balancer still considers this instance ready, and a portion of them land straight back on the container that is trying to empty.

Media

SFUs break two more assumptions

Everything above applies to an SFU, plus two problems that are specific to media.

  1. UDP media ports are not HTTP and your ingress does not understand them

    An SFU needs a wide UDP range reachable directly, and clients negotiate a specific address during ICE. Standard HTTP ingress does nothing for this: the media path is hostPort or host networking, with the node's public address advertised explicitly in the SFU config. Getting this wrong produces the classic symptom — signalling connects, participants appear, nobody can hear anyone.

  2. A room cannot move

    Everyone in a call is pinned to one SFU instance. There is no draining a room to another node without renegotiating every participant. So the deploy unit is effectively the room: stop assigning new rooms to an instance, wait for its existing ones to end, then replace it.

What does not work

Waiting for every room to empty naturally.

On a busy service that can be hours, and one long meeting blocks the deploy indefinitely. We had a container refuse to exit for most of a working day.

What works

A drain window with an announced migration.

Stop assigning new rooms; wait up to 30 minutes; then notify remaining rooms in-app that they will be briefly reconnected, and move them. Deploy during a low-traffic window and most rooms end on their own.

For a stateful media service, "zero downtime" is not achievable in the way it is for HTTP. What is achievable is that nobody is disconnected without warning, and that is a completely different promise — one you can actually keep.

Pipeline

What CI has to do differently

  • Tag by commit SHA, never latest. With connection draining a deploy spans minutes and both generations run simultaneously; latest makes it impossible to say which container is which, exactly when you need to.
  • Set maxSurge: 1, maxUnavailable: 0 so the new instance is serving before the old one starts emptying. This needs headroom for both — draining costs capacity as well as time.
  • Grace period longer than the drain deadline, always. If your drain waits 90 seconds, the grace period is 120. Getting this backwards silently reverts you to hard kills.
  • Gate the rollout on connection count, not on HTTP checks. A new instance that is up but has not yet accepted a socket has proved nothing.
  • Deploy on a schedule for media services. Continuous deployment and hour-long calls are in genuine tension. A nightly window is not a failure of engineering, it is reading the constraint correctly.

Honesty about what is left

  • Deploys are slow now. Minutes rather than seconds. That is the cost of not disconnecting people, and it makes "deploy twenty times a day" impractical for these services specifically.
  • Reconnecting clients lose ephemeral state. Anything held only in the socket's memory is gone. It has to be recoverable from the server after a reconnect, which is a design requirement that reaches into the application.
  • Autoscaling still needs custom metrics. Connections per instance, exported and scaled on. CPU-based autoscaling will not do anything useful.
  • This does not survive a node failure. Draining is graceful shutdown; a machine that vanishes takes its connections with it, and the client's reconnect logic is the only mitigation.
  • Test the drain path in CI. Start the container, open sockets, docker stop, assert the reconnect frames were sent and the process exited zero. It is the code most likely to rot silently, because nothing else exercises it.

The checklist

  • Exec-form CMD so the process is PID 1 and receives SIGTERM.
  • Separate /healthz and /readyz — liveness stays green while draining.
  • Readiness also red at the connection ceiling, not only when draining.
  • Pause after going unready before notifying clients.
  • Jittered reconnect instructions, never a mass disconnect.
  • Grace period > drain deadline > jitter window.
  • maxSurge: 1, maxUnavailable: 0, with capacity for both generations.
  • Image tags by commit SHA.
  • SFU media ports on host networking, with the public address advertised explicitly.
  • A CI test for the drain path, because nothing else will catch it breaking.

Where to start on Monday

Run docker stop on your real-time container and watch the logs. If your drain handler does not log anything, you have the shell-form CMD problem, and every deploy you have ever done was a hard kill. That is a five-minute check and a one-line fix.

  • exec-form CMD
  • liveness vs readiness
  • connection draining
  • jittered reconnect
  • grace period
  • maxSurge
  • SHA tags
  • host networking

This closes the deployment series: the environments, securing the box, the pipeline, the atomic swap, and now the stateful case. The architecture this deploys is in adding real-time to microservices, and the drain pattern first appears in the WebSocket pillar.

#WebSocket #Real-Time #Scalability #LiveKit #CI/CD #Docker

Comments (0)

No comments yet

Be the first to share a thought on this article.

Join the conversation

Comments are moderated before they appear.

Keep reading

Related articles