Scaling WebSocket Connections: What Breaks Past a Few Thousand Concurrent Users

WebSockets don't scale like REST APIs, they're stateful, long-lived connections, and that single difference breaks chat and notifications past a few thousand concurrent users. A breakdown of what actually needs to change: shared pub/sub state, connection-aware load balancing.
Scaling WebSocket Connections: What Breaks Past a Few Thousand Concurrent Users
A chat feature built on a single WebSocket server works perfectly in every demo, every staging test, and the first few weeks in production. Then usage grows past a few thousand concurrent connections, and things that were never a problem, one server holding all the connection state, a broadcast loop that was fine at hundreds of users, a load balancer that was never told these connections are different from normal HTTP traffic, start quietly breaking message delivery, presence status, and reconnection behavior, usually without a clear error pointing at the actual cause.
WebSockets are not a scaling problem in the way REST APIs are. A REST API is stateless, any server can answer any request, so you scale by adding more servers behind a load balancer and calling it done. A WebSocket connection is a long-lived, stateful pipe between one specific client and one specific server, and that single difference is what makes everything about scaling them past a single machine genuinely different from scaling stateless HTTP traffic.
Tip
TL;DR, A single well-tuned server can hold far more idle WebSocket connections than most teams assume, hundreds of thousands in the right conditions. The real ceiling isn't connection count, it's message throughput and broadcast fan-out on top of those connections. Past a few thousand concurrent users, the fixes that matter are: connection-aware load balancing, moving broadcast/presence state out of individual server memory and into a shared pub/sub layer, and heartbeat-based reconnection handling, all of it planned before you need it, because retrofitting a stateful system under live traffic is far harder than retrofitting a stateless one.
Why WebSockets scale differently than everything else in your stack
| Stateless HTTP/REST | WebSocket connections | |
|---|---|---|
| Server affinity | None, any server can handle any request | The client stays connected to one specific server for the connection's lifetime |
| Scaling method | Add servers behind a load balancer, done | Requires connection-aware routing and shared state across servers |
| What "load" means | Requests per second | Both connection count and message throughput per connection |
| Failure mode when overloaded | Slow or dropped individual requests | An entire server's worth of connections can drop at once |
| Load balancer behavior needed | Simple round-robin works fine | Needs sticky sessions or connection-aware routing, round-robin creates uneven load because connections are long-lived, not one-and-done |
That last row is where a lot of early scaling attempts go wrong: a load balancer configured the same way it handles REST traffic will happily route WebSocket upgrade requests round-robin, which looks fine until connections pile up unevenly across servers because some clients simply stay connected far longer than others.
The real ceiling isn't connection count
The most common misconception is that WebSocket scaling is primarily about how many connections a server can hold open. In practice, idle connections are cheap, a single well-tuned event-driven server can hold hundreds of thousands of idle connections in memory, limited mainly by file descriptor limits and a few kilobytes of memory per connection, not by anything exotic. The ceiling that actually matters is what happens when those connections start being active at the same time: broadcasting a message to ten thousand clients means running that many send operations, and if broadcasts happen frequently, a live dashboard, a busy group chat, a multiplayer session, that fan-out work is what saturates a server's CPU long before connection count does.
This is the distinction that trips up teams scaling a chat or notification feature for the first time: the app "worked" up to a few thousand users because connections were mostly idle, and it starts degrading not because a connection limit was hit, but because message volume finally caught up to what a single event loop can actually push out per second.
The first real fix: get connection state out of individual servers
On a single server, "who's online" and "which room is this connection part of" can live in a plain in-memory map, fast, simple, and completely fine until you need a second server. The moment you add a second server to handle more connections, that in-memory state becomes a trap: a message meant for a user connected to server B has no way of reaching them if it was published from server A, because server A has no idea server B even exists, let alone who's connected to it.
The standard fix is a shared pub/sub layer, Redis pub/sub is the common starting point, sitting between your WebSocket servers. Instead of a server trying to deliver a message directly to a connection it doesn't own, it publishes the message to a shared channel; every server subscribed to that channel receives it and forwards it only to the connections it actually holds. This one change is what turns a fleet of independent WebSocket servers into something that behaves like a single logical system to your users, regardless of which physical server they happen to be connected to.
A companion piece to this is a connection registry, a shared, fast-lookup record of which server currently holds which user's connection. Without it, "is this user online right now, and if so, which server do I ask" has no answer once you're past a single machine, which breaks accurate presence status and targeted, one-to-one messaging at the exact moment your app has grown enough to need both reliably.
Load balancing that respects long-lived connections
Round-robin load balancing assumes every unit of work is short and roughly equal, true for HTTP requests, false for WebSocket connections that can stay open for hours. A better default for WebSocket traffic is least-connections balancing, which routes new connections to whichever server currently holds the fewest, naturally evening out load over time instead of assuming every connection costs the same regardless of how long it's already been open elsewhere.
Sticky sessions, pinning a client to the same server across reconnects, typically via IP hash or a cookie, are a reasonable starting point at moderate scale. The tradeoff shows up later: sticky sessions mean a server going down takes all of its pinned clients with it, and they have to reconnect and get re-pinned elsewhere. Combining sticky routing with externally-stored session state (in Redis, not just server memory) is what lets a reconnecting client land on any available server without losing context, instead of being fragile to exactly which machine it was previously attached to.
Heartbeats and reconnection: the part that's invisible until it isn't
A WebSocket connection can silently die, a phone switches from Wi-Fi to cellular, a laptop goes to sleep, a network middlebox drops an idle connection without telling either side. Without a way to detect this, a server can hold thousands of connections it believes are healthy that are actually already gone, which wastes resources and, worse, makes presence status lie to your users about who's actually online.
The standard fix is a heartbeat: the server pings each connection at a regular interval, commonly every 30 seconds, and a connection that doesn't respond within a reasonable window is considered dead and cleaned up. On the client side, reconnection logic needs the same discipline as any retry behavior on an unreliable network, exponential backoff with a bit of randomness (jitter) so that a server recovering from an outage isn't immediately hit with every disconnected client retrying at the exact same instant, which just recreates the outage it's trying to recover from.
When you need message durability, not just delivery
Pub/sub alone delivers a message to whoever happens to be connected right now, if a client is offline or briefly disconnected, that message is simply gone. That's fine for ephemeral events like a "user is typing" indicator, and not fine at all for a chat message or a notification a user genuinely needs to see later. The moment your real-time feature needs message history, guaranteed delivery, or the ability to replay missed events to a client that just reconnected, plain pub/sub needs to be paired with (or replaced by) something with actual persistence and replay capability, a durable message log rather than a fire-and-forget broadcast.
This distinction is easy to miss early on because a live demo never has anyone disconnect at the wrong moment. In production, at real scale, someone is disconnecting at the wrong moment constantly, and whether your architecture treats that as normal or as data loss is a decision made far earlier than most teams realize.
A simple decision model
Stay on a single server, with in-memory state, if: you're comfortably under a few thousand concurrent connections and message volume is light, this covers more early-stage products than teams expect, and premature distributed-systems complexity here usually isn't worth it yet.
Add a shared pub/sub layer and connection registry once: you're adding, or planning to add, a second server for capacity or redundancy, this is the point where in-memory state stops working correctly, not a nice-to-have for later.
Move load balancing to least-connections with externally-stored session state once: connections are unevenly distributed across servers, or a single server restart is visibly disrupting a noticeable share of active users.
Add message durability and replay once: users are reporting missed messages or notifications after a disconnect, or the feature has grown from "live indicator" into something people rely on not to lose data.
Wrapping up
The mistake most teams make with WebSockets isn't a technology choice, it's assuming a pattern that works cleanly for stateless HTTP traffic will translate directly to a stateful, long-lived connection model. It doesn't, and the gap between the two is exactly where chat features and live notifications quietly start failing somewhere between a few hundred and a few thousand concurrent users. Planning for a shared pub/sub layer, connection-aware load balancing, and proper heartbeat/reconnection handling from the start costs very little when the system is small, and saves months of retrofitting once it isn't.
The same underlying discipline, protecting shared backend capacity from uneven client behavior, is what I covered in rate limiting and caching at the API gateway layer, and the connection-state and session concerns here echo the multi-device and session-handling issues in handling users smoothly from the Flutter client to the backend.
Related Posts
Related Articles

From 0 to 100K Users: Handling Users Smoothly from the Mobile App Client to the Backend
What actually breaks as a Flutter app scales isn't the UI — it's how sessions, token refresh, and concurrent writes are handled between the client and backend. A technical breakdown of the user-lifecycle failures that show up between 1K and 100K users.

API Security Checklist for SaaS Builders
Most SaaS breaches aren't zero-days, they're missing auth checks, leaked keys, and endpoints nobody remembered to lock down. Here's the checklist that actually prevents them.

Next.js SEO: What Actually Moves the Needle
Rebuilt your site in React and watched traffic drop? Here's what Google actually sees, and the three fixes that matter more than any plugin.
