From 0 to 100K Users: Handling Users Smoothly from the Mobile App Client to the Backend

From 0 to 100K Users: Handling Users Smoothly from the Mobile App Client to the Backend

What actually breaks as a Flutter app scales isn't the UI — it's how sessions, token refresh, and concurrent writes are handled between the client and backend. A technical breakdown of the user-lifecycle failures that show up between 1K and 100K users.

From 0 to 100K Users: Handling Users Smoothly from the Flutter Client to the Backend

Most "scaling" content for mobile apps is marketing advice wearing a technical costume — ASO tips, retention hooks, growth loops. Almost none of it answers the question an engineer actually has: what happens to a user as they move through your system — from the moment the app opens on their device to the moment your backend reads or writes something on their behalf — and where does that journey start to wobble as the number of concurrent users climbs?

I've shipped and operated Flutter apps in production long enough to watch that curve happen more than once — most notably with Rewaytk, which crossed 100K+ active users on a stack that was never rewritten from scratch to get there. What changed wasn't the app's features. It was how carefully a single user's session, identity, and requests were handled end to end, at every layer between the client and the database. This is the technical account of that, not the pitch-deck version.

Tip

TL;DR — A user is not just "logged in" or "logged out." Between those two states sits token refresh, session expiry, retry behavior, request ordering, and backend concurrency — and each one behaves fine at low volume and degrades silently as volume grows. Design the client-to-backend user lifecycle to be idempotent, stateless where possible, and resilient to duplicate or out-of-order requests before you need it to be, because retrofitting it under real traffic is where teams lose months.

The stages, and what actually changes at each one

Stage User count What breaks What doesn't matter yet
Prototype 0–1K Nothing — every session/auth shortcut "works" Rate limiting, request queuing
Early traction 1K–10K Token refresh races cause random logouts; duplicate requests from retry logic start corrupting data Multi-region session storage
Growth 10K–50K Concurrent writes from the same user (multi-device) collide; backend session lookups slow down under load Horizontal auto-scaling
Scale 50K–100K+ Stale sessions pile up and bloat the session store; abusive or misbehaving clients degrade service for everyone else Rewriting the auth system from scratch (usually a mistake)

The pattern across every stage: nothing catastrophic happens overnight. A handful of users get logged out randomly, someone reports "my data duplicated," a support ticket mentions the app being slow on one device but not another — and each of those is a symptom of the same underlying gap between how the client assumes a user's session behaves and how the backend actually handles it under concurrency.

Identity and session handling: the decision that compounds

At under 1,000 users, storing an access token in memory and refreshing it "when a request fails" feels adequate. At 50,000 users, that pattern reliably produces two symptoms: users randomly bounced to the login screen because two requests raced to refresh the same expired token at once, and duplicate writes because a client retried a request it assumed had failed, when the backend had actually already processed it.

The fix isn't a bigger auth library. It's making two guarantees explicit and enforcing them consistently everywhere the client talks to the backend:

  • Token refresh must be serialized on the client, so that if five requests discover an expired token simultaneously, only one refresh call happens and the other four wait on its result instead of each firing their own refresh and racing to write the new token back.
  • Every state-changing request must be safely repeatable. If a client can't tell whether a request succeeded before retrying it, the backend has to be able to tell the difference between "this is a new action" and "this is the same action arriving twice."

This second guarantee is the one that quietly saves a growing app: the backend accepts a client-generated identifier for each meaningful action (a message send, a payment, a profile update) and treats a repeated identifier as "already handled" rather than "do it again." At low traffic this never gets tested because retries are rare. At scale, on flaky mobile networks, retries are constant.

The backend usually breaks before the app does

Flutter apps get blamed for lag that's actually a backend problem with how it handles concurrent users, not raw request volume. At low traffic, looking up a user's session and permissions on every request is cheap enough that nobody profiles it. At 100K users with any real daily-active fraction, that same per-request lookup — hitting the database fresh every time instead of a fast session cache — becomes the dominant cost on every endpoint in the app, and it shows up as generalized slowness that's hard to pin on any one feature.

The three user-handling issues that show up in this exact order as concurrent users grow:

  1. Session and permission lookups that hit the primary database on every request instead of a cache layer, which is fine at low concurrency and becomes the bottleneck on every endpoint once enough users are active at once.
  2. No idempotency guarantee on write endpoints, so a client retry — normal and expected on mobile networks — creates a duplicate order, duplicate message, or duplicate account action instead of being recognized as a repeat.
  3. No per-user rate limiting, so one misbehaving client — a bug in a specific app version, or a user with a script — can consume enough backend capacity to visibly slow the experience for every other concurrent user, even though nothing is technically "down."

This is the same category of fix as the MongoDB indexing and connection-pooling work I've covered before — the failure isn't in a single feature's logic, it's in an implicit assumption about how many users are hitting the same code path at the same time, an assumption that was true on day one and stopped being true somewhere on the way to 100K.

Multi-device and out-of-order requests

A user with the app open on a phone and a tablet at the same time isn't an edge case once you're past a few thousand users — it's routine. The failure mode is two devices sending conflicting updates for the same user record within the same second, and the backend accepting whichever one happens to arrive last, silently overwriting the other. At low scale this basically never happens because so few users are on two devices simultaneously. At scale it becomes a recurring support ticket that looks like "the app lost my changes," when the actual cause is the backend never having an explicit rule for resolving concurrent writes to the same user's data — no version check, no last-write timestamp comparison, nothing.

The fix is deciding, deliberately, what "the newest update wins" means for each type of user data, and enforcing that check on the write path instead of just accepting whatever arrives.

Handling the misbehaving minority

At scale, a small fraction of clients will always behave unexpectedly — an old app version with a bug that hammers an endpoint in a loop, a user on a bad connection whose client retries far more aggressively than intended, or occasionally deliberate abuse. At low user counts this fraction is a handful of sessions and the backend absorbs it without anyone noticing. At 100K users, that same small percentage is thousands of sessions generating enough load to degrade response times for everyone else sharing the same backend capacity.

The habit that actually works: rate limit and throttle per user identity, not just per IP address, since mobile users routinely share IPs (carrier NAT, public Wi-Fi) and a per-IP limit either blocks innocent users sharing a network or fails to catch a single abusive user rotating networks.

A simple decision model

Serialize token refresh on the client now if:
  - You're still under ~10K users
  - Retrofitting later means auditing every network call site
    for how it handles a session expiring mid-request

Add idempotency keys to write endpoints if:
  - Any endpoint changes data as a result of a client action
    (not just reads)
  - Support tickets mention duplicated actions on flaky connections

Add explicit conflict resolution for multi-device writes if:
  - Users can plausibly be logged in on more than one device
  - "The app lost my changes" tickets are starting to appear

Rate limit per user identity, not per IP, the moment your
concurrent user count is high enough that one misbehaving
client's load is noticeable against the total.

Wrapping up

None of this is exotic. Serialized token refresh, idempotent writes, explicit conflict resolution, and per-user rate limiting are all well-understood practices — the actual skill is noticing when each one stops being optional, because a growing user base rarely announces the problem directly. It shows up as a handful of random logouts, an occasional duplicated record, or a general sense that the app is "a bit slower than it used to be" while every individual dashboard still looks fine on average.

The backend half of this story — how session and permission data is stored and read reliably under concurrent load — sits on the same infrastructure I've covered before: the MongoDB disaster recovery strategy I run in production, and the underlying self-hosted infrastructure choices laid out in building my own Firebase alternative.

Related Posts

From 0 to 100K Users: Handling Users Smoothly from the Flutter Client to the Backend | Khalid Arafa