Redis Pub/Sub vs Caching vs Rate Limiting: When to Add Each Layer

Redis can cache reads, enforce shared limits, or broadcast events—but each solves a different problem. Learn when to add each layer and when not to.
Redis Pub/Sub vs Caching vs Rate Limiting: When to Add Each Layer
Your database is getting slower, one client is flooding an endpoint, and your WebSocket servers need to share events. Redis can help with all three, but not with the same feature, and not necessarily by adding three pieces of infrastructure at once.
Redis Pub/Sub vs caching vs rate limiting is really a choice between three different kinds of shared state: copies of data, counters that shape traffic, and short-lived messages. By the end of this guide, you’ll know which problem each solves, what to measure first, and when a Redis layer is worth operating.
Tip
TL;DR , Add caching when repeated reads are measurably expensive and bounded staleness is acceptable. Add Redis-backed rate limiting when multiple app instances must enforce the same quota. Add Pub/Sub when live consumers on different processes need the same transient event. If you need replay or guaranteed processing, use a durable queue or Redis Streams instead.
Redis Pub/Sub vs caching vs rate limiting: three different jobs
These patterns all use Redis, but their data and failure behavior are different. Treating them as interchangeable is how a cache quietly turns into a queue or a notification channel becomes a source of truth.
| Layer | What Redis holds | Add it when… | It is not… |
|---|---|---|---|
| Cache | A temporary copy of data | Repeated reads make your database or downstream API slow or expensive | The canonical database |
| Rate limiter | Counters or token state | Several processes must make one consistent allow/deny decision | A cache of user data |
| Pub/Sub | No durable history; messages are sent to current subscribers | One event needs to reach multiple live processes quickly | A job queue or event log |
A quick way to tell whether you need Redis is to name the state you need to share. Is it a reusable answer, a quota, or a notification? If you can’t name the state and the failure it prevents, instrument your current system before adding another service.
This is related to, but different from, rate limiting and caching at the API gateway layer. A gateway can reject or serve a request before it reaches your app; this guide focuses on when Redis should provide shared state behind those decisions.
Redis caching: add it when repeated reads are the bottleneck
Caching is usually the first Redis feature teams consider. It’s also the easiest one to add for the wrong reason. A fast cache doesn’t help much if most requests miss, the database is already comfortably below its latency target, or the cached data changes too often to serve safely.
Look for repeated reads of the same relatively stable data: product details, public profiles, configuration, or an expensive report. Measure database query volume, connection-pool pressure, and p95/p99 latency first. If the same query dominates those numbers, a cache-aside layer can avoid doing that work on every request. Redis’s cache-aside guide describes this pattern.
The application checks Redis, falls back to the source on a miss, and stores the result for a bounded time:
async function getProduct(id: string) {
const key = `product:${id}`;
const cached = await redis.get(key);
if (cached !== null) return JSON.parse(cached);
const product = await database.products.findById(id);
if (product) {
await redis.set(key, JSON.stringify(product), { EX: 60 });
}
return product;
}The 60-second TTL is an example, not a universal setting. Pick a staleness window the product can tolerate, then invalidate a key after a successful write when users should see changes sooner. Add jitter to expiration or a stampede-control strategy if many requests can miss on the same popular key at once.
A cache is a copy, so decide what happens when Redis is unavailable. Usually the app should fall back to the database with a short timeout, but only if the database can handle that traffic. Track hit rate, miss rate, Redis latency, evictions, and database load together; a high hit rate is not a win if users are seeing stale results or evictions are constant.
Redis rate limiting: add it when quotas must be shared
A rate limiter answers a different question: should this request be allowed now? A counter stored in one application process is enough for a single server, but it gives each instance its own quota. Behind a load balancer, a client can spread requests across instances and exceed the limit.
Use Redis when that inconsistency matters, for example, per-account API quotas, login-attempt controls, tenant usage caps, or protection for a costly third-party API. Choose the identity carefully: an IP can be useful for anonymous traffic, while authenticated routes usually need an account, API key, or tenant identifier.
A simple fixed-window counter can be made atomic with a Redis transaction. This Node.js example counts requests for one subject in one 60-second bucket:
const windowSeconds = 60;
const bucket = Math.floor(Date.now() / (windowSeconds * 1000));
const key = `rate-limit:${subjectId}:${bucket}`;
const result = await redis
.multi()
.incr(key)
.expire(key, windowSeconds)
.exec();
const count = Number(result[0]);
const allowed = count <= 100;Fixed windows are easy to explain, but traffic can burst at the boundary: a client may use nearly the full quota just before the clock rolls over, then use it again immediately after. For smoother limits or controlled bursts, use a sliding-window or token-bucket algorithm. Redis’s rate-limiter documentation covers the common patterns and atomicity concerns.
Decide what a Redis outage means before enabling enforcement. Some endpoints should fail open to preserve availability; expensive or abuse-sensitive operations may need to fail closed. Return a clear 429 response and, where possible, a retry hint. Monitor allowed and rejected requests, key cardinality, and Redis latency, not just the total request count.
Redis Pub/Sub: add it when live processes need to hear about an event
Pub/Sub is for fan-out. One process publishes an event to a channel, and processes currently subscribed to that channel receive it. This is useful for live notifications, presence updates, WebSocket broadcasts, or telling several app instances that a cache entry changed.
It becomes useful when an in-process event emitter stops being enough: the publisher and consumer may run on different machines, or a WebSocket client may be connected to a different server from the one handling the update. That cross-instance coordination is one of the scaling problems covered in what breaks as WebSocket connections grow.
The tradeoff is delivery. Redis Pub/Sub is at-most-once: a subscriber that is disconnected or unable to process a message misses it, and Redis doesn’t store the message for later. That is fine for a transient “refresh this view” hint. It is not fine for “charge this order,” “send this invoice,” or any job that must survive a restart. For those, use a durable queue or Redis Streams with an acknowledgement and retry strategy.
Cache invalidation is a tempting Pub/Sub use case, but account for missed messages. If a subscriber disconnects, it may keep stale local data after reconnecting. A short TTL, reconnect-triggered cache flush, version check, or Redis client-side tracking can provide a recovery path. Don’t rely on a best-effort notification as the only thing keeping business data correct.
When should you add each Redis layer?
Add the smallest layer that addresses a measured problem. The order below is a useful starting point, not a rule that every application must follow.
- Start without Redis. Keep the database as the source of truth and use an in-process event for work that stays inside one process. Record latency and load so you know what actually needs attention.
- Add a cache for hot reads. Use it when repeated reads are consuming meaningful database capacity or missing your latency target, and you can define acceptable staleness plus an invalidation plan.
- Add shared rate-limit state when you scale out or need consistent quotas. If one process can enforce the policy reliably, a remote store may be unnecessary. If requests hit multiple instances, local counters no longer represent one shared limit.
- Add Pub/Sub when independent live consumers need fan-out. Keep it for transient signals. If consumers must resume after downtime, use a durable messaging feature instead.
If you’re also deciding where Redis should run, start with the operational basics in how to Dockerize an application: health checks, configuration, and service boundaries matter just as much as the Redis command you choose.
The operational cost is part of the decision
Redis adds a network hop and another dependency to deploy, monitor, secure, and pay for. It also introduces memory limits, key naming, eviction policy, backup questions, and behavior during failover. A layer that reduces database calls can still make the overall system less reliable if every request now depends on an overloaded Redis instance.
Before shipping, write down four things: what data or message is stored, how long it lives, what happens when Redis is unavailable, and which metric tells you the layer is helping. Set connection and command timeouts, use TLS and authentication on reachable deployments, avoid unbounded keys, and test the fallback path. Don’t put secrets or high-risk authorization decisions in a cache without an explicit freshness and invalidation policy.
The right monitoring differs by layer: cache hit rate and source load for caching; allow/deny counts and limit-key growth for rate limiting; subscriber health and publish volume for Pub/Sub. At the Redis level, keep an eye on memory, evictions, command latency, rejected connections, and replication/failover health.
Warning
Pub/Sub is not durable just because Redis is durable for other data types. A persisted key or stream and a Pub/Sub message have different guarantees. Pick the primitive based on what must happen when a consumer is offline.
Wrapping Up
Redis caching, rate limiting, and Pub/Sub solve three separate problems: repeated work, shared quotas, and live fan-out. Add each only when you can point to the bottleneck or coordination requirement it addresses, and define its TTL, failure behavior, and success metric before it becomes part of the request path.
Start with one layer, measure its effect, and add the next only when the system gives you a reason.
Building a system that needs these tradeoffs? Browse my projects or get in touch.
Related Posts
Related Articles

Rate Limiting and Caching at the API Gateway Layer: Protecting Your Backend From Its Own Traffic
Most outages aren't attacks, they're your own traffic overwhelming itself. A concept-first look at how rate limiting and caching work together at the API gateway to keep one misbehaving client, integration, or traffic spike from degrading service for everyone else.

The Backup Strategy That Actually Protects Your VPS Data (And Why Automation Isn't Optional)
Manual backups fail under pressure, and unmonitored automated ones fail silently. A layered VPS backup strategy, the 3-2-1 rule, application-aware dumps, encrypted off-site transfer, dead-man's-switch monitoring.

Backup Strategy for a Dockerized MongoDB Replica Set: Disaster Recovery You've Actually Tested
A replica set isn't a backup — it protects against a dead node, not a dropped collection or a bad migration. Here's the mongodump/oplog setup I actually run against a Dockerized MongoDB replica set, why restore testing is the step everyone skips, and when to graduate to filesystem snapshots or PBM i
