Redis Performance Tuning: From Configuration to Monitoring
Redis processes commands in a single-threaded event loop, which means its performance ceiling is determined by how efficiently that single thread handles each operation. While this architecture eliminates concurrency bugs and lock contention, it also means that a single slow command blocks every other client. Tuning Redis for production workloads requires understanding memory management, persistence trade-offs, command efficiency, and the operational patterns that protect that single thread from bottlenecks.
Most Redis performance problems in production are not caused by Redis itself but by how applications interact with it. Oversized values, missing key expiration, unoptimized data structures, and naive command patterns are responsible for the majority of latency spikes and memory exhaustion events.
Memory Architecture and Configuration
Redis stores all data in memory, making memory management the most critical tuning dimension. The maxmemory directive sets the hard upper bound on Redis memory usage. Without this setting, Redis grows until the operating system kills the process via the OOM killer, typically at the worst possible time during peak traffic.
# redis.conf - Memory configuration
maxmemory 4gb
maxmemory-policy allkeys-lfu
# Memory optimization for small values
hash-max-listpack-entries 128
hash-max-listpack-value 64
list-max-listpack-size -2
set-max-listpack-entries 128
zset-max-listpack-entries 128
zset-max-listpack-value 64
The eviction policy determines what happens when Redis reaches maxmemory. The allkeys-lfu (Least Frequently Used) policy provides the best cache hit ratio for most workloads because it evicts keys that are accessed least often, regardless of how recently they were last accessed. The older allkeys-lru policy evicts the least recently used key, which can prematurely evict frequently accessed keys that happen to have a brief idle period.
Memory Optimization Techniques
Redis internal data structures have significant per-key overhead. A simple string key-value pair consumes approximately 90 bytes of overhead beyond the key and value data. For workloads with millions of small values, this overhead dominates total memory usage.
Hash packing reduces per-key overhead by storing multiple fields under a single Redis key. Instead of storing 1 million individual keys like user:1234:name, user:1234:email, user:1234:role, store a single hash user:1234 with fields for name, email, and role. This reduces the number of top-level keys by a factor of 3 and leverages the listpack encoding for small hashes, which is dramatically more memory-efficient.
Persistence Trade-offs
Redis offers two persistence mechanisms, each with distinct performance characteristics. RDB (Redis Database) persistence creates periodic point-in-time snapshots by forking the process. AOF (Append-Only File) persistence logs every write command, providing stronger durability guarantees at higher I/O cost.
| Parameter | RDB | AOF | RDB + AOF |
|---|---|---|---|
| Durability | Minutes of loss possible | 1 second or less | Best of both |
| Write performance | No impact between saves | fsync per second (~5% overhead) | AOF-limited |
| Fork overhead | Large (copies page tables) | Smaller rewrites | Both fork types |
| Recovery speed | Fast (load snapshot) | Slow (replay commands) | Uses RDB for speed |
| Disk usage | Compact | Larger (command log) | Both files |
For pure caching workloads where data can be regenerated from the database, disable both persistence mechanisms entirely. This eliminates fork overhead and disk I/O, allowing Redis to dedicate all resources to serving requests. Set save "" and appendonly no in the configuration.
For workloads requiring persistence, appendfsync everysec provides a good balance between durability and performance. Redis buffers AOF writes and fsyncs to disk once per second, limiting the potential data loss window to approximately one second while avoiding the severe performance penalty of appendfsync always, which fsyncs after every write command.
Pipeline Batching
Each Redis command incurs a network round-trip between the client and server. For a workload that sends 100 sequential commands, the total latency includes 100 round-trip times. On a network with 0.5 millisecond RTT, the network overhead alone is 50 milliseconds, far exceeding the actual command processing time.
Pipelining sends multiple commands to Redis without waiting for individual responses, then reads all responses at once. This collapses 100 round trips into a single round trip, reducing network overhead by two orders of magnitude.
# Without pipeline: 100 round trips
for key in keys:
value = redis.get(key) # 0.5ms RTT each
# Total: ~50ms
# With pipeline: 1 round trip
pipe = redis.pipeline(transaction=False)
for key in keys:
pipe.get(key)
values = pipe.execute() # Single round trip
# Total: ~0.5ms + processing time
Set transaction=False when using pipelines purely for batching without atomicity requirements. Transactional pipelines (the default in many client libraries) use MULTI/EXEC wrapping, which adds overhead and prevents Redis from interleaving pipelined commands with requests from other clients. For application caching patterns where individual cache lookups are independent, non-transactional pipelines provide better throughput.
Command Optimization
Certain Redis commands have O(N) time complexity and can block the single-threaded event loop for extended periods when operating on large data structures.
KEYS *— Scans every key in the database. UseSCANwith a cursor for iterative key discovery in production.SMEMBERSon large sets — Returns all members at once. UseSSCANfor large sets.HGETALLon large hashes — Retrieves all fields. UseHSCANorHMGETfor specific fields.SORT— Sorts entire collections in memory. Pre-sort using sorted sets instead.DELon large keys — Synchronously frees memory. UseUNLINKfor asynchronous deletion of large keys.
The SLOWLOG command identifies operations that exceed a configured time threshold. Set slowlog-log-slower-than 10000 (10 milliseconds in microseconds) to capture commands that could cause noticeable latency. Review the slow log regularly to catch emerging performance problems before they affect users.
# Configure slow log
CONFIG SET slowlog-log-slower-than 10000
CONFIG SET slowlog-max-len 128
# Review slow commands
SLOWLOG GET 10
Cluster Mode Architecture
Redis Cluster distributes data across multiple nodes using hash slot assignment. The 16,384 hash slots are divided among primary nodes, and each key maps to a specific slot based on a CRC16 hash of the key name. This architecture enables horizontal scaling beyond the memory and throughput limits of a single node.
Cluster mode introduces several performance considerations. Cross-slot operations (commands that reference keys on different nodes) require client-side coordination or the use of hash tags to force related keys onto the same slot. The {tag} syntax in key names ensures that all keys sharing the same hash tag map to the same slot.
# These keys hash to different slots - cross-slot operations fail
SET user:1234:profile "..."
SET user:1234:sessions "..."
# Hash tags force same slot assignment
SET {user:1234}:profile "..."
SET {user:1234}:sessions "..."
# Both keys hash based on "user:1234" → same slot
Cluster resharding (adding or removing nodes) moves hash slots between nodes, which triggers key migration. During migration, accessing a key in a migrating slot results in an ASK redirect, adding one extra round trip. Plan resharding operations during low-traffic periods and monitor latency during the process. The connection pooling configuration should account for the additional connections required to all cluster nodes.
Pub/Sub Performance
Redis Pub/Sub enables event-driven architectures where publishers broadcast messages to channels and subscribers receive them in real time. Pub/Sub operates outside the key-value data model and does not persist messages. A subscriber that disconnects misses all messages published during the disconnection.
Pub/Sub throughput depends on the number of subscribers per channel and the message size. Each published message is copied to every subscriber's output buffer, so publishing a 1 KB message to 100 subscribers requires 100 KB of buffer memory. Large subscriber counts with frequent messages can overwhelm the output buffer limits, causing Redis to disconnect slow subscribers.
For workloads requiring reliable message delivery with persistence and acknowledgment, Redis Streams provide a more robust alternative. Streams support consumer groups, message acknowledgment, and automatic claim of unprocessed messages from failed consumers, making them suitable for task queues and event processing pipelines.
Monitoring and Alerting
Effective Redis monitoring tracks both operational health and performance efficiency. The INFO command provides comprehensive statistics across multiple categories. Integrate these metrics with your application performance monitoring stack for correlation with application-level performance data.
- used_memory vs maxmemory — Memory utilization percentage. Alert at 80% to allow time for scaling before eviction begins.
- evicted_keys — Rate of key evictions due to memory pressure. Any non-zero eviction rate in a non-cache workload indicates insufficient memory.
- keyspace_hits vs keyspace_misses — Cache hit ratio. Calculate as
hits / (hits + misses). Below 90% warrants investigation into TTL values and access patterns. - connected_clients — Current client connection count. Sudden spikes may indicate connection leaks in application code.
- instantaneous_ops_per_sec — Current command throughput. Compare against baseline to detect traffic anomalies.
- latest_fork_usec — Duration of the most recent fork operation for persistence. Fork times exceeding 100 milliseconds on large datasets cause latency spikes visible to clients.
Set up alerting on used_memory_rss (resident set size) rather than used_memory because RSS includes memory fragmentation and overhead not captured by the logical memory counter. Redis memory fragmentation ratio (RSS / used_memory) above 1.5 indicates significant fragmentation that can be addressed by restarting the instance or enabling activedefrag.
everysec fsync as the baseline and adjust based on durability requirements.
Frequently Asked Questions
How much memory does Redis actually use per key?
Each top-level key in Redis consumes approximately 90 bytes of overhead beyond the key name and value data. This overhead includes the dictionary entry, the key string object, the value object metadata, and expiration tracking if a TTL is set. For small values, the overhead can exceed the data size. Hash packing, where multiple fields are stored under a single hash key, reduces this overhead by sharing a single top-level key entry across multiple data fields.
Should I use RDB or AOF persistence?
Use both RDB and AOF together for production workloads requiring persistence. RDB provides fast recovery by loading a compact snapshot, while AOF captures commands written since the last snapshot, minimizing data loss. For pure caching where data can be regenerated, disable both for maximum performance. If choosing only one, AOF with everysec fsync provides the best durability-to-performance ratio, limiting potential data loss to approximately one second.
When should I use Redis Cluster versus a single instance with replicas?
Use Redis Cluster when your dataset exceeds the memory capacity of a single server or when your command throughput saturates a single core. A single Redis instance with replicas provides read scaling but the primary handles all writes. Redis Cluster shards the keyspace across multiple primaries, enabling both memory and write throughput scaling. Most applications with less than 25 GB of data and fewer than 100,000 operations per second work well with a single primary and replicas.
How do I diagnose Redis latency spikes?
Start with the slow log to identify commands taking longer than expected. Check the latest_fork_usec metric to determine if persistence forks are causing pauses. Use the LATENCY command to track different latency event sources. Monitor the operating system for swap activity, which causes severe Redis performance degradation because any swapped memory page access blocks the event loop. Disable transparent huge pages and ensure the Redis process has sufficient resident memory.
What is the maximum recommended number of keys in a Redis instance?
Redis can handle hundreds of millions of keys as long as sufficient memory is available. The practical limit is determined by memory, not key count. However, operations that scan all keys, such as KEYS, DBSIZE-related bookkeeping, and expiration sampling, become slower with very high key counts. For instances exceeding 100 million keys, consider partitioning across multiple instances or using Redis Cluster to distribute the keyspace and reduce per-instance operational overhead.