Home›Database & Query Performance›Redis Performance Tuning
Database & Query Performance

Redis Performance Tuning: From Configuration to Monitoring

Redis processes commands in a single-threaded event loop, which means its performance ceiling is determined by how efficiently that single thread handles each operation. While this architecture eliminates concurrency bugs and lock contention, it also means that a single slow command blocks every other client. Tuning Redis for production workloads requires understanding memory management, persistence trade-offs, command efficiency, and the operational patterns that protect that single thread from bottlenecks.

Most Redis performance problems in production are not caused by Redis itself but by how applications interact with it. Oversized values, missing key expiration, unoptimized data structures, and naive command patterns are responsible for the majority of latency spikes and memory exhaustion events.

Memory Architecture and Configuration

Redis stores all data in memory, making memory management the most critical tuning dimension. The maxmemory directive sets the hard upper bound on Redis memory usage. Without this setting, Redis grows until the operating system kills the process via the OOM killer, typically at the worst possible time during peak traffic.

# redis.conf - Memory configuration
maxmemory 4gb
maxmemory-policy allkeys-lfu

# Memory optimization for small values
hash-max-listpack-entries 128
hash-max-listpack-value 64
list-max-listpack-size -2
set-max-listpack-entries 128
zset-max-listpack-entries 128
zset-max-listpack-value 64

The eviction policy determines what happens when Redis reaches maxmemory. The allkeys-lfu (Least Frequently Used) policy provides the best cache hit ratio for most workloads because it evicts keys that are accessed least often, regardless of how recently they were last accessed. The older allkeys-lru policy evicts the least recently used key, which can prematurely evict frequently accessed keys that happen to have a brief idle period.

Memory Optimization Techniques

Redis internal data structures have significant per-key overhead. A simple string key-value pair consumes approximately 90 bytes of overhead beyond the key and value data. For workloads with millions of small values, this overhead dominates total memory usage.

Hash packing reduces per-key overhead by storing multiple fields under a single Redis key. Instead of storing 1 million individual keys like user:1234:name, user:1234:email, user:1234:role, store a single hash user:1234 with fields for name, email, and role. This reduces the number of top-level keys by a factor of 3 and leverages the listpack encoding for small hashes, which is dramatically more memory-efficient.

Redis Memory: Individual Keys vs Hash Packing Individual Keys (270 bytes overhead) user:1234:name → "Alice" ~90B overhead user:1234:email → "a@b.com" ~90B overhead user:1234:role → "admin" ~90B overhead Hash Packing (90 bytes overhead) user:1234 → {name:"Alice", email:"a@b.com", role:"admin"} ~90B overhead 3× savings
Hash packing reduces per-key memory overhead by consolidating related data under a single key

Persistence Trade-offs

Redis offers two persistence mechanisms, each with distinct performance characteristics. RDB (Redis Database) persistence creates periodic point-in-time snapshots by forking the process. AOF (Append-Only File) persistence logs every write command, providing stronger durability guarantees at higher I/O cost.

ParameterRDBAOFRDB + AOF
DurabilityMinutes of loss possible1 second or lessBest of both
Write performanceNo impact between savesfsync per second (~5% overhead)AOF-limited
Fork overheadLarge (copies page tables)Smaller rewritesBoth fork types
Recovery speedFast (load snapshot)Slow (replay commands)Uses RDB for speed
Disk usageCompactLarger (command log)Both files

For pure caching workloads where data can be regenerated from the database, disable both persistence mechanisms entirely. This eliminates fork overhead and disk I/O, allowing Redis to dedicate all resources to serving requests. Set save "" and appendonly no in the configuration.

For workloads requiring persistence, appendfsync everysec provides a good balance between durability and performance. Redis buffers AOF writes and fsyncs to disk once per second, limiting the potential data loss window to approximately one second while avoiding the severe performance penalty of appendfsync always, which fsyncs after every write command.

Pipeline Batching

Each Redis command incurs a network round-trip between the client and server. For a workload that sends 100 sequential commands, the total latency includes 100 round-trip times. On a network with 0.5 millisecond RTT, the network overhead alone is 50 milliseconds, far exceeding the actual command processing time.

Pipelining sends multiple commands to Redis without waiting for individual responses, then reads all responses at once. This collapses 100 round trips into a single round trip, reducing network overhead by two orders of magnitude.

# Without pipeline: 100 round trips
for key in keys:
    value = redis.get(key)  # 0.5ms RTT each
# Total: ~50ms

# With pipeline: 1 round trip
pipe = redis.pipeline(transaction=False)
for key in keys:
    pipe.get(key)
values = pipe.execute()  # Single round trip
# Total: ~0.5ms + processing time

Set transaction=False when using pipelines purely for batching without atomicity requirements. Transactional pipelines (the default in many client libraries) use MULTI/EXEC wrapping, which adds overhead and prevents Redis from interleaving pipelined commands with requests from other clients. For application caching patterns where individual cache lookups are independent, non-transactional pipelines provide better throughput.

Command Optimization

Certain Redis commands have O(N) time complexity and can block the single-threaded event loop for extended periods when operating on large data structures.

The SLOWLOG command identifies operations that exceed a configured time threshold. Set slowlog-log-slower-than 10000 (10 milliseconds in microseconds) to capture commands that could cause noticeable latency. Review the slow log regularly to catch emerging performance problems before they affect users.

# Configure slow log
CONFIG SET slowlog-log-slower-than 10000
CONFIG SET slowlog-max-len 128

# Review slow commands
SLOWLOG GET 10

Cluster Mode Architecture

Redis Cluster distributes data across multiple nodes using hash slot assignment. The 16,384 hash slots are divided among primary nodes, and each key maps to a specific slot based on a CRC16 hash of the key name. This architecture enables horizontal scaling beyond the memory and throughput limits of a single node.

Cluster mode introduces several performance considerations. Cross-slot operations (commands that reference keys on different nodes) require client-side coordination or the use of hash tags to force related keys onto the same slot. The {tag} syntax in key names ensures that all keys sharing the same hash tag map to the same slot.

# These keys hash to different slots - cross-slot operations fail
SET user:1234:profile "..."
SET user:1234:sessions "..."

# Hash tags force same slot assignment
SET {user:1234}:profile "..."
SET {user:1234}:sessions "..."
# Both keys hash based on "user:1234" → same slot

Cluster resharding (adding or removing nodes) moves hash slots between nodes, which triggers key migration. During migration, accessing a key in a migrating slot results in an ASK redirect, adding one extra round trip. Plan resharding operations during low-traffic periods and monitor latency during the process. The connection pooling configuration should account for the additional connections required to all cluster nodes.

Pub/Sub Performance

Redis Pub/Sub enables event-driven architectures where publishers broadcast messages to channels and subscribers receive them in real time. Pub/Sub operates outside the key-value data model and does not persist messages. A subscriber that disconnects misses all messages published during the disconnection.

Pub/Sub throughput depends on the number of subscribers per channel and the message size. Each published message is copied to every subscriber's output buffer, so publishing a 1 KB message to 100 subscribers requires 100 KB of buffer memory. Large subscriber counts with frequent messages can overwhelm the output buffer limits, causing Redis to disconnect slow subscribers.

For workloads requiring reliable message delivery with persistence and acknowledgment, Redis Streams provide a more robust alternative. Streams support consumer groups, message acknowledgment, and automatic claim of unprocessed messages from failed consumers, making them suitable for task queues and event processing pipelines.

Monitoring and Alerting

Effective Redis monitoring tracks both operational health and performance efficiency. The INFO command provides comprehensive statistics across multiple categories. Integrate these metrics with your application performance monitoring stack for correlation with application-level performance data.

Set up alerting on used_memory_rss (resident set size) rather than used_memory because RSS includes memory fragmentation and overhead not captured by the logical memory counter. Redis memory fragmentation ratio (RSS / used_memory) above 1.5 indicates significant fragmentation that can be addressed by restarting the instance or enabling activedefrag.

Key Takeaway: Redis performance tuning focuses on protecting the single-threaded event loop from blocking operations, right-sizing memory with appropriate eviction policies, using pipeline batching to reduce network overhead, and monitoring key health metrics. For pure caching, disable persistence entirely. For data that matters, use AOF with everysec fsync as the baseline and adjust based on durability requirements.

Frequently Asked Questions

How much memory does Redis actually use per key?

Each top-level key in Redis consumes approximately 90 bytes of overhead beyond the key name and value data. This overhead includes the dictionary entry, the key string object, the value object metadata, and expiration tracking if a TTL is set. For small values, the overhead can exceed the data size. Hash packing, where multiple fields are stored under a single hash key, reduces this overhead by sharing a single top-level key entry across multiple data fields.

Should I use RDB or AOF persistence?

Use both RDB and AOF together for production workloads requiring persistence. RDB provides fast recovery by loading a compact snapshot, while AOF captures commands written since the last snapshot, minimizing data loss. For pure caching where data can be regenerated, disable both for maximum performance. If choosing only one, AOF with everysec fsync provides the best durability-to-performance ratio, limiting potential data loss to approximately one second.

When should I use Redis Cluster versus a single instance with replicas?

Use Redis Cluster when your dataset exceeds the memory capacity of a single server or when your command throughput saturates a single core. A single Redis instance with replicas provides read scaling but the primary handles all writes. Redis Cluster shards the keyspace across multiple primaries, enabling both memory and write throughput scaling. Most applications with less than 25 GB of data and fewer than 100,000 operations per second work well with a single primary and replicas.

How do I diagnose Redis latency spikes?

Start with the slow log to identify commands taking longer than expected. Check the latest_fork_usec metric to determine if persistence forks are causing pauses. Use the LATENCY command to track different latency event sources. Monitor the operating system for swap activity, which causes severe Redis performance degradation because any swapped memory page access blocks the event loop. Disable transparent huge pages and ensure the Redis process has sufficient resident memory.

What is the maximum recommended number of keys in a Redis instance?

Redis can handle hundreds of millions of keys as long as sufficient memory is available. The practical limit is determined by memory, not key count. However, operations that scan all keys, such as KEYS, DBSIZE-related bookkeeping, and expiration sampling, become slower with very high key counts. For instances exceeding 100 million keys, consider partitioning across multiple instances or using Redis Cluster to distribute the keyspace and reduce per-instance operational overhead.