Home›Database & Query Performance›Caching Layers Architecture
Database & Query Performance

Multi-Layer Caching Architecture for Web Applications

Caching is the most effective tool for reducing latency and increasing throughput in web applications. A well-designed caching architecture eliminates redundant computation, reduces database load, and moves data closer to users. The challenge lies not in implementing a single cache but in designing a coherent multi-layer system where each tier serves a specific purpose and the invalidation strategy keeps data consistent across all layers.

Modern web applications typically employ four distinct caching layers: browser cache, CDN edge cache, application-level cache, and database query cache. Each layer intercepts requests at a different point in the stack, and the cumulative effect of a high cache hit ratio across all layers can reduce origin server load by 95% or more.

The Four-Layer Cache Stack

Multi-Layer Cache Architecture User Layer 1 Browser Cache Layer 2 CDN Edge Cache Layer 3 App Cache (Redis) Layer 4 Query Cache / DB Typical Latency by Layer Browser Cache <1ms CDN Edge 5-30ms App Cache (Redis) 1-5ms Database Query 10-500ms Higher cache layers intercept requests before they reach slower tiers
Each caching layer intercepts requests closer to the user, with dramatically lower latency at higher tiers

Layer 1: Browser Cache

The browser cache is the fastest cache layer because it requires zero network requests. When a resource is served from the browser cache, the response time is effectively zero. Browser caching is controlled entirely through HTTP response headers sent by the origin server or CDN.

The Cache-Control header is the primary directive. For static assets like CSS, JavaScript, and images that use content-hashed filenames, Cache-Control: public, max-age=31536000, immutable instructs the browser to cache the resource for one year without revalidation. The immutable directive prevents unnecessary conditional requests during page navigation.

# Nginx: Static assets with hashed filenames
location ~* \.(css|js|woff2|png|jpg|webp|svg)$ {
    add_header Cache-Control "public, max-age=31536000, immutable";
}

# HTML pages: Always revalidate
location ~* \.html$ {
    add_header Cache-Control "no-cache";
    add_header ETag "";
}

For HTML documents and API responses, Cache-Control: no-cache does not prevent caching. It instructs the browser to store the response but validate it with the server before using it, using ETag or Last-Modified conditional headers. This pattern ensures users always see fresh content while avoiding full downloads when the content has not changed.

Layer 2: CDN Edge Cache

CDN edge caches store content at points of presence (PoPs) geographically distributed near users. When a CDN cache hit occurs, the response travels a short network distance from the nearest PoP rather than traversing the internet to the origin server. This reduces latency from hundreds of milliseconds to single-digit milliseconds for users near a PoP.

CDN cache behavior is controlled through Cache-Control headers directed specifically at shared caches. The s-maxage directive overrides max-age for shared caches like CDNs while allowing different browser cache durations. The stale-while-revalidate directive enables edge servers to serve stale content immediately while refreshing the cache in the background, eliminating the latency penalty of cache misses for subsequent users.

# Cache at CDN for 5 min, browser for 60s, stale serve for 1 hour
Cache-Control: public, max-age=60, s-maxage=300,
    stale-while-revalidate=3600

Cache key design at the CDN layer requires careful attention. The default cache key includes the full URL with query parameters, which means /api/products?sort=price and /api/products?sort=name are cached as separate entries. Normalize query parameter ordering, strip tracking parameters, and configure the CDN to include only relevant parameters in the cache key to improve hit ratios. For details on CDN-level optimization, see our guide on CDN performance fundamentals.

Layer 3: Application Cache (Redis and Memcached)

Application-level caching stores computed results, database query results, and serialized objects in an in-memory data store. Redis and Memcached are the two dominant technologies, each with distinct architectural characteristics.

FeatureRedisMemcached
Data structuresStrings, hashes, lists, sets, sorted sets, streamsStrings only
PersistenceRDB snapshots + AOF logNone (pure cache)
Memory efficiencyHigher per-key overhead (~90 bytes)Lower overhead (~48 bytes per slab)
ClusteringNative cluster with hash slotsClient-side consistent hashing
EvictionMultiple policies (LRU, LFU, volatile-*)LRU only
Pub/subBuilt-inNot supported

Redis has become the dominant choice for application caching because its rich data structure support enables caching patterns that Memcached cannot express efficiently. Storing a user profile as a Redis hash allows updating individual fields without deserializing and re-serializing the entire object. Sorted sets enable cached leaderboards that can be updated incrementally. For in-depth configuration guidance, see Redis performance tuning.

Cache-Aside Pattern

The cache-aside (or lazy-loading) pattern is the most widely used application caching strategy. The application first checks the cache. On a hit, it returns the cached data directly. On a miss, it queries the database, writes the result to the cache, and returns it to the caller.

async function getProduct(productId) {
  const cacheKey = `product:${productId}`;

  // Check cache first
  const cached = await redis.get(cacheKey);
  if (cached) {
    return JSON.parse(cached);
  }

  // Cache miss: query database
  const product = await db.query(
    'SELECT * FROM products WHERE id = $1',
    [productId]
  );

  // Write to cache with 5 minute TTL
  await redis.setex(cacheKey, 300, JSON.stringify(product));

  return product;
}

Write-Through and Write-Behind Patterns

Write-through caching updates the cache synchronously when the database is updated, ensuring the cache always contains the latest data. This eliminates the stale data window that exists in cache-aside patterns between a database write and cache expiration. The tradeoff is added latency on write operations because both the database and cache must be updated before the write completes.

Write-behind (also called write-back) caching writes to the cache immediately and asynchronously flushes changes to the database. This reduces write latency at the cost of potential data loss if the cache server fails before the flush completes. Write-behind is appropriate for data that can tolerate small amounts of loss, such as view counts, analytics events, or non-financial session data.

Cache Invalidation

Phil Karlton's observation that cache invalidation is one of the two hard problems in computer science remains accurate decades later. Invalidation complexity grows with the number of cache layers and the relationships between cached entities.

TTL-Based Expiration

Time-to-live (TTL) expiration is the simplest and most predictable invalidation strategy. Each cached entry has a fixed lifetime after which it is automatically removed. TTL values should reflect the data's tolerance for staleness: product catalog data might use a 5-minute TTL, while user session data might use 30 minutes.

The key advantage of TTL expiration is that it requires no coordination between the write path and the cache. Data eventually becomes consistent without explicit invalidation logic. The disadvantage is that data can be stale for up to the full TTL duration after a change, and cache misses cluster at TTL boundaries, creating periodic load spikes on the database.

Event-Driven Invalidation

For data that requires near-immediate consistency, event-driven invalidation explicitly removes or updates cache entries when the underlying data changes. Database triggers, change data capture (CDC) streams, or application-level events publish invalidation messages that cache subscribers act on.

# Event-driven invalidation via pub/sub
async def update_product(product_id, data):
    # Update database
    await db.execute(
        "UPDATE products SET ... WHERE id = %s", data, product_id
    )

    # Invalidate cache
    await redis.delete(f"product:{product_id}")

    # Invalidate related caches
    category_id = data.get('category_id')
    await redis.delete(f"category_products:{category_id}")
    await redis.delete("featured_products")

The challenge with event-driven invalidation is maintaining a complete mapping of which cache keys are affected by each data change. A product update might invalidate the product detail cache, the category listing cache, the search results cache, and the homepage featured products cache. Missing any of these mappings results in stale data that persists until TTL expiration.

Cache Stampede Prevention

A cache stampede occurs when a frequently accessed cache entry expires and many concurrent requests simultaneously discover the cache miss, all attempting to regenerate the cached value from the database. If the underlying query takes 500 milliseconds and 200 requests arrive during that window, all 200 execute the same expensive query concurrently, potentially overwhelming the database.

Locking

Distributed locking ensures that only one request regenerates the cache entry while others wait or receive the stale value. Redis SET NX EX provides atomic lock acquisition with automatic expiration. The winning request rebuilds the cache while other requests either wait briefly and retry or serve a stale cached copy.

Probabilistic Early Expiration

Rather than all entries expiring at exactly the same time, probabilistic early expiration adds jitter to TTL values. Each request that reads a cached value computes whether to proactively refresh it based on how close the entry is to expiring. As the entry approaches its TTL, the probability of a proactive refresh increases, ensuring that regeneration happens before expiration and only one or a few requests trigger it.

def should_refresh(cached_at, ttl, beta=1.0):
    """XFetch algorithm for probabilistic early refresh"""
    import random, math, time
    age = time.time() - cached_at
    remaining = ttl - age
    if remaining <= 0:
        return True
    # Higher probability as remaining time shrinks
    return remaining < beta * ttl * -math.log(random.random())

Layer 4: Database Query Cache

Database-level query caching stores the results of recently executed queries and returns them directly for identical subsequent queries without re-executing the query plan. PostgreSQL does not have a built-in query cache in the MySQL sense, instead relying on the operating system's filesystem cache and its own shared buffer pool to cache frequently accessed data pages.

MySQL's query cache was deprecated in version 5.7 and removed in 8.0 because it became a scalability bottleneck under concurrent write workloads. The query cache required a global lock for every read and was fully invalidated for any write to a table referenced in cached queries, making it counterproductive for write-heavy workloads.

For PostgreSQL, the shared buffer pool is the primary database-level cache. Sizing it correctly, typically 25% of total system memory, ensures that frequently accessed data pages and index pages remain in memory. Monitor the buffer cache hit ratio through pg_stat_database and target a hit ratio above 99% for OLTP workloads. The query optimization guide covers how properly indexed queries maximize buffer cache effectiveness.

Cache Observability

Each cache layer needs its own monitoring to diagnose performance issues and capacity planning. Track these metrics per layer:

Correlate cache metrics with server response time to understand how cache performance affects end-user latency. A 5% drop in CDN cache hit ratio can translate to a 50% increase in origin traffic and corresponding response time degradation.

Key Takeaway: Effective caching architecture layers multiple cache tiers, each serving a specific latency and consistency requirement. Browser and CDN caches handle static and semi-static content closest to the user. Application caches store computed results and database query outputs. The critical design decisions are TTL values, invalidation strategy, and stampede prevention, not which technology to use.

Frequently Asked Questions

Should I use Redis or Memcached for application caching?

Redis is the better default choice for most applications because it supports complex data structures, persistence, and pub/sub for cache invalidation. Memcached has a slight advantage in raw throughput for simple key-value workloads with very high concurrency because its multi-threaded architecture scales better on multi-core servers. Choose Memcached only when you need pure caching without persistence and your workload is exclusively simple string storage.

How long should cache TTLs be?

TTL values depend on data volatility and staleness tolerance. Static assets with hashed filenames can use TTLs of one year. Product catalog data typically works well with 5 to 15 minute TTLs. User session data uses 15 to 30 minutes. Real-time data like stock prices or live scores should use TTLs of 1 to 5 seconds or event-driven invalidation. Start with conservative TTLs and increase them as you confirm that staleness is acceptable.

How do I prevent stale data in a multi-layer cache?

Use event-driven invalidation to clear cache entries across all layers when data changes. For CDN caches, issue purge API calls for the specific URLs affected. For application caches, delete the relevant Redis keys in the same transaction or event handler that modifies the database. For browser caches, use content-hashed filenames for static assets and short max-age values with revalidation for dynamic content.

What is the difference between no-cache and no-store?

Cache-Control no-cache allows the browser and CDN to store the response but requires revalidation with the origin before serving it to subsequent requests. Cache-Control no-store prohibits any caching entirely: the browser must fetch the resource from the server on every request. Use no-cache for content that can be cached but must be validated for freshness. Use no-store only for truly sensitive data that should never persist in any cache, such as authentication tokens or personal financial information.

How do I measure the business impact of caching?

Measure the cache hit ratio at each layer and calculate the equivalent database load and latency reduction. A 95% CDN cache hit ratio means the origin server handles only 5% of total traffic. Quantify this as server cost savings and latency improvement for end users. Track the correlation between cache hit ratio changes and business metrics like page load time, conversion rate, and bounce rate to demonstrate the direct revenue impact of cache performance.