Home›Performance›Server Response Time

Server Response Time Optimization: Reducing TTFB for Faster Delivery

Tom LindgrenSeptember 22, 202614 min read

Time to First Byte (TTFB) measures the elapsed time from when a client sends an HTTP request to when it receives the first byte of the response. It is the most direct indicator of server-side performance and a critical component of Largest Contentful Paint (LCP) — every millisecond of TTFB delays the start of content rendering in the browser.

Google considers a TTFB under 800ms acceptable for web pages, but competitive applications target sub-200ms for origin responses and sub-100ms for CDN-cached content. This guide breaks down the components of TTFB, identifies the most common server-side bottlenecks, and provides concrete optimization strategies from database tuning to edge caching.

Anatomy of TTFB

TTFB is not a single operation — it is the sum of multiple sequential operations, each of which can become a bottleneck.

TTFB Breakdown DNS Lookup TCP Handshake TLS Negotiation Request Transfer Server Processing App logic + DB + APIs Response First byte ~5-50ms ~10-50ms ~10-80ms ~1-5ms 50-2000ms+ ~1-5ms Server Processing is almost always the largest contributor to TTFB Network phases (DNS + TCP + TLS) can be minimized with CDN and connection reuse

Database Query Optimization

Database queries are the single most common cause of high TTFB. A page that executes 15 queries averaging 40ms each adds 600ms of server processing time before network latency even enters the picture.

Identifying Slow Queries

-- PostgreSQL: Find top slow queries
SELECT
    calls,
    round(total_exec_time::numeric, 2) AS total_ms,
    round(mean_exec_time::numeric, 2) AS avg_ms,
    round(max_exec_time::numeric, 2) AS max_ms,
    query
FROM pg_stat_statements
ORDER BY mean_exec_time DESC
LIMIT 20;

-- MySQL: Enable slow query log
SET GLOBAL slow_query_log = 'ON';
SET GLOBAL long_query_time = 0.1;  -- 100ms threshold
SET GLOBAL log_queries_not_using_indexes = 'ON';

Common Query Optimizations

Caching Strategies

Caching eliminates repeated work by storing computed results and serving them directly on subsequent requests. Effective caching operates at multiple layers, each with different trade-offs between freshness and performance.

Cache LayerHit TimeBest ForInvalidation
Application memory (in-process)<1msConfiguration, static reference data, session dataTTL or application restart
Distributed cache (Redis/Memcached)1-5msDatabase query results, computed API responses, rate limit countersTTL, explicit invalidation, pub/sub
HTTP cache (reverse proxy)1-10msFull page responses, API responses, static assetsCache-Control headers, purge API
CDN edge cache5-50msStatic assets, cacheable pages, API responses with geographic distributionTTL, purge API, stale-while-revalidate

Cache-Aside Pattern

# Python: Cache-aside with Redis
import redis
import json

cache = redis.Redis(host='redis', port=6379, db=0)

def get_product(product_id):
    cache_key = f"product:{product_id}"

    # Try cache first
    cached = cache.get(cache_key)
    if cached:
        return json.loads(cached)

    # Cache miss: query database
    product = db.query(
        "SELECT * FROM products WHERE id = %s", product_id
    )

    # Store in cache with TTL
    cache.setex(
        cache_key,
        timedelta(minutes=15),
        json.dumps(product)
    )

    return product

def update_product(product_id, data):
    db.execute(
        "UPDATE products SET ... WHERE id = %s", product_id
    )
    # Invalidate cache immediately
    cache.delete(f"product:{product_id}")

Connection Pooling

Opening a new database connection for every request is expensive — TCP handshake, TLS negotiation, authentication, and protocol setup can add 20-100ms per connection. Connection pooling maintains a reservoir of pre-established connections that requests can borrow and return.

# Python: Connection pooling with SQLAlchemy
from sqlalchemy import create_engine

engine = create_engine(
    "postgresql://user:pass@db:5432/myapp",
    pool_size=20,          # Maintained connections
    max_overflow=10,       # Temporary connections above pool_size
    pool_timeout=30,       # Seconds to wait for a connection
    pool_recycle=1800,     # Recycle connections after 30 min
    pool_pre_ping=True,    # Verify connection health before use
)

Connection Pool Sizing

Pool size should match your application's concurrency, not your peak traffic. A common formula: pool_size = (number_of_CPU_cores * 2) + effective_spindle_count. For cloud-managed databases, consult the provider's connection limit documentation — exceeding the database's maximum connections causes connection errors that are worse than the latency a smaller pool introduces.

CDN and Edge Delivery

Content Delivery Networks (CDNs) fundamentally reshape the TTFB equation by serving cached content from edge servers geographically close to the user. Instead of every request traveling to the origin server (potentially across continents), the CDN serves responses from the nearest Point of Presence (PoP).

How CDNs Reduce TTFB

CDN impact on TTFB: For static content, a well-configured CDN typically reduces TTFB from 200-800ms (origin) to 10-50ms (edge cache hit). For dynamic content using edge compute, TTFB reductions of 60-80% are achievable compared to centralized origin processing. The Asia-Pacific region sees the most dramatic improvements due to the long network distances to US/EU origin servers — organizations serving APAC markets should prioritize CDN providers with dense PoP coverage in the region.

Cache-Control Headers

# Nginx: Cache-Control configuration for CDN edge caching
location /static/ {
    # Static assets: cache aggressively at CDN and browser
    add_header Cache-Control "public, max-age=31536000, immutable";
}

location /api/products {
    # API responses: cache at CDN, short TTL
    add_header Cache-Control "public, max-age=60, s-maxage=300";
    add_header CDN-Cache-Control "max-age=300";
    # s-maxage overrides max-age for shared caches (CDN)
    # CDN-Cache-Control is CDN-specific (Cloudflare, Fastly)
}

location /api/user {
    # Per-user content: no CDN caching
    add_header Cache-Control "private, no-store";
}

location /api/feed {
    # Stale-while-revalidate: serve stale for 60s while
    # fetching fresh content in the background
    add_header Cache-Control "public, max-age=10, stale-while-revalidate=60";
}

Application-Level Optimizations

Async and Non-Blocking I/O

Synchronous request processing blocks the thread while waiting for database queries, API calls, or file I/O. Async frameworks (Node.js, Python asyncio, Go goroutines) handle I/O concurrently, allowing a single process to handle thousands of concurrent requests.

# Python: Async handler with concurrent database queries
import asyncio
import asyncpg

async def get_dashboard_data(user_id):
    async with pool.acquire() as conn:
        # Execute queries concurrently instead of sequentially
        profile, orders, notifications = await asyncio.gather(
            conn.fetchrow(
                "SELECT * FROM users WHERE id = $1", user_id
            ),
            conn.fetch(
                "SELECT * FROM orders WHERE user_id = $1 "
                "ORDER BY created_at DESC LIMIT 10", user_id
            ),
            conn.fetch(
                "SELECT * FROM notifications WHERE user_id = $1 "
                "AND read = false LIMIT 5", user_id
            ),
        )
        # Three queries in ~40ms (parallel)
        # instead of ~120ms (sequential)
        return {
            "profile": dict(profile),
            "orders": [dict(o) for o in orders],
            "notifications": [dict(n) for n in notifications],
        }

Response Compression

Compressing HTTP responses reduces transfer size by 60-80%, cutting both server egress bandwidth and client download time. Brotli offers 15-20% better compression than gzip for text content at a modest CPU cost.

# Nginx: Brotli and gzip compression
brotli on;
brotli_comp_level 4;  # 1-11, higher = better ratio, more CPU
brotli_types text/html text/plain text/css text/xml
             application/json application/javascript
             application/xml application/rss+xml
             image/svg+xml;

gzip on;
gzip_comp_level 5;
gzip_min_length 256;
gzip_types text/html text/plain text/css text/xml
           application/json application/javascript
           application/xml application/rss+xml
           image/svg+xml;

Profiling Server Performance

Optimization without measurement is guesswork. Profiling tools identify the actual bottlenecks in your request pipeline, often revealing surprises — the slow middleware, the chatty ORM, the blocking DNS lookup that nobody expected.

Request Profiling Checklist

  1. Measure total TTFB: Use synthetic monitoring from multiple locations to establish baseline TTFB across regions
  2. Profile the request lifecycle: Instrument middleware, routing, authentication, business logic, and template rendering as individual spans in your distributed tracing system
  3. Identify database bottlenecks: Log slow queries and count queries per request. Lighthouse audits flag high TTFB as a diagnostic opportunity
  4. Check for connection overhead: Monitor connection pool utilization — high wait times indicate pool exhaustion
  5. Validate cache effectiveness: Track cache hit ratios. A cache with a 40% hit ratio is doing more harm (complexity and memory cost) than good
  6. Monitor under load: TTFB at idle is meaningless. Profile under realistic traffic patterns using load testing tools

TTFB Targets by Content Type

Content TypeGood TTFBAcceptable TTFBPoor TTFB
CDN-cached static assets<50ms50-100ms>100ms
CDN-cached HTML pages<100ms100-200ms>200ms
Server-rendered pages (origin)<200ms200-500ms>500ms
API responses (simple query)<100ms100-300ms>300ms
API responses (complex aggregation)<500ms500-1000ms>1000ms

Common TTFB Pitfalls

Key Takeaways