Home›Network Performance & Latency›Network Latency
Network Performance & Latency

Network Latency: Understanding and Reducing Round-Trip Time

Network latency is the time elapsed between sending a request and receiving the first byte of the response. For web applications, latency determines how fast pages load, how responsive interactive features feel, and how many concurrent users a system can serve. While bandwidth governs how much data can flow through a connection, latency governs how quickly that flow begins, and for the small request/response exchanges that dominate web traffic, latency is the dominant performance factor.

Understanding the components of network latency reveals optimization opportunities at each layer. A complete round trip includes DNS resolution, TCP connection establishment, TLS negotiation, request transmission, server processing, and response transmission. Each component contributes measurable delay, and reducing any single component improves the total time to first byte.

Anatomy of a Network Round Trip

When a browser requests a resource from a server for the first time, the total latency is the sum of multiple sequential steps. On a typical connection between a user in Tokyo and a server in Virginia, the physics of light traveling through fiber optic cable impose a minimum one-way latency of approximately 70 milliseconds. The complete first-request sequence adds significant overhead on top of this physical minimum.

For a first-time HTTPS connection with TLS 1.2, the minimum latency before any application data arrives is approximately 4 round trips: DNS + TCP + TLS + HTTP. At 140 ms RTT between Tokyo and Virginia, this translates to 560 ms before the first byte of the response reaches the browser. Optimizing each component systematically reduces this total.

Geographic Latency and CDN Impact

The speed of light in fiber optic cable is approximately 200,000 km/s, or about two-thirds the speed of light in vacuum. This physical constant creates a hard floor for latency based on distance. The great-circle distance between network endpoints, combined with the actual cable routing (which is rarely a straight line), determines the minimum achievable one-way latency.

RouteDistanceTheoretical Min RTTTypical RTT
Same data center<1 km<0.1 ms0.2 - 0.5 ms
Same region (NYC ↔ DC)~350 km3.5 ms5 - 10 ms
Cross-continent (NYC ↔ LA)~4,000 km40 ms60 - 80 ms
Transpacific (LA ↔ Tokyo)~8,800 km88 ms100 - 140 ms
Transatlantic (NYC ↔ London)~5,500 km55 ms70 - 90 ms

Content delivery networks reduce geographic latency by serving content from edge servers located close to users. Instead of every request traveling to the origin server, cached content is served from the nearest CDN point of presence. A user in Singapore accessing a CDN-cached resource from a Singapore PoP experiences 2 to 5 ms latency instead of the 180 ms round trip to a US-based origin server. The CDN performance fundamentals article covers edge caching architecture in detail.

TCP Optimization

TCP's congestion control algorithm directly impacts transfer performance on high-latency connections. The default initial congestion window (initcwnd) of 10 segments (approximately 14.6 KB) limits how much data the server can send in the first round trip after the connection is established. For web pages where the critical rendering path fits within this window, a single round trip delivers enough data to begin rendering. For larger resources, additional round trips are needed as TCP slowly increases the congestion window.

# Linux: increase initial congestion window
# Check current value
ip route show | grep initcwnd

# Set initcwnd to 10 (recommended default per RFC 6928)
ip route change default via 10.0.0.1 initcwnd 10 initrwnd 10

# Verify TCP settings
sysctl net.ipv4.tcp_congestion_control
# Should output: bbr (or cubic)

TCP BBR (Bottleneck Bandwidth and Round-trip propagation time) is a congestion control algorithm developed by Google that achieves higher throughput and lower latency than the traditional CUBIC algorithm, especially on high-bandwidth, high-latency connections. BBR models the network's bottleneck bandwidth and propagation delay to pace packets optimally, rather than relying on packet loss as a congestion signal.

Connection reuse through HTTP keep-alive and HTTP/2 multiplexing eliminates the TCP and TLS handshake overhead for subsequent requests to the same server. A single persistent connection can carry hundreds of requests without repeating the setup cost. The HTTP/2 and HTTP/3 performance article explores protocol-level optimizations in detail.

DNS Optimization

DNS resolution adds 20 to 100 milliseconds to the first request for each unique domain. Web pages that load resources from multiple domains (CDN, analytics, fonts, APIs) accumulate DNS resolution overhead for each distinct hostname.

DNS prefetching instructs the browser to resolve domain names proactively before the user clicks a link or the page references the domain. The dns-prefetch resource hint triggers background DNS resolution without establishing a connection.

<!-- Prefetch DNS for domains used on the page -->
<link rel="dns-prefetch" href="//cdn.example.com">
<link rel="dns-prefetch" href="//api.example.com">

<!-- Preconnect establishes TCP + TLS proactively -->
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>

On the server side, DNS TTL values control how long resolvers cache DNS records. Short TTLs (30 to 60 seconds) enable fast failover but increase DNS query volume and resolution latency for users whose cache has expired. Longer TTLs (300 to 3600 seconds) reduce lookup frequency but slow down DNS-based failover. For most web applications, a TTL of 300 seconds (5 minutes) provides a good balance. The DNS performance optimization guide covers resolver configuration and anycast strategies.

Measuring Network Latency

Accurate latency measurement requires distinguishing between network latency and server processing time. The Time to First Byte (TTFB) metric reported by browsers includes both network latency and server processing. To isolate network latency, compare TTFB measurements from different geographic locations for the same request, or use the TCP connection time (available in the Resource Timing API as connectEnd - connectStart) as a proxy for pure network RTT.

Traceroute reveals the path packets take between the client and server, including the latency contribution of each network hop. The mtr (My Traceroute) tool combines traceroute with continuous ping to show both the path and the latency stability of each hop over time, making it easier to identify intermittent routing problems.

# mtr provides continuous traceroute with statistics
mtr --report --report-cycles 100 api.example.com

# Output shows each hop with loss%, avg, and worst latency
# Hop 1:  10.0.0.1     0.0%  0.5ms
# Hop 2:  172.16.0.1   0.0%  1.2ms
# Hop 3:  203.0.113.1  0.1%  12.5ms  ← ISP gateway
# Hop 4:  198.51.100.1 0.0%  35.2ms  ← Peering point
# ...
# Hop 9:  93.184.216.1 0.0%  68.3ms  ← Destination

For production latency monitoring, synthetic monitoring from multiple geographic locations provides consistent baseline measurements independent of real user traffic patterns. Combine synthetic monitoring with Real User Monitoring (RUM) to understand the actual latency distribution experienced by your user population, including the long tail of high-latency connections from mobile networks and distant regions.

Key Takeaway: Network latency is a sum of components: DNS, TCP, TLS, and processing time. Each component offers optimization opportunities. CDNs eliminate geographic latency for cached content, TLS 1.3 reduces handshake overhead, connection reuse amortizes setup costs, and DNS prefetching overlaps resolution with page processing. Measure latency end-to-end and per-component to identify the highest-impact optimizations for your specific user base.
First HTTPS Request: Round-Trip Breakdown DNS Lookup 20-100ms TCP Handshake 1 RTT TLS Handshake 1-2 RTT HTTP Request 1 RTT Response + transfer Total: 4+ round trips before first byte At 140ms RTT (US ↔ Asia): ~560ms minimum latency With CDN: 1-2 RTTs at 5ms = 5-10ms total
Each step in the connection sequence adds one or more round trips before application data flows

Frequently Asked Questions

What is the difference between latency and bandwidth?

Latency is the time delay for a single packet to travel from source to destination and back, measured in milliseconds. Bandwidth is the maximum data transfer rate of the connection, measured in megabits per second. For web applications with many small requests, latency is the dominant performance factor because each request must wait for a round trip regardless of bandwidth. High bandwidth helps only when transferring large files where the transfer time exceeds the latency overhead.

How much does TLS add to connection latency?

TLS 1.2 adds 1 to 2 round trips for the handshake, depending on whether session resumption is available. TLS 1.3 reduces this to 1 round trip for new connections and supports 0-RTT resumption for repeat connections, eliminating the handshake latency entirely for returning visitors. On a 100ms RTT connection, upgrading from TLS 1.2 to TLS 1.3 saves 100 to 200 ms on first connections and up to 200ms on subsequent connections.

Can I reduce latency below the speed-of-light limit?

The speed of light in fiber sets a hard floor for network latency based on distance. You cannot go below this limit on a given path. However, you can reduce the effective distance by serving content from servers closer to users via CDN edge caching, or by using shorter network paths through peering agreements and anycast routing. Some financial trading firms use microwave links for shorter paths than fiber because microwave travels closer to the speed of light in vacuum.

How does HTTP/2 reduce latency compared to HTTP/1.1?

HTTP/2 multiplexes multiple requests over a single TCP connection, eliminating the need to open parallel connections for concurrent requests. It also uses header compression to reduce per-request overhead and supports server push to send resources before the client requests them. These features reduce the number of round trips and connections needed to load a page, with typical latency improvements of 20 to 50 percent for complex pages with many resources.

What tools should I use to diagnose latency problems?

Use mtr or traceroute to identify which network hops contribute the most latency. Use curl with timing output to measure DNS, TCP, TLS, and TTFB separately. Use browser DevTools Network panel waterfall to visualize request timing. For production monitoring, use synthetic monitoring tools that measure latency from multiple geographic locations and Real User Monitoring to capture actual user experience data across diverse network conditions.