Frontend Performance

Streaming Server-Side Rendering: Progressive HTML Delivery

Traditional server-side rendering waits for all data to be fetched and all HTML to be generated before sending a single byte to the browser. If a page requires three API calls that take 200ms, 500ms, and 800ms, the TTFB is at least 800ms plus rendering time — the entire page is blocked by its slowest data dependency. Streaming SSR eliminates this bottleneck by sending HTML to the browser as each section becomes ready, starting with the shell and progressively filling in data-dependent sections.

How Streaming SSR Works

Streaming SSR uses HTTP chunked transfer encoding to send the response in pieces. The server flushes HTML as soon as it is available — the document head, navigation, and static content arrive first, while data-dependent sections are sent later as their API responses complete.

Traditional SSR vs Streaming SSR Timeline Traditional Server: Fetch all data + render (800ms) Send Browser renders Streaming Shell TTFB: 50ms API 1 (200ms) Flush API 2 (500ms) Flush API 3 (800ms) — runs in parallel Flush Browser renders progressively Result: User sees content 750ms earlier

The Browser's Incremental Parser

Browsers are designed to handle incomplete HTML. When a chunk of HTML arrives, the browser's incremental parser processes it immediately — constructing DOM nodes, discovering CSS and font resources, and rendering visible content. This means the user sees the page skeleton and above-the-fold content while the server is still fetching data for below-the-fold sections.

This behavior directly improves Largest Contentful Paint (LCP). If the LCP element (a hero image or main heading) is in the initial chunk, LCP fires at the first flush time rather than at full page completion time.

Streaming Architecture Patterns

Shell-First Pattern

The most common streaming pattern sends the document shell — head, navigation, layout skeleton, and CSS — in the first chunk. This chunk is typically ready in under 50ms because it requires no data fetching. The browser begins rendering immediately: loading CSS, resolving fonts, and painting the page skeleton.

Subsequent chunks fill in the content sections as their data becomes available. Each section wraps in a container that the initial shell defines, so the page layout remains stable as content arrives — avoiding layout shift.

Out-of-Order Streaming

What if the footer data (50ms fetch) is ready before the main content (500ms fetch)? Out-of-order streaming solves this by sending each section as soon as its data is ready, regardless of document order. The technique sends placeholder elements in document order, then streams completed sections with inline script tags that move the content into the correct placeholder.

React 18's Suspense-based streaming SSR implements this natively. Each <Suspense> boundary defines an independently streamable section. When a suspended component's data resolves, React streams the HTML along with a small inline script that replaces the fallback content — all without any client-side JavaScript framework code having loaded yet.

Performance Insight: Out-of-order streaming ensures no section's TTFB is penalized by a sibling section's slow data fetch. Each section appears at its own data latency, not at the maximum data latency of the page.

Impact on Core Web Vitals

MetricTraditional SSRStreaming SSRImprovement
TTFBBlocked by slowest APIShell-only latency (50-100ms)Often 500ms+ reduction
FCP (First Contentful Paint)After full renderAfter first chunkProportional to data wait
LCPAfter full renderAfter hero chunkMajor if LCP is above fold
CLSZero (full page sent)Risk of shift if not plannedRequires skeleton layout
INPNo differenceNo difference (hydration controls this)Neutral

CLS Management in Streaming

Streaming introduces a layout stability risk: as new content chunks arrive, they could push existing content down. Prevention requires the initial shell to define the exact dimensions of each streamable section using CSS. Height reservations (explicit height or min-height) and aspect-ratio containers prevent content shift when the actual content replaces the skeleton.

Server Infrastructure Requirements

Streaming SSR requires specific infrastructure support that traditional SSR does not need:

CDN Compatibility

Edge CDNs interact with streaming in important ways. Some CDNs buffer the entire response before forwarding, negating the streaming benefit. Others (Cloudflare Workers, Fastly Compute) support edge streaming, forwarding chunks to the client as they arrive from the origin. Verify your CDN's behavior with streaming responses — the TTFB at the client should closely match the TTFB at the origin, not the total response time.

Error Handling in Streams

Once the HTTP response starts streaming (status 200 sent), the server cannot change the status code if a later data fetch fails. Error handling strategies for streaming include:

Measuring Streaming Performance

Traditional TTFB measurement reports the time to the first byte of the response — which for streaming SSR is the first chunk, not the complete page. Additional metrics are needed to fully understand streaming performance:

Track these with server-side timing headers (Server-Timing) that label each chunk's timing. Client-side, use the PerformanceObserver API to measure when resources discovered in each chunk begin loading.

When Not to Stream

Streaming adds complexity. It is not always the right choice:

Key Takeaways

Streaming SSR decouples TTFB from your slowest data dependency by sending HTML progressively as each section's data becomes available. The browser's incremental parser renders content immediately, improving FCP and LCP without waiting for the full page to generate. Use the shell-first pattern to deliver navigation and layout instantly, out-of-order streaming to avoid head-of-line blocking between sections, and proper height reservations to prevent layout shift. Verify that your CDN and reverse proxy support chunked transfer without buffering. Measure per-section latency, not just overall TTFB, to identify optimization targets within the streaming pipeline.