Streaming Server-Side Rendering: Progressive HTML Delivery
Traditional server-side rendering waits for all data to be fetched and all HTML to be generated before sending a single byte to the browser. If a page requires three API calls that take 200ms, 500ms, and 800ms, the TTFB is at least 800ms plus rendering time — the entire page is blocked by its slowest data dependency. Streaming SSR eliminates this bottleneck by sending HTML to the browser as each section becomes ready, starting with the shell and progressively filling in data-dependent sections.
How Streaming SSR Works
Streaming SSR uses HTTP chunked transfer encoding to send the response in pieces. The server flushes HTML as soon as it is available — the document head, navigation, and static content arrive first, while data-dependent sections are sent later as their API responses complete.
The Browser's Incremental Parser
Browsers are designed to handle incomplete HTML. When a chunk of HTML arrives, the browser's incremental parser processes it immediately — constructing DOM nodes, discovering CSS and font resources, and rendering visible content. This means the user sees the page skeleton and above-the-fold content while the server is still fetching data for below-the-fold sections.
This behavior directly improves Largest Contentful Paint (LCP). If the LCP element (a hero image or main heading) is in the initial chunk, LCP fires at the first flush time rather than at full page completion time.
Streaming Architecture Patterns
Shell-First Pattern
The most common streaming pattern sends the document shell — head, navigation, layout skeleton, and CSS — in the first chunk. This chunk is typically ready in under 50ms because it requires no data fetching. The browser begins rendering immediately: loading CSS, resolving fonts, and painting the page skeleton.
Subsequent chunks fill in the content sections as their data becomes available. Each section wraps in a container that the initial shell defines, so the page layout remains stable as content arrives — avoiding layout shift.
Out-of-Order Streaming
What if the footer data (50ms fetch) is ready before the main content (500ms fetch)? Out-of-order streaming solves this by sending each section as soon as its data is ready, regardless of document order. The technique sends placeholder elements in document order, then streams completed sections with inline script tags that move the content into the correct placeholder.
React 18's Suspense-based streaming SSR implements this natively. Each <Suspense> boundary defines an independently streamable section. When a suspended component's data resolves, React streams the HTML along with a small inline script that replaces the fallback content — all without any client-side JavaScript framework code having loaded yet.
Performance Insight: Out-of-order streaming ensures no section's TTFB is penalized by a sibling section's slow data fetch. Each section appears at its own data latency, not at the maximum data latency of the page.
Impact on Core Web Vitals
| Metric | Traditional SSR | Streaming SSR | Improvement |
|---|---|---|---|
| TTFB | Blocked by slowest API | Shell-only latency (50-100ms) | Often 500ms+ reduction |
| FCP (First Contentful Paint) | After full render | After first chunk | Proportional to data wait |
| LCP | After full render | After hero chunk | Major if LCP is above fold |
| CLS | Zero (full page sent) | Risk of shift if not planned | Requires skeleton layout |
| INP | No difference | No difference (hydration controls this) | Neutral |
CLS Management in Streaming
Streaming introduces a layout stability risk: as new content chunks arrive, they could push existing content down. Prevention requires the initial shell to define the exact dimensions of each streamable section using CSS. Height reservations (explicit height or min-height) and aspect-ratio containers prevent content shift when the actual content replaces the skeleton.
Server Infrastructure Requirements
Streaming SSR requires specific infrastructure support that traditional SSR does not need:
- Chunked transfer encoding: The HTTP connection must support
Transfer-Encoding: chunkedand the server must flush chunks without buffering the entire response. Many reverse proxies (Nginx, Cloudflare) buffer responses by default — disable response buffering for streaming routes. - Long-lived connections: Streaming keeps the HTTP connection open until all chunks are sent. Connection timeout settings must accommodate the slowest data dependency plus rendering time.
- Memory pressure: Each concurrent streaming response holds a render context in server memory. Under high concurrency, this can exhaust server memory faster than traditional SSR, which releases the context after the full response is sent.
CDN Compatibility
Edge CDNs interact with streaming in important ways. Some CDNs buffer the entire response before forwarding, negating the streaming benefit. Others (Cloudflare Workers, Fastly Compute) support edge streaming, forwarding chunks to the client as they arrive from the origin. Verify your CDN's behavior with streaming responses — the TTFB at the client should closely match the TTFB at the origin, not the total response time.
Error Handling in Streams
Once the HTTP response starts streaming (status 200 sent), the server cannot change the status code if a later data fetch fails. Error handling strategies for streaming include:
- Fallback content: Replace the failed section with a static fallback message rendered in the same layout slot. The user sees content degradation, not a broken page.
- Client-side recovery: Send a script tag in the error chunk that triggers a client-side retry for the failed section, fetching the data from an API endpoint after the page loads.
- Timeout boundaries: Set maximum wait times for each data dependency. If a fetch exceeds its timeout, stream the fallback immediately rather than keeping the connection open indefinitely.
Measuring Streaming Performance
Traditional TTFB measurement reports the time to the first byte of the response — which for streaming SSR is the first chunk, not the complete page. Additional metrics are needed to fully understand streaming performance:
- Time to first chunk: When the shell arrives. This should be under 100ms for most pages.
- Time to last chunk: When the final section is flushed. This is the equivalent of traditional SSR's TTFB.
- Per-section latency: How long each streamable section takes to resolve and flush. This identifies the slowest data dependency.
- Stream utilization: The ratio of time spent sending data vs. waiting for data. High utilization means the stream is consistently delivering content; low utilization means long pauses between chunks.
Track these with server-side timing headers (Server-Timing) that label each chunk's timing. Client-side, use the PerformanceObserver API to measure when resources discovered in each chunk begin loading.
When Not to Stream
Streaming adds complexity. It is not always the right choice:
- Simple pages with fast data: If all data fetches complete in under 100ms, streaming adds infrastructure complexity with minimal TTFB improvement.
- Pages that need complete data for rendering: An analytics dashboard that cannot show any section until all data is cross-referenced gains nothing from streaming individual sections.
- Static pages: Statically generated pages are pre-rendered — there is nothing to stream. They are already the fastest possible delivery mechanism.
Key Takeaways
Streaming SSR decouples TTFB from your slowest data dependency by sending HTML progressively as each section's data becomes available. The browser's incremental parser renders content immediately, improving FCP and LCP without waiting for the full page to generate. Use the shell-first pattern to deliver navigation and layout instantly, out-of-order streaming to avoid head-of-line blocking between sections, and proper height reservations to prevent layout shift. Verify that your CDN and reverse proxy support chunked transfer without buffering. Measure per-section latency, not just overall TTFB, to identify optimization targets within the streaming pipeline.