PWA Performance: Offline-First and Fast Loading
Progressive Web Applications promise native-like reliability from web technology. The core of that promise is performance — specifically, the ability to load instantly on repeat visits, function without a network connection, and remain responsive under degraded conditions. Achieving this requires deliberate architecture choices around caching, asset delivery, and lifecycle management. This guide examines the service worker caching strategies, app shell patterns, and background synchronization techniques that make PWAs genuinely fast.
The App Shell Architecture
The app shell model separates an application's infrastructure (navigation, layout, shared UI) from its dynamic content. The shell loads from the service worker cache on repeat visits — typically in under 100 milliseconds — while content loads from the network or a content cache. This separation produces the instant-load experience users associate with native applications.
Defining the Shell Boundary
The app shell includes everything that remains constant across navigations:
- HTML skeleton: The document structure, navigation bars, footer, and layout containers. This is typically a single HTML file that serves as the entry point for all routes.
- Core CSS: The stylesheet covering layout, typography, and component styles. Critical CSS should be inlined; non-critical CSS loaded asynchronously.
- Application JavaScript: The routing logic, state management, and core UI components. This should be the minimum JavaScript needed to render the shell and initiate content loading.
- Static assets: Icons, logos, placeholder images, and fonts that appear on every page.
Everything else — article content, user-specific data, dynamic lists, API responses — is content that loads separately and caches independently.
Design Rule: The app shell should render a meaningful, interactive skeleton within 200 milliseconds from cache. If it takes longer, the shell includes too much. Audit it ruthlessly — every asset in the shell is an asset that must be cached, versioned, and updated.
Shell Caching Strategy
The service worker caches the app shell during installation and serves it using a cache-first strategy. When the shell's assets change (a new deployment), the service worker detects the update, downloads the new shell in the background, and activates it on the next navigation.
This lifecycle — install, activate, fetch — determines how quickly users see updates. A common pitfall is configuring the service worker to wait for user action before activating a new version. While this prevents mid-session disruption, it means users may run stale code for days. The right balance depends on the application: content sites can activate immediately; transactional applications should prompt users and activate on navigation.
Service Worker Caching Strategies
Different resource types require different caching strategies. The four canonical strategies — cache-first, network-first, stale-while-revalidate, and network-only — each optimize for a different tradeoff between speed and freshness.
Cache-First for Static Assets
Cache-first serves the cached version immediately and never hits the network for that request. It is appropriate for versioned assets — files whose URLs change when their content changes. Content-hashed filenames (style.a3f9b2.css) are ideal for cache-first because the URL itself guarantees freshness: if the content changes, the URL changes, and the new URL fetches from the network.
For assets without content hashing (such as /fonts/main.woff2), pair cache-first with a maximum cache age. When the service worker installs a new version, it re-fetches these assets and updates the cache.
Stale-While-Revalidate for Content
Stale-while-revalidate returns the cached version immediately (fast) while simultaneously fetching an updated version from the network (fresh). The updated version replaces the cached version for the next request. This strategy delivers near-instant content loads with eventual consistency — users always see content, and it is usually current within one visit cycle.
This strategy works well for content that changes periodically but where showing a slightly stale version is acceptable: blog posts, product descriptions, documentation pages, and feed items. It is not appropriate for content where staleness causes confusion — stock prices, availability status, or real-time chat messages.
Network-First for Dynamic Data
Network-first tries the network and falls back to cache on failure. This preserves freshness when the network is available and provides offline functionality when it is not. The tradeoff is speed: every request waits for the network before falling back, which means slower responses than cache-first or stale-while-revalidate.
Add a network timeout (typically 3 to 5 seconds) to prevent indefinite waiting on slow connections. If the network does not respond within the timeout, serve the cached version. This prevents the worst-case scenario: a slow but not disconnected network that makes every request take 30 seconds.
Precaching and Runtime Caching
Precaching fills the cache during service worker installation. Runtime caching fills the cache as users navigate. Both are necessary — precaching ensures the critical path loads from cache on the first revisit, while runtime caching captures content the user has actually visited.
Precache Manifest Design
The precache manifest lists every URL that should be cached during installation. Keep it focused on the app shell and critical assets. Common mistakes include precaching too much (slowing installation and wasting bandwidth) or too little (missing assets that leave the shell incomplete).
A well-designed precache manifest for a content site includes:
- The HTML shell (index.html or the app's entry point)
- Critical CSS and JavaScript bundles
- Web fonts used in the shell
- The offline fallback page
- Key navigation icons and the app icon set
It does not include content pages, large images, video, or third-party scripts. These are candidates for runtime caching with appropriate strategies.
Cache Storage Limits
Browsers impose storage quotas that vary by platform and available disk space. Chrome typically allows an origin to use up to 60 percent of the total disk space, but this is shared across all storage APIs (Cache API, IndexedDB, localStorage). Mobile devices have tighter practical limits because disk space is scarcer.
Implement cache eviction policies for runtime caches: limit the number of entries, the total size, or the age of cached responses. A common policy caps runtime caches at 50 to 100 entries and evicts the least recently used entry when the cap is reached.
Background Sync for Offline Actions
Caching solves read operations — users can view previously loaded content offline. Background sync solves write operations — users can submit forms, post comments, or save changes offline, and the service worker replays those actions when connectivity returns.
Sync Registration
The Background Sync API lets the service worker register a sync event that fires when the browser detects a usable network connection. The application stores the pending action (typically in IndexedDB), registers the sync tag, and displays a confirmation to the user. When connectivity returns, the service worker processes the queue.
Key design considerations for background sync:
- Idempotency: The sync handler may fire multiple times for the same event. Every action in the queue must be safe to replay — use unique identifiers and server-side deduplication.
- Conflict resolution: If the user modifies the same resource online and offline, the sync handler must resolve the conflict. Last-write-wins is simplest but loses data; three-way merge is correct but complex. Choose based on data sensitivity.
- User feedback: Show clear indicators for pending, syncing, and synced states. Users need to know whether their offline action has been committed to the server.
- Retry limits: The browser retries failed sync events with exponential backoff, but eventually gives up. Monitor sync failures and surface them to users before the retry limit is reached.
Performance Measurement for PWAs
PWA performance measurement adds complexity because the same page may load from the network (first visit), the service worker cache (repeat visit), or a mix of both (shell cached, content from network). Each path has different performance characteristics and requires separate measurement.
Distinguishing Cache and Network Loads
Tag RUM events with the load source: was the navigation served from the service worker cache, the HTTP cache, or the network? This segmentation reveals whether caching improvements are actually reaching users and how much of your traffic benefits from the service worker.
The Navigation Timing API provides the workerStart timestamp, which is nonzero when a service worker handled the navigation. Use this to segment performance data into service-worker-controlled and uncontrolled loads. The PerformanceObserver API provides resource-level timing that can distinguish cached from network-fetched assets.
Measuring Install and Activation
Service worker installation and activation are invisible to users but critical to PWA performance. Track these lifecycle events to understand how long the service worker takes to become operational and how many users complete installation:
| Event | What to Measure | Healthy Target |
|---|---|---|
| Install duration | Time to download and cache precache manifest | < 5 seconds |
| Activation duration | Time to clean old caches and take control | < 500ms |
| Claim time | Time until SW controls the page | < 100ms after activation |
| Install failure rate | Percentage of failed installations | < 1% |
| Update frequency | How often new SW versions deploy | Matches app deploy cadence |
Push Notifications and Performance Impact
Push notifications wake the service worker to process incoming messages. Each wake-up consumes battery and may trigger network activity. Poorly implemented push handling — large payload processing, unnecessary cache updates, or heavy UI rendering in the notification — degrades the device experience and can lead users to disable notifications entirely.
Performance guidelines for push notification handling:
- Keep notification payloads small — under 4 KB. Use the payload as a pointer to content, not the content itself.
- Show the notification quickly. Browsers enforce a time limit on notification display — if the service worker takes too long, the browser shows a generic notification.
- Defer heavy work. If the notification triggers data synchronization or cache updates, queue that work for when the user next opens the application.
- Batch notification processing. If multiple notifications arrive while the app is closed, consolidate them into a summary rather than firing individual handlers for each.
Installability and Performance
The web app manifest controls how the PWA appears when installed — its name, icon, display mode, and theme color. Installability criteria require a service worker, a manifest, and HTTPS. But the manifest also affects performance through launch behavior.
When installed in standalone display mode, the PWA launches without browser chrome, which reduces the visual complexity of the initial frame. The start_url determines which page loads on launch — set it to a lightweight entry point, not a content-heavy page. If the home feed requires API calls, the launch experience is a blank screen until data arrives. Instead, start with the cached app shell and progressively load content.
The theme_color and background_color in the manifest affect the splash screen that appears between launch and first paint. Set the background color to match the app shell's background — a white splash screen followed by a dark app shell creates a jarring flash. This is a perceived performance detail, but perceived performance is real performance to users.
Offline Fallback Patterns
A robust offline experience requires planned fallbacks for every resource type:
- Navigation requests: Serve a custom offline page that explains the situation, lists recently cached content the user can still access, and provides a retry mechanism.
- Image requests: Serve a placeholder SVG or a low-resolution cached version. Avoid showing broken image icons — they signal abandonment, not intention.
- API requests: Return cached data with a "last updated" timestamp so users know the data's age. For write operations, queue them for background sync.
- Font requests: Fall back to system fonts. CSS
font-display: swapensures text remains visible even when the custom font is unavailable.
Common PWA Performance Mistakes
Several patterns appear repeatedly in underperforming PWAs:
- Caching too aggressively: Caching every network request fills storage quickly and serves stale content from URLs that should always be fresh (auth endpoints, real-time data).
- Not versioning the service worker: Without cache-busting the service worker registration, browsers may serve stale service worker code from the HTTP cache, preventing updates from reaching users.
- Ignoring cache invalidation: Adding to the cache without removing old entries leads to unbounded storage growth. Implement max entries, max age, or explicit invalidation on deployment.
- Blocking install on large precache: A precache manifest with 50 MB of assets makes installation take minutes on slow connections. Users abandon the page before installation completes.
- No measurement segmentation: Treating all Core Web Vitals data as a single population hides the performance difference between cached and uncached visits. Segment to understand actual impact.
Key Takeaways
PWA performance is fundamentally about caching architecture. The app shell model provides instant repeat loads. Service worker strategies — cache-first for static assets, stale-while-revalidate for content, network-first for dynamic data — match each resource type to the right speed-freshness tradeoff. Background sync extends reliability to write operations. And measurement must segment cached from uncached visits to reveal actual user experience.
The investment pays off in measurable ways: repeat visit load times under 200 milliseconds, offline functionality that prevents complete service loss, and perceived reliability that keeps users engaged even on unreliable networks.