Web Workers: Offloading Heavy Computation from the Main Thread
The browser's main thread handles layout, painting, event handling, and JavaScript execution in a single-threaded event loop. When computation-heavy JavaScript runs on this thread, everything else stops — the page becomes unresponsive, animations stutter, and INP scores degrade. Web Workers solve this by executing JavaScript in background threads, completely separate from the main thread. This guide covers when to use workers, how to structure communication between threads, and the performance patterns that maximize their benefit.
How Web Workers Execute
A Web Worker runs in its own global context (DedicatedWorkerGlobalScope) with its own event loop. It has no access to the DOM, window, or document. Communication with the main thread happens exclusively through postMessage, which serializes data using the structured clone algorithm.
This isolation is both the strength and the constraint. The worker cannot directly manipulate the UI, which prevents thread-safety issues. But it means all data crossing the thread boundary must be serializable and the serialization cost itself can become a bottleneck for large payloads.
When to Use Web Workers
Not every computation benefits from a worker. The overhead of thread creation, message serialization, and context switching means workers are only worthwhile when the computation would otherwise block the main thread for a noticeable duration — typically more than 50 milliseconds (one long task).
Good Candidates for Workers
- Data transformation: Sorting, filtering, or aggregating large datasets (10,000+ rows). A spreadsheet component processing formulas, a charting library preparing render data from raw API responses, or CSV/JSON parsing of large files.
- Image processing: Applying filters, resizing, or format conversion. Canvas pixel manipulation at scale without blocking scrolling or animation frames.
- Cryptographic operations: Hashing, encryption, decryption, and key derivation. These are CPU-intensive by design and have no DOM dependency.
- Text processing: Search indexing, diff computation, markdown rendering, syntax highlighting of large files. Any operation that traverses a large string.
- Compression: Gzip/brotli compression or decompression of payloads before sending or after receiving from APIs.
Poor Candidates for Workers
- Simple DOM updates: Any operation that needs DOM access. Workers cannot read or modify the DOM at all.
- Short computations: Operations under 16ms. The serialization overhead may exceed the computation time, resulting in net slower performance.
- Frequent small messages: Patterns that send hundreds of tiny messages per second. Each
postMessagehas overhead; batch operations instead.
Worker Types and Their Trade-offs
| Worker Type | Scope | Use Case | Browser Support |
|---|---|---|---|
| Dedicated Worker | Single page | Page-specific computation | All modern browsers |
| Shared Worker | Multiple tabs/pages | Cross-tab state, connection pooling | Chrome, Firefox, Safari 16+ |
| Service Worker | Origin-wide | Network proxy, caching, offline | All modern browsers |
For performance offloading, Dedicated Workers are the primary tool. Service Workers serve a different purpose (network interception), and Shared Workers add complexity that is rarely justified unless you specifically need cross-tab coordination.
Optimizing Message Passing
The structured clone algorithm copies data between threads. For large payloads, this copy dominates the total worker operation time. Two techniques reduce this cost dramatically.
Transferable Objects
Instead of copying an ArrayBuffer, you can transfer ownership from one thread to another. The buffer becomes unusable in the sending thread (its byteLength goes to zero) but arrives in the receiving thread with zero copy cost. This is ideal for image data, audio buffers, and any binary payload.
Transfer vs. Copy: A 10MB ArrayBuffer takes approximately 10-15ms to clone via structured clone. Transferring the same buffer takes under 0.1ms — a 100x improvement. Always transfer ArrayBuffers when the sender no longer needs them.
SharedArrayBuffer
SharedArrayBuffer allows multiple threads to read and write the same memory region without copying. Combined with Atomics for synchronization, this enables true shared-memory parallelism. However, it requires cross-origin isolation headers (Cross-Origin-Opener-Policy: same-origin and Cross-Origin-Embedder-Policy: require-corp), which can break third-party embeds.
Architecture Patterns
Worker Pool Pattern
Creating a new worker for each task adds startup overhead (typically 5-20ms per worker creation). A worker pool maintains a fixed number of pre-created workers and distributes tasks among them using a queue. This amortizes creation cost across many operations.
Pool sizing depends on the hardware. navigator.hardwareConcurrency reports the number of logical processors. A pool of hardwareConcurrency - 1 workers leaves one core for the main thread. On devices with fewer cores, a smaller pool prevents contention.
Request-Response Pattern
The most common pattern wraps postMessage in a Promise-based API. Each request includes a unique ID. The worker includes the ID in its response. The main thread resolves the correct Promise when the response arrives. Libraries like Comlink abstract this entirely, exposing worker functions as if they were local async functions.
Streaming Pattern
For operations that produce results incrementally — like parsing a large file or running a progressive search — the worker sends partial results as they become available. The main thread renders updates without waiting for the full computation to complete. This keeps the UI responsive and gives users immediate feedback.
Measuring Worker Performance
Worker performance measurement requires tracking time across thread boundaries. The performance.now() timer runs in both the main thread and worker thread but starts from the same origin (page navigation time), making cross-thread timing straightforward.
Key metrics to track:
- Message serialization time: Time between calling
postMessageand the message arriving in the worker (measured using timestamps in the message payload). - Computation time: Time the worker spends processing, independent of communication overhead.
- Round-trip time: Total time from sending a request to receiving the response. This is what the user experiences.
- Main thread blocking: Any time the main thread spends preparing data for the worker or processing its response. Use Long Task detection to identify these.
Common Pitfalls
Several mistakes undermine the performance benefits of workers:
- Over-serialization: Sending entire application state to the worker when only a subset is needed. Send the minimum data required for the computation.
- Worker per task: Creating and destroying workers for each operation. Use persistent workers or pools instead.
- Ignoring transfer: Cloning large ArrayBuffers when they could be transferred. Always use the transfer list parameter for binary data.
- Blocking worker thread: Running synchronous XMLHttpRequest or compute loops that prevent the worker from processing other messages. Workers have their own event loop — keep it responsive.
- Excessive granularity: Splitting work into too many tiny messages. Batch operations to reduce message overhead. A single message with 1000 items processes faster than 1000 messages with 1 item each.
Browser Compatibility and Fallbacks
Dedicated Workers are supported in all modern browsers, including mobile Safari and Android WebView. However, module workers (type: "module") have narrower support — Safari only added support in version 15. If your build system can bundle worker scripts, use classic workers for maximum compatibility.
Always implement a synchronous fallback. If typeof Worker === 'undefined', run the computation on the main thread. This ensures the feature works everywhere, even in restricted environments that disable workers.
Key Takeaways
Web Workers eliminate main-thread blocking for CPU-intensive operations, directly improving INP and perceived responsiveness. Use them when computation exceeds 50ms — below that threshold, the message passing overhead may negate the benefit. Transfer ArrayBuffers instead of cloning them for binary data. Use worker pools to amortize creation costs. Wrap communication in Promise-based APIs for clean integration with async application code. Measure round-trip time, not just computation time, to understand the true performance impact.