Engineering
Geometric abstraction of a segmented circular ring buffer and synchronizing pointer nodes in teal and violet
• • 12 min read

SharedArrayBuffer and Atomics: Lock-Free Ring Buffers Between Web Workers

Shared-memory concurrency in the browser: COOP/COEP isolation, SharedArrayBuffer, atomic operations, and a lock-free SPSC ring buffer walkthrough.

Our earlier deep dive on Web Workers and transferable objects ended at the edge of what message passing can do. Transferring an ArrayBuffer across threads costs nothing per crossing—but every crossing still is a crossing: a queued message, a scheduling hop, a fresh event-loop task between producer and consumer. For exports and parses that latency is invisible. For workloads where threads exchange data dozens or hundreds of times per frame—an audio callback pulling samples, a physics step feeding a renderer, a capture pipeline draining a sensor stream—the per-message overhead becomes the architecture’s ceiling.

The browser has a second model for exactly those workloads: genuine shared memory. SharedArrayBuffer gives two threads access to the same bytes; Atomics provides the ordering guarantees that make concurrent access defined rather than chaotic. Between them lives systems programming in JavaScript—rewarding when the budget demands it, overkill otherwise, and gated behind deployment requirements that surprise nearly everyone the first time. This article covers all three honestly, then builds the canonical first structure: a single-producer, single-consumer ring buffer.

The latency floor of postMessage

Where does message-passing overhead actually go? Not in copying—transferables solved that—but in sequencing. Each postMessage serializes a small envelope into an internal queue, wakes the receiving thread’s event loop, waits for its current task to finish, and delivers the payload as a new task. That pipeline is robust and fair, and each individual hop is fast by web standards. But it inherits the grain of the event loop: delivery happens between tasks, never inside them, and the receiving thread must poll its queue on some cadence to notice arrivals.

Budget arithmetic makes the ceiling visible. An audio callback running every 128 samples at 48 kHz fires roughly 375 times per second with hard real-time expectations; a 120 Hz input sampler wants fresh data eight milliseconds apart; a game loop sharing state with a worker twice per frame spends its coordination budget in four hops. None of these numbers is exotic, yet each sits uncomfortably close to the granularity of task-based messaging once queue depth and scheduler jitter join the picture. When two threads need to behave less like correspondents and more like two halves of one machine, the design moves from messages to memory.

Cross-origin isolation: the ticket you must buy first

SharedArrayBuffer was briefly available everywhere until 2018, when the Spectre vulnerability family demonstrated that speculative execution could, in principle, read memory a page should never touch. Browsers responded by gating high-resolution timers and shared memory behind a document-level security posture called cross-origin isolation. Two response headers switch it on:

Cross-Origin-Opener-Policy: same-origin
Cross-Origin-Embedder-Policy: require-corp

The first isolates your browsing context from any window opened by another origin; the second declares that every subresource you load must consent to being embedded (via its own CORP/COEP headers or explicit CORS approval). Set both, and the document gains access to SharedArrayBuffer, high-resolution timing, and related capabilities.

The honest operational cost sits with embedded third parties. Fonts, scripts, and images pulled from CDNs must send appropriate CORS headers or the isolation policy blocks them—and you control neither their servers nor their timelines. Analytics snippets and ad embeds are frequent casualties. Before committing, inventory every third-party resource your pages load and verify each one cooperates; several popular hosts already do, but “several” is not “all.”

Detection belongs in code, not assumptions, because the gate can vary by browser version and embedding context:

export function sharedMemoryAvailable() {
  try {
    new SharedArrayBuffer(8);              // throws if not isolated
    return typeof Atomics !== "undefined";
  } catch {
    return false;                          // missing COOP/COEP, or blocked context
  }
}

Shared memory semantics

A SharedArrayBuffer is a fixed-length block of bytes whose storage is reachable from multiple threads simultaneously. Typed-array views over it behave like ordinary views—one thread writes, another reads the same physical cells:

const sab = new SharedArrayBuffer(16);
const main = new Int32Array(sab);

const worker = new Worker("peer.js");
worker.postMessage(sab);                   // SAB passes by reference, not clone

That last line hides the semantic break with everything else in postMessage land: structured clone copies most objects, but a SharedArrayBuffer crosses by reference. Both sides now see one memory.

With shared access comes the precise definition of a data race: two threads accessing the same location without synchronization, at least one writing. In languages with formal memory models, races are undefined behavior—reads may observe torn values, stale values, or values reordered in ways source order never suggested. JavaScript’s memory model stops short of UB for plain accesses but guarantees almost nothing useful about them across threads: no ordering relative to other threads’ operations, no cross-thread visibility timing. Ordinary reads and writes are simply the wrong tool for coordination. Coordination is what Atomics exists for.

The Atomics toolkit

The namespace exposes small, uninterruptible operations on typed-array views backed by shared buffers:

OperationPurpose
Atomics.load(ta, i)Atomic read of cell i
Atomics.store(ta, i, v)Atomic write of v into cell i
Atomics.add/sub(ta, i, v)Fetch-and-modify counters
Atomics.compareExchange(ta, i, exp, v)Write v only if current value equals exp; returns prior value
Atomics.exchange(ta, i, v)Swap in v, return prior value
Atomics.wait(ta, i, v)Sleep until cell i leaves value v (workers only)
Atomics.notify(ta, i, n)Wake up to n threads waiting on cell i

Two properties matter more than the full list. First, atomic operations are sequentially consistent by default: all threads observe all atomic operations in one global order. Practically, that hands you a powerful rule of thumb—if a producer writes payload data first and then atomically stores a publish index, any consumer that atomically loads that index will observe every byte written before it. The atomic store acts as the publication barrier; the matching load as the acquisition point. That pairing is the entire foundation of the ring buffer below.

Second, wait/notify implement futex-style parking: a thread sleeps in kernel space instead of spinning, waking only when notified. The long-standing restriction still holds—the main thread cannot call Atomics.wait (browsers throw, since blocking the UI thread defeats its purpose). Waiting belongs in workers; the main thread coordinates via events or polls.

Code walkthrough: an SPSC ring buffer

Single-producer/single-consumer is the sweet spot where lock-free structures stay comprehensible: exactly one writer advances the head, exactly one reader advances the tail, and each index has exactly one owner. No compare-and-exchange loops needed—atomic loads and stores suffice.

Design choices worth naming before the code: indices are monotonic (they count up forever; only the slot position wraps), which removes the ambiguity of reusing wrapped-around index values; capacity is a power of two, so modulo becomes a bitwise mask; slots are fixed size, trading generality for predictable memory and branch-free indexing—ideal for audio frames, telemetry samples, or pixel rows.

// ring.js — lock-free SPSC queue over shared memory.
// Layout: [ head:i32 ][ tail:i32 ][ slot₀ … slotN ]   slots: SLOT bytes wide.
const HEAD = 0, TAIL = 1, HEADER = 8;

export class SpscRing {
  #ctrl;     // Int32Array over the two index cells
  #slots;    // Uint8Array over the data region

  constructor(sab, { capacity = 4096, slotSize = 256 } = {}) {
    if ((capacity & (capacity - 1)) !== 0)
      throw new Error("capacity must be a power of two");
    if (!sab) sab = new SharedArrayBuffer(HEADER + capacity * slotSize);
    this.capacity = capacity;
    this.slotSize = slotSize;
    this.#ctrl  = new Int32Array(sab, 0, 2);
    this.#slots = new Uint8Array(sab, HEADER);
    return this;                          // constructor returns the SAB-backed ring
  }

  get buffered() {                        // items currently in flight
    const c = this.#ctrl;
    return Atomics.load(c, HEAD) - Atomics.load(c, TAIL);
  }

  // Producer: write payload FIRST, then publish the new head.
  push(frame) {                           // frame: Uint8Array(slotSize)
    const c = this.#ctrl;
    const head = Atomics.load(c, HEAD);
    if (head - Atomics.load(c, TAIL) === this.capacity) return false;  // full
    this.#slots.set(frame, (head & (this.capacity - 1)) * this.slotSize);
    Atomics.store(c, HEAD, head + 1);     // publication barrier
    return true;
  }

  // Consumer: acquire head, copy out, then release the slot via tail.
  shift(out) {                            // out: Uint8Array(slotSize)
    const c = this.#ctrl;
    const tail = Atomics.load(c, TAIL);
    if (tail === Atomics.load(c, HEAD)) return false;                  // empty
    const start = (tail & (this.capacity - 1)) * this.slotSize;
    out.set(this.#slots.subarray(start, start + this.slotSize));
    Atomics.store(c, TAIL, tail + 1);     // release barrier
    return true;
  }
}

Correctness rests on the ordering argument from the previous section, applied twice. The producer’s payload write precedes its atomic store of HEAD; the consumer observes the incremented HEAD only through an atomic load, so it can never see the index before the bytes it announces. Symmetrically, the consumer finishes copying before storing TAIL, so the producer—which checks TAIL before overwriting a slot—can never reclaim a cell still being drained. Each side treats the other’s index as immutable-on-read, writable-by-one. No locks exist anywhere in the file because none are required by that discipline.

Both threads construct views over the same buffer; ownership of the indices, not the memory, defines who calls what:

// Main thread — producer side
const shared = new SharedArrayBuffer();
const ring = new SpscRing(shared);
worker.postMessage(shared);

setInterval(() => ring.push(nextSample()), 4);

// Worker thread — consumer side
self.onmessage = ({ data }) => {
  const ring = new SpscRing(data);
  const out = new Uint8Array(ring.slotSize);
  (function drain() {
    while (ring.shift(out)) process(out);
    setTimeout(drain, 2);                 // yield; park harder with wait/notify
  })();
};

Production refinements come later, and they rhyme across implementations: cache loaded indices in local variables inside loops (the other side changes them asynchronously anyway), batch multiple items per reserve to amortize atomic traffic, add wait/notify around an empty check to stop busy-polling, and consider BigInt64Array indices for sessions long enough to exhaust Int32 monotonic counts.

An honest scope check

Shared-memory concurrency buys throughput at a real price in complexity, and pretending otherwise produces the worst of both worlds:

  • Debugging loses its safety nets. Race conditions manifest as occasional wrong pixels or dropped frames, far from their cause. Traditional breakpoint debugging across two live threads sharing mutable bytes is limited; expect to reason about interleavings on paper.
  • Performance does not transfer. The win depends on cache behavior, core topology, and OS scheduling. Numbers measured on one machine mean little elsewhere—profile on target-class devices, qualitatively.
  • Most pipelines do not need it. An OffscreenCanvas transferred to a worker, fed by a few large transfers per second, renders games and editors smoothly with none of this machinery. Message passing remains the correct default; shared memory is the exception you earn.

Questions people often ask

Is SharedArrayBuffer supported everywhere now?

Desktop evergreen browsers support it behind the cross-origin-isolation gate described above. Mobile support is broad but historically patchier, and some embedded webviews ignore COOP/COEP entirely. Treat availability as a runtime capability, not a given.

Do I need any of this for JSON APIs or request/response work?

No. Structured cloning plus transferables handles request/response patterns comfortably, and their costs scale acceptably down to tens of crossings per second. The ring buffer earns its complexity above roughly hundreds of exchanges per second with strict freshness requirements.

Can Atomics operate on non-shared ArrayBuffers?

Yes—all atomic operations work on ordinary buffers too, though there they mainly serve as documentation of intent, since no other thread can touch the memory. Atomics.pause()-style backoff hints and future shared-nothing uses aside, the practical home for Atomics is shared memory.

How do garbage collection and SharedArrayBuffer interact?

Carefully but transparently. The GC understands shared buffers and will not free memory another realm still references; the buffer lives while either side holds a view. What the GC cannot fix is logic-level leaks—a producer pushing into a ring nobody drains fills real memory, regardless of collection.

The takeaway

Shared memory converts inter-thread communication from correspondence into coexistence: one buffer, two threads, and a pair of monotonically advancing indices whose atomic updates carry every ordering guarantee the design needs. The SPSC ring buffer is the pattern’s cleanest expression—single writer, single reader, power-of-two masking, publish-after-write, release-after-read—and the right first structure whenever measurement shows message passing standing between your threads and their deadline. Deploy it with open eyes about the COOP/COEP tax and debugging’s lost comforts, and keep message passing as the default architecture everywhere the budget forgives it.

ADVERTISEMENT
SPREAD THE WORD

Found this guide helpful? Share it with your team & network.

ADVERTISEMENT
Author

Author

Verified

Engineer at Anirone, building Awesome Crate — free browser tools that keep your files on your device — and writing about how they work.