Engineering
Bauhaus modernist modular grid blocks representing zero-copy FlatBuffer memory partitions
• • 10 min read

Zero-Copy Binary Formats: Reading Structured Data Without Deserializing It

Why JSON.parse taxes the garbage collector, how zero-copy deserialization reads fields straight from raw bytes, and when binary protocols earn their cost.

JSON.parse is the quietest heavy operation in web development. A ten-megabyte response arrives, one line of code consumes it, everything feels fine—and underneath, the runtime has tokenized every character, built an abstract syntax tree, allocated an object for every brace-level, a string for every key, boxed numbers where they fit awkwardly, and handed the garbage collector a mountain of short-lived allocations to reclaim seconds later during your demo. Parse cost is not just CPU time; it is allocation pressure, and allocation pressure becomes jank on a schedule you do not control.

Zero-copy serialization formats attack the problem at the root: instead of transforming bytes into objects, they arrange the bytes so the raw buffer is the structure—fields read directly from their wire locations through offsets, with no intermediate representation ever allocated. This article makes the hidden costs visible, demystifies the offset-and-vtable mechanics with a hand-built example, and closes with the decision framework the topic deserves—because most applications should keep JSON, and knowing precisely when not to is the valuable part.

What JSON.parse actually does to your heap

Consider what must exist after parsing a modest document—a user record with nested preferences. The parser walks tokens and constructs a live object graph: one object per {}, one array per [], one string per key and per value (keys duplicated across records rarely intern), numbers materialized as doubles or integers depending on shape. Every one of those is an individual heap allocation with headers and pointers between them.

Three costs follow, none visible in a function profile:

  • Parse-time CPU scales with document complexity, not just size—deep nesting and many keys mean many small allocations rather than few large ones.
  • GC pressure arrives deferred. Young-generation collection sweeps the freshly allocated graph shortly after; large survivors get promoted and eventually cost major collections. The pause lands wherever the schedule puts it—often mid-interaction.
  • Memory footprint multiplies. The parsed graph routinely occupies several times the raw byte count once object headers, property storage, and string metadata are counted.
  • Keys repeat ruthlessly. Every object re-allocates its property-name strings; a thousand-record array with eight fields allocates eight thousand identical key strings unless the engine happens to intern them—and interning behavior varies by engine and context.

The pattern shows up in predictable places: configuration blobs parsed at startup, save files and telemetry replays measured in tens of megabytes, high-frequency messages where parse cost repeats per message. Wherever large payloads meet repeated parsing, the same tax accrues.

The zero-copy idea

A zero-copy format inverts construction: the writer arranges bytes so that reading requires no rearrangement. Scalars sit at fixed-width positions; variable-length data (strings, nested structures) lives elsewhere in the buffer, referenced by integer offsets; and a small indirection table—the vtable—maps each logical field to wherever its bytes happen to be.

Reading becomes pointer arithmetic. Fetching user.age means: consult the vtable entry for that field, add its offset to the table base, read four bytes. No object was created, no string decoded unless requested, nothing allocated. Access latency approaches raw memory reads, and the entire document occupies exactly its wire size in memory.

The vtable indirection sounds like overhead but pays double duty as the schema-evolution mechanism. Fields added in a newer writer produce extra vtable entries; old readers consult only the indices they know and skip the rest cleanly. Fields removed leave entries pointing nowhere, which readers interpret as “absent—use default.” Versions interoperate without negotiation, the property that made this family of formats popular across service boundaries.

The contenders in one table

The ecosystem offers several answers along this spectrum, differing mainly in how much decoding they still perform and what tooling surrounds them:

FormatDecode modelRandom access after loadSchema neededErgonomics
JSONFull parse → object graphOnly after full parseNone (by convention)Universal
Protocol BuffersFull decode → objectsNo—works on decoded copy.proto fileMature, ubiquitous
CBORStreaming decode → objectsNoOptionalJSON’s compact cousin
FlatBuffersNavigation over raw bytesYes—direct into bufferSchema + generated codeSolid, more verbose

No benchmark numbers appear in that table deliberately. Relative performance depends overwhelmingly on payload shape, document depth, string density, and the engine’s current GC mood—and published comparisons age badly. Measure on your payloads instead, with a harness that defeats dead-code elimination:

function bench(label, make, runs = 200) {
  make();                                          // warm caches and JIT
  performance.mark("start");
  let sink = 0;
  for (let i = 0; i < runs; i++) sink += make().byteLength ?? 1;
  const ms = (performance.now() - performance.measure("m", "start").end) / runs;
  console.log(`${label}: ${ms.toFixed(3)} ms/op (sink ${sink})`);
}

bench("json roundtrip", () =>
  new TextEncoder().encode(JSON.stringify(JSON.parse(rawJson))));

Pair timing with allocation awareness—long tasks observed via the Performance panel tell you whether parse work lands inside interactions, which matters more than raw throughput for UI responsiveness.

Code walkthrough: navigating a flat buffer by hand

Real FlatBuffers ship with code generators, alignment rules, and union support. To expose the mechanism, the snippets below implement a simplified cousin of the layout—not the production wire format, but the same three ideas: a root pointer, a vtable of field offsets, and length-prefixed strings.

First, a writer that lays out one record:

// write.js — encode {name, age} into our teaching wire format (little-endian).
// Layout: [u32 root][table: u16→vtable][vtable: u16,u16][string: u16+len][u32 age]
const LE = true;

export function writeUser(name, age) {
  const nb = new TextEncoder().encode(name);
  const table = 4, vtable = 6, str = vtable + 4;      // fixed prefix geometry
  const ageAt = str + 2 + nb.length;
  const buf = new ArrayBuffer(ageAt + 4);
  const dv = new DataView(buf), u8 = new Uint8Array(buf);

  dv.setUint32(0, table, LE);                         // root points at table
  dv.setUint16(table, vtable - table, LE);            // table → vtable (relative)
  dv.setUint16(vtable,      str - table, LE);         // field 0: name offset
  dv.setUint16(vtable + 2, ageAt - table, LE);        // field 1: age offset
  dv.setUint16(str, nb.length, LE);                   // length-prefixed string…
  u8.set(nb, str + 2);                                // …then its utf-8 bytes
  dv.setUint32(ageAt, age, LE);                       // scalar stored in place
  return buf;
}

Then the reader—which is the point of the exercise. Notice that reaching either field allocates nothing except the final result object:

// read.js — field access via vtable indirection; no intermediate tree exists.
const LE = true;

export function readUser(buf) {
  const dv = new DataView(buf);
  const root = dv.getUint32(0, LE);                   // navigate: root → table
  const vt = root + dv.getUint16(root, LE);           // table → vtable

  const fieldPos = (i) => {                           // i-th entry → absolute
    const rel = dv.getUint16(vt + i * 2, LE);         // position; 0 = absent
    return rel === 0 ? null : root + rel;
  };

  const namePos = fieldPos(0);
  let name = "";
  if (namePos !== null) {
    const len = dv.getUint16(namePos, LE);            // decode strings lazily…
    name = new TextDecoder().decode(
      new Uint8Array(buf, namePos + 2, len));         // …only when asked
  }

  const agePos = fieldPos(1);                         // scalars read in place
  return { name, age: agePos === null ? 0 : dv.getUint32(agePos, LE) };
}

Trace one access mentally: readUser fetches the root pointer (one read), hops to the table, hops to the vtable, reads two bytes naming the age’s position, then reads four bytes there. Four navigational reads, zero construction. A real FlatBuffer accessor performs the identical walk behind a generated method—the generated code is this paragraph, formalized.

Schema evolution falls out for free. Extend the writer with a third field appended after the existing ones and a third vtable entry; the old reader above keeps working untouched—it never consults index two. Remove the age field (entry set to zero) and the reader’s null branch supplies the default. Versioning without registry negotiation, from six lines of offset logic.

Choosing deliberately

Binary formats trade universal readability and frictionless tooling for speed and footprint. That trade resolves differently per scenario:

ScenarioSensible default
Config files, typical API responses (< ~1 MB)JSON — clarity wins; costs are invisible
Cross-team / cross-language servicesProtocol Buffers — mature contracts, wide toolchain
Large local artifacts: saves, replays, exportsCBOR or custom binary — compact and cheap to load
High-frequency in-browser pipelines (workers, shared memory)FlatBuffers-style flat data — zero allocation per access

That last row connects to architecture choices covered earlier: flat buffers pair naturally with transferable ArrayBuffers crossing worker boundaries and with shared-memory ring buffers, where avoiding per-message deserialization is often the entire reason shared memory exists. The mindset overlaps with other compact-data techniques too—Bloom filters trade exactness for bytes in the same spirit of paying attention to representation.

The rows also mix happily inside one system, and most serious adopters end up hybrid: JSON at trust boundaries where humans inspect traffic, binary inside hot loops where allocation shows up in profiles. Conversion points stay few and explicit—decode once at ingress, navigate flat structures thereafter—which contains the debugging cost while capturing nearly all of the performance.

Questions people often ask

Is JSON.parse actually slow enough to matter?

For sub-megabyte payloads consumed occasionally: no—readability and ubiquity win outright. The tax becomes real around large documents, startup-critical config, or per-message parsing at high frequency. Measure those specific paths before restructuring anything.

How do compression and binary formats interact?

Wire compression (gzip, Brotli) narrows the size gap considerably—JSON compresses excellently because of its repetitive keys. Binary formats compress too, but their advantage shifts from transfer size to parse avoidance: even identically-sized payloads differ sharply in decode cost and allocation pressure.

Do these formats require schema registries?

Protocol Buffers effectively wants centralized schema management—that is part of its discipline. FlatBuffers schemas live with generated code and evolve via vtable tolerance. Schema-less JSON avoids the ceremony and accumulates silent drift instead; pick which failure mode suits your team.

What is the debugging experience like?

Honest answer: worse. Hex dumps replace pretty-printed JSON, and while inspector tools exist for the major formats, nothing matches opening a response in devtools and reading it. Common mitigations: debug endpoints that emit JSON alongside, golden-file tests against known encodings, and reserving binary for internal boundaries where tooling is controlled.

The takeaway

JSON’s convenience is funded by an invisible factory: every parse constructs a full object graph, and the garbage collector invoices for it later. Zero-copy formats dissolve that factory by storing data in navigated form—offsets instead of pointers, vtables instead of constructors—so reading costs arithmetic and allocating nothing. The mechanics fit in forty lines, which is the real revelation: there is no magic in FlatBuffers, only disciplined byte arrangement. Adopt that discipline where measurement demands it—large local artifacts, hot in-browser pipelines—and keep JSON everywhere else, because readability is also a feature with real returns.

ADVERTISEMENT
SPREAD THE WORD

Found this guide helpful? Share it with your team & network.

ADVERTISEMENT
Author

Author

Verified

Engineer at Anirone, building Awesome Crate — free browser tools that keep your files on your device — and writing about how they work.