Skip to content
HN On Hacker News ↗

Cache-Conscious Data Layout in Rust: Field Zoning, False Sharing, and the 128-Byte Rule

▲ 39 points • 24 comments • by eigenBasis • 3mo ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this document is fully AI-generated

99 %

AI likelihood · overall

AI
0% human-written 100% AI-generated
SEGMENTS · HUMAN 0 of 5
SEGMENTS · AI 5 of 5
WORD COUNT 1,622
PEAK AI % 99% · §1
Analyzed
Jul 10
backend: pangram/v3.3
Segments scanned
5 windows
avg 324 words each
Distribution
0 / 100%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 1,622 words · 5 segments analyzed

Human AI-generated
§1 AI · 99%

Cache-Conscious Data Layout: Field Zoning, False Sharing, and the 128-Byte Rule Part 1 of Low-Level Systems Design in Rust - a series on writing high-throughput, low-latency systems code, using a single-producer / single-consumer (SPSC) ring buffer as the running example. Part 0 - Architectural Decomposition made the highest-leverage decision (remove contention structurally, so every writer gets its own ring) and holds the full series index and guiding principles. This post starts the micro-level work: given one such SPSC ring, how should it be laid out in memory? The patterns aren't specific to ring buffers - they apply to any structure more than one core touches.

The thing nobody tells you about multi-threaded structs When you write a single-threaded data structure, the only layout question that matters is "does it fit in cache." When you write a multi-threaded one, a second, relevant and a bit subtle question appears:

For each field, which core touches it, and how often?

Getting this wrong will make you pay in a way that no profiler flags as a single hot line. Two cores writing to different fields that happen to share a 64-byte cache line will quietly serialize each other through the hardware coherence protocol - a pathology called false sharing. The code looks lock-free. The benchmark says otherwise. This post is about designing the layout deliberately. It comes in two parts:

Field zoning - grouping fields by (write owner, frequency) so that one core's hot writes don't evict another core's hot working set. Alignment and padding - the mechanics that make zoning real: why #[repr(C)] is load-bearing, why the magic number is often 128 and not 64, and why adding prefetch hints can make things worse.

The running example is a single-producer / single-consumer (SPSC) ring buffer. One core (the producer) appends, another core (the consumer) drains. They coordinate through two monotonic cursors - tail (where the producer writes next) and head (where the consumer reads next). Though the following discussion and the design principles are general enough, you can refer to this ring buffer implementation as a reference point that honors these guiding principles.

§2 AI · 99%

Part 1 - Field zoning: design based on "who touches what" The cache line - commonly 64 bytes on x86-64 and many AArch64 cores - is the unit of currency for inter-core communication. Coherence traffic is accounted per line, not per byte or per field. So the first design move for any shared structure is to sort its fields into zones by access pattern: ZoneFieldsWrite ownerCross-core access Producer-hottail, cached_headproducerconsumer samples tail Consumer-hothead, cached_tailconsumerproducer samples head Coldclosed, config, metricslifecycle/observabilityrare, not hot-path polling

"Producer-hot" does not mean "producer-only." The producer owns writes to tail, but the consumer still reads tail when its cached view runs dry. The important distinction is write ownership: the producer is the core that keeps making the line exclusive, so the fields near that write must be chosen deliberately. First, the vocabulary, since the rest of the post leans on it. The hot path is the code that runs on (almost) every operation - here, the send and receive loops that each core executes millions of times a second. Hot fields are the ones those loops touch, like tail, head, and the two cursor caches. The cold path, by contrast, is everything that runs rarely - construction, shutdown, an occasional metrics read - and cold fields are the ones only the cold path touches, like config and metrics. "Hot" and "cold" are about frequency of access, and they're what the zones in the table above are sorted by. The rule is then simple to state: place each zone on its own set of cache lines. A hot field written by one core must not share a line with hot fields another core reads or writes frequently - a producer store to tail should not invalidate the line holding the consumer's head or cached_tail, and vice versa. Cold fields, since nobody contends over them on the hot path, can stay packed together at the back. Here is a ring buffer that encodes those zones directly in its declaration.

§3 AI · 99%

The type parameter A is just a pluggable allocator for the backing buffer, ignore it if you like - what matters is the order and grouping of the fields: #[repr(C)] pub struct Ring<T, A: BufferAllocator = HeapAllocator> { // PRODUCER HOT tail: CacheAligned<AtomicU64>, cached_head: CacheAligned<UnsafeCell<u64>>,

// CONSUMER HOT head: CacheAligned<AtomicU64>, cached_tail: CacheAligned<UnsafeCell<u64>>,

// COLD closed: AtomicBool, metrics: Metrics, config: Config, buffer: UnsafeCell<A::Buffer<T>>, } A few things are worth reading carefully here, because each one is a deliberate choice rather than an accident of style. The two cached_* fields live in the hot zones, not the cold one. A producer that wants to know "is there space?" must compare tail against head. But head is written by the other core, so reading it directly is a potential cross-core coherence miss - tens to hundreds of cycles. Instead, the producer keeps cached_head: a single-writer, producer-owned snapshot of where the consumer was last time we looked. Because a stale view of head can only ever under-report available space (never over-report it), the cache is always safe to trust on the fast path, and the expensive Acquire read of the real head fires only when the cache says we're out of room. That snapshot is touched on every send, so it belongs in the producer-hot zone - and it is deliberately a plain UnsafeCell<u64>, not an atomic, precisely because only one core ever writes it. The consumer has the mirror image: cached_tail lets it avoid reading the producer-owned tail until the local snapshot is exhausted. Cold fields share lines on purpose. closed, metrics, and config are packed together with no padding between them. That's not laziness - it's the point. The goal of zoning is not to align everything; it's to align only the things that ping between cores. Padding a cold field just wastes a cache line you could have spent keeping hot data resident.

§4 AI · 99%

If a lifecycle flag is polled on every send or receive, it is no longer cold; promote it into its own hot or lifecycle zone instead of hiding it behind the word "shutdown." Why #[repr(C)] is doing real work Notice the #[repr(C)] on the struct. It is not decorative, and removing it would silently undermine the entire layout strategy. By default, Rust uses repr(Rust), and the Rust Reference is explicit that repr(Rust) guarantees only field alignment, that fields don't overlap, and that the aggregate is sufficiently aligned. It explicitly does not guarantee that fields are laid out in declaration order - the compiler is free to reorder them (typically to minimize padding). For a normal struct that's a feature. For a struct whose correctness-of-performance depends on tail and head sitting in different zones, a reorder could collapse your carefully separated zones back onto shared lines. #[repr(C)] pins fields in declaration order, so the zone layout you wrote is the zone layout you get. One subtlety that's easy to miss: repr(C) on the outer struct does not recursively freeze the layout of nested repr(Rust) fields. It fixes the order of Ring's own fields, but the internal ordering of, say, Metrics or Config is still repr(Rust) unless those types carry their own repr(C). If a nested struct has its own hot/cold split that matters, give it repr(C) too. Here it doesn't matter, because Metrics and Config are entirely cold.

Part 2 - Alignment and false sharing: making zoning real Zoning is the intent. Alignment is the mechanism. Declaring that tail and head belong to different zones accomplishes nothing unless the bytes actually land on different cache lines. That's the job of CacheAligned<T>. The one-line helper that pays for itself /// Aligns a value to a 128-byte boundary so that cross-core fields don't /// share a cache line. The 128 (vs. 64) is a target policy: it also leaves /// room for adjacent-line/spatial-prefetch effects on CPUs where those matter.

§5 AI · 99%

#[repr(C, align(128))] struct CacheAligned<T> { value: T, }

impl<T> CacheAligned<T> { const fn new(value: T) -> Self { Self { value } } }

impl<T> std::ops::Deref for CacheAligned<T> { type Target = T; fn deref(&self) -> &Self::Target { &self.value } } That's the whole thing - a #[repr(C, align(128))] newtype with a Deref so it's transparent at the call site (self.tail.load(...) just works). No library dependency required. Because Ring is repr(C), the compiler lays fields out in declaration order and inserts whatever padding is needed to satisfy each field's alignment. Because a Rust type's size is a multiple of its alignment, each small CacheAligned<T> also consumes a full 128-byte slot. No two hot fields can share a line, and - critically - neither can they share a line with the cold fields that follow. Why 128 bytes, when the cache line is 64? This is the part that trips people up. If the cache line is 64 bytes, why pad to 128? Because strict line-level separation can still be too optimistic. Some CPUs have adjacent-line or spatial prefetch behavior: when you touch one 64-byte line, the core may speculatively pull in its neighbor. If producer-hot data sits on the line immediately adjacent to consumer-hot data, the prefetcher can drag the two into each other's caches even though, strictly speaking, they never shared a line. The result is prefetcher-induced false sharing: the lines are technically separate, the contention is real, and line-granularity tools may not point directly at it. You see it as throughput that's lower than the algorithm predicts. For small cursor fields, aligning each hot field to 128 bytes makes it start on an even 64-byte-line boundary, so an adjacent-line prefetch usually pulls in padding rather than another core's hot data. Treat the exact width as a target-specific policy, not a universal constant.