The state of SIMD in Rust in 2026 | Sergey "Shnatsel" Davidoff
Pangram verdict · v3.3
We believe that this entire text is human-written.
AI likelihood · overall
HumanArticle text · 1,493 words · 1 segments analyzed
A lot of progress was made since last year, and I made some of it!After last year's survey I started contributing to the SIMD library that seemed the most promising. One thing led to another, and now I'm a maintainer of Fearless SIMD.To avoid a conflict of interest, I invited authors of other libraries (std::simd, wide, pulp, macerator) to review and provide feedback on a draft of this article. However, I retained editorial control, and all mistakes are my own.This year's survey is more in-depth than my previous one. So buckle up, and let's take it... from the top!What’s SIMD? Why SIMD?Hardware that does arithmetic is cheap, so any CPU made this century has plenty of it. But you still only have one instruction decoding block and it is hard to get it to go fast, so the arithmetic hardware is vastly underutilized.To get around the instruction decoding bottleneck, you can feed the CPU a batch of numbers all at once for a single arithmetic operation like addition. Hence the name: “single instruction, multiple data,” or SIMD.Instead of adding two numbers together, you can add two batches or “vectors” of numbers and it takes about the same amount of time as doing just one addition.On recent x86 chips these batches can be up to 512 bits in size, so in theory you can get an 8x speedup for math on f64 or a 64x speedup on u8. In practice it can run both slower and faster.Instruction setsHistorically, SIMD instructions were added after the CPU architecture was already designed, so SIMD is an extension with its own marketing name on each architecture.ARM calls theirs “NEON”, and all 64-bit ARM CPUs have it.WebAssembly doesn’t have a marketing department, so they just call theirs “WebAssembly 128-bit packed SIMD extension”.64-bit x86 shipped with one called “SSE2” which has basic instructions for 128-bit vectors, but later they added a whole menagerie of extensions on top of that, with SSE 4.2 adding more operations, AVX and AVX2 adding 256-bit vectors and AVX-512 adding 512-bit vectors and even more operations.The word “later” in the above paragraph creates a problem.Does this CPU have that instruction?If you’re running a program on an x86_64 CPU, it’s not a given that the CPU has any particular SIMD extension. So by default the compiler isn’t allowed to use instructions beyond SSE2 because that won’t work on all x86_64 CPUs.There are two ways around this problem.If you work for a company that only ever runs their binaries on their own servers or on a public cloud, you can just assert that they’re all recent enough to at least have AVX2 that was introduced over 10 years ago, and have the program crash or misbehave if it ever runs on anything without AVX2:RUSTFLAGS='-C target-cpu=x86-64-v3' cargo build --releaseHowever, if you are distributing the binaries for other people to run, that’s not really an option.Instead you can do something called function multiversioning: compile the same function multiple times for different SIMD extensions, and when the program actually runs, check what features the CPU supports and select the appropriate version based on that.Fortunately, this problem only exists on x86.ARM made NEON mandatory on its 64-bit CPUs and hasn't really added useful SIMD extensions after that (more on that later).WebAssembly makes you compile two different binaries, one with SIMD and one without, and use JavaScript to check if the browser supports SIMD.How do I SIMD?There are three ways to leverage SIMD:Automatic vectorization: &[i32].sum()Portable SIMD abstractions: i32x4 + i32x4Platform-specific intrinsics - hang on, we're gonna need a bigger code block:#[cfg(all(any(target_arch = "x86", target_arch = "x86_64"),target_feature = "sse2"))] _mm_add_epi32(__m128i , __m128i) #[cfg(all(target_arch = "aarch64", target_feature = "neon"))] vaddq_u32(int32x4_t, int32x4_t)Let's look at what each one entails and what the state of each programming model is.Automatic vectorizationJust write plain Rust and let the compiler heuristics do the work!You can get it to work quite well, if you are careful to write code in a way that the compiler can reliably(ish) vectorize. This usually involves iterating over &[i32].as_chunks() instead of &[i32] and benchmarking or staring at the assembly to verify it worked. See Can You Trust a Compiler to Optimize Your Code? for details.This is the easiest option to use, requires no dependencies, and automatically supports all instruction sets the compiler supports, no matter how obscure.The downside is that this method is not very reliable. The larger and more complex your function is, the greater is the chance that the compiler will not be able to vectorize it. Performance can also swing wildly depending on the compiler version or due to changes to the surrounding code.Floating-point types also need special care.Floats are weird. Even something as trivial as summing an array of floats with reasonable precision gets surprisingly involved, see Taming Floating-Point Sums.Previously automatic vectorization didn't work with floating-point types because it would change the precision of the result (often for the better, but the compiler is not permitted to change any observable results).This changed in Rust 1.98 which stabilized algebraic ops such as algebraic_add() that let the compiler change the observable result, like a less dangerous -ffast-math. You still have to rewrite your code to use them for it to be eligible for vectorization in most cases.And you still need to get multiversioning somehow. So while we're at it...The 'multiversion' crateThe all-in-one SIMD crates discussed below also provide multiversioning, but let's take a look at multiversion real quick since it's most useful for automatic vectorization.It's very easy to use: you add the #[multiversion(targets = "simd")] annotation to your function and that's it.But that ease hides an undocumented pitfall: calling a function annotated with #[multiversion] has a little bit of overhead. It is very small - under a dozen instructions, but it shows up as significant overhead if the function you put it on is itself tiny.As a rule of thumb, if your function has a loop in it, add #[multiversion]; if it processes a handful of values add #[inline(always)], so long as there is #[multiversion] somewhere up the call chain.The other crates listed below don't have this pitfall and don't make you think about the sizes of functions, at the cost of more boilerplate.multiversion is the only crate that allows you to list the exact CPU extensions you require, as opposed to opting in to a predefined SIMD level. So if your code happens to benefit from some very recent instruction, you can opt in to it. But in my experience this hardly ever comes up for autovectorized code.For AVX-512 multiversion checks if it's present, not whether it's actually fast, which may hurt performance in practice (more on that below). You can work around that at the cost of boilerplate - you have to put this on every function:#[multiversion::multiversion(targets( "x86_64+cmpxchg16b+popcnt+sse3+sse4.1+sse4.2+ssse3", // x86_64-v2 "x86_64+avx+avx2+bmi1+bmi2+cmpxchg16b+f16c+fma+lzcnt+movbe+popcnt+sse3+sse4.1+sse4.2+ssse3+xsave", // x86_64-v3 "x86_64+fxsr,adx,avx512bitalg,avx512bw,avx512cd,avx512dq,avx512f,avx512ifma,avx512vbmi,avx512vbmi2,avx512vl,avx512vnni,avx512vpopcntdq,bmi1,bmi2,cmpxchg16b,fma,gfni,lzcnt,movbe,pclmulqdq,popcnt,vpclmulqdq,xsave,xsavec,xsaveopt,xsaves", // Ice Lake and later )]]Portable SIMD abstractionsThere are several production-ready ones. The desirable features are:Fixed-width vectors: write code in terms of f32x4, u8x16, etc (known size)Hardware-width vectors: use the largest vector size the hardware supports, without knowing it in advanceGeneric over element type: write code that works on both f32x4 and f64x2Generic over vector width: write code that works on all of f32x4, f32x8, f32x16The TL;DR table:std::simd (nightly)fearless simdwidepulpmaceratormultiversioning📦/🛠✅❌✅✅fixed-width vectors✅✅✅☑❌hardware-width vectors📦/🛠✅🛠✅✅generic over element type✅✅🛠🛠✅generic over vector width✅✅🛠✅✅safe access to intrinsics🛠✅☑✅🛠trigonometry📦/🛠🛠☑🛠🛠✅ Yes☑ Yes, with caveats📦 Yes, with a third-party crate🛠 Build it yourself❌ Absolutely notAnd the instruction set support:std::simdfearless simdwidepulpmaceratorSSE2✅✅✅☑🐌SSE4.x✅✅✅☑✅AVX2✅✅✅✅✅AVX-512✅✅✅✅✅NEON✅✅✅✅✅WASM✅✅✅✅✅All the rest✅🐌🐌🐌🐌*✅ Has optimized routines☑ Implemented but not used. Requires writing a custom dispatch to opt in.🐌 Reliant on autovectorization, often slow* macerator also supports LoongArch because the author was, and I quote, "bored".std::simdstd::simd is not a complete solution for SIMD. It's more of a set of building blocks that absolutely has to be in the standard library, while everything else is left up to the ecosystem crates.The largest drawback is that it's nightly-only, and still undergoes infrequent breaking API changes. So one day you update the compiler and your code stops compiling, and you have to go and fix it. But so long as you're OK with that, and only need fixed-width vectors and maybe multiversioning, it's pretty great!std::simd's raison d'être is that it sits directly on top of LLVM and can target any platform LLVM can target, including weird CPUs that only large banks use or that only the Chinese government uses. On the flip side, if LLVM doesn't have a perfectly matching operation inside it for std::simd to make use of, there is no plan B and no SIMD is actually used.This happens disturbingly often. Its sin() could not be more apt: shipping scalar implementations in a SIMD guise is the cardinal sin. And reduce_sum() is somehow the worst case for both performance and accuracy. So don't bother using any non-trivial functions on floats.The closest thing we have to proper trigonometry is the sleef crate, a partial port of SLEEF to std::simd that's only a little buggy. And that's the best trigonometry I have in this whole article!std::simd is uniquely flexible when it comes to multiversioning.