Skip to content
HN On Hacker News ↗

Performance Improvements in .NET 11

▲ 349 points • 112 comments • by soheilpro • 3w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this text is a mix of AI and human-written content.

33 %

AI likelihood · overall

Mixed
72% human-written 28% AI-generated
SEGMENTS · HUMAN 4 of 8
SEGMENTS · AI 2 of 8
WORD COUNT 1,502
PEAK AI % 88% · §2
Analyzed
Sep 16
backend: pangram/v3.3
Segments scanned
8 windows
avg 188 words each
Distribution
72 / 28%
human / AI fraction
Verdict
Mixed
Pangram v3.3

Article text · 1,502 words · 8 segments analyzed

Human AI-generated
§1 Human · 2%

Before television shows like The Office and Parks and Recreation cemented the mockumentary in the minds of millions, there was Christopher Guest. He didn’t invent the genre, but he’s widely recognized as one of its most influential practitioners, and for my money, there’s none better. I’ve watched Waiting for Guffman and Best in Show more times than I can count. But the one that has stuck with me the most, the one I quote at the slightest provocation, is This Is Spinal Tap. If you’ve seen it you already know where this is going (and if you haven’t, you now have weekend plans). The film is a fictional documentary about an aging English rock band named Spinal Tap, whose members are everything we picture when we picture over-the-top rock stars. In one of its more memorable scenes, the guitarist (Nigel) gives the filmmaker (Marty) a tour of his most prized gear, in particular showing off an amplifier unlike any other: its dials don’t stop at ten. That leads to what might be the single most quoted exchange in the entire movie: Nigel: “You see, most blokes, you know, will be playing at ten. You’re on ten here, all the way up, all the way up, all the way up, you’re on ten on your guitar. Where can you go from there? Where?” Marty: “I don’t know.” Nigel: “Nowhere. Exactly. What we do is, if we need that extra push over the cliff, you know what we do?” Marty: “Put it up to eleven?” Nigel: “Eleven. Exactly. One louder.” This is .NET 11. It’s one louder, with another year’s worth of performance work having gone into making the runtime and libraries that much faster. Of course, the premise of Nigel’s special amplifier is ludicrous, as is exemplified in the subsequent few lines of dialog: Marty: “Why don’t you just make ten louder and make ten be the top number and make that a little louder?” Nigel: (pauses) “…these go to eleven.” In contrast, .NET 11 is actually one higher, one louder. The sections that follow are full of real improvements.

§2 AI · 88%

A bounds check removed, an allocation that no longer happens, a lock that isn’t taken, a loop that runs in fewer cycles than it did a year ago, a comparison folded to a constant here, a redundant check hoisted out of a loop there, a couple of instructions fused into one, a syscall sidestepped, an array copy handed off to SIMD, and on and on. That’s how real performance work goes, accumulating gain after gain, each compounding on the last, until the whole thing is measurably, provably louder.

§3 Human · 12%

And so, in this post, as I’ve done in past years with .NET 10, .NET 9, .NET 8, .NET 7, .NET 6, .NET 5, .NET Core 3.0, .NET Core 2.1, and .NET Core 2.0 before it, we’ll take an unhurried tour through hundreds of them. This is a long one. It’s meant to be. Grab your hot beverage of choice, settle in, and let’s turn it up. Benchmarking Setup As in previous years, the post is chock full of micro-benchmarks that demonstrate the individual improvements. Almost all of them use BenchmarkDotNet, and each is written to be self-contained so you can try it out yourself. Start by ensuring you have both .NET 10 and .NET 11 installed (most of the benchmarks compare the same code running on both versions) and create a new console project in a fresh benchmarks directory: dotnet new console -o benchmarks cd benchmarks Replace the contents of the generated benchmarks.csproj with the following, which multi-targets both versions so that BenchmarkDotNet can build for each: <Project Sdk="Microsoft.NET.Sdk"> <PropertyGroup> <OutputType>Exe</OutputType> <TargetFrameworks>net11.0;net10.0</TargetFrameworks> <LangVersion>preview</LangVersion> <ImplicitUsings>enable</ImplicitUsings> <Nullable>enable</Nullable> <AllowUnsafeBlocks>true</AllowUnsafeBlocks> <ServerGarbageCollection>true</ServerGarbageCollection> <SystemPackageVersion Condition="'$(TargetFramework)' == 'net10.0'">10.0.12</SystemPackageVersion> <SystemPackageVersion Condition="'$(TargetFramework)' == 'net11.0'">11.0.0-rc.1.26425.128</SystemPackageVersion> </PropertyGroup> <ItemGroup> <PackageReference Include="BenchmarkDotNet" Version="0.16.0-preview.1" /> <PackageReference Include="System.IO.Hashing" Version="$(SystemPackageVersion)" /> <PackageReference Include="System.Runtime.Caching" Version="$(SystemPackageVersion)" /> <PackageReference Include="System.Numerics.Tensors" Version="$(SystemPackageVersion)" /> </ItemGroup> </Project> For a given benchmark to test, copy its complete contents over everything in Program.cs and then run it. Each benchmark includes as a comment at the top the exact command to use. In most cases, it’s: dotnet run -c Release -f net10.0 --filter "*" --runtimes net10.0 net11.0 which builds in Release and runs the benchmark against both .NET 10 and .NET 11, emitting a side-by-side comparison.

§4 Mixed · 60%

The other common form, used when a benchmark is comparing two coding approaches on a single runtime (rather than the same code across two runtimes) is: dotnet run -c Release -f net11.0 --filter "*" The usual disclaimer applies: these are micro-benchmarks, many measuring operations so short that a blink would miss them. Your results will vary with your hardware, OS, runtime configuration, what else your machine happens to be doing at that exact moment, and whether Mercury is in retrograde.

§5 Human · 16%

Every line of managed code ultimately ends up at the just-in-time compiler, so let’s start there. JIT Of all the places to improve .NET’s performance, few have as broad an impact as the just-in-time (JIT) compiler. C#, F#, and Visual Basic are typically compiled first to intermediate language (IL), and the JIT ultimately turns that IL into the native instructions the CPU executes.

§6 Mixed · 53%

A JIT improvement can therefore benefit application and library code wherever the optimized pattern occurs, often with no source changes or recompilation of the application itself. Even removing a single instruction or proving one check unnecessary can add up when the code is on a very hot path. Deabstraction We as developers love our abstractions. They let us write clean, reusable, object-oriented code, but we don’t want to pay for every abstraction at run time. The runtime can often undo an abstraction when it proves the effects aren’t observable.

§7 Human · 25%

It can look at a virtual call and determine which concrete method it’ll invoke, look at a heap allocation and recognize that the object never leaves the current stack frame, or look at an interface cast and reuse a type fact already established earlier in the method. This process is called “deabstraction.” .NET has improved steadily in this area for years, and that continues in .NET 11. Every time you write interface in C#, you’re creating a contract, a promise that any type implementing that interface can be substituted for any other. That flexibility is enormously valuable because, for example, it’s what lets us write IEnumerable<T> and have it work equally well over arrays, lists, other collections, LINQ, custom iterators, and so on. But the CPU doesn’t know anything about these contracts; it just knows how to execute instructions. Turning “call whatever method this interface reference points to” into actual machine instructions requires special machinery. Consider this example: // dotnet run -c Release -f net11.0 --filter "*" using BenchmarkDotNet.Attributes; using BenchmarkDotNet.Running; using System.Runtime.CompilerServices; BenchmarkSwitcher.FromAssembly(typeof(Benchmarks).Assembly).Run(args); [DisassemblyDiagnoser, HideColumns("Job", "Error", "StdDev", "Median", "RatioSD")] public class Benchmarks { private Animal _animal = Environment.TickCount >= 0 ? new Dog() : new Cat(); [Benchmark] public int Speak() => _animal.Speak(); public abstract class Animal { public abstract int Speak(); } private sealed class Dog : Animal { [MethodImpl(MethodImplOptions.NoInlining)] public override int Speak() => 1; } private sealed class Cat : Animal { [MethodImpl(MethodImplOptions.NoInlining)] public override int Speak() => 2; } } At compile time, all else equal, the JIT doesn’t know whether _animal is a Dog or a Cat. It generates code that loads the instance’s “method table pointer” (its object type handle), sometimes called a “vtable pointer”, stored at the beginning of every .NET object, indexes into the method table at the known slot for Speak, and calls the function pointer found there: ; x64 mov rcx, [rcx+8] ; load _animal mov rax, [rcx] ; load method table mov rax, [rax+40] ; load vtable chunk call qword ptr [rax+20] For this one call to Speak, we pay three dependent memory dereferences and an indirect call because the processor doesn’t know for certain in advance where the call is going (it might guess, or “speculatively execute”, but it has to be prepared for the possibility it was wrong), and because the call target is indirect, the JIT can’t inline the callee.

§8 AI · 79%

Whatever Speak does, its code can’t be folded into the calling method. That’s a performance problem. Those indirections have overhead, but the bigger cost is the lost opportunity to inline. Inlining not only saves function call overhead, more importantly it opens the callee’s code up to the same optimizations that are operating on the caller, such as constant propagation, dead code elimination, bounds check elimination, further devirtualization, etc. That means a series of small virtual calls that each look innocent can, when devirtualized and inlined, collapse into a handful of instructions that would be unrecognizable and way cheaper when compared to the original source code. Without inlining, each callee is an opaque box; with it, the JIT can see through the layers. We as .NET developers constantly rely on the JIT’s sophisticated heuristics for inlining that weigh the IL size of the callee, the exact work the callee is performing, the call frequency of the method, the expected benefit from constant arguments, and dozens of other factors.