Pangram verdict · v3.3
We believe that this entire text is human-written.
AI likelihood · overall
HumanArticle text · 1,510 words · 1 segments analyzed
The PC market is one of the toughest areas for a CPU designer to compete in. Consumers in the PC segment expect high performance across a wide range of applications, and famously expect their devices to reach beyond a constrained, curated set of use cases. Microsoft’s own Surface RT prominently failed more than a decade ago because it could not run the programs that PC users expect to just work on Windows. Recent 64-bit Arm cores from Qualcomm and others are fast enough to satisfy the performance component of the PC equation, but software compatibility presents a tougher roadblock. PC software is traditionally built for x86-64. Getting developers to offer 64-bit Arm (aarch64) versions of their programs is a slow and gradual process. Some programs may never get aarch64 ports because they’re no longer under active development and were only distributed in binary form. Arm’s PC market chances therefore ride on Microsoft’s efforts to ensure x86-64 binaries can run seamlessly on aarch64 hosts.Windows 11’s newest binary translator, dubbed Prism, enables this by translating x86-64 instructions to aarch64 ones. Binary translation is challenging because x86-64 instructions sometimes don’t map to aarch64 ones in a straightforward manner. Translation also has to be fast to minimize program launch delays and avoid consuming excessive CPU time for doing the translation. All of this means binary translation comes with a performance penalty compared to running equivalent native code.Here, I’m checking out what that penalty looks like using Geekbench 7. Geekbench 7 comes in both aarch64 and x86-64 versions, providing an opportunity to compare performance with binary translation against native aarch64 performance. John Poole (Founder of Primate Labs, Creator of Geekbench) has kindly provided a pro key, which lets me profile individual workloads. Of course, one benchmark suite can’t reflect behavior across a wide range of applications, but I consider this a good starting point. As for hardware, I’m testing with the Snapdragon X2 Elite Extreme X2E-96-100 in an Asus Zenbook A16 laptop that was sampled by Asus as well as Arm instances available on Microsoft Azure.Like Geekbench 6, Geekbench 7 has a number of workloads that leverage vector extensions when available. The suite is slanted towards high IPC, compute-bound workloads. Running workloads through Intel’s Software Development Emulator (SDE) set to expose Haswell’s feature set shows more than half the workloads using AVX and 256-bit vector width. I’m running each workload with 200 iterations to ensure the vast majority of counted instructions come from the workload rather than Geekbench’s test harness.With 200 iterations, each workload executes roughly a few trillion instructions. Performance counter data from various aarch64 cores show roughly similar instruction counts when executing Geekbench 7’s native aarch64 binary, with occasional exceptions. Qualcomm’s cores curiously report higher retired instruction counts than Neoverse N1 and N2 in many tests, which suggests there may be some inaccuracy in hardware performance monitoring.Executing Geekbench 7’s x86-64 version under binary translation results in dramatically inflated instruction counts across the board, except in PDF Viewer. A Geekbench 7 workload will execute roughly twice as many aarch64 instructions as x86-64 ones when going through binary translation, as a rule of thumb.Microsoft hints that binary translation doesn’t work the same way across all aarch64 CPUs. Different CPUs support different instruction set extensions, some of which may let Prism more closely map x86-64 instructions to aarch64 ones. An aarch64 CPU can also hypothetically provide guarantees beyond what the ISA requires, like maintaining store ordering even though aarch64 lets one core observe stores from another out of program order.Prism is optimized and tuned specifically for Qualcomm Snapdragon processors. Some performance features within Prism require hardware features only available in the Snapdragon X series, but Prism is available for all supported Windows 11 on Arm devices with Windows 11 24H2.https://learn.microsoft.com/en-us/windows/arm/apps-on-arm-x86-emulationIf Prism is optimized specifically for Qualcomm’s cores, it doesn’t make a difference from the instruction count side. Ampere Altra’s Neoverse N1 is the oldest core I’ve tested on, and performance counters show very similar executed instruction counts.X-Axis is Number of Executed Instruction Count in TrillionsSimilar instruction counts can of course hide large differences. Performance is ultimately what matters in the end, and that’s affected both by the core architecture and the nature of the generated instructions.Geekbench 7’s score shows harsh penalties from binary translation. Every core, including Qualcomm’s, suffers several generations worth of performance loss. However, the cores in the Snapdragon X2 Elite have much higher baseline performance than Neoverse N1 and N2. The Snapdragon X2 Elite has a 6-core cluster of “performance” cores, and two 6-core cluster of “prime” cores. “Performance” cores are 6-wide and run at 3.6 GHz. Since they fill the same role as efficiency cores in other chips, I’ll call them E-Cores for simplicity. “Prime” cores are 9-wide and run at 5 GHz. They aim for maximum performance, so I’ll call them P-Cores.Qualcomm’s E-Cores manage to outperform Neoverse N1 even when taking a binary translation penalty. The same applies to the P-Cores against Neoverse N2. There’s really no arguing with much wider cores running at very high clock speeds, even with binary translation penalties in play.Qualcomm’s cores also take a slightly lower penalty from binary translation. It’s hard to tell whether this is down to Qualcomm-specific optimizations in Prism, or whether it’s down to Qualcomm’s cores having much higher throughput. Instruction count increases from binary translation remind me of playing with Claude’s C Compiler (CCC). CCC’s case involved creating a lot of extra instructions off the critical path, which an out-of-order CPU can often absorb. Recent high performance cores usually leave most of their core width unused, so extra core width can mitigate higher instruction counts to some extent. But I’m not sure that’s the case here, because Qualcomm’s 4-wide E-Core comes off with a lighter penalty than 5-wide Neoverse N2.Individual workloads suffer to varying degrees when run under binary translation. Navigation is a low IPC workload that’s heavily bound by branch prediction and backend memory latency, and takes a comparatively low penalty when run under binary translation. At the other end of the spectrum, Video Player is a well vectorized workload that takes advantage of AVX instructions. Score differences in that test are absolutely massive. Photo Editor and Photo Library are in a similar situation.Workload characteristics under binary translation stay mostly the same on Arm’s Neoverse N1 and N2.Performance counters show Qualcomm’s cores putting a dent in binary translation overhead by dipping into unused core width. However, the IPC increase is minor compared to the instruction count overhead, which explains the large binary translation penalty. Snapdragon X2 Elite Extreme’s P-Core manages a huge IPC increase in the Video Player workload, but gains elsewhere tend to be muted.Qualcomm’s E-Core is in a similar situation, though as a narrower core it’s able to make much better use of its core width. Text Processing, HDR, and Video Player basically have the core running up against its 4-wide width limitation when running under binary translation. However, the meager IPC increases seen on Qualcomm’s P-Core in Text Processing and HDR suggest that widening the E-Core might do little, as the workload might soon run into other limitations.Neoverse N2 starts off at lower IPC and curiously fails to gain much IPC when running x86-64 under binary translation. The core performs fine in an absolute sense when looking at IPC for native workloads, but it doesn’t look great next to Qualcomm’s latest efficiency optimized core. No workload on Neoverse N2 can average above 3 IPC, with or without binary translation in play.Neoverse N1 is the oldest core in this set, and unsurprisingly has the hardest time. N1 in Ampere Altra spearheaded Arm’s push to get a foothold in the server market, and succeeded because the core could deliver adequate performance while being more density optimized than its x86-64 contemporaries. But that was years ago. Qualcomm’s E-Core is also 4-wide, and leaves it in the dust.Comparatively good performance from Qualcomm’s E-Core serves as a reminder that core width means relatively little compared to other architecture characteristics. Qualcomm gives their modern E-Core a much larger out-of-order engine than Neoverse N1. Doing so lets the core keep more instructions in flight, reducing the impact of short duration delays on individual instructions.Accounting for dispatch-stage pipeline throughput shows a largely similar picture for workloads running natively and under binary translation, which isn’t a surprise because they should be doing the same high level work. On Qualcomm’s P-Core, the biggest difference is more utilized pipeline slots, along with some movement towards being less frontend bound.Qualcomm’s narrower E-Core uses a huge portion of its core width in many tests, especially the x86-64 versions running under binary translation. A narrower core is much easier to feed, and pipeline slots lost to frontend reasons are far fewer compared to on the wider P-Core.Detailed performance monitoring is harder on Azure because the platform doesn’t pass through top-down counters that account at the pipeline slot granularity. Roughly sketching things out with cycle-level accounting (STALL_FRONTEND and STALL_BACKEND) however shows that Neoverse N2 loses more pipeline slots to both backend and frontend reasons than Qualcomm’s E-Core. On one hand, feeding a