Writing Rust code that's faster than state-of-the-art libraries by asking agents to make the code faster
Pangram verdict · v3.3
We believe that this entire text is human-written.
AI likelihood · overall
HumanArticle text · 1,597 words · 1 segments analyzed
In January 2025, I had a fun hypothesis for a blog post: can LLMs write better code if you keep asking them to “write better code”? That was prior to the advent of robust agentic coding, but Claude Sonnet 3.5 was still able to iteratively improve on algorithmic Python code. The “better” instruction turned out to be underspecified: Sonnet abused that ambiguity to instead add a ton of useless features but the code was indeed faster. Even with the rise of agentic LLMs specifically RLHFed to handle solving common pass/fail coding problems, optimization is generally not a part of that suite.At the end of that blog post, I opined on a hypothetical future where LLMs could be able to write superfast Python code by instead writing Rust code and using PyO3 to bridge the languages to get Python’s ergonomics with Rust’s speed. An earlier draft of that post asserted that the same “write code better” instruction could instead be applied to the base Rust code and drastically improve its speed which would then propagate down to the Python code: however, back then I did not know enough about Rust and making such a claim would be too spicy without evidence.After months of testing and experimenting since the release of Opus 4.5 made agentic coding more viable, I can confidently confirm that modern agentic LLMs can indeed write Rust code that is significantly faster than current state-of-the-art approaches if given appropriate guardrails and constraints. Additionally, as LLMs have made drastic improvements in coding in each successive frontier model release since Opus 4.5, the optimizations have become even better, cumulatively resulting in anywhere from 2x-20x speedup depending on the domain.More importantly, this blog post is not a vaguepost and I am including both the prompts I used and the benchmark results. You’ve been warned.Iterative “Benchmaxxing”#At first, making software faster was a good quantitative way for me to test and compare these new agentic models. I used Rust as the target language primarily due to the Python integration and speed, but there are other aspects of the Rust language that make it particularly useful if a fast implementation is indeed discovered, such as memory safety and the ability to compile it to WebAssembly/WASM so it can run in a web browser without much effort. However, one important constraint I will follow that technically may not result in the fastest code is to forbid unsafe code whenever possible.My first test case was reimplementing machine learning algorithms in Rust, which would lead to a meaningful productivity increase for me as a data scientist if I had faster and scalable tooling. At the time, it was arrogant to assume that I could beat battle-tested algorithms that have been iterated on for over a decade and are already written in C so Rust’s low-level benefits are not as pronounced. The algorithm I wanted to optimize the most was UMAP, which is a valuable algorithm for dimensionality reduction I used in my work, but scales poorly to big data and is very slow, with alternatives such as cuML being time-consuming to set up. UMAP Rust crates such as umap-rs already exist where I could just fork them and prompt Claude Opus 4.5 to add Python/PyO3 support, but as an experiment and learning experience I wanted to have the agent write the algorithm from scratch with minimal Rust dependencies in order to make optimizations at as low of a level as possible.Rust has a comprehensive benchmarking tool with the criterion crate which all agents know how to leverage. criterion will run the benchmarks, track results across iterations to see if performance improved or regressed, and can calculate if this change is statistically significant or just noise.Typical criterion output, depicting a 3.5x speedup relative to the previous run of the benchmark.First, in the initial prompt for creating a Rust crate for UMAP, I asked Opus 4.5 to create benchmarks with different input data sizes since optimizations for small datasets may not work for large datasets and vice versa.Afterwards, create and run a benchmark suite which stores the results as a Markdown file. The benchmark suite MUST include inputs up to 100000x768 and be tested in both CPU and GPU modes.This approach created benchmarks using criterion and I manually reran the benchmark suite after prompting performance feature improvements such as using faer for faster linear algebra and using simsimd for faster SIMD operations. This quickly became cumbersome as I had to manually rerun each benchmark after each change to verify there are no speed regressions.A sidenote on prompting styleThe way I prompt agentic LLMs is unusual: I typically provide the agents very long prompts prewritten in a Markdown document with the additional use of ALL CAPS and **bolding** for emphasis. This is to ensure I capture all nuances through the use of prompt engineering, along with several other tricks as detailed in this blog post. Although some may argue prompt engineering is dead as the latest models have become smart enough to correctly handle ambiguity, I strongly disagree as LLMs have also become much better at following said nuances.Running a prompt from a Markdown document (right) by tagging the file (left), using Zed Agent.After having enough confidence that the agent will not accidentally rm -rf the repo, I experimented with letting the agent be autonomous, giving them permission to iterate until they achieve a speed increase, hopefully.**YOU MUST KEEP ITERATING OPTIMIZATIONS AND SOLVING ISSUES UNTIL THE BENCHMARK RESULTS STOP IMPROVING AND THE CRATE IS AS FAST AS IT CAN BE**. You have permission to keep iterating until you run out of ideas.It turned out “fast as it can be” is too ambiguous and Opus 4.5 was lazy so it tweaked a few hyperparameters without much of an actual speed increase and called it a day. What I needed was a clear target goal that can be pass/failed, so I refined the prompt:First, **without making any futher changes**, run the CPU Rust benchmarks to establish a True Performance Baseline. Then, optimize the crate code to make it such that ALL CPU benchmarks run **atleast 1.2x faster** than the True Performance Baseline; ideally as fast as possible. NEVER hack the benchmarks to accomplish this runtime reduction, only iterate on the library code. You may use ANY techniques to do so (e.g. import new crates) other than adding `unsafe` code. **REPEAT THIS PROCESS UNTIL BENCHMARK PERFORMANCE CONVERGES AND YOU ARE OUT OF OPTIMIZATION IDEAS.** You have permission to keep iterating. After each benchmark iteration, report the relative results to the True Performance Baseline to console. Prioritize making quick/high-impact wins iteratively and making changes accordingly. Do not overthink the necessary changes.This worked very well and not only did I get a 1.2x speed up on the benchmarks, but the agent continued after hitting the metric constraint and only stopped if a metric constraint was infeasible; in this instance, the agent hit 1.5x-2.0x speedups. The low-level Rust optimizations centered around a number of techniques including but not limited to: leveraging SIMD operations more aggressively, fusing functions, unrolling loops, creating intermediate caches, using Arc instead of borrowing wherever possible, and creating performance profiles based on input data (e.g. if the data is small, don’t use rayon data parallelism as the overhead erases gains).I chose “1.2x faster” as a sanity test: if the goal is too high, the agent may cheat to achieve it through risky/verbose rewrites. Smaller changes are better since the agent can more easily isolate the cause of a speedup/regression, hence the note about iteration. After new frontier LLMs released such as GPT-5.3 Codex and Opus 4.6, I repeated this prompt unchanged for every new LLM and each were able to achieve a cumulative 1.5x-2.0x speedup over the previous pass. Going all the way to GPT-6 Astra over many months, that’s around 7.5x-32x faster than the initial implementation baseline.This approach is hyperoptimizing for given benchmarks and therefore it could be considered benchmaxxing: a derogatory term for frontier LLMs that are only oriented to getting the high score on a benchmark which generalizes poorly to real-world use. However, if the benchmarks are sufficiently heterogeneous and truly representative of real-world use cases, then this is less of a concern. For this type of project, there are two ways to address concerns of benchmaxxing: 1) have the agent design diverse/unusual/adversarial input datasets instead of the generic “inputs up to 100000x768” and 2) enforce a quality gate on the output by comparing the output to a known correct implementation. In the case of machine learning algorithms, there is always a tradeoff between speed and quality, but in this instance it’s surprisingly easier to get the model fast, then make it correct. That is not how scientific engineering typically works, but it’s unlikely for a new implementation to match a known good implementation across many different benchmarks in an apples-to-apples comparison unless it’s truly correct.An agent-optimized gradient boosted decision tree implementation which beats xgboost significantly in speed, but also sometimes quality! (MSE: lower is better; other metrics, higher is better)Fortunately, there is a canonical implementation of UMAP with the Python package umap-learn and Python bindings to the Rust crate were already trivially added, so the new objective is simultaneous constraints: improve the code’s quality while capping the speed loss.Create a Python Jupyter Notebook comparing the performance of the Python bindings with `umap-learn`, including a check to confirm where the outputs and UMAP losses are as similar. Use diverse datasets with different matrix sizes than the benchmarks. If the outputs are not sufficiently similar, investigate methods to fix it **without causing more than a 5% speed regression**.Indeed, the