Pangram verdict · v3.3
We believe this text is mainly human-written, with some AI content.
AI likelihood · overall
HumanArticle text · 1,457 words · 1 segments analyzed
October 2026Correspondence to [email protected]·Code· TL;DR We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models. Dust perturbs activations (node perturbation) independently at every token, so each token is a virtual population member and one forward pass evaluates them all in parallel. Dust approximates backprop closely at large population (i.e. substantially more compute) and in multiple settings even exceeds it. This hints that in a compute-rich regime we might be able to surpass backprop. Dust is orders of magnitude more efficient than weight-space ES. From 1M tokens up, Dust is on the order of $10^3$ to $10^4$ times more efficient than a transformer implementation of EGGROLL, a state-of-the-art ES method, based on our extrapolations. Zeroth-order methods are widely believed not to scale to large networks. Strikingly, we find larger models are more population-efficient, not less: a 243M-parameter model outperforms a $120\times$ smaller model at most population sizes. Dust’s gradient estimates align better with backprop’s as population grows, and stay well aligned at every scale we test, up to 1B tokens, which is encouraging for scaling. Contents TL;DR 1 Introduction 2 Method 2.1 Activation-Space Perturbation 2.2 Credit Assignment 2.3 Interference and Tuning 3 Pretraining Without a Backward Pass 3.1 Setup 3.2 Main Results 3.3 Dust Under Adam 4 Search in High-Dimensional Space 4.1 Overparameterization 4.2 Emergence of Backprop-Like Gradients 5 Conclusion 6 Related Work References Appendix 1 Introduction Deep learning has been built around backprop, the only credit assignment algorithm capable of training modern neural nets, including transformer-based language models. Backprop requires differentiability and produces first-order gradients, and deep learning’s architectures, optimizers, and hardware have co-evolved around this constraint. However, as the amount of compute available in the world increases, we might prefer more generic and brute-force learning algorithms based on search over inductive biases like differentiability, backprop, and approximations of higher-order gradients. The bitter lesson (Sutton, 2019Richard S. Sutton. The bitter lesson. http://www.incompleteideas.net/IncIdeas/BitterLesson.html, 2019. Blog post.) is that general methods that scale with compute eventually win, and AlphaGo Zero (Silver et al., 2017David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. Mastering the game of Go without human knowledge. Nature, 550 (7676): 354–359, 2017. doi: 10.1038/nature24270.) is the obvious example. Bootstrapping AlphaGo on human data helped the network learn faster initially, but with a lot of computation the purely self-play network overtook it. Similarly, differentiability and backprop might be good inductive biases in the low-compute regime, where they make learning efficient, but in the high-compute regime they limit the space of architectures that work. Even within an architecture, gradient-based methods fail to explore the loss landscape optimally (Liu et al., 2020Shengchao Liu, Dimitris Papailiopoulos, and Dimitris Achlioptas. Bad global minima exist and SGD can reach them. In Advances in Neural Information Processing Systems, volume 33, 2020.). This might also explain why current neural nets require massive amounts of data to generalize. A more flexible credit assignment algorithm based on search is likely an important step towards much better generalization. In this paper, we aim to replace backprop with a learning algorithm based much more on brute-force computation and much less on analytic structure. We call it Dust. Dust is a zeroth-order optimization algorithm that perturbs activations, rewards each perturbation by how much it lowers the loss, and averages the reward-weighted perturbations over a population to estimate the gradient. Traditional ES methods that perturb weights (Salimans et al., 2017Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever. Evolution strategies as a scalable alternative to reinforcement learning. arXiv preprint arXiv:1703.03864, 2017.), like EGGROLL (Sarkar et al., 2025Bidipta Sarkar, Mattie Fellows, Juan Agustin Duque, Alistair Letcher, Antonio León Villares, Anya Sims, Clarisse Wibault, Dmitry Samsonov, Dylan Cope, Jarek Liesen, Kang Li, Lukas Seier, Theo Wolf, Uljad Berdica, Valentin Mohl, Alexander David Goldie, Aaron Courville, Karin Sevegnani, Shimon Whiteson, and Jakob Nicolaus Foerster. Evolution strategies at the hyperscale. arXiv preprint arXiv:2511.16652, 2025.), scale with population, but scaling the population is costly because each member must be materialized and evaluated. We remove both costs with the concept of virtual population, where we avoid materializing every member by bypassing weight space entirely and instead perturb activations, as in node perturbation (Werfel et al., 2003Justin Werfel, Xiaohui Xie, and H. Sebastian Seung. Learning curves for stochastic gradient descent in linear feedforward networks. In Advances in Neural Information Processing Systems, volume 16, 2003.; Widrow and Lehr, 1990Bernard Widrow and Michael A. Lehr. 30 years of adaptive neural networks: Perceptron, Madaline, and backpropagation. Proceedings of the IEEE, 78 (9): 1415–1442, 1990. doi: 10.1109/5.58323.). We do so independently at every token, so each token is a member and one forward pass evaluates them all in parallel. Activations are a more interesting space to search over than weights. Mechanistic interpretability has shown that reasoning, whether verbalizable or not, lives in the activations (Gurnee et al., 2026Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, Ben Thompson, David Abrahams, Subhash Kantamneni, Emmanuel Ameisen, Joshua Batson, and Jack Lindsey. Verbalizable representations form a global workspace in language models. arXiv preprint arXiv:2607.15495, 2026.; Lindsey et al., 2025Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, Brian Chen, Adam Pearce, Nicholas L. Turner, Craig Citro, et al. On the biology of a large language model. Transformer Circuits Thread, 2025.), which means this approach could turn training into a search over latent reasoning (Vegesna and Dahal, 2025Akshay Vegesna and Samip Dahal. Decoupling search and learning in neural net training. arXiv preprint arXiv:2509.10973, 2025.). We then pair the activation-space perturbation with a very generic credit assignment rule that assigns different token-level rewards to different layer types in a transformer block. Those two biases, along with a few implementation details and efficiency measures, like avoiding interference between perturbed modules, are the whole algorithm. We make the following contributions. We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models. At large populations Dust exceeds backprop in multiple settings, which suggests that in a compute-rich regime we might be able to surpass backprop. Dust is orders of magnitude more efficient than weight-space ES. From 1M tokens up, Dust is on the order of $10^3$ to $10^4$ times more efficient than a transformer implementation of EGGROLL, based on our extrapolations. Contrary to conventional wisdom, larger models are often more population-efficient, not less, and can make use of larger populations. This gives a new view of overparameterization as a larger search space with potentially better geometry. Dust’s gradient estimates align better with backprop’s as the population grows, and the alignment holds up at every scale we test, up to 1B tokens, which is encouraging for scaling. The goal of this paper is to lay the foundations of a search-based credit assignment algorithm that is competitive with backprop on the hardest task we could think of: pretraining transformers. We do not attempt to make it compute-efficient enough to replace backprop today. We also do not train the new kinds of neural nets it makes accessible, like nets with an external program in the loop or transformers looped over many steps that backpropagation through time struggles to train. Both are left to future work. 2 Method Dust works as follows. We add Gaussian noise to the output of each linear layer, independently at every token, run a forward pass, and reward each token’s noise by the change in loss at that token. The reward-weighted noise, averaged over draws, is the estimated error at the layer’s output, and its outer product with the layer’s input is the weight gradient. Attention internals get a variant of it: they are credited through the estimated error at the attention output over current and future tokens, instead of the tokens’ loss directly. The core intuition is that while weight-space ES evaluates one population member per forward pass, we evaluate one per token, in parallel, and a member is materialized by adding noise to a hidden state, which is cheap. On a modern transformer a single forward pass therefore evaluates a population at least three orders of magnitude larger than weight-space ES. We describe each component in detail below. 2.1 Activation-Space Perturbation The bottleneck of evolution strategies is population size. Every member needs its own perturbed copy of the weights and its own forward pass. EGGROLL (Sarkar et al., 2025Bidipta Sarkar, Mattie Fellows, Juan Agustin Duque, Alistair Letcher, Antonio León Villares, Anya Sims, Clarisse Wibault, Dmitry Samsonov, Dylan Cope, Jarek Liesen, Kang Li, Lukas Seier, Theo Wolf, Uljad Berdica, Valentin Mohl, Alexander David Goldie, Aaron Courville, Karin Sevegnani, Shimon Whiteson, and Jakob Nicolaus Foerster. Evolution strategies at the