Introducing Beam: Reflection's 501B open-weight model — Reflection
Pangram verdict · v3.3
We believe this text is mainly human-written, with some AI content.
AI likelihood · overall
HumanArticle text · 1,376 words · 1 segments analyzed
We are introducing Beam, Reflection’s first open-weight model. Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.Beam’s capabilities come from major investments in both pretraining and reinforcement learning (RL). We pretrained the model on 23.8 trillion diverse, curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming available similar-sized open base models. In parallel, we developed the algorithms, training environments, and infrastructure needed to sustain high-compute RL at exceptional scale. Our high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks of training.Together, these efforts produced competitive open-weight performance with frontier inference compute efficiency.Beam is undergoing final red-teaming and evaluations. You can sign up here for early access to the model. We will release the weights, technical report, model card, and developer artifacts later this month.Model CapabilityWe trained Beam with a particular focus on coding and agentic performance. Beam advances the Western open-weight frontier and is competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks. Where frontier open models like Kimi K3 remain ahead on raw capability, Beam's advantage is efficiency at inference time.The below figure shows Beam's performance across a range of coding, agentic, reasoning, and STEM benchmarks. NR denotes scores that have not been reported. Agentic Coding/TerminalRow labelBeamInklingNemotron 3 UltraGLM 5.2GLM 5.3Kimi K3Qwen 3.8 MaxDeepSeek V4.1 FlashDeepSWE v1.144.4NRNR44.061.068.051.074.2SWE Bench Pro v2-Hard77.256.9NRNR84.388.2NRNRSWE Bench Pro v165.554.346.462.1NRNR67.7NRTerminal Bench v2.1 80.163.856.481.088.288.386.690.6SWE Atlas Codebase QnA34.6NRNRNR61.068.0NRNRSWEBench Multilingual78.0NR67.7NRNRNRNRNRSWEBench Verified80.977.670.7NRNRNRNRNRBeam pairs coding and agentic capabilities with highly efficient reasoning. On advanced reasoning benchmarks, it achieves scores comparable to GLM-5.2 while using 3–4× less inference compute. Efficiency gains are even more pronounced when comparing to models in the 2T+ parameter family like Qwen 3.8-Max, which require significantly more inference compute per token.These results translate into more intelligence per token, delivering strong model capabilities at lower cost, making Beam a powerful workhorse model for enterprise coding and agentic workloads.Figure 2: Beam demonstrates frontier-level inference efficiency, both when measured in terms of FLOPS and token count across DeepSWE, Humanity’s Last Exam (HLE), and Terminal Bench 2.1. We used data from Artificial Analysis and DataCurve, estimating generation forward-pass compute as FLOPs ≈ 2 × active parameter count × mean generated tokens per attempt, counting each multiply-add as two operations. Generated tokens include both reasoning and the final answer. For mixture-of-experts models, we used the parameters activated per token rather than the total model size. These estimates exclude prompt prefill, context-dependent attention operations, and serving overhead, so they represent an approximate compute comparison rather than measured inference cost. We use Artificial Analysis and DataCurve as sources for other model’s evals.High-Compute Reinforcement LearningWe made high-compute reinforcement learning a central scaling axis for Beam, investing in RL science, data, and infrastructure to turn more compute into stronger capabilities. Scaling RL enables more extensive exploration of problem-solving strategies, while longer rollouts support multi-step reasoning, tool use, and adaptation to environment feedback.To scale reinforcement learning, we deployed 10.5K NVIDIA GB300 GPUs for four weeks generating more than 100 million rollouts with a maximum context length of 256K tokens. Training and grading used approximately 1.3 billion sandboxes. To sustain a run of this magnitude, we sourced one million high-quality coding, agentic, and STEM environments. We believe this is one of the largest scale RL runs conducted by any open lab to date. Across our evaluation suite, capabilities continued to improve as we increased RL compute, with no sign of a plateau.Figure 3: Terminal-Bench 2.1, HLE, and DeepSWE scores as a function of cumulative RL rollouts during training of Beam’s reasoning expert, accounting for 80M of the over 100M rollouts generated across the full RL campaign. For comparison, Inkling was trained on 30M rollouts and MiMo on 753K.We trained Beam with asynchronous policy gradients. At scale, policy staleness becomes a major source of instability for these methods. Long running rollouts have tokens that are generated by multiple model checkpoints, with earlier tokens becoming increasingly stale relative to the current policy. Numerical mismatch between training and inference engines further compounds this challenge.We developed new algorithms to maintain stable learning under these conditions while systematically reducing training–inference mismatch throughout our pipeline. These advances enable fully asynchronous RL at scale that remains stable, even when learning from interactions generated more than a day earlier.Figure 4: Stable learning continues as staleness builds up over time. The top plot shows the oldest sample in the batch, with the bottom plot demonstrating stable numerics. Even when training Beam with one-day staleness—107 weight versions behind the current policy—the numerics remain stable.Learning to reason efficientlyWe trained Beam with a controllable length penalty that rewards successful solutions while discouraging unnecessary tokens. Early in RL, performance improved even as completion lengths fell: the model learned to solve tasks more effectively with less reasoning. Later, as Beam developed stronger agentic capabilities, completion lengths grew again, but those additional tokens supported further gains in performance. Throughout training, RL improved the tradeoff between capability and token usage.Figure 5: DeepSWE scores during part of the RL run. Each point corresponds to a different reasoning effort. The Pareto frontier moves in two phases. First, it contracts as the policy learns to be more token efficient. Then, the higher reasoning efforts expand outwards to achieve high performance. Users can control this tradeoff through Beam’s reasoning effort parameter: lower settings favor shorter responses, while higher settings allow longer reasoning to improve performance on demanding tasks. This gives users the flexibility to match reasoning effort to their task and compute budget.How behavior generalizes with RLWe designed Beam’s RL training to develop reasoning and agentic capabilities that generalize beyond its training tasks. During a phase of training on reasoning, software engineering, and terminal tasks, we saw consistent gains in browsing despite the absence of browsing tasks from the RL mixture. This transfer suggests that Beam was learning broader agentic capabilities that generalize across domains. When given web access, it organically learned to search for and query other large language models, and to use OCR APIs to read documents.The demos below showcase Beam applying these capabilities across research, application development, gameplay, and machine learning workflows. The examples range from building a live NYC subway dashboard using public data to creating interactive applications and preparing model fine-tuning notebooks. Although Beam is text-only, it can work with information from other modalities when represented as text. In another out of distribution domain, Beam also created a fine-tuning notebook for the latest and smallest Gemma-4 model on a Text2SQL task.Together, these demos illustrate the breadth of tasks Beam can tackle by combining reasoning, coding, and tool use. Each example includes the initial request and resulting output, along with relevant setup and user iterations.Scaling reinforcement learning environmentsFrontier-scale reinforcement learning requires a large volume of difficult, high-quality tasks. We built a pool of nearly one million environments, primarily through synthetic data pipelines, supplemented by proprietary vendor data and open-source sources.We relied heavily on an iterative curation process. First, we synthesized or sourced environments across a broad set of domains including software engineering, terminal use, competitive coding, STEM, web search, tool use, and general knowledge work. Second, we heavily filtered tasks for difficulty (ensuring they were neither consistently solvable, nor impossible for the model) and quality (e.g., not underspecified, misleading, guessable, hackable, or otherwise broken or noisy). Third, we tested the tasks through RL, which allowed us to identify further quality or difficulty issues and inform the next iteration of sourcing and filtering.Throughout Beam’s development, we found that compromises in data quality led to capability plateaus and other training issues. Systematic improvements to task quality were essential to sustaining capability gains throughout the run, which ended with no sign of saturation.Frontier RL infrastructureHigh-compute agentic RL requires generating rollouts, executing tools, evaluating outcomes, and updating the model at scale. We built an asynchronous platform that lets these processes run independently while coordinating the flow of experience and model updates.During Beam’s training, we sustained an average of 110K concurrent rollouts. Seven capabilities made this practical:Fully asynchronous execution: Agents generate rollouts while the trainer learns and publishes new model versions. Each token is tagged with the version that produced it, allowing the training algorithm to account for policy staleness as completed rollouts flow into