Skip to content
HN On Hacker News ↗

Introducing Laguna S 2.1

▲ 416 points 89 comments by rexledesma 5w ago HN discussion ↗

Pangram verdict · v3.3

We believe that this document is fully human-written

0 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 7 of 7
SEGMENTS · AI 0 of 7
WORD COUNT 1,629
PEAK AI % 1% · §4
Analyzed
Jul 21
backend: pangram/v3.3
Segments scanned
7 windows
avg 233 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,629 words · 7 segments analyzed

Human AI-generated
§1 Human · 0%

Today we’re releasing Laguna S 2.1, a significant step forward in our development of models that pursue longer horizon work and make effective use of reasoning.Laguna S 2.1 is a 118B total parameter Mixture-of-Experts model with 8B activated parameters per token and supports a context window of up to 1M tokens in thinking and no-thinking modes. It went from the start of training to launch in under nine weeks, and on long-horizon coding benchmarks it holds its own against models many times its size. For every benchmark score we publish today, we are releasing full trajectories for every trial in the final evaluation set at trajectories.poolside.ai. Laguna S 2.1 118B-A8B Tencent Hy3 295B-A21B Inkling 975B-A41B Nemotron 3 Ultra 550B-A55B DeepSeek-V4-Pro Max 1.6T-A49B Kimi K3 2.8T-A50B Qwen 3.7 Max — Muse Spark 1.1 — Claude Fable 5 — Terminal-Bench 2.1 Resolved tasks on Terminal-Bench 2.1. SWE-Bench Multilingual Resolved tasks on SWE-Bench Multilingual. SWE-Bench Pro (Public Dataset) Resolved tasks on SWE-Bench Pro (Public Dataset). DeepSWE Resolved tasks on DeepSWE. SWE Atlas (Codebase QnA) Resolved tasks on SWE Atlas (Codebase QnA). Toolathlon Verified Resolved tasks on Toolathlon Verified. Benchmarks as of 21 July 2026. pass@1 averaged over 4 attempts per task, except DeepSWE, SWE Atlas (Codebase QnA) and Toolathlon Verified that had 3 attempts per task. For all benchmarks we take the maximum of the vendor self-reported score, benchmark author leaderboard or third-party leaderboard (Artificial Analysis), except SWE Atlas (Codebase QnA) where we do not use third-party leaderboard figures.

§2 Human · 1%

Laguna S 2.1 (118B-A8B) Tencent Hy3 (295B-A21B) Inkling (975B-A41B) Nemotron 3 Ultra (550B-A55B) DeepSeek-V4-Pro Max (1.6T-A49B) Kimi K3 (2.8T-A50B) Qwen 3.7 Max (—) Muse Spark 1.1 (—) Claude Fable 5 (—)

Terminal-Bench 2.1 70.2 71.7 63.8 56.4 64.0 88.3 74.5 80 88.0

SWE-Bench Multilingual 78.5 75.8 - 67.7 76.2 - 78.3 - -

SWE-Bench Pro (Public Dataset) 59.4 57.9 54.3 - 55.4 - 60.6 61.5 80.3

DeepSWE 40.4 - - - 9.0 69.0 - 53.3 70.0

SWE Atlas (Codebase QnA) 46.2 - - - 27.2 - - 42.2 -

Toolathlon Verified 49.7 - 45.5 34.3 55.9 - - 75.6 -

Punching above its weight classLaguna S 2.1 is, as far as we can measure, the most capable agentic coding model in its weight class by a wide margin.S 2.1 scores 70.2% on Terminal-Bench 2.1 in our agent harness with thinking enabled. Its compact size makes it uniquely suitable for complex work on local machines.

§3 Human · 1%

Benchmark Open weights Closed / size undisclosed 1 GPT-5.6 Sol 88.82 Kimi K3 2.8T-A50B 88.33 Claude Fable 5 88.04 GPT-5.6 Terra 87.45 GPT-5.6 Luna 84.76 Claude Opus 4.8 84.67 Claude Sonnet 5 80.48 Muse Spark 1.1 80.09 Qwen-3.7 Max 74.510 Hy3 295B-A21B 71.711 Laguna S 2.1 118B-A8B 70.212 MiniMax M3 428B-A23B 66.013 DeepSeek-V4-Pro-Max 1600B-A49B 64.014 Inkling 975B-A41B 63.815 DeepSeek-V4-Flash-Max 284B-A13B 61.816 Nemotron 3 Ultra 550B-A55B 56.417 Inkling-Small 276B-A12B 52.718 Qwen3.6-27B 27B 51.319 Qwen3.6-35B-A3B 35B-A3B 44.920 Nemotron 3 Super 120B-A12B 38.621 Laguna XS 2.1 33B-A3B 33.422 Mistral Small 4 119B 21.4 Benchmarks as of 21 July 2026. pass@1 averaged over 4 attempts per task, except DeepSWE, SWE Atlas (Codebase QnA) and Toolathlon Verified that had 3 attempts per task.

§4 Human · 1%

For all benchmarks we take the maximum of the vendor self-reported score, benchmark author leaderboard or third-party leaderboard (Artificial Analysis), except SWE Atlas (Codebase QnA) where we do not use third-party leaderboard figures.Terminal-Bench 2.1 evaluates a wide, high-quality set of long-horizon tasks where an agent model is connected to its environment through a terminal. Laguna S 2.1 is a standout model in its size category on this benchmark. Benchmark Laguna S 2.1 Other Laguna Other disclosed models Total parameters, log scale. Models with undisclosed total parameter counts are omitted.A closer look at DeepSWEThe benchmarks above are all meaningful, and we're glad to be close to the frontier on them. But part of that closeness is a property of maturing benchmarks: as the frontier advances, top scores cluster in the 70-90% range and models that behave very differently end up no more than a few points apart. Datacurve’s DeepSWE still has significant headroom. Its tasks are longer-horizon and hard to partially solve, and the scores actually spread: frontier models range from 54% to 73% on the v1.1 variant, with some 1T+ parameter open models scoring below 10%.On DeepSWE v1.1, Laguna S 2.1 scores 40.4 in thinking mode in pool harness. Open weights Closed / size undisclosed 1 GPT-5.6 Sol 73.02 Claude Fable 5 70.03 GPT-5.6 Terra 70.04 Kimi K3 2.8T-A50B 69.05 GPT-5.6 Luna 67.26 GPT-5.5 67.07

§5 Human · 0%

Claude Opus 4.8 59.08 Claude Sonnet 5 54.09 Grok 4.5 54.010 Muse Spark 1.1 53.311 GPT-5.4 52.012 GLM 5.2 753B-A40B 44.013 Laguna S 2.1 118B-A8B 40.414 Gemini 3.5 Flash 37.015 Kimi K2.7 Code 31.016 Claude Sonnet 4.6 30.017 Gemini 3.1 Pro 12.018 DeepSeek-V4-Pro-Max 1600B-A49B 9.019 Laguna XS 2.1 33B-A3B 0.3 pass@1 averaged over 3 attempts per task, harnesses vary (DeepSWE's leaderboard uses mini-swe-agent, model providers may report in their own harnesses and we report in pool, our agent harness). For all benchmarks we take the maximum of the vendor self-reported score, benchmark author leaderboard or third-party leaderboard (Artificial Analysis).It is worth noting that Laguna S 2.1 scored 40.4% in our agent harness, pool, not mini-swe-agent which DeepSWE’s leaderboard uses. For other models we report maximal over reported scores which for most models are the official leaderboard results reported by Datacurve. While this makes scores less comparable, we don’t believe it puts us in a particularly advantageous position as it’s been reported that many of the models score the same or better in mini-swe-agent compared to their native harnesses. Every trajectory in the final evaluation run is available here.Evaluation methodologyEvaluation of agent models is notoriously difficult due to prevalence of reward hacking. We have previously written about reward hacking in leading benchmarks and our evaluations system and rigor as part of the technical report on our Laguna M.1 and XS.2 models. Recent work has focused on adversarial judging to increase reward hacking detection.

§6 Human · 0%

With this release, we are making all trajectories from our final evaluations of the published Laguna S 2.1 checkpoint available to view and download at trajectories.poolside.ai.Seeing the model workBenchmark scores give a quantitative view into the model behavior, but to get a better intuitive understanding of how the model works it’s useful to look into runs on real world tasks. We share three such tasks with unedited trajectories and commentary.Case study 1 A browser engine from a blank folder One of our favorite things about Laguna S 2.1 is its resourcefulness: It will find clever ways to get to the goal even if the direct path is not available. We saw a great demonstration of this when we asked it to build a browser engine from scratch; knowing it would be a challenge for Laguna to verify its work given its lack of vision capabilities. In one 50-minute session of 181 steps, with no human intervention, Laguna S 2.1 built a working HTML/CSS rendering engine from an empty folder, then proved it renders like a real browser by measuring itself against one. Throughout its work, the model found increasingly complex ways to validate its work despite its limitations, leading to running headless Chromium to read canvases back and comparing screenshots numerically. See the full trajectory here. Read the full case study // the verbatim prompt · reproduce it yourself your job is it to build a simple browser engine (just html/css) in javascript to demonstrate the capabilities of poolsides new "Laguna S" model. the goal is to take render html snippets in a canvas like a real browser. to demonstrate it the engine, build a self-contained single page app that showcases a gallery of multiple html snippets and renders them side by side (canvas with our render engine + iframe letting the hosting browser render it for real for comparison).

§7 Human · 0%

support for most common layout and styling elements Over the session the model built the full pipeline, parser → cascade → layout → renderer, in vanilla JavaScript: an HTML tokenizer and DOM tree, a CSS parser with selector specificity, a cascade engine with inheritance, box-model layout, and a canvas-2D renderer, wrapped in an app that shows nine snippets on its own canvas beside the same markup in an iframe, so the hosting browser sits right there as the reference. The model's engine rendering on a canvas (left) beside the hosting browser's own rendering of the same markup (right). The header reports pixel dimensions and the measured difference between the two.Case study 2 Optimizing our own harness Laguna S 2.1 is capable of pursuing meaningful engineering and research work. In one example, one of our researchers pointed it at our agent harness, used for training/evaluation and user interaction with our models. In an automated loop, Laguna S 2.1 made our harness 5.2% faster with ~70% lower memory allocation. See the full trajectory here. Read the full case study For this task, we instrumented the harness with benchmarks so the model could see exactly where the time and memory went. We set strict rules: one approach at a time, benchmark after every change, keep only what measurably wins. We then ran Laguna S 2.1 in an automated research loop that fed each result back to it and pushed it to keep improving. Results. Over multiple hours of work, Laguna S 2.1 found and implemented multiple different optimizations in our agent harness, resulting in an overall speedup of 5.2%, and reducing memory allocation by ~70%. The plot below shows the progression of the optimization, with the insights and discoveries the model made along its way. Laguna S 2.1 found that streaming-token accumulation used O(n^2) string concatenation and replaced it with buffers. It also found several instances of redundant copying and over-allocation during trajectory materialization, which it resolved by memoizing materializations and pre-allocating slices to their exact sizes.