Skip to content
HN On Hacker News ↗

Introducing Grok 4.6

▲ 632 points 616 comments by iLuddite 2w ago HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

5 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 675
PEAK AI % 5% · §1
Analyzed
Aug 12
backend: pangram/v3.3
Segments scanned
1 windows
avg 675 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 675 words · 1 segments analyzed

Human AI-generated
§1 Human · 5%

Today we are releasing Grok 4.6. Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work. It stays with complex tasks across many steps, whether researching a topic, analyzing information, working across a codebase, or turning an idea into a polished application or work artifact. Grok 4.6 achieves frontier intelligence across several agentic coding and knowledge work benchmarks. It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, which is a composite score of nine benchmarks. 0204060AA Intelligence Index62Fable 5 Max61Grok 4.661GPT-5.6 Sol Max56Grok 4.5 HighCompetitor figures are drawn from the respective developers’ published system cards or benchmark leaderboardsBenchmark bar charts comparing Grok 4.6 with other leading models across AA Intelligence, GDPVal-AA, DeepSWE 1.1, CursorBench 3.2, and FrontierCode 1.1. Competitor figures are drawn from the respective developers’ published system cards or benchmark leaderboards. Grok 4.6 is available today in Cursor and Grok Build. We’re offering 2x included usage inside Grok Build and Cursor for the first week so you can start trying 4.6 immediately. Training Grok 4.6 Grok 4.6 underwent a longer supplemental training run than Grok 4.5, with curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe. This produced a stronger foundation for the SFT and RL stages that followed. We then used Grok 4.5 to regenerate the SFT trajectories across reasoning efforts, agent harnesses, and domains such as STEM, software engineering, and knowledge work, and filtered out problematic traces with model-based checks. The resulting SFT checkpoint shows strong performance and improved behavior. Grok 4.6 is trained on a wide range of agentic RL tasks, including knowledge work, general coding, and domain-specific environments for kernel optimization, web development, computer-aided design, and more. Turning ambitious ideas into working projects We tested Grok 4.6 on projects designed to stretch its range and ability to sustain work over many steps. We found the model is especially strong at turning a broad product idea into a working first version. It can research unfamiliar domains, structure the application, implement the core interactions, and continue refining the result through several rounds of feedback. On longer trajectories, we also started to see more self-testing and verification, with the model checking its own work before moving on. Grok 4.6 produces stronger first passes on visual and interactive projects than we typically saw with Grok 4.5. Given a concrete product idea, it is able to establish structure and visual language for an application in one pass. This has made it especially useful for projects where the fastest route to a good result was to begin with something substantial and then iterate in the loop. Safety and capabilities Grok 4.6’s safeguards have been improved and calibrated in line with the model’s capabilities. Our safety stack is designed to maximize utility and security across legitimate use cases, allowing Grok 4.6 to be helpful and safe in domains such as vulnerability patching, accelerating the engineering design cycle, and augmenting AI research. Our safeguard evaluation work reflects Grok 4.6’s expanded capabilities, with our widest-ever suite of pre-deployment testing for capabilities and safeguard calibration, as well as extensive post-deployment and third-party testing. Evals Grok 4.6 HighGrok 4.5 HighGPT-5.6 Sol MaxFable 5 MaxAA Intelligence Index61566162GDPVal-AA v21753152617281741CursorBench v3.269.9%66.7%67.2%70.5%DeepSWE v1.165.9%54%73%70%FrontierCode v1.1 (Extended)61.3%56.6%60.6%63.6%APEX-Agents57.5%47.1%56.7%59.2%Terminal-Bench v3.026%15.7%34.6%34.1%APEX-SWE56.4%53.6%—58.8%AA-Briefcase1577131315021574Harvey LAB (Vals)15.8%12.9%2.5%11.3%Best score per evaluation in bold. Third-party model scores are the best of self-reported or publicly available results. Get started with Grok 4.6 Grok 4.6 is available today in Cursor and Grok Build. It’s also available in the API and other partners like OpenRouter, Vercel, and Cloudflare. Pricing starts at $2 per million input tokens and $6 per million output tokens. Additionally, there is a fast variant which is twice the price. We’re offering 2x included usage inside Grok Build and Cursor for the first week so you can start trying 4.6 immediately. Create an API KeyStart building with Grok 4.6 today via the SpaceXAI API.API DocsRead the docs and integrate Grok 4.6 into your stack. Try it in Grok Build for freeGet started today at x.ai/build.