Skip to content
HN On Hacker News ↗

Inkling: Our Open-Weights Model

▲ 1227 points 294 comments by vimarsh6739 1mo ago HN discussion ↗

Pangram verdict · v3.3

We believe that this document is fully human-written

3 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 6 of 6
SEGMENTS · AI 0 of 6
WORD COUNT 1,361
PEAK AI % 16% · §2
Analyzed
Jul 15
backend: pangram/v3.3
Segments scanned
6 windows
avg 227 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,361 words · 6 segments analyzed

Human AI-generated
§1 Human · 2%

Our mission is to build AI that extends human will and judgment. We have developed a platform that lets anyone customize models, previewed an AI system built for interactive collaboration, and published novel research. Today we are advancing our mission by releasing a model we trained from scratch with the full weights available, so that people can make it their own. Our model, called Inkling, is a Mixture-of-Experts transformer with 975B total parameters, 41B active. It supports a context window of up to 1M tokens. It was pretrained on 45 trillion tokens of text, images, audio and video. It is the first in a family of models of different sizes: alongside it we are sharing a preview of Inkling-Small, a lighter-weight model with 12B active parameters, trained with a similar recipe, that achieves strong performance with even lower cost and latency. Inkling reasons natively over text, images, and audio, and balances cost with performance through efficient and controllable thinking effort. We trained it to be a broad, balanced foundation model: strong across many domains, flexible enough to adapt. Inkling is not the strongest overall model available today, open or closed. Instead, a combination of qualities makes it a good open-weights base for customization: multimodal capabilities, efficient thinking, and availability on Tinker for fine-tuning. Inkling is just the start: our first release in a model family we will continue to build on. We want to make customization accessible for more use cases, so Inkling is available for fine-tuning on Tinker today. Picking the right base model to fine-tune is a qualitative judgment that combines measurable benchmarks with the unique feel of a model that comes from playing with it. To enable the latter we’re adding the Inkling Playground in the Tinker console: a developer-facing interface for chatting with Inkling. To show what customization means in practice, we asked Inkling to fine-tune itself.

§2 Human · 16%

Using Tinker, the model wrote its own fine-tuning job, ran it, and evaluated the result:

Build · inkling · tinker-prod ~/news/introducing-inkling/ 1.33.7

Start in OpenCode: Inkling runs inside the OpenCode harness.

Capabilities Real-world applications require models with a wide range of capabilities that can be combined and improved with fine-tuning. We showcase what Inkling can do, and how it measures on important qualities such as trustworthiness and safety. Generalist model Inkling is designed to be broad. We trained it across agentic, reasoning, coding, instruction-following, factuality, vision, and audio tasks, rather than narrowly optimizing for one domain. That breadth matters for customization and real-world use: different users need models that can adapt to very different workflows, not just excel on benchmarks.

Spider chart comparing Inkling, Nemotron 3 Ultra, GLM 5.2, GPT 5.6 Sol, and Claude Fable 5 on ten evaluations scored from zero to one hundred. Inkling is shown with the heavier cobalt line. Evaluations without a reported model score are plotted at zero. Hover an evaluation to compare every model's score.

Inkling is a broad, balanced generalist model. Benchmark scores are shown on a shared 0–100 scale; higher is better. The results show competitive performance across text, agentic, multimodal, and audio evaluations, rather than a model narrowly optimized for one benchmark family. This breadth reflects Inkling’s intended role: a practical multimodal foundation model for customization across domains, workflows, and products.

Agentic coding and tool use A strong base for fine-tuning needs to flexibly solve a wide variety of tasks with agentic tool use. Inkling scores well among open-weights models on most agentic benchmarks without any significant weak points.

§3 Human · 3%

We trained Inkling to run inside a variety of coding and agent harnesses, and randomized the tool set and schema during training to reduce sensitivity to any particular one. Inkling’s controllable thinking effort, described in the next section, can be set from within the harness. Below are a few demos showcasing Inkling’s agentic coding and tool use, and the artifacts it creates. One-shot web app with embedded browser use Inkling builds a functional web app in a single shot, then powers an embedded AI assistant that can operate the web app interface through natural language instructions.

Inkling one-shots a job-application web app from the prompt in the second tab, then a browser-use agent fills out the form from a saved profile.

Design Arena Inkling was evaluated on Design Arena’s Agentic Web Dev leaderboard, where blinded human evaluators compare generated web apps head to head. It ranks among the strongest open-weight models.

Claude Sonnet 5 1333

Claude Fable 5 1329

Claude Opus 4.8 1285

GLM 5.2 1275

Grok 4.5 1271

GPT-5.6 Sol 1260

Inkling 1257

Claude Opus 4.6 1257

Gemini 3.5 Flash 1254

Kimi K2.6 1249

§4 Human · 1%

Claude Sonnet 4.6 1237

Kimi K2.7 Code 1234

GLM 5.1 1233

Claude Opus 4.5 1212

Grok 4.20 Reasoning 1203

Gemini 3.1 Pro Preview 1187

Grok 4.3 1185

Kimi K2.5 (Thinking) 1185

Inkling’s position on Design Arena’s Agentic Web Dev leaderboard, a blinded human evaluation of generated apps.

Cohesively styled artifacts Inkling creates multi-page artifacts with precise instruction following, cohesive styling and design throughout.

From the prompt in the second tab, Inkling produces a polished nine-page PDF food and travel journal — browse the actual document above.

Multiplayer game created through long refinement loop Inkling refined an online snake game through 40 iterations of feedback from GPT Codex serving as a reviewer. The ability to sustain a long process of refinement and improve from feedback is crucial to creating the best collaborative work.

A multiplayer snake game generated by Inkling from the prompt in the second tab — real-time server, bots, leaderboard and all.

Controllable thinking effort Test-time scaling and problem-solving are the core capability of every model, but that capacity is hard to capture with a single number. Developers fine-tuning models for a specialized task care as much about efficiency as about the max-effort performance on a public benchmark. Cost and latency are often binding constraints in real-world applications, and low latency in particular is crucial for enabling collaboration and improvement through iteration. Inkling supports controllable thinking effort, allowing you to balance performance with token efficiency.

§5 Human · 6%

The chart below shows the effort/performance curve of Inkling as well as other open weight models on a range of benchmarks: Terminal Bench 2.1 for agentic coding, HLE for advanced reasoning, and IFBench for instruction following. Inkling spends 1/3 as many tokens to achieve the same performance as Nemotron 3 Ultra on Terminal Bench. Cost and latency matter for a model that you run millions of times and as part of longer workflows; looking at the full cost curve allows developers to choose the best model for each use case.

Inkling (effort sweep) GLM-5.2 Kimi K2.6 Nemotron 3 Ultra Kimi K2.5 GPT-OSS (high)

Sweeping Inkling's effort setting from 0.2 to 0.99 traces its performance against mean generated tokens on Terminal Bench 2.1, HLE, and IFBench; competing models are shown at their default operating point. Inkling reaches a given score at fewer tokens — for example, it matches Nemotron 3 Ultra on Terminal Bench 2.1 at roughly a third of the tokens. *Humanity's Last Exam scores reflect an earlier checkpoint and run slightly below the final release.

Multimodality A major goal of Inkling’s design is to serve as the background reasoning model in the interaction models system we recently introduced. Interaction models enable the user to collaborate naturally, using voice and vision in real-time. This requires a model natively trained for broad multimodal capabilities.

Open weights Closed weights Inklingeffort=0.99 Qwen3-Omni Nemotron-3Nano-Omni Kimi K2.5 Kimi K2.6 Qwen3.5Omni-Plus Gemini 3.1 Pro(high)

§6 Human · 6%

Audio Audio MC 56.6% 24.3% 23.2% – – 37.6% 66.8% MMAU 77.2% 77.5% 76.7% – – 81.1% 82.5% VoiceBench 91.4% 88.8% 89.4% – – 92.4% 94.3% Vision MMMU Pro(Standard 10) 73.5% 60.0% 53.0% 75.0% 79.0% 71.0% 82.0% Charxiv RQ 78.1% 61.1% 63.6% 77.5% 80.4% 72.5% 80.2% Charxiv RQwith python 82.0% – – 78.7% 86.7% – 89.9%

Audio and vision benchmarks against specialist omni models (open- and closed-weight), reported at effort=0.99. The multimodal components are trained from scratch on general-domain data. We opted for an encoder-free architecture for audio and vision inputs, consistent with the interaction model design. Audio signals are input as discrete dMel spectrogramsBai et al., 2024., while images are encoded as patches of 40×40 pixels using a four-layer hMLPTouvron et al.,