Skip to content
HN On Hacker News ↗

Open-sourcing AstaBrief, the fast report-generation model in Asta | Ai2

▲ 26 points • 3 comments • by malshe • 1w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

9 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 1,454
PEAK AI % 9% · §1
Analyzed
Oct 2
backend: pangram/v3.3
Segments scanned
1 windows
avg 1454 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,454 words · 1 segments analyzed

Human AI-generated
§1 Human · 9%

October 2, 2026Ai2ModelDataLanguage models can already help researchers search the literature, synthesize evidence, and work through complex questions. But scientific work places particular demands on these models—answers need to stay grounded in evidence, the models need to preserve what the evidence actually supports rather than quietly broadening a study’s conclusions, and researchers need to be able to verify the final outputs.We see that in how scientists use Asta, our agentic platform for scientific work. Instead of simple keyword searches, users often bring substantial context and many constraints—for example, asking Asta to compare approaches across a body of literature while accounting for a particular method, population, or setting. Many also return to generated reports later, treating them as working research artifacts rather than one-off answers.We wanted to help scientists generate cited reports faster, with a model they could download and run themselves. To do that, we tested whether a small, open model trained specifically for scientific report generation could match the report quality of the proprietary models we were using, while reducing generation time and serving costs. We built AstaBrief 8B, a model that turns a research question and retrieved literature excerpts into a cited report. AstaBrief is available in Asta’s Generate a report feature today as Fast mode alongside Claude-powered Thinking mode, and we’re also open-sourcing it and the training data so others can study, reproduce, and build on our approach.Developing AstaBrief required tens of thousands of real research queries, citation-focused filtering, preference data, and a redesigned report-generation pipeline that writes the full report in one pass rather than section by section. The result is nearly an order-of-magnitude reduction in report generation time compared to the proprietary models we tracked—across the full Asta pipeline, Fast mode averages 51.1 seconds per report compared with 178.5 seconds for Thinking mode, about 3.5× faster. Together, those efficiency gains made AstaBrief a useful test case for a broader goal: building open language models that can be adapted to the specific demands of scientific work. Open weights will also let institutions run AstaBrief on their own infrastructure, which is necessary when research questions reveal sensitive or unpublished work. Alongside the model weights, we’re releasing an example workflow that researchers can adapt to create reports from their own PDFs, providing a starting point for local report generationThis post covers how we trained AstaBrief, what we learned about grounding it in scientific evidence, and which parts of our approach we think can carry forward to future models for science. Most of the training and evaluation described was completed in 2025, so the proprietary models used to generate training data and as comparison points reflect the frontier at the time. We haven’t rerun the full evaluation against today’s frontier models; the results below are best read as evidence about the particular training and system design choices we tested.Training the modelOur goal with AstaBrief was to build an open-weights model with all the qualities that matter most for long-form scientific synthesis: answer quality, relevance, structure, and citation grounding. We started from Qwen3-8B and focused most of our effort on the post-training data, evaluation, and surrounding report-generation scaffolding.Adapting general-purpose models for scientific work – and training new scientific models from scratch – is something we're exploring broadly across Ai2. Through NSF OMAI, a U.S. national initiative led by Ai2 to build fully open AI infrastructure and models for scientific discovery, our researchers are working directly with scientific communities to understand what they need from future open models and where today's general-purpose models fall short. That includes studying how needs differ across scientific fields and workflows, with more findings from that research to share in the future.Recent work, including our DR Tulu, has shown that reinforcement-learning-based (RL) methods can improve long-form report generation for open-weights models, especially when judge models are involved in the training loop. We considered that path for AstaBrief, but ultimately focused on a simpler recipe built around supervised fine-tuning (SFT) and direct preference optimization (DPO).RL-based training can be unstable and expensive. We wanted to see how far we could push report generation quality with a cheaper, more operationally manageable setup—one that's also easier to debug and iterate on.That made the quality of the training data especially important. Rather than relying on a more complex optimization method to compensate for noisy examples, we spent much of the project figuring out how to generate, select, and filter examples that actually demonstrated the report-writing behavior we wanted. We also wanted AstaBrief to be faster so that users could get preliminary reports quickly that they could then iterate over in subsequent turns. For speed improvements, we decided to train AstaBrief to directly generate the final report in one pass given a user query and relevant retrieved snippets, bypassing the expensive snippet summarization and clustering stages our Claude-based Thinking mode uses and not writing out the answer section-by-section. Interestingly, we found it was possible to do so without sacrificing performance. Collecting SFT training dataThe training pipeline began with real user queries submitted through the system described in our paper “Synthesizing scientific literature with retrieval-augmented LMs” and ScholarQA, the framework that now underpins Asta’s Generate a report feature. Rather than training only on synthetic prompts or benchmark-style tasks, we wanted AstaBrief to learn from real queries from real scientists.Our research suggests that scientists often ask different things of language models than users do of general-purpose chatbots or traditional search tools. In our analysis of hundreds of thousands of Asta queries, expert researchers frequently supplied substantial context, multiple constraints, and relationships between concepts rather than relying on short, keyword-style prompts.More recent Asta user studies have also surfaced differences in how researchers want AI involved in their work—some are comfortable using models for ideation or experimentation, while others prefer a narrower role in synthesis, literature surveillance, or pattern-finding. Across those differences, participants want clearer source traceability, more visibility into what a model is doing, and greater control over the context it uses.We filtered the user logs we collected for quality, relevance, and privacy, stripping out beta-tester and bot traffic, dropping queries that were too short to be meaningful, and using an LLM-based filtering pass to catch non-English queries, non-scientific requests, and prompts containing personal information. That left a pool of 90K research-focused queries.For SFT, we generated full-report target outputs from the filtered queries using the multi-step ScholarQA pipeline behind Asta's report generation. The pipeline retrieved relevant literature, organized the material into sections, and used a backing report-generating model to synthesize the evidence into a cited report. We drew on a mix of proprietary systems: Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1. After quality filtering, this yielded 47K usable training examples.Creating DPO pairsDPO required a different kind of training data. Instead of a single target report per query, we needed pairs of reports with one preferred over the other.We built those pairs from a separate subset of queries not used during SFT data generation. One report per query came from the existing ScholarQA pipeline, typically backed by Claude 3.5 Sonnet or 3.7 Sonnet. The competing report was generated by feeding ScholarQA's retrieved literature excerpts to a different model: o3, o4-mini, DeepSeek-V3, or DeepSeek-R1, depending on the example.Two judge models – GPT-4.1 and DeepSeek-R1 – compared each pair and picked a winner. We ensured that LLM judges were aligned with human preferences (95% agreement) and only kept pairs where both judges agreed, which gave us a cleaner preference set and cut much of the noise that typically shows up in preference data generated at scale.After quality filtering, the final DPO dataset came to about 6K examples. Using multiple generators and requiring agreement between two judges gave us a relatively simple way to construct preference data without treating any single model’s output or judgment as ground truth.Filtering data for better attributionOur main evaluation target was SQABench-CS2, a set of 200 user-written computer science research questions. We tracked four metrics throughout the development of AstaBrief:Rubric score, which measures how much necessary content is covered by the report.Answer precision, which measures whether each paragraph is relevant to the question.Citation precision, which measures whether each citation supports the claim it's attached to. Citation recall, which measures whether the report's claims are fully supported by the citations provided.For our final model, we also ran secondary evaluations: DeepScholarBench, a 63-query benchmark for long-form research synthesis built from recent ArXiv papers, and two separate pairwise evaluations against reports generated by the Claude-powered pipeline—an LLM-judged comparison on SQABench-CS2 and a small human study. A report can sound polished and complete while meandering from the question or attaching citations to claims from which the underlying evidence doesn't follow. For scientific synthesis, we needed to measure those behaviors separately. But citation support is only part of scientific faithfulness—a