A cited-report model, now downloadable
The Allen Institute for AI (Ai2) released the open weights and training data for AstaBrief 8B, the model that generates cited research reports inside the company's Asta platform, in a Hugging Face blog post published on October 2, 2026. Anyone can now download the model and run it on their own computer or server.
AstaBrief turns a research question plus retrieved literature excerpts into a fully cited report in a single pass. Inside Asta, it powers the "Generate a report" feature's Fast mode, alongside a Claude-powered Thinking mode. Speed is the headline claim: across the full Asta pipeline, Fast mode averages 51.1 seconds per report versus 178.5 seconds for Thinking mode, which Ai2 describes as about 3.5 times faster.
That pipeline figure should not be confused with a separate claim on the generation step itself. The blog states there was a "nearly an order-of-magnitude reduction in report generation time compared to the proprietary models we tracked." That order-of-magnitude figure applies to the report-generation step compared to the proprietary models Ai2 tracked, not to the whole pipeline.
Ai2 also framed the release as a matter of privacy as much as speed. "Open weights will also let institutions run AstaBrief on their own infrastructure, which is necessary when research questions reveal sensitive or unpublished work," the post says. Alongside the weights, Ai2 published an example workflow that researchers can adapt to generate reports from their own PDFs locally.
How the model was trained
AstaBrief is a fine-tune of Qwen3-8B, a general-purpose open model, rather than a model trained from scratch. Ai2 concentrated its effort on the post-training data, the evaluation, and the report-generation scaffolding around the model.
The training recipe is deliberately simple: supervised fine-tuning (SFT) followed by direct preference optimization (DPO). Ai2 writes that it considered reinforcement-learning-based training, which its own DR Tulu work has shown can help long-form report generation, but that RL methods "can be unstable and expensive." The team wanted a cheaper, easier-to-debug setup, which shifted the burden onto data quality.
The SFT data started from roughly 90,000 real research queries, filtered from user logs for quality, relevance, and privacy, with bot traffic, too-short queries, non-English prompts, and prompts containing personal information removed. Full-report target outputs were generated using a mix of proprietary systems: Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1. After quality filtering, this yielded about 47,000 usable SFT examples.
The DPO stage required pairs of reports with one preferred over the other, built from a separate subset of queries. Competing reports came from o3, o4-mini, DeepSeek-V3, or DeepSeek-R1. Two judge models, GPT-4.1 and DeepSeek-R1, compared each pair; Ai2 reports 95 percent agreement between LLM judges and human preferences and kept only pairs where both judges agreed. The final DPO set came to about 6,000 examples.
What was measured, and how
The primary benchmark was SQABench-CS2, a set of 200 user-written computer science research questions. Ai2 tracked four metrics: rubric score (how much necessary content is covered), answer precision (whether each paragraph is relevant to the question), citation precision (whether each citation supports the claim attached to it), and citation recall (whether the report's claims are fully supported by the citations). Secondary evaluations included DeepScholarBench, a 63-query benchmark for long-form research synthesis built from recent arXiv papers, and two pairwise comparisons against the Claude-powered pipeline, one LLM-judged and one a small human study.
One design choice stands out for readers: the model writes the full report in one pass, given the query and relevant retrieved snippets, rather than summarizing and clustering snippets first and writing section by section. Ai2 says it found this did not sacrifice performance, and it is a major reason the Fast mode is faster.
It is worth being clear about what these numbers do and do not establish. The comparison is an internal engineering evaluation by Ai2 about its own design choices, published by the same team that made the choices. Independent replication is now possible because the weights and data are open, but it has not happened yet.
The caveats Ai2 itself puts first
Ai2's most important statements in the post are its own caveats, and they belong in the main story rather than a footnote. "Most of the training and evaluation described was completed in 2025, so the proprietary models used to generate training data and as comparison points reflect the frontier at the time," the post states. "We haven't rerun the full evaluation against today's frontier models; the results below are best read as evidence about the particular training and system design choices we tested."
In other words, the quality comparisons involve Claude 3.5 and 3.7 Sonnet, o3, o4-mini, and GPT-4.1 as they existed in 2025. Whether AstaBrief matches 2026-era frontier systems on report quality is, per Ai2's own statement, an open question.
Ai2 also names a gap in its own metrics. "Our development metrics focused primarily on relevance, coverage, and citation grounding; a richer evaluation of scientific report writers should also test whether they preserve the scope and strength of the claims in their sources," the post says.
That failure mode is easy to describe in plain language: a report can cite the right study and still quietly overstate it. A finding measured on a particular sample becomes a generic claim about an entire population. A result reported in the past tense becomes a present-tense statement that sounds universally true. A descriptive finding becomes a recommendation for what clinicians, policymakers, or researchers should do. Each step can broaden the apparent scope of the evidence without introducing an obviously false statement, so the citation still looks correct while the claim is stronger than what the researchers established.
Ai2 cites published research on generalization bias in large language models as an example of how subtle this can be. The point matters more, not less, as report-writing models become faster and cheaper to run at scale.
What changes for scientists and readers
For working scientists, the concrete change is that the model behind a cited literature report is now a file they can run themselves. Three practical consequences follow.
First, speed and cost of preliminary reports. A first-pass synthesis of a literature question can be produced in under a minute on the Asta pipeline, and institutions running the open weights locally avoid per-report API costs. Ai2 positions these reports as starting points to iterate on in subsequent turns, not finished reviews.
Second, privacy. When a research question itself contains sensitive or unpublished material, sending it to a hosted service may be unacceptable. Local execution keeps both the query and the retrieved documents on institutional infrastructure.
Third, a new open question about quality. Faster and open is not the same as verified. Open weights enable reproducibility: other groups can now rerun Ai2's evaluation, test the model against current frontier systems, and build better measures of whether AI-written syntheses preserve the scope and strength of their sources. But none of that verification has happened yet, and Ai2's own metrics do not yet capture the overstatement failure mode it describes.
My assessment, as distinct from the reported facts: the release is most valuable as infrastructure for study rather than as a proven drop-in replacement for careful human literature review. A 51-second cited report is useful precisely because a researcher will check it. The honest framing of this release is that it makes checking both more necessary and, because everything is open, more possible.
