Next upSF Pitch Night by the AI Collective - #SFTechWeek
News

Ai2 releases open-weights AstaBrief 8B for scientific reports

Ai2 released open-weights AstaBrief 8B, two training datasets and workflow code for turning retrieved scientific literature into cited reports.

D
Oct 2, 2026 · 2 min read

Ai2 has released AstaBrief 8B, an open-weights model designed to turn a research question and retrieved scientific-literature excerpts into a cited report. The organization also published two training datasets and report-generation code, and added the model to Asta’s “Generate a report” feature as its Fast mode.

The release gives research institutions a downloadable report-writing component that can be studied, adapted or run on their own infrastructure. AstaBrief does not retrieve literature by itself: Ai2 says in the model card that it expects both a question and already-retrieved excerpts as input, then produces a report with citations.

Ai2 says Fast mode generates an entire report in one pass. That replaces the separate snippet-summarization, clustering and section-by-section writing stages used by Asta’s Claude-powered Thinking mode. In the team’s full-pipeline measurements, Fast averaged 51.1 seconds per report, compared with 178.5 seconds for Thinking—about 3.5 times faster. The release does not disclose the hardware, serving costs or report-length distribution behind those averages; the opened release materials do not cite an independent reproduction.

The project team also reports that AstaBrief scored 87.0 on the 100-question ScholarQA-CS2 test split, including 90.5 for citation precision and 78.2 for citation recall. In the same table, Asta ScholarQA scored 86.2 and DR-Tulu scored 88.8. AstaBrief’s 53.50 score on DeepScholarBench trailed Asta ScholarQA at 60.25 and DR-Tulu at 56.26. Those are Ai2’s evaluation results; the opened release materials do not cite independent verification.

Ai2 cautions that most training and evaluation took place in 2025 and that it did not rerun the complete evaluation against 2026 frontier models. The announcement also describes a separate human study covering 14 questions from three scientific researchers. The team presents the findings as support for its engineering approach rather than as a current frontier-model ranking.

Ai2 published a 39,500-example supervised fine-tuning mix. Its dataset card says Ai2 started with 90,000 filtered research queries and generated 47,000 usable examples before applying a citation-density cutoff; it also says retained queries came only from users who opted into data sharing. Ai2 also published a 6,622-example preference dataset, which it used for direct preference optimization. Ai2 says GPT-4.1 and DeepSeek-R1 compared two reports for each query, with pairs retained only when both judges agreed after a check that found 95% agreement with human preferences.

Early usage figures are also internal. Ai2 says 374 Asta users had tried Fast mode, with 29.1% using it on at least two days and users averaging 3.67 report threads. The team says 23% continued using Fast without switching back for later threads, while another 18% alternated between Fast and Thinking. It did not disclose the observation window, cohort-selection method or statistical uncertainty.

Ai2 labels the AstaBrief 8B weights Apache-2.0. Both training datasets are licensed CC BY-NC 4.0 and warn that synthetic outputs from third-party models remain subject to those providers’ terms. Ai2 also published its ScholarQA workflow code, but the release materials do not specify minimum hardware requirements for running AstaBrief locally.

More news