OFICIAL Hugging Face Blog

Open-sourcing AstaBrief, the fast report-generation model in Asta

What happened
Based on Hugging Face Blog · Oct 02, 2026

Hugging Face open-sources AstaBrief 8B, a fast report-generation model for scientific literature, reducing generation time by 3.5× while maintaining cited accuracy.

Open-sourcing AstaBrief, the fast report-generation model in Asta
Hugging Face Blog — Hugging Face
Key points
·
AstaBrief 8B reduces report generation time from 178.5 to 51.1 seconds, a 3.5× speed improvement over proprietary models
·
Model trained using 90,000 filtered research queries and 47,000 SFT examples derived from ScholarQA pipeline outputs
·
Open-sourced alongside example workflow for local report generation from PDFs under Ai2’s NSF OMAI initiative
Key numbers
·
This approach reduced average report generation time from 178.
·
5 seconds to 51.
·
1 seconds, a 3.

Hugging Face has released AstaBrief 8B, an open-weights model designed to generate cited scientific reports from research questions and retrieved literature excerpts. The model is available in Asta’s Generate a report feature as Fast mode, alongside a slower proprietary Thinking mode. Researchers can now download and run AstaBrief locally, enabling faster preliminary reports for iterative refinement.

AstaBrief was trained using tens of thousands of real research queries and citation-focused filtering to ensure outputs remain grounded in evidence. The model generates full reports in a single pass, bypassing multi-stage summarization used in proprietary systems. This approach reduced average report generation time from 178.5 seconds to 51.1 seconds, a 3.5× speed improvement across the full Asta pipeline.

The training pipeline relied on supervised fine-tuning (SFT) and direct preference optimization (DPO) using data derived from real user queries. Over 90,000 research-focused queries were filtered for quality and privacy, with 47,000 used for SFT after generating full-report targets via a multi-step ScholarQA pipeline. DPO training required pairs of reports per query, with preferences determined by comparative evaluations.

Hugging Face also released an example workflow for local report generation from PDFs, alongside the open-sourced model weights. The initiative aligns with broader efforts to build open AI infrastructure for scientific discovery, including work under the NSF OMAI initiative led by Ai2. The team explored reinforcement-learning methods but ultimately prioritized a simpler, more stable training setup focused on high-quality data.

Original source → Deals on Clipraptor.com →