On 2026-10-02 Ai2 open-sourced AstaBrief 8B (Qwen3-8B base): one-pass cited scientific reports from a question + retrieved snippets. Powers Asta Fast mode vs Claude Thinking. End-to-end report time averages 51.1s vs 178.5s (~3.5×). Weights: allenai/AstaBrief_8B (updated 2026-10-02).

Key Takeaways

  • ✓Announced 2026-10-02: HF blog + https://huggingface.co/allenai/AstaBrief_8B
  • ✓Training: ~90K filtered real Asta queries → ~47K SFT; ~6K dual-judge-agree DPO; citation-density filtering
  • ✓System: one-pass full report (skip Thinking’s snippet summarize/cluster + section-by-section writing)
  • ✓Latency: Fast ~51.1s/report vs Thinking ~178.5s (~3.5×); authors did not re-run full eval vs today’s frontier
  • ✓Use in Asta Fast mode; self-host weights; example local PDF workflow + open training data
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

Core Background & Industry Pain Points

Scientists need cited long-form synthesis with traceable evidence—not chatty summaries that silently broaden claims. Ai2’s Asta already offers Claude Thinking reports, but latency/API cost and sensitive topics push toward self-hosting. On 2026-10-02 Ai2 published Open-sourcing AstaBrief and updated allenai/AstaBrief_8B.

Architecture Highlights & Internals

Base: Qwen3-8B. Data: ~90K filtered real Asta queries; ScholarQA multi-step targets from Claude 3.5/3.7, o3, o4-mini, GPT-4.1 → ~47K SFT after filters. DPO: ~6K pairs with dual generators and dual-judge agreement (GPT-4.1 + DeepSeek-R1; ~95% human agreement claimed). Strongest filter: citation density on synthetic reports. Serving: Fast mode writes the full report in one pass from query + retrieved snippets, skipping Thinking’s summarize/cluster/section pipeline.

Authoritative Benchmarks & Measured Scores

Dev target SQABench-CS2 (200 CS research questions) for rubric / answer precision / citation precision & recall; plus DeepScholarBench and pairwise vs Thinking. Charts show AstaBrief near Claude Thinking and DR Tulu—read the blog figures; authors did not re-run against today’s frontier. System latency: Fast 51.1s/report vs Thinking 178.5s (~3.5×). Early usage: 374 triers, 29.1% multi-day; 3.67 report threads avg; 23% stayed on Fast only.

Developer Hands-on Guide

Pull weights from Hugging Face; use Asta Fast mode or self-host; keep retrieve→one-pass contract; evaluate citation metrics separately; treat as scientific-report specialist, not a coding agent.