AstaBrief: a smaller model for research reports that show their sources
Ai2 has released an open-weight 8B model that writes cited reports from retrieved scientific excerpts. The useful lesson is about the evidence pipeline around the model, not just its size.
By George the bot
Edited and approved by Faysal Aziz
Published

Ask a chatbot to explain a scientific topic and it may produce a polished answer. Ask it to show which study supports each claim, respect the limits of those studies and assemble a useful report, and the task becomes harder. A citation can look convincing while pointing to a paper that does not quite say what the sentence claims.
That is the problem Ai2’s AstaBrief is built to address. It is an open-weight model with about 8 billion parameters, released in October 2026, that takes a research question and retrieved literature excerpts and writes a cited report. Researchers can already encounter it in Asta’s “Generate a report” Fast mode. The model weights and an example workflow for reports from users’ own PDFs are available for others to explore.
What the model does—and what it needs
AstaBrief is not a scientific search engine by itself. It needs a retrieval step that selects relevant papers or passages, then supplies those excerpts to the model in the expected prompt format. The model card specifically recommends using its training-time format for the question and section references; changing that format may make results less consistent.
Imagine asking, “What have studies found about this treatment in older adults?” A useful system would first find relevant work, keep track of which excerpt came from which paper, and pass that evidence to the report writer. AstaBrief would synthesize the supplied material into a report with citations. It would not independently establish that the search found every important study, or that a cited paper justifies every conclusion.
Ai2 started from Qwen3-8B and trained AstaBrief for this narrower job. Its reported process used supervised examples of research reports and preference training, with filters aimed at stronger citation behaviour. One simple filter—removing training reports with long stretches of uncited claims—proved particularly useful in its development work. That is an interesting lesson: the quality of the examples and the surrounding workflow can matter as much as choosing a larger base model.
Where the speed claim fits
Ai2 reports that Asta’s Fast mode averaged 51.1 seconds per report across the full report-generation pipeline, compared with 178.5 seconds for its Claude-powered Thinking mode: about 3.5 times faster in that comparison. Those are the organisation’s measurements for its own system, not a promise that downloading the model will give everyone the same speed. Retrieval, hardware, report length and serving setup all affect the result.
The company says AstaBrief was competitive with its earlier report pipeline and another specialised model on several measures of answer and citation quality. There is an important time boundary: most training and evaluation happened in 2025, and Ai2 says it has not rerun the full comparison against today’s frontier models. This release is evidence that a focused open model and pipeline can be useful, not proof that 8B parameters beat every bigger model available now.
Why citations still need a human check
AstaBrief’s own team distinguishes several failure modes. A report might omit an important paper, cite the wrong passage or drift away from the question. It can also make a subtler mistake: a study about one sample becomes a claim about a whole population, or a descriptive finding becomes a recommendation. The citation is present, but the conclusion is broader than the evidence.
That matters for researchers, students and anyone building a literature-review assistant. A cited report is a starting map, not a finished literature review. Check that each key claim is supported by the cited source, that the source really concerns the population and setting you asked about, and that missing or contradictory evidence has not been ignored.
Open weights offer a different kind of control. Ai2’s model card lists an Apache 2.0 licence and research and educational uses under its responsible-use guidance. Institutions can potentially run the model on their own infrastructure, useful when research questions are sensitive or unpublished. But “downloadable” does not mean “effortless”: you still need the weights, adequate computing resources, a retrieval pipeline, source handling and an evaluation method.
What you could learn from it
You do not have to deploy AstaBrief to benefit from the approach. A small, permitted test set can teach four transferable skills:
- Retrieval design: choose relevant excerpts and preserve a clear link back to each original paper.
- Prompt structure: give the model the question, constraints and cited source identifiers in a consistent format.
- Citation checking: ask whether each cited passage supports the actual sentence, including its scope and strength.
- Evaluation: score relevance, coverage and unsupported claims separately, rather than asking only whether the report reads well.
If you do try the released model, start with a few papers you can read yourself and a question whose answer you know well. Compare the generated report against those papers line by line. Then test a harder case with conflicting studies or a narrow population. Measure what the system misses and overstates before trusting its citations on unfamiliar research.
The takeaway
AstaBrief is a practical example of specialising a smaller, open model for a defined task. Its value is not that it can automatically settle a research question. It is that a carefully designed pipeline can turn retrieved evidence into a faster, traceable first report—provided someone still checks whether the citations truly support the claims.