Science and research
ArxivDIGESTables: Synthesizing Scientific Literature into Tables using Language Models
Language models generated comparison tables from research papers. Supplying captions and surrounding text helped them identify which features to compare.

Researchers often compare earlier studies in a table: one row per paper, one column per feature that matters to the review. Building that table takes judgment. The writer must choose what to compare, find the evidence in many articles and fill the cells consistently. “ArxivDIGESTables” studies whether language models can help with this work while keeping the two parts of the task—choosing the columns and filling in their values—separate.
The table is two problems
The authors describe table creation as schema generation followed by value generation. The schema is the set of comparison aspects, such as a paper’s data, method or evaluation. Value generation then identifies how each selected aspect applies to each paper. This decomposition gives researchers a way to ask whether a model struggles to decide what matters, extract information, or both.
To support evaluation, the team created ARXIVDIGESTABLES, a dataset of 2,228 literature-review tables extracted from arXiv papers. The tables synthesize information from 7,542 research papers. They also developed DECONTEXT EVAL, an automatic measure that attempts to align columns that refer to the same underlying aspect even when their wording differs. That matters because two authors can ask about the same feature using different labels.
The authors then tested language models on reconstructing reference tables. They found that adding context—such as captions and references in the text—helped models recover schema information. Models still did not perfectly reconstruct the reference tables. The researchers also asked people to assess newly generated aspects that did not match a column in the reference. In that evaluation, the novel aspects were judged to have usefulness, specificity and insightfulness comparable to the human-authored reference aspects.
A benchmark for both missing and new ideas
The dataset offers a practical resource for studying how literature synthesis might be supported at scale. A reference table gives a target for evaluating a system, but a strict match can overlook a genuinely useful new comparison. The paper’s human evaluation addresses that tension: a model can differ from the reference and still suggest a useful dimension, though that requires additional judgment.
This does not make a generated table a substitute for reading the cited work. A concise cell can hide nuance, omit qualifications or incorrectly generalize a result. In a review, the table is valuable because it helps a scholar think across papers. If an automated system makes the comparison look complete, it can also create misplaced confidence in its summaries.
What the data cannot represent
The dataset is built from English-language arXiv papers and is weighted toward computer science. Its results do not establish how table generation performs in fields with different publishing conventions or article structures. The proposed evaluation metric improves alignment, but the authors report that many generated columns still fail to match the corresponding reference columns. Human evaluation of novel aspects is also costly, which limits how easily that quality check can scale.
The authors explicitly note a further risk: generated tables could misrepresent original work or discourage readers from consulting the papers themselves. This is not only a model-accuracy problem. Literature tables compress claims from many sources, so a mistake can affect how readers understand an entire research area.
Possible implications for practice
For researchers and knowledge teams, the study suggests a measured role for language models: help draft candidate schemas and locate evidence, then keep the original papers close at hand while checking the results. A table that exposes its source trail could speed up early exploration without disguising the work still required to verify it. More diverse datasets and affordable human assessment would help establish whether these systems can support reviews beyond the arXiv and computer-science setting.
“ArxivDIGESTables: Synthesizing Scientific Literature into Tables using Language Models” appeared in the 2024 EMNLP proceedings. Its contribution includes both a benchmark and an evaluation method, along with evidence that added context helps without eliminating reconstruction errors.
Bibliography & sources
- Benjamin Newman, Yoonjoo Lee, Aakanksha Naik, Pao Siangliulue, Raymond Fok, Juho Kim, Daniel S. Weld, Joseph Chee Chang, Kyle Lo. 2024. ArxivDIGESTables: Synthesizing Scientific Literature into Tables using Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 9612–9631. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.538.
