ArxivDIGESTables: Synthesizing Scientific Literature into Tables using Language Models
Benjamin Newman, Yoonjoo Lee, Aakanksha Naik, Pao Siangliulue, Raymond Fok, Juho Kim, Daniel S. Weld, Joseph Chee Chang, Kyle Lo
Abstract
When conducting literature reviews, scientists often create literature review tablestables whose rows are publications and whose columns constitute a schema, a set of aspects used to compare and contrast the papers. Can we automatically generate these tables using language models (LMs)? In this work, we introduce a framework that leverages LMs to perform this task by decomposing it into separate schema and value generation steps. To enable experimentation, we address two main challenges: First, we overcome a lack of high-quality datasets to benchmark table generation by curating and releasing ARXIVDIGESTABLES, a new dataset of 2,228 literature review tables extracted from ArXiv papers that synthesize a total of 7,542 research papers. Second, to support scalable evaluation of model generations against humanauthored reference tables, we develop DECON-TEXTEVAL, an automatic evaluation method that aligns elements of tables with the same underlying aspects despite differing surface forms. Given these tools, we evaluate LMs' abilities to reconstruct reference tables, finding this task benefits from additional context to ground the generation (e.g. table captions, in-text references). Finally, through a human evaluation study we find that even when LMs fail to fully reconstruct a reference table, their generated novel aspects can still be useful. blnewman/arxivDIGESTables bnewm0609/arxivDIGESTables * Equal contributions. Dataset size Annotation method Intended Application Evaluation Metric Paper 1 1,200 video sequences Subjectively annotated Objective VQA method development Subjective Mean Opinion Score Paper 2 585 videos Subjective video quality scores via crowdsourcing NR video quality prediction advancement Subjective video quality scores Paper 3 153,841 videos Coarsely annotated set with five quality ratings each Deep-learning VQA model training Spearman rank-order correlation coefficient Paper 4 1 million YouTube videos N/A Large-scale video classification and action recognition Performance improvements over baselines Dataset Size Task Annotations
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59b3267a-e53c-469c-b72e-a5a0053503ebCited by top-tier papers8
- AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research SuiteJonathan Bragg, Mike D'Arcy, Nishant Balepur, Dan Bareket et al.ICLR 2026 · 51 citations
- The Nature of NLP: Analyzing Contributions in NLP PapersAniket Pramanick, Yufang Hou, Saif M. Mohammad, Iryna GurevychACL 2025 · 9 citations
- arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table GenerationWeiqi Wang, Jiefu Ou, Yangqiu Song, Benjamin Van Durme et al.ACL 2026 · 8 citations
- From Automation to Autonomy: A Survey on Large Language Models in Scientific DiscoveryTianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang et al.EMNLP 2025 · 5 citations
- Toward Living Narrative Reviews: An Empirical Study of the Processes and Challenges in Updating Survey Articles in Computing ResearchRaymond Fok, Alexa F. Siu, Daniel S. WeldCHI 2025 · 4 citations
Builds on12
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis EvaluationYujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang et al.NeurIPS 2023 · 119 citations
- MS2: Multi-Document Summarization of Medical StudiesJay DeYoung, Iz Beltagy, Madeleine van Zuylen, Bailey Kuehl et al.EMNLP 2021 · 83 citations
- QASA: Advanced Question Answering on Scientific ArticlesYoonjoo Lee, Kyungjae Lee, Sunghyun Park, Dasol Hwang et al.ICML 2023 · 76 citations
- Coarse-to-Fine Query Focused Multi-Document SummarizationYumo Xu, Mirella LapataEMNLP 2020 · 76 citations
Related papers
- SciTables : A Dataset and Evaluation Framework for Complex Table-to-Text GenerationMehrnoush Alizade, Tengrui Kong, Suman Kalyan MaityVLDB 2026
- AutoDDG: Automated Dataset Description Generation using Large Language ModelsHaoxiang Zhang, Yurong Liu, Aécio S. R. Santos, Wei-Lun Hung et al.SIGMOD 2026 · 17 citations
- LiveXiv - A Multi-Modal live benchmark based on Arxiv papers contentNimrod Shabtay, Felipe Maia Polo, Sivan Doveh, Wei Lin et al.ICLR 2025
- Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review CompositionXuemei Tang, Xufeng Duan, Zhenguang G. CaiEMNLP 2025 · 5 citations
- Evaluating LLM-Generated Diagrams as GraphsChumeng Liang, Jiaxuan YouEMNLP 2025
