ArxivDIGESTables: Synthesizing Scientific Literature into Tables using Language Models
Benjamin Newman, Yoonjoo Lee, Aakanksha Naik, Pao Siangliulue, Raymond Fok, Juho Kim, Daniel S. Weld, Joseph Chee Chang, Kyle Lo
摘要
When conducting literature reviews, scientists often create literature review tablestables whose rows are publications and whose columns constitute a schema, a set of aspects used to compare and contrast the papers. Can we automatically generate these tables using language models (LMs)? In this work, we introduce a framework that leverages LMs to perform this task by decomposing it into separate schema and value generation steps. To enable experimentation, we address two main challenges: First, we overcome a lack of high-quality datasets to benchmark table generation by curating and releasing ARXIVDIGESTABLES, a new dataset of 2,228 literature review tables extracted from ArXiv papers that synthesize a total of 7,542 research papers. Second, to support scalable evaluation of model generations against humanauthored reference tables, we develop DECON-TEXTEVAL, an automatic evaluation method that aligns elements of tables with the same underlying aspects despite differing surface forms. Given these tools, we evaluate LMs' abilities to reconstruct reference tables, finding this task benefits from additional context to ground the generation (e.g. table captions, in-text references). Finally, through a human evaluation study we find that even when LMs fail to fully reconstruct a reference table, their generated novel aspects can still be useful. blnewman/arxivDIGESTables bnewm0609/arxivDIGESTables * Equal contributions. Dataset size Annotation method Intended Application Evaluation Metric Paper 1 1,200 video sequences Subjectively annotated Objective VQA method development Subjective Mean Opinion Score Paper 2 585 videos Subjective video quality scores via crowdsourcing NR video quality prediction advancement Subjective video quality scores Paper 3 153,841 videos Coarsely annotated set with five quality ratings each Deep-learning VQA model training Spearman rank-order correlation coefficient Paper 4 1 million YouTube videos N/A Large-scale video classification and action recognition Performance improvements over baselines Dataset Size Task Annotations
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research SuiteJonathan Bragg, Mike D'Arcy, Nishant Balepur, Dan Bareket 等ICLR 2026 · 被引用 51 次
- The Nature of NLP: Analyzing Contributions in NLP PapersAniket Pramanick, Yufang Hou, Saif M. Mohammad, Iryna GurevychACL 2025 · 被引用 9 次
- arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table GenerationWeiqi Wang, Jiefu Ou, Yangqiu Song, Benjamin Van Durme 等ACL 2026 · 被引用 8 次
- From Automation to Autonomy: A Survey on Large Language Models in Scientific DiscoveryTianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang 等EMNLP 2025 · 被引用 5 次
- Toward Living Narrative Reviews: An Empirical Study of the Processes and Challenges in Updating Survey Articles in Computing ResearchRaymond Fok, Alexa F. Siu, Daniel S. WeldCHI 2025 · 被引用 4 次
它引用的顶会 Paper12
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis EvaluationYujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang 等NeurIPS 2023 · 被引用 119 次
- MS2: Multi-Document Summarization of Medical StudiesJay DeYoung, Iz Beltagy, Madeleine van Zuylen, Bailey Kuehl 等EMNLP 2021 · 被引用 83 次
- QASA: Advanced Question Answering on Scientific ArticlesYoonjoo Lee, Kyungjae Lee, Sunghyun Park, Dasol Hwang 等ICML 2023 · 被引用 76 次
- Coarse-to-Fine Query Focused Multi-Document SummarizationYumo Xu, Mirella LapataEMNLP 2020 · 被引用 76 次
相关 Paper
- SciTables : A Dataset and Evaluation Framework for Complex Table-to-Text GenerationMehrnoush Alizade, Tengrui Kong, Suman Kalyan MaityVLDB 2026
- AutoDDG: Automated Dataset Description Generation using Large Language ModelsHaoxiang Zhang, Yurong Liu, Aécio S. R. Santos, Wei-Lun Hung 等SIGMOD 2026 · 被引用 17 次
- LiveXiv - A Multi-Modal live benchmark based on Arxiv papers contentNimrod Shabtay, Felipe Maia Polo, Sivan Doveh, Wei Lin 等ICLR 2025
- Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review CompositionXuemei Tang, Xufeng Duan, Zhenguang G. CaiEMNLP 2025 · 被引用 5 次
- Evaluating LLM-Generated Diagrams as GraphsChumeng Liang, Jiaxuan YouEMNLP 2025
