SciTables : A Dataset and Evaluation Framework for Complex Table-to-Text Generation
Mehrnoush Alizade, Tengrui Kong, Suman Kalyan Maity
摘要
Generating coherent and factually grounded text from structured data is a core challenge in natural language generation, with applications in scientific communication, medical documentation, and automated reporting. Existing datasets primarily focus on open-domain or simplified table formats, limiting progress in more complex, high-stakes domains. We present SciTables , a new dataset and evaluation framework for scientific table-to-text generation, addressing the gap in existing resources that focus largely on open-domain or simplified tables. Our dataset is constructed from Computer Science papers on arXiv (2017–2023) and features complex tables rich in numeric, symbolic, and mathematical content paired with naturally occurring textual descriptions. We develop a scalable, semi-automated pipeline to extract, clean, and align tables with their associated text, preserving domain-specific language while minimizing annotation cost. The resulting benchmark poses realistic challenges for current models and supports evaluation beyond semantic similarity, including factual accuracy, relevance, and multiple forms of reasoning. We conduct extensive experiments with state-of-the-art generation models and show that while current models achieve strong semantic alignment with reference descriptions, they struggle with higher-order reasoning, aggregation, and factual grounding as table complexity increases. Our work provides a realistic and scalable benchmark for advancing faithful, informative, and reasoning-aware table-to-text generation in scientific domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Prometheus: Inducing Fine-Grained Evaluation Capability in Language ModelsSeungone Kim, Jamin Shin, Yejin Choi, Joel Jang 等ICLR 2024 · 被引用 468 次
- Logical Natural Language Generation from Open-Domain TablesWenhu Chen, Jianshu Chen, Yu Su, Zhiyu Chen 等ACL 2020 · 被引用 116 次
- ToTTo: A Controlled Table-To-Text Generation DatasetAnkur P. Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui 等EMNLP 2020 · 被引用 69 次
- Text-Tuple-Table: Towards Information Integration in Text-to-Table Generation via Global Tuple ExtractionZheye Deng, Chunkit Chan, Weiqi Wang, Yuxi Sun 等EMNLP 2024 · 被引用 2 次
相关 Paper
- Towards Table-to-Text Generation with Numerical ReasoningLya Hulliyyatus Suadaa, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura 等ACL 2021
- SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific TablesXinyuan Lu, Liangming Pan, Qian Liu, Preslav Nakov 等EMNLP 2023 · 被引用 7 次
- A Multi-Task Learning Framework for Reading Comprehension of Scientific Tabular DataXu Yang, Meihui Zhang, Ju Fan, Zeyu Luo 等ICDE 2024 · 被引用 1 次
- Text2Tabular - Reconstructing Tabular Research Data from Scientific PublicationsJonas Gottal, Florian MatthesACL 2026
- HiTab: A Hierarchical Table Dataset for Question Answering and Natural Language GenerationZhoujun Cheng, Haoyu Dong, Zhiruo Wang, Ran Jia 等ACL 2022
