ToTTo: A Controlled Table-To-Text Generation Dataset
Ankur P. Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, Dipanjan Das
Abstract
We present TOTTO, an open-domain English table-to-text dataset with over 120,000 training examples that proposes a controlled generation task: given a Wikipedia table and a set of highlighted table cells, produce a one-sentence description. To obtain generated targets that are natural but also faithful to the source table, we introduce a dataset construction process where annotators directly revise existing candidate sentences from Wikipedia. We present systematic analyses of our dataset and annotation process as well as results achieved by several state-of-the-art baselines. While usually fluent, existing methods often hallucinate phrases that are not supported by the table, suggesting that this dataset can serve as a useful research benchmark for high-precision conditional text generation. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08305df8-af1a-4acd-860e-fee3576a8154Cited by top-tier papers90
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen et al.ICLR 2024 · 699 citations
- LLMs Get Lost In Multi-Turn ConversationPhilippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer NevilleICLR 2026 · 491 citations
- ExT5: Towards Extreme Multi-Task Scaling for Transfer LearningVamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao et al.ICLR 2022 · 237 citations
- UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language ModelsTianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong et al.EMNLP 2022 · 222 citations
- TableBench: A Comprehensive and Complex Benchmark for Table Question AnsweringXianjie Wu, Jian Yang, Linzheng Chai, Ge Zhang et al.AAAI 2025 · 138 citations
Builds on3
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
- Logical Natural Language Generation from Open-Domain TablesWenhu Chen, Jianshu Chen, Yu Su, Zhiyu Chen et al.ACL 2020 · 116 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
Related papers
- VisToT: Vision-Augmented Table-to-Text GenerationPrajwal Gatti, Anand Mishra, Manish Gupta, Mithun Das GuptaEMNLP 2022 · 4 citations
- SciTables : A Dataset and Evaluation Framework for Complex Table-to-Text GenerationMehrnoush Alizade, Tengrui Kong, Suman Kalyan MaityVLDB 2026
- Towards Faithfulness in Open Domain Table-to-text Generation from an Entity-centric ViewTianyu Liu, Xin Zheng, Baobao Chang, Zhifang SuiAAAI 2021 · 37 citations
- TempTabQA: Temporal Question Answering for Semi-Structured TablesVivek Gupta, Pranshu Kandoi, Mahek Bhavesh Vora, Shuo Zhang et al.EMNLP 2023 · 4 citations
- Towards Table-to-Text Generation with Numerical ReasoningLya Hulliyyatus Suadaa, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura et al.ACL 2021
