INFOTABS: Inference on Tables as Semi-structured Data
Vivek Gupta, Maitrey Mehta, Pegah Nokhiz, Vivek Srikumar
Abstract
In this paper, we observe that semi-structured tabulated text is ubiquitous; understanding them requires not only comprehending the meaning of text fragments, but also implicit relationships between them. We argue that such data can prove as a testing ground for understanding how we reason about information. To study this, we introduce a new dataset called INFOTABS, comprising of human-written textual hypotheses based on premises that are tables extracted from Wikipedia info-boxes. Our analysis shows that the semi-structured, multi-domain and heterogeneous nature of the premises admits complex, multi-faceted reasoning. Experiments reveal that, while human annotators agree on the relationships between a table-hypothesis pair, several standard modeling strategies are unsuccessful at the task, suggesting that reasoning about tables can pose a difficult modeling challenge.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f353f293-aae8-42c9-a0d7-2d7329cc3377Cited by top-tier papers34
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training RecipeTianyu Yu, Zefan Wang, Chongyi Wang, Fuwei Huang et al.CVPR 2026 · 179 citations
- Large Language Models are Versatile Decomposers: Decomposing Evidence and Questions for Table-based ReasoningYunhu Ye, Binyuan Hui, Min Yang, Binhua Li et al.SIGIR 2023 · 75 citations
- Turning Tables: Generating Examples from Semi-structured Tables for Endowing Language Models with Reasoning SkillsOri Yoran, Alon Talmor, Jonathan BerantACL 2022 · 57 citations
- MUSTIE: Multimodal Structural Transformer for Web Information ExtractionQifan Wang, Jingang Wang, Xiaojun Quan, Fuli Feng et al.ACL 2023 · 16 citations
- DiSCoMaT: Distantly Supervised Composition Extraction from Tables in Materials Science ArticlesTanishq Gupta, Mohd Zaki, Devanshi Khatsuriya, Kausik Hira et al.ACL 2023 · 15 citations
Builds on3
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi et al.ICLR 2020 · 521 citations
Related papers
- TempTabQA: Temporal Question Answering for Semi-Structured TablesVivek Gupta, Pranshu Kandoi, Mahek Bhavesh Vora, Shuo Zhang et al.EMNLP 2023 · 4 citations
- Right for the Right Reason: Evidence Extraction for Trustworthy Tabular ReasoningVivek Gupta, Shuo Zhang, Alakananda Vempala, Yujie He et al.ACL 2022
- HiTab: A Hierarchical Table Dataset for Question Answering and Natural Language GenerationZhoujun Cheng, Haoyu Dong, Zhiruo Wang, Ran Jia et al.ACL 2022
- Toward a Unified Framework for Unsupervised Complex Tabular ReasoningZhenyu Li, Xiuxing Li, Zhichao Duan, Bowen Dong et al.ICDE 2023 · 4 citations
- ToTTo: A Controlled Table-To-Text Generation DatasetAnkur P. Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui et al.EMNLP 2020 · 69 citations
