Open Question Answering over Tables and Text
Wenhu Chen, Ming-Wei Chang, Eva Schlinger, William Yang Wang, William W. Cohen
摘要
In open question answering (QA), the answer to a question is produced by retrieving and then analyzing documents that might contain answers to the question. Most open QA systems have considered only retrieving information from unstructured text. Here we consider for the first time open QA over both tabular and textual data and present a new large-scale dataset Open Table-and-Text Question Answering (OTT-QA) to evaluate performance on this task. Most questions in OTT-QA require multi-hop inference across tabular data and unstructured text, and the evidence required to answer a question can be distributed in different ways over these two types of input, making evidence retrieval challenging---our baseline model using an iterative retriever and BERT-based reader achieves an exact match score less than 10%. We then propose two novel techniques to address the challenge of retrieving and aggregating evidence for OTT-QA. The first technique is to use ``early fusion'' to group multiple highly relevant tabular and textual units into a fused block, which provides more context for the retriever to search for. The second technique is to use a cross-block reader to model the cross-dependency between multiple retrieved evidence with global-local sparse attention. Combining these two techniques improves the score significantly, to above 27%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper47
- MultiModalQA: complex question answering over text, tables and imagesAlon Talmor, Ori Yoran, Amnon Catav, Dan Lahav 等ICLR 2021 · 被引用 229 次
- TableBench: A Comprehensive and Complex Benchmark for Table Question AnsweringXianjie Wu, Jian Yang, Linzheng Chai, Ge Zhang 等AAAI 2025 · 被引用 138 次
- MATE: Multi-view Attention for Table Transformer EfficiencyJulian Martin Eisenschlos, Maharshi Gor, Thomas Müller, William W. CohenEMNLP 2021 · 被引用 62 次
- DeepAnalyze: Agentic Large Language Models for Autonomous Data ScienceShaolei Zhang, Ju Fan, Meihao Fan, Yizhe Liu 等ICML 2026 · 被引用 48 次
- Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical ReasoningPan Lu, Liang Qiu, Kai-Wei Chang, Ying Nian Wu 等ICLR 2023 · 被引用 41 次
它引用的顶会 Paper11
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 等ICLR 2020 · 被引用 674 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher 等ICLR 2020 · 被引用 322 次
相关 Paper
- RINK: Reader-Inherited Evidence Reranker for Table-and-Text Open Domain Question AnsweringEunhwan Park, Sung-Min Lee, Daeryong Seo, Seonhoon Kim 等AAAI 2023 · 被引用 4 次
- Multi-Row, Multi-Span Distant Supervision For Table+Text Question AnsweringVishwajeet Kumar, Yash Gupta, Saneem A. Chemmengath, Jaydeep Sen 等ACL 2023 · 被引用 3 次
- Dual Reader-Parser on Hybrid Textual and Tabular Evidence for Open Domain Question AnsweringAlexander Hanbo Li, Patrick Ng, Peng Xu, Henghui Zhu 等ACL 2021
- Open Domain Question Answering with A Unified Knowledge InterfaceKaixin Ma, Hao Cheng, Xiaodong Liu, Eric Nyberg 等ACL 2022 · 被引用 45 次
- CRAFT: Training-Free Cascaded Retrieval for Tabular QAAdarsh Singh, Kushal Raj Bhandari, Jianxi Gao, Soham Dan 等ACL 2026 · 被引用 2 次
