SPARTA: Scalable and Principled Benchmark of Tree-Structured Multi-hop QA over Text and Tables
Sungho Park, Jueun Kim, Wook-Shin Han
摘要
Real-world Table–Text question answering (QA) tasks require models that can reason across long text and source tables, traversing multiple hops and executing complex operations such as aggregation. Yet existing benchmarks are small, manually curated—and therefore error-prone—and contain shallow questions that seldom demand more than two hops or invoke aggregations, grouping, or other advanced analytical operations expressible in natural-language queries. We present SPARTA, an end-to-end construction framework that automatically generates large-scale Table–Text QA benchmarks with lightweight human validation, requiring only one quarter of the annotation time of HybridQA. The framework first constructs a reference fact database by enriching each source table with grounding tables whose tuples are atomic facts automatically extracted from the accompanying unstructured passages, then synthesizes nested queries whose number of nested predicates matches the desired hop count. To ensure that every SQL statement is executable and that its verbalization yields a fluent, human-sounding question, we propose two novel techniques: provenance-based refinement, which rewrites any syntactically valid query that returns a non-empty result, and realistic-structure enforcement, which confines generation to post-order traversals of the query graph. The resulting pipeline produces thousands of high-fidelity question–answer pairs covering aggregations, grouping, and deep multi-hop reasoning across text and tables. On SPARTA, state-of-the-art models that reach over 70 F1 on HybridQA or over 50 F1 on OTT-QA drop by more than 30 F1 points, exposing fundamental weaknesses in current cross-modal reasoning. We will release the benchmark, construction code, and baseline results to spur progress toward robust, realistic Table–Text QA models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- MultiHiertt: Numerical Reasoning over Multi Hierarchical Tabular and Textual DataYilun Zhao, Yunxiang Li, Chenying Li, Rui ZhangACL 2022 · 被引用 168 次
- Open Question Answering over Tables and TextWenhu Chen, Ming-Wei Chang, Eva Schlinger, William Yang Wang 等ICLR 2021 · 被引用 76 次
- CABINET: Content Relevance-based Noise Reduction for Table Question AnsweringSohan Patnaik, Heril Changwal, Milan Aggarwal, Sumit Bhatia 等ICLR 2024 · 被引用 34 次
- Chain-of-Skills: A Configurable Model for Open-Domain Question AnsweringKaixin Ma, Hao Cheng, Yu Zhang, Xiaodong Liu 等ACL 2023 · 被引用 12 次
相关 Paper
- MultiTabQA: Generating Tabular Answers for Multi-Table Question AnsweringVaishali Pal, Andrew Yates, Evangelos Kanoulas, Maarten de RijkeACL 2023 · 被引用 7 次
- CompTab: A Comprehensive Benchmark for Real-World TableQA with Complex Reasoning and Irregular TablesZhen Yang, Wei Du, Jie Wang, Wenze Zhou 等ACL 2026
- TableRAG: A Retrieval Augmented Generation Framework for Heterogeneous Document ReasoningXiaohan Yu, Pu Jian, Chong ChenEMNLP 2025 · 被引用 4 次
- Same Content, Different Representations: A Controlled Study for Table QAYue Zhang, Seiji Maekawa, Nikita BhutaniICLR 2026 · 被引用 5 次
- Multi-Row, Multi-Span Distant Supervision For Table+Text Question AnsweringVishwajeet Kumar, Yash Gupta, Saneem A. Chemmengath, Jaydeep Sen 等ACL 2023 · 被引用 3 次
