How Far Can LLM Agents Reason with Tables? Benchmarking Multi-Turn Agentic Table Question Answering in the Wild
Jingwang Huang, Jie Zhang, Haoyang Zeng, Changzai Pan, Xianjie Wu, Guanting Dong, Jiaheng Liu, Wei Zhang, Mingyu Zheng, Chunxiao Liu, Kaiwen Wei, Jiang Zhong, Jian Yang
Abstract
Recent advances in large language models (LLMs) have substantially expanded the scope of Table Question Answering (TableQA). However, existing benchmarks primarily treat TableQA as a passive, single-turn natural language understanding task, lacking the capacity to evaluate autonomous reasoning and tool-call trajectories in realistic, multi-turn scenarios. To bridge this gap, we introduce TableAgent-Bench, a large-scale bilingual benchmark that reformulates TableQA as proactive, agentic interactions over structurally complex, multi-table environments. With a topology-aware construction strategy, TableAgent-Bench captures dynamic intent evolution through 1,310 multi-turn dialogues grounded in 2,275 industrial tables. Furthermore, we propose the Table-centric Agent Evaluation Framework (TAEF) to assess agent interactions with complex table structures. Specifically, TAEF integrates a specialized agent toolset and 4 metric categories to systematically diagnose intermediate failure modes, assessing performance across table localization, tool-invocation rationality, and trajectory-level pass rate. Extensive experiments with 25 state-of-the-art LLM agents reveal a substantial capability gap, with even the strongest model Gemini-3-Pro-Preview achieving only 53.4% information coverage. We expect TableAgent-Bench to serve as a rigorous testbed for developing and evaluating agents capable of robust table-centric reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5407166-5f07-40f2-80a2-586a7e369060Builds on12
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
- TableBench: A Comprehensive and Complex Benchmark for Table Question AnsweringXianjie Wu, Jian Yang, Linzheng Chai, Ge Zhang et al.AAAI 2025 · 138 citations
- ToTTo: A Controlled Table-To-Text Generation DatasetAnkur P. Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui et al.EMNLP 2020 · 69 citations
- Tool Learning in the Wild: Empowering Language Models as Automatic Tool AgentsZhengliang Shi, Shen Gao, Lingyong Yan, Yue Feng et al.WWW 2025 · 59 citations
- SheetAgent: Towards a Generalist Agent for Spreadsheet Reasoning and Manipulation via Large Language ModelsYibin Chen, Yifu Yuan, Zeyu Zhang, Yan Zheng et al.WWW 2025 · 13 citations
Related papers
- TopBench: A Benchmark for Implicit Predictive Reasoning in Tabular Question AnsweringAn-Yang Ji, Jun-Peng Jiang, De-Chuan Zhan, Han-Jia YeICML 2026 · 1 citation
- CompTab: A Comprehensive Benchmark for Real-World TableQA with Complex Reasoning and Irregular TablesZhen Yang, Wei Du, Jie Wang, Wenze Zhou et al.ACL 2026
- T2R-BENCH: A Benchmark for Real World Table-to-Report TaskJie Zhang, Changzai Pan, Sishi Xiong, Kaiwen Wei et al.EMNLP 2025 · 2 citations
- ODUTQA-MDC: A Task for Open-Domain Underspecified Tabular QA with Multi-turn Dialogue-based ClarificationZhensheng Wang, ZhanTeng Lin, Wenmian Yang, Kun Zhou et al.ACL 2026
- TALON: A Multi-Agent Framework for Long-Table Exploration and Question AnsweringRuochun Jin, Xiyue Wang, Dong Wang, Haoqi Zheng et al.EMNLP 2025
