TableEval: A Real-World Benchmark for Complex, Multilingual, and Multi-Structured Table Question Answering
Junnan Zhu, Jingyi Wang, Bohan Yu, Xiaoyu Wu, Junbo Li, Lei Wang, Nan Xu
Abstract
LLMs have shown impressive progress in natural language processing. However, they still face significant challenges in TableQA, where real-world complexities such as diverse table structures, multilingual data, and domain-specific reasoning are crucial. Existing TableQA benchmarks are often limited by their focus on simple flat tables and suffer from data leakage. Furthermore, most benchmarks are monolingual and fail to capture the crosslingual and cross-domain variability in practical applications. To address these limitations, we introduce TableEval, a new benchmark designed to evaluate LLMs on realistic TableQA tasks. Specifically, TableEval includes tables with various structures (such as concise, hierarchical, and nested tables) collected from four domains (including government, finance, academia, and industry reports). Additionally, TableEval features cross-lingual scenarios with tables in Simplified Chinese, Traditional Chinese, and English. To reduce potential data leakage, we curate data from recent real-world documents. Considering that existing TableQA metrics fail to capture semantic accuracy, we further propose SEAT, a new evaluation framework that assesses the alignment between model responses and reference answers at the sub-question level. Experimental results have shown that SEAT achieves high agreement with human judgment. Extensive experiments on TableEval reveal critical gaps in the ability of state-of-the-art LLMs to handle these complex, real-world TableQA tasks, offering insights for future improvements. We make our dataset available here: https://github.com/wenge-research/TableEval .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7903c883-abc5-4565-8d7d-4820b60ac10fCited by top-tier papers6
- Same Content, Different Representations: A Controlled Study for Table QAYue Zhang, Seiji Maekawa, Nikita BhutaniICLR 2026 · 5 citations
- MIRAGE: Scaling Test-Time Inference with Parallel Graph-Retrieval-Augmented Reasoning ChainsKaiwen Wei, Rui Shan, Dongsheng Zou, Jianzhong Yang et al.AAAI 2026 · 5 citations
- Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and EvaluationWei Zhou, Bolei Ma, Annemarie Friedrich, Mohsen MesgarACL 2026 · 3 citations
- ASCIIEval: Benchmarking Models' Visual Perception in Text Strings via ASCII ArtQi Jia, Xiang Yue, Shanshan Huang, Ziheng Qin et al.ICLR 2026 · 2 citations
- Replacing Multi-Step Assembly of Data Preparation Pipelines with One-Step LLM Pipeline Generation for Table QAFengyu Li, Junhao Zhu, Kaishi Song, Lu Chen et al.VLDB 2026 · 1 citation
Builds on14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi et al.ICLR 2022 · 347 citations
- Chain-of-Table: Evolving Tables in the Reasoning Chain for Table UnderstandingZilong Wang, Hao Zhang, Chun-Liang Li, Julian Martin Eisenschlos et al.ICLR 2024 · 244 citations
Related papers
- CompTab: A Comprehensive Benchmark for Real-World TableQA with Complex Reasoning and Irregular TablesZhen Yang, Wei Du, Jie Wang, Wenze Zhou et al.ACL 2026
- MMQA: Evaluating LLMs with Multi-Table Multi-Hop Complex QuestionsJian Wu, Linyi Yang, Dongyuan Li, Yuliang Ji et al.ICLR 2025
- TableBench: A Comprehensive and Complex Benchmark for Table Question AnsweringXianjie Wu, Jian Yang, Linzheng Chai, Ge Zhang et al.AAAI 2025 · 138 citations
- MMTableBench: A Multi-level Multimodal Benchmark for Reasoning and Layout Complexity in Table QAXianjie Wu, Xiaohang Xu, Tingyu Jiang, Jian Yang et al.WWW 2026 · 3 citations
- T2R-BENCH: A Benchmark for Real World Table-to-Report TaskJie Zhang, Changzai Pan, Sishi Xiong, Kaiwen Wei et al.EMNLP 2025 · 2 citations
