ACL2026

M3TQA: Massively Multilingual Multitask Table Question Answering

Daixin Shu, Jian Yang, Zhenhe Wu, Xianjie Wu, Xianfu Cheng, Xiangyuan Guan, Yanghai Wang, Pengfei Wu, Tingyang Yang, Hualei Zhu, Wei Zhang, Ge Zhang, Jiaheng Liu, Zhoujun Li

2 citations

Abstract

Tabular data is a fundamental component of real-world information systems, yet most research in table understanding remains confined to English, leaving multilingual comprehension significantly underexplored. Existing multilingual table benchmarks suffer from geolinguistic imbalance-overrepresenting Indo-European and high-resource languages-and insufficient scale for rigorous cross-lingual analysis. To address these limitations, we introduce a complete framework to enhance the massively multilingual multitask table question answering by introducing the instruction and reinforcement learning dataset M 3 TQA-INSTRUCT, a large-scale benchmark spanning 97 languages across diverse language families, including underrepresented and lowresource ones. Further, we construct M 3 TQA by curating 50 real-world tables in Chinese and English, then applying a robust, six-step LLM-based translation pipeline powered by DeepSeek and GPT-4o, achieving high translation fidelity (median BLEU: 60.19) as validated by back-translation. The benchmark includes 2,916 professionally annotated questionanswering pairs across four tasks designed to evaluate nuanced table reasoning capabilities. Experiments on state-ofthe-art LLMs reveal critical insights into cross-lingual generalization, showing that synthetically generated, unannotated QA data can significantly boost performance-especially for low-resource languages. M3T-Bench thus establishes a new standard for multilingual table understanding, offering both a challenging evaluation platform and a scalable methodology for future research.