Table Question Answering for Low-resourced Indic Languages
Vaishali Pal, Evangelos Kanoulas, Andrew Yates, Maarten de Rijke
Abstract
TableQA is the task of answering questions over tables of structured information, returning individual cells or tables as output. TableQA research has focused primarily on high-resource languages, leaving medium-and low-resource languages with little progress due to scarcity of annotated data and neural models. We address this gap by introducing a fully automatic large-scale table question answering (tableQA) data generation process for low-resource languages with limited budget. We incorporate our data generation method on two Indic languages, Bengali and Hindi, which have no tableQA datasets or models. TableQA models trained on our large-scale datasets outperform stateof-the-art LLMs. We further study the trained models on different aspects, including mathematical reasoning capabilities and zero-shot cross-lingual transfer. Our work is the first on low-resource tableQA focusing on scalable data generation and evaluation procedures. Our proposed data generation method can be applied to any low-resource language with a web presence. We release datasets, models, and code. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed847a15-b065-4cf1-a671-98ccc261ef58Cited by top-tier papers1
Ask how each one uses itBuilds on16
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi et al.ICLR 2022 · 347 citations
- MultiHiertt: Numerical Reasoning over Multi Hierarchical Tabular and Textual DataYilun Zhao, Yunxiang Li, Chenying Li, Rui ZhangACL 2022 · 168 citations
Related papers
- One Question Answering Model for Many Languages with Cross-lingual Dense Passage RetrievalAkari Asai, Xinyan Yu, Jungo Kasai, Hanna HajishirziNeurIPS 2021 · 86 citations
- Cross-lingual Transfer for Automatic Question Generation by Learning Interrogative Structures in Target LanguagesSeonjeong Hwang, Yunsu Kim, Gary Geunbae LeeEMNLP 2024 · 2 citations
- MultiTabQA: Generating Tabular Answers for Multi-Table Question AnsweringVaishali Pal, Andrew Yates, Evangelos Kanoulas, Maarten de RijkeACL 2023 · 7 citations
- POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question AnsweringYichen Xu, Liangyu Chen, Liang Zhang, Zihao Yue et al.ACL 2026 · 2 citations
- RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial PerturbationsYilun Zhao, Chen Zhao, Linyong Nan, Zhenting Qi et al.ACL 2023 · 7 citations
