LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake
Haonan Wang, Jiaxiang Liu, Yurong Liu, Austin Wijaya, Tianle Zhou, Yifan Wu, Yijia Chen, Wanting You, Reya Vir, Daniela Pinto Veizaga, Grace Fan, Yusen Zhang
Abstract
Recent large language models (LLMs) have shown rapid progress on reading-based question answering (QA), where the evidence is explicitly provided or trivially retrievable. In contrast, real-world questions are often not paired with accurate evidence documents. The useful evidence resides in a massive collection of data lakes, necessitating searching as a prerequisite for answering. However, there is a lack of a comprehensive benchmark that requires searching and reasoning over a large collection of data lakes. To this end, we introduce LakeQA, a comprehensive benchmark for search-centric question answering over data lakes that jointly emphasizes searching and reasoning capabilities. LakeQA is built on a heterogeneous collection of 9.5 TB text resources from Wikipedia and open-source government data, spanning structured and unstructured data. To ensure the quality of LakeQA's tasks, each sample is annotated by at least one Ph.D level expert. Each task requires long-horizon multi-hop reasoning with implicit intermediate steps: agents need to discover the correct document(s) and then compose evidence across sources to produce the answer. Intensive experiment results on seven frontier LLMs have demonstrated that LakeQA is challenging. For instance, GPT-5.2 only obtains an exact matching score of 14.73% on LakeQA. Overall LakeQA provides a realistic testbed for developing LLM agents that can both find and analyze data in modern data lakes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce096002-41b6-461c-aa58-c383f4cd57a5Builds on7
- MultiModalQA: complex question answering over text, tables and imagesAlon Talmor, Ori Yoran, Amnon Catav, Dan Lahav et al.ICLR 2021 · 229 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation LearningGrace Fan, Jin Wang, Yuliang Li, Dan Zhang et al.VLDB 2023 · 139 citations
- Open Question Answering over Tables and TextWenhu Chen, Ming-Wei Chang, Eva Schlinger, William Yang Wang et al.ICLR 2021 · 76 citations
- LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data LakesYuhao Deng, Chengliang Chai, Lei Cao, Qin Yuan et al.VLDB 2024 · 36 citations
Related papers
- LiveNewsBench: Evaluating Web Search Agents with Freshly Curated NewsYunfan Zhang, Kathleen McKeown, Smaranda MuresanICML 2026 · 2 citations
- MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Long-tail KnowledgeJie He, Nan Hu, Wanqiu Long, Jiaoyan Chen et al.ACL 2026 · 1 citation
- SealQA: Raising the Bar for Reasoning in Search-Augmented Language ModelsThinh Pham, Nguyen Phan Nguyen, Pratibha Zunjare, Weiyuan Chen et al.ICLR 2026 · 69 citations
- CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmarkJian Wu, Linyi Yang, Zhen Wang, Manabu Okumura et al.ICLR 2025
- MMQA: Evaluating LLMs with Multi-Table Multi-Hop Complex QuestionsJian Wu, Linyi Yang, Dongyuan Li, Yuliang Ji et al.ICLR 2025
