Revisiting Single-Table Retrieval: An Open Problem Under 360° Stress Tests
Chenyu Yang, Ziyu Jiang, Junhao Li, Yuyu Luo, Ju Fan, Nan Tang
Abstract
Single-table retrieval (STR)-selecting the most relevant table from a data lake to answer a natural language question-remains an open challenge despite recent progress in dense retrieval. A conclusive assessment remains out of reach because (i) publicly available datasets are limited to fully reflect real-world complexity, (ii) experimental pipelines lack standardization, preventing fair comparison, and (iii) evaluations focus narrowly on standard performance on a fixed data set, overlooking robustness, generalization, and near-duplicate handling. We conduct the first systematic, large-scale investigation of dense retrieval for STR and contribute four advances: (1) a unified pipeline that factorizes STR into four design spaces: encoder architecture, model training, table structure encoding, and linearization; (2) TR360, a modular testbed that instantiates this pipeline and enables multi-angle stress testing under realistic, evolving data-lake scenarios; (3) TRBench, a benchmark that fuses six heterogeneous corpora, supplies realistic questions with varying reasoning depth, and includes a synthetic-data generator for rapid domain adaptation; and (4) the largest empirical study to date, spanning over 10 models, which distills actionable guidelines while revealing persistent weaknesses in generalization, robustness, and near-duplicate discrimination. Our results demonstrate that, despite notable advances, singletable retrieval remains unsolved, with persistent weaknesses in generalization, robustness, and the discrimination of nearduplicate tables. All code, data, and evaluation scripts are at https://github.com/CTXY/table_retrieval.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get fd23c303-92a2-4560-901c-12fad16aa280Related papers
- T2R-BENCH: A Benchmark for Real World Table-to-Report TaskJie Zhang, Changzai Pan, Sishi Xiong, Kaiwen Wei et al.EMNLP 2025 · 2 citations
- REaR : Retrieve, Expand and Refine for Effective Multitable RetrievalRishita Agarwal, Himanshu Singhal, Peter Baile Chen, Manan Roy Choudhury et al.ACL 2026 · 2 citations
- How Far Can LLM Agents Reason with Tables? Benchmarking Multi-Turn Agentic Table Question Answering in the WildJingwang Huang, Jie Zhang, Haoyang Zeng, Changzai Pan et al.ICML 2026
- CRAFT: Training-Free Cascaded Retrieval for Tabular QAAdarsh Singh, Kushal Raj Bhandari, Jianxi Gao, Soham Dan et al.ACL 2026 · 2 citations
- SURE or Not? Investigating Semantic Understanding in Dense Retrieval ModelsLingdi Kong, Xuanang Chen, Ben He, Le SunACL 2026
