Human-Centered Exploration of Table Unionability
Nina Klimenkova, Sreeram Marimuthu, Roee Shraga
摘要
Table union search (TUS) identifies tables that can be meaningfully combined with a given query table and is a core task in data discovery over data lakes. Yet, what it means for two tables to be "unionable" is inherently ambiguous: domain experts may disagree even on seemingly simple cases, and existing benchmarks collapse this disagreement into binary labels, omitting the behavioral context behind human decisions. We take a human-centered view of table unionability and study how humans, traditional TUS methods, and large language models (LLMs) interact on this task. We introduce TUNE (Table UNionability with human Evaluation), a benchmark of 464 expert judgments over 26 table pairs that records binary decisions, confidence scores, decision times, interaction traces, textual explanations, and post-survey reflections. Using TUNE, we (i) characterize human performance, overconfidence, and metacognitive quality (calibration and resolution); (ii) benchmark state-of-the-art TUS methods (Starmie, SANTOS, D3L), revealing complementary strengths and systematic misalignment with human judgments; and (iii) evaluate four experimental scenarios that combine human behavioral signals and TUS features using classical ML models and LLMs. The best configuration we experimented with reaches 84% accuracy, improving over both human majority vote and the strong standalone TUS method, while LLMs act as useful second opinions but are sensitive to conflicting signals. Overall, our results suggest that unionability labels reflect a structured yet imperfect human decision process and that hybrid human-model (may it be traditional classifiers or LLMs) pipelines provide more reliable and interpretable unionability assessments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation LearningGrace Fan, Jin Wang, Yuliang Li, Dan Zhang 等VLDB 2023 · 被引用 139 次
- Dataset Discovery in Data LakesAlex Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos KonstantinouICDE 2020 · 被引用 118 次
- SANTOS: Relationship-based Semantic Table Union SearchAamod Khatiwada, Grace Fan, Roee Shraga, Zixuan Chen 等SIGMOD 2023 · 被引用 61 次
- ADnEV: Cross-Domain Schema Matching using Deep Similarity Matrix Adjustment and EvaluationRoee Shraga, Avigdor Gal, Haggai RoitmanVLDB 2020 · 被引用 37 次
- LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data LakesYuhao Deng, Chengliang Chai, Lei Cao, Qin Yuan 等VLDB 2024 · 被引用 36 次
相关 Paper
- CHORUS: Foundation Models for Unified Data Discovery and ExplorationMoe Kayali, Anton Lykov, Ilias Fountalis, Nikolaos Vasiloglou 等VLDB 2024 · 被引用 33 次
- T2R-BENCH: A Benchmark for Real World Table-to-Report TaskJie Zhang, Changzai Pan, Sishi Xiong, Kaiwen Wei 等EMNLP 2025 · 被引用 2 次
- Integrating Data Lake TablesAamod Khatiwada, Roee Shraga, Wolfgang Gatterbauer, Renée J. MillerVLDB 2023 · 被引用 59 次
- TopBench: A Benchmark for Implicit Predictive Reasoning in Tabular Question AnsweringAn-Yang Ji, Jun-Peng Jiang, De-Chuan Zhan, Han-Jia YeICML 2026 · 被引用 1 次
- LIFTus: An Adaptive Multi-Aspect Column Representation Learning for Table Union SearchErmu Qiu, Jun Gao, Yaofeng Tu, Jingru YangICDE 2025
