Data Ambiguity Profiling for the Generation of Training Examples
Enzo Veltri, Gilbert Badaro, Mohammed Saeed, Paolo Papotti
摘要
Several applications, such as text-to-SQL and computational fact checking, exploit the relationship between relational data and natural language text. However, state of the art solutions simply fail in managing "data-ambiguity", i.e., the case when there are multiple interpretations of the relationship between text and data. Given the ambiguity in language, text can be mapped to different subsets of data, but existing training corpora only have examples in which every sentence/question is annotated precisely w.r.t. the relation. This unrealistic assumption leaves the target applications unable to handle ambiguous cases. To tackle this problem, we present an end-to-end solution that, given a table D, generates examples that consist of text, annotated with its data evidence, with factual ambiguities w.r.t. D. We formulate the problem of profiling relational tables to identify row and attribute data ambiguity. For the latter, we propose a deep learning method that identifies every pair of data ambiguous attributes and a label that describes both columns. Such metadata is then used to generate examples with data ambiguities for any input table. To enable scalability, we finally introduce a SQL approach that can generate millions of examples in seconds. We show the high accuracy of our solution in profiling relational tables and report on how our automatically generated examples lead to drastic quality improvements in two fact-checking applications, including a website with thousands of users, and in a text-to-SQL system.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Logical and Physical Optimizations for SQL Query Execution over Large Language ModelsDario Satriani, Enzo Veltri, Donatello Santoro, Sara Rosato 等SIGMOD 2025 · 被引用 7 次
- Generation of Training Examples for Tabular Natural Language InferenceJean-Flavien Bussotti, Enzo Veltri, Donatello Santoro, Paolo PapottiSIGMOD 2024 · 被引用 7 次
- PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?Jingzhe Xu, Rui Wang, Jiannan Wang, Guoliang LiVLDB 2026 · 被引用 3 次
- Unknown Claims: Generation of Fact-Checking Training Examples from Unstructured and Structured DataJean-Flavien Bussotti, Luca Ragazzi, Giacomo Frisoni, Gianluca Moro 等EMNLP 2024 · 被引用 3 次
- SQUAB: Evaluating LLM robustness to Ambiguous and Unanswerable Questions in Semantic ParsingSimone Papicchio, Luca Cagliero, Paolo PapottiEMNLP 2025
相关 Paper
- Benchmarking and Improving Text-to-SQL Generation under AmbiguityAdithya Bhaskar, Tushar Tomar, Ashutosh Sathe, Sunita SarawagiEMNLP 2023 · 被引用 13 次
- Test Data Generation for Complex SQL QueriesSunanda Somwase, Parismita Das, S. SudarshanSIGMOD 2026 · 被引用 1 次
- Text2sql-Flow: a Robust Sql-Aware Data Augmentation Framework for Text-To-SqlQifeng Cai, Hao Liang, Chang Xu, Tao Xie 等ICDE 2026 · 被引用 1 次
- DBPal: A Fully Pluggable NL2SQL Training PipelineNathaniel Weir, Prasetya Ajie Utama, Alex Galakatos, Andrew Crotty 等SIGMOD 2020 · 被引用 36 次
- SQL-Checker: Error Detection and Labeling for Text-to-SQL with Interpretability AnalysisXingyu Ma, Xin Tian, Lingxiang Wu, Xuepeng Wang 等WWW 2026
