Data Ambiguity Profiling for the Generation of Training Examples
Enzo Veltri, Gilbert Badaro, Mohammed Saeed, Paolo Papotti
Abstract
Several applications, such as text-to-SQL and computational fact checking, exploit the relationship between relational data and natural language text. However, state of the art solutions simply fail in managing "data-ambiguity", i.e., the case when there are multiple interpretations of the relationship between text and data. Given the ambiguity in language, text can be mapped to different subsets of data, but existing training corpora only have examples in which every sentence/question is annotated precisely w.r.t. the relation. This unrealistic assumption leaves the target applications unable to handle ambiguous cases. To tackle this problem, we present an end-to-end solution that, given a table D, generates examples that consist of text, annotated with its data evidence, with factual ambiguities w.r.t. D. We formulate the problem of profiling relational tables to identify row and attribute data ambiguity. For the latter, we propose a deep learning method that identifies every pair of data ambiguous attributes and a label that describes both columns. Such metadata is then used to generate examples with data ambiguities for any input table. To enable scalability, we finally introduce a SQL approach that can generate millions of examples in seconds. We show the high accuracy of our solution in profiling relational tables and report on how our automatically generated examples lead to drastic quality improvements in two fact-checking applications, including a website with thousands of users, and in a text-to-SQL system.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6e034a2c-cfb1-49fd-8399-e43c2fdaa084Cited by top-tier papers5
- Logical and Physical Optimizations for SQL Query Execution over Large Language ModelsDario Satriani, Enzo Veltri, Donatello Santoro, Sara Rosato et al.SIGMOD 2025 · 7 citations
- Generation of Training Examples for Tabular Natural Language InferenceJean-Flavien Bussotti, Enzo Veltri, Donatello Santoro, Paolo PapottiSIGMOD 2024 · 7 citations
- PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?Jingzhe Xu, Rui Wang, Jiannan Wang, Guoliang LiVLDB 2026 · 3 citations
- Unknown Claims: Generation of Fact-Checking Training Examples from Unstructured and Structured DataJean-Flavien Bussotti, Luca Ragazzi, Giacomo Frisoni, Gianluca Moro et al.EMNLP 2024 · 3 citations
- SQUAB: Evaluating LLM robustness to Ambiguous and Unanswerable Questions in Semantic ParsingSimone Papicchio, Luca Cagliero, Paolo PapottiEMNLP 2025
Related papers
- Benchmarking and Improving Text-to-SQL Generation under AmbiguityAdithya Bhaskar, Tushar Tomar, Ashutosh Sathe, Sunita SarawagiEMNLP 2023 · 13 citations
- Test Data Generation for Complex SQL QueriesSunanda Somwase, Parismita Das, S. SudarshanSIGMOD 2026 · 1 citation
- Text2sql-Flow: a Robust Sql-Aware Data Augmentation Framework for Text-To-SqlQifeng Cai, Hao Liang, Chang Xu, Tao Xie et al.ICDE 2026 · 1 citation
- DBPal: A Fully Pluggable NL2SQL Training PipelineNathaniel Weir, Prasetya Ajie Utama, Alex Galakatos, Andrew Crotty et al.SIGMOD 2020 · 36 citations
- SQL-Checker: Error Detection and Labeling for Text-to-SQL with Interpretability AnalysisXingyu Ma, Xin Tian, Lingxiang Wu, Xuepeng Wang et al.WWW 2026
