Explaining Dataset Changes for Semantic Data Versioning with Explain-Da-V
Roee Shraga, Renée J. Miller
Abstract
In multi-user environments in which data science and analysis is collaborative, multiple versions of the same datasets are generated. While managing and storing data versions has received some attention in the research literature, the semantic nature of such changes has remained under-explored. In this work, we introduce Explain-Da-V, a framework aiming to explain changes between two given dataset versions. Explain-Da-V generates explanations that use data transformations to explain changes. We further introduce a set of measures that evaluate the validity, generalizability, and explainability of these explanations. We empirically show, using an adapted existing benchmark and a newly created benchmark, that Explain-Da-V generates better explanations than existing data transformation synthesis methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3087eccd-fbb9-464f-9133-07f68fbc4bcbCited by top-tier papers6
- Discovering Functional Dependencies through Hitting Set EnumerationTobias Bleifuß, Thorsten Papenbrock, Thomas Bläsius, Martin Schirneck et al.SIGMOD 2024 · 9 citations
- Weak-to-Strong Prompts with Lightweight-to-Powerful LLMs for High-Accuracy, Low-Cost, and Explainable Data TransformationChanglun Li, Chenyu Yang, Yuyu Luo, Ju Fan et al.VLDB 2025 · 6 citations
- Gen-T: Table Reclamation in Data LakesGrace Fan, Roee Shraga, Renée J. MillerICDE 2024 · 5 citations
- TabulaX: Leveraging Large Language Models for Multi-Class Table TransformationsArash Dargahi Nobari, Davood RafieiVLDB 2025 · 3 citations
- Retrieve-and-Verify: A Table Context Selection Framework for Accurate Column AnnotationsZhihao Ding, Yongkang Sun, Jieming ShiSIGMOD 2026 · 2 citations
Builds on12
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan et al.CHI 2021 · 663 citations
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- How do Data Science Workers Collaborate? Roles, Workflows, and ToolsAmy X. Zhang, Michael J. Muller, Dakuo WangCSCW 2020 · 260 citations
- Understanding and Visualizing Data Iteration in Machine LearningFred Hohman, Kanit Wongsuphasawat, Mary Beth Kery, Kayur PatelCHI 2020 · 114 citations
- Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science NotebooksCong Yan, Yeye HeSIGMOD 2020 · 64 citations
Related papers
- Diff in the Loop: Supporting Data Comparison in Exploratory Data AnalysisApril Yi Wang, Will Epperson, Robert A. DeLine, Steven Mark DruckerCHI 2022 · 34 citations
- SDEcho: Efficient Explanation of Aggregated Sequence DifferenceFei Ye, Zikang Liu, Xi Zhang, Yinan Jing et al.VLDB 2025
- Compact, Tamper-Resistant Archival of Fine-Grained ProvenanceNan Zheng, Zack IvesVLDB 2021 · 6 citations
- Explain the Synth: Interpretable Evaluation of LLM Data SynthesisYue Yang, Fan Yang, Yu Bai, Hao WangACL 2026
- Datamations: Animated Explanations of Data Analysis PipelinesXiaoying Pu, Sean Kross, Jake M. Hofman, Daniel G. GoldsteinCHI 2021 · 33 citations
