Explaining Dataset Changes for Semantic Data Versioning with Explain-Da-V
Roee Shraga, Renée J. Miller
摘要
In multi-user environments in which data science and analysis is collaborative, multiple versions of the same datasets are generated. While managing and storing data versions has received some attention in the research literature, the semantic nature of such changes has remained under-explored. In this work, we introduce Explain-Da-V, a framework aiming to explain changes between two given dataset versions. Explain-Da-V generates explanations that use data transformations to explain changes. We further introduce a set of measures that evaluate the validity, generalizability, and explainability of these explanations. We empirically show, using an adapted existing benchmark and a newly created benchmark, that Explain-Da-V generates better explanations than existing data transformation synthesis methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Discovering Functional Dependencies through Hitting Set EnumerationTobias Bleifuß, Thorsten Papenbrock, Thomas Bläsius, Martin Schirneck 等SIGMOD 2024 · 被引用 9 次
- Weak-to-Strong Prompts with Lightweight-to-Powerful LLMs for High-Accuracy, Low-Cost, and Explainable Data TransformationChanglun Li, Chenyu Yang, Yuyu Luo, Ju Fan 等VLDB 2025 · 被引用 6 次
- Gen-T: Table Reclamation in Data LakesGrace Fan, Roee Shraga, Renée J. MillerICDE 2024 · 被引用 5 次
- TabulaX: Leveraging Large Language Models for Multi-Class Table TransformationsArash Dargahi Nobari, Davood RafieiVLDB 2025 · 被引用 3 次
- Retrieve-and-Verify: A Table Context Selection Framework for Accurate Column AnnotationsZhihao Ding, Yongkang Sun, Jieming ShiSIGMOD 2026 · 被引用 2 次
它引用的顶会 Paper12
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan 等CHI 2021 · 被引用 663 次
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- How do Data Science Workers Collaborate? Roles, Workflows, and ToolsAmy X. Zhang, Michael J. Muller, Dakuo WangCSCW 2020 · 被引用 260 次
- Understanding and Visualizing Data Iteration in Machine LearningFred Hohman, Kanit Wongsuphasawat, Mary Beth Kery, Kayur PatelCHI 2020 · 被引用 114 次
- Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science NotebooksCong Yan, Yeye HeSIGMOD 2020 · 被引用 64 次
相关 Paper
- Diff in the Loop: Supporting Data Comparison in Exploratory Data AnalysisApril Yi Wang, Will Epperson, Robert A. DeLine, Steven Mark DruckerCHI 2022 · 被引用 34 次
- SDEcho: Efficient Explanation of Aggregated Sequence DifferenceFei Ye, Zikang Liu, Xi Zhang, Yinan Jing 等VLDB 2025
- Compact, Tamper-Resistant Archival of Fine-Grained ProvenanceNan Zheng, Zack IvesVLDB 2021 · 被引用 6 次
- Explain the Synth: Interpretable Evaluation of LLM Data SynthesisYue Yang, Fan Yang, Yu Bai, Hao WangACL 2026
- Datamations: Animated Explanations of Data Analysis PipelinesXiaoying Pu, Sean Kross, Jake M. Hofman, Daniel G. GoldsteinCHI 2021 · 被引用 33 次
