SDEcho: Efficient Explanation of Aggregated Sequence Difference
Fei Ye, Zikang Liu, Xi Zhang, Yinan Jing, Zhenying He, Yuxin Che, Haoran Xiong, Kai Zhang, X. Sean Wang
摘要
Understanding the reasons behind differences between aggregated sequences derived from SQL queries is crucial for data scientists. However, existing methods often suffer from being labor-intensive, lacking scalability, providing only approximate solutions, and inadequately supporting sequence difference explanations. In response, we introduce SDEcho, a novel framework designed to automate the explanation searching for sequence differences in high-dimensional and high-volume datasets. SDEcho utilizes advanced pruning techniques, considering pattern, order, and dimension perspectives, as well as their interactions, to prune the entire explanation space while maintaining explanations accurate and concise. This hybrid pruning approach significantly accelerates the explanation searching process, making SDEcho a valuable tool for data analysis tasks. Extensive experiments on synthetic and real-world datasets, along with a case study, demonstrate that SDEcho outperforms existing methods in terms of both effectiveness and efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- What went wrong and when? Instance-wise feature importance for time-series black-box modelsSana Tonekaboni, Shalmali Joshi, Kieran Campbell, David Duvenaud 等NeurIPS 2020 · 被引用 94 次
- Computing Local Sensitivities of Counting Queries with JoinsYuchao Tao, Xi He, Ashwin Machanavajjhala, Sudeepa RoySIGMOD 2020 · 被引用 37 次
- Table2Charts: Recommending Charts by Learning Shared Table RepresentationsMengyu Zhou, Qingtao Li, Xinyi He, Yuejiang Li 等KDD 2021 · 被引用 35 次
- Guided Exploration of Data SummariesBrit Youngmann, Sihem Amer-Yahia, Aurélien PersonnazVLDB 2022 · 被引用 22 次
- XInsight: eXplainable Data Analysis Through The Lens of CausalityPingchuan Ma, Rui Ding, Shuai Wang, Shi Han 等SIGMOD 2023 · 被引用 20 次
相关 Paper
- Summarized Causal Explanations For Aggregate ViewsBrit Youngmann, Michael J. Cafarella, Amir Gilad, Sudeepa RoySIGMOD 2024 · 被引用 12 次
- "What makes my queries slow?": Subgroup Discovery for SQL Workload AnalysisYoucef Remil, Anes Bendimerad, Romain Mathonat, Philippe Chaleat 等ASE 2021 · 被引用 12 次
- On Explaining Confounding BiasBrit Youngmann, Michael J. Cafarella, Yuval Moskovitch, Babak SalimiICDE 2023 · 被引用 7 次
- SAGA: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning ApplicationsShafaq Siddiqi, Roman Kern, Matthias BoehmSIGMOD 2024 · 被引用 24 次
- On Detecting Cherry-picked GeneralizationsYin Lin, Brit Youngmann, Yuval Moskovitch, H. V. Jagadish 等VLDB 2022 · 被引用 18 次
