SDEcho: Efficient Explanation of Aggregated Sequence Difference
Fei Ye, Zikang Liu, Xi Zhang, Yinan Jing, Zhenying He, Yuxin Che, Haoran Xiong, Kai Zhang, X. Sean Wang
Abstract
Understanding the reasons behind differences between aggregated sequences derived from SQL queries is crucial for data scientists. However, existing methods often suffer from being labor-intensive, lacking scalability, providing only approximate solutions, and inadequately supporting sequence difference explanations. In response, we introduce SDEcho, a novel framework designed to automate the explanation searching for sequence differences in high-dimensional and high-volume datasets. SDEcho utilizes advanced pruning techniques, considering pattern, order, and dimension perspectives, as well as their interactions, to prune the entire explanation space while maintaining explanations accurate and concise. This hybrid pruning approach significantly accelerates the explanation searching process, making SDEcho a valuable tool for data analysis tasks. Extensive experiments on synthetic and real-world datasets, along with a case study, demonstrate that SDEcho outperforms existing methods in terms of both effectiveness and efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5294ef9e-22a1-467d-bd8c-60ec7ddee34cBuilds on13
- What went wrong and when? Instance-wise feature importance for time-series black-box modelsSana Tonekaboni, Shalmali Joshi, Kieran Campbell, David Duvenaud et al.NeurIPS 2020 · 94 citations
- Computing Local Sensitivities of Counting Queries with JoinsYuchao Tao, Xi He, Ashwin Machanavajjhala, Sudeepa RoySIGMOD 2020 · 37 citations
- Table2Charts: Recommending Charts by Learning Shared Table RepresentationsMengyu Zhou, Qingtao Li, Xinyi He, Yuejiang Li et al.KDD 2021 · 35 citations
- Guided Exploration of Data SummariesBrit Youngmann, Sihem Amer-Yahia, Aurélien PersonnazVLDB 2022 · 22 citations
- XInsight: eXplainable Data Analysis Through The Lens of CausalityPingchuan Ma, Rui Ding, Shuai Wang, Shi Han et al.SIGMOD 2023 · 20 citations
Related papers
- Summarized Causal Explanations For Aggregate ViewsBrit Youngmann, Michael J. Cafarella, Amir Gilad, Sudeepa RoySIGMOD 2024 · 12 citations
- "What makes my queries slow?": Subgroup Discovery for SQL Workload AnalysisYoucef Remil, Anes Bendimerad, Romain Mathonat, Philippe Chaleat et al.ASE 2021 · 12 citations
- On Explaining Confounding BiasBrit Youngmann, Michael J. Cafarella, Yuval Moskovitch, Babak SalimiICDE 2023 · 7 citations
- SAGA: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning ApplicationsShafaq Siddiqi, Roman Kern, Matthias BoehmSIGMOD 2024 · 24 citations
- On Detecting Cherry-picked GeneralizationsYin Lin, Brit Youngmann, Yuval Moskovitch, H. V. Jagadish et al.VLDB 2022 · 18 citations
