SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging
Svetlana Sagadeeva, Matthias Boehm
摘要
Slice finding---a recent work on debugging machine learning (ML) models---aims to find the top-K data slices (e.g., conjunctions of predicates such as gender female and degree PhD), where a trained model performs significantly worse than on the entire training/test data. These slices may be used to acquire more data for the problematic subset, add rules, or otherwise improve the model. In contrast to decision trees, the general slice finding problem allows for overlapping slices. The resulting search space is huge as it covers all subsets of features and their distinct values. Hence, existing work primarily relies on heuristics and focuses on small datasets that fit in memory of a single node. In this paper, we address these scalability limitations of slice finding in a holistic manner from both algorithmic and system perspectives. We leverage monotonicity properties of slice sizes, errors and resulting scores to facilitate effective pruning. Additionally, we present an elegant linear-algebra-based enumeration algorithm, which allows for fast enumeration and automatic parallelization on top of existing ML systems. Experiments with different real-world regression and classification datasets show that effective pruning and efficient sparse linear algebra renders exact enumeration feasible, even for datasets with many features, correlations, and data sizes beyond single node memory.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Domino: Discovering Systematic Errors with Cross-Modal EmbeddingsSabri Eyuboglu, Maya Varma, Khaled Kamal Saab, Jean-Benoit Delbrouck 等ICLR 2022 · 被引用 178 次
- Mandoline: Model Evaluation under Distribution ShiftMayee F. Chen, Karan Goel, Nimit Sharad Sohoni, Fait Poms 等ICML 2021 · 被引用 84 次
- Interpretable Data-Based Explanations for Fairness DebuggingRomila Pradhan, Jiongli Zhu, Boris Glavic, Babak SalimiSIGMOD 2022 · 被引用 53 次
- SAGA: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning ApplicationsShafaq Siddiqi, Roman Kern, Matthias BoehmSIGMOD 2024 · 被引用 24 次
- Towards Observability for Production Machine Learning Pipelines [Vision]Shreya Shankar, Aditya G. ParameswaranVLDB 2022 · 被引用 21 次
它引用的顶会 Paper5
- Towards Automated Neural Interaction Discovery for Click-Through Rate PredictionQingquan Song, Dehua Cheng, Hanning Zhou, Jiyan Yang 等KDD 2020 · 被引用 63 次
- Cerebro: A Data System for Optimized Deep Learning Model SelectionSupun Nakandala, Yuhao Zhang, Arun KumarVLDB 2020 · 被引用 61 次
- A Tensor Compiler for Unified Machine Learning Prediction ServingSupun Nakandala, Karla Saur, Gyeong-In Yu, Konstantinos Karanasos 等OSDI 2020 · 被引用 60 次
- Identifying Insufficient Data Coverage in Databases with Multiple RelationsYin Lin, Yifan Guan, Abolfazl Asudeh, H. V. JagadishVLDB 2020 · 被引用 50 次
- Getting Swole: Generating Access-Aware Code with Predicate PullupsAndrew Crotty, Alex Galakatos, Tim KraskaICDE 2020 · 被引用 11 次
相关 Paper
- SliceTeller: A Data Slice-Driven Approach for Machine Learning Model ValidationXiaoyu Zhang, Jorge Piazentin Ono, Huan Song, Liang Gou 等IEEE VIS 2022 · 被引用 43 次
- Error Discovery By Clustering Influence EmbeddingsFulton Wang, Julius Adebayo, Sarah Tan, Diego Garcia-Olano 等NeurIPS 2023 · 被引用 10 次
- What Is Wrong with My Model? Identifying Systematic Problems with Semantic Data SlicingChenyang Yang, Yining Hong, Grace A. Lewis, Tongshuang Wu 等ASE 2024 · 被引用 2 次
- Explaining mispredictions of machine learning models using rule inductionJürgen Cito, Isil Dillig, Seohyun Kim, Vijayaraghavan Murali 等FSE 2021 · 被引用 26 次
- FACTS: First Amplify Correlations and Then Slice to Discover BiasSriram Yenamandra, Pratik Ramesh, Viraj Prabhu, Judy HoffmanICCV 2023 · 被引用 33 次
