Domino: Discovering Systematic Errors with Cross-Modal Embeddings
Sabri Eyuboglu, Maya Varma, Khaled Kamal Saab, Jean-Benoit Delbrouck, Christopher Lee-Messer, Jared Dunnmon, James Zou, Christopher Ré
Abstract
Machine learning models that achieve high overall accuracy often make systematic errors on important subsets (or slices) of data. Identifying underperforming slices is particularly challenging when working with high-dimensional inputs (e.g. images, audio), where important slices are often unlabeled. In order to address this issue, recent studies have proposed automated slice discovery methods (SDMs), which leverage learned model representations to mine input data for slices on which a model performs poorly. To be useful to a practitioner, these methods must identify slices that are both underperforming and coherent (i.e. united by a human-understandable concept). However, no quantitative evaluation framework currently exists for rigorously assessing SDMs with respect to these criteria. Additionally, prior qualitative evaluations have shown that SDMs often identify slices that are incoherent. In this work, we address these challenges by first designing a principled evaluation framework that enables a quantitative comparison of SDMs across 1,235 slice discovery settings in three input domains (natural images, medical images, and time-series data). Then, motivated by the recent development of powerful cross-modal representation learning approaches, we present Domino, an SDM that leverages cross-modal embeddings and a novel error-aware mixture model to discover and describe coherent slices. We find that Domino accurately identifies 36% of the 1,235 slices in our framework - a 12 percentage point improvement over prior methods. Further, Domino is the first SDM that can provide natural language descriptions of identified slices, correctly generating the exact name of the slice in 35% of settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 035116cf-590e-4c19-9bcd-2e1eb915a87dCited by top-tier papers47
- Zoology: Measuring and Improving Recall in Efficient Language ModelsSimran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson et al.ICLR 2024 · 140 citations
- MetaShift: A Dataset of Datasets for Evaluating Contextual Distribution Shifts and Training ConflictsWeixin Liang, James ZouICLR 2022 · 103 citations
- Assaying Out-Of-Distribution Generalization in Transfer LearningFlorian Wenzel, Andrea Dittadi, Peter V. Gehler, Carl-Johann Simon-Gabriel et al.NeurIPS 2022 · 93 citations
- LANCE: Stress-testing Visual Models by Generating Language-guided Counterfactual ImagesViraj Prabhu, Sriram Yenamandra, Prithvijit Chattopadhyay, Judy HoffmanNeurIPS 2023 · 59 citations
- Zeno: An Interactive Framework for Behavioral Evaluation of Machine LearningÁngel Alexander Cabrera, Erica Fu, Donald Bertucci, Kenneth Holstein et al.CHI 2023 · 51 citations
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li et al.NeurIPS 2020 · 390 citations
- No Subclass Left Behind: Fine-Grained Robustness in Coarse-Grained Classification ProblemsNimit Sharad Sohoni, Jared Dunnmon, Geoffrey Angus, Albert Gu et al.NeurIPS 2020 · 316 citations
Related papers
- What Is Wrong with My Model? Identifying Systematic Problems with Semantic Data SlicingChenyang Yang, Yining Hong, Grace A. Lewis, Tongshuang Wu et al.ASE 2024 · 2 citations
- Error Slice Discovery via Manifold CompactnessHan Yu, Hao Zou, Jiashuo Liu, Renzhe Xu et al.AAAI 2026 · 2 citations
- CB-SLICE: Concept-Based Interpretable Error Slice DiscoveryYael Konforti, Mateo Espinosa Zarlenga, Elaf Almahmoud, Mateja JamnikICML 2026
- Error Discovery By Clustering Influence EmbeddingsFulton Wang, Julius Adebayo, Sarah Tan, Diego Garcia-Olano et al.NeurIPS 2023 · 10 citations
- VLSlice: Interactive Vision-and-Language Slice DiscoveryEric Slyman, Minsuk Kahng, Stefan LeeICCV 2023 · 11 citations
