Bridging Explainability and Embeddings: BEE Aware of Spuriousness
Cristian Daniel Paduraru, Antonio Barbalau, Radu Filipescu, Andrei Liviu Nicolicioiu, Elena Burceanu
Abstract
Current methods for detecting spurious correlations rely on analyzing dataset statistics or error patterns, leaving many harmful shortcuts invisible when counterexamples are absent. We introduce BEE (Bridging Explainability and Embeddings), a framework that shifts the focus from model predictions to the weight space, and to the embedding geometry underlying decisions. By analyzing how fine-tuning perturbs pretrained representations, BEE uncovers spurious correlations that remain hidden from conventional evaluation pipelines. We use linear probing as a transparent diagnostic lens, revealing spurious features that not only persist after full fine-tuning but also transfer across diverse state-of-the-art models. Our experiments cover numerous datasets and domains: vision (Waterbirds, CelebA, ImageNet-1k), language (CivilComments, MIMIC-CXR medical notes), and multiple embedding families (CLIP, CLIP-DataComp.XL, mGTE, BLIP2, SigLIP2). BEE consistently exposes spurious correlations: from concepts that slash the ImageNet accuracy by up to 95%, to clinical shortcuts in MIMIC-CXR notes that induce dangerous false negatives. Together, these results position BEE as a general and principled tool for diagnosing spurious correlations in weight space, enabling principled dataset auditing and more trustworthy foundation models. Code publicly available HERE. INTRODUCTION AND BACKGROUND Deep neural networks, and especially fine-tuned versions of foundation models, are increasingly deployed in critical areas such as healthcare, finance, and criminal justice, where decisions based on spurious correlations (SCs) can have severe societal consequences (Angwin et al., 2016; Caliskan et al., 2017) . Even if a pretrained model has been validated by the community, the dataset leveraged
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 31759657-e617-4498-88a1-20310e1a7488Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- Leveraging Sparse Linear Layers for Debuggable Deep NetworksEric Wong, Shibani Santurkar, Aleksander MadryICML 2021 · 101 citations
- NeuronTune: Towards Self-Guided Spurious Bias MitigationGuangtao Zheng, Wenqian Ye, Aidong ZhangICML 2025
- Mitigating Spurious Correlations in Multi-modal Models during Fine-tuningYu Yang, Besmira Nushi, Hamid Palangi, Baharan MirzasoleimanICML 2023 · 65 citations
- Diagnosing Transformers: Illuminating Feature Spaces for Clinical Decision-MakingAliyah R. Hsu, Yeshwanth Cherapanamjeri, Briton Park, Tristan Naumann et al.ICLR 2024 · 1 citation
- Common Sense Bias Modeling for Classification TasksMiao Zhang, Zee Fryer, Ben Colman, Ali Shahriyari et al.AAAI 2025
