Pruning the Paradox: How CLIP's Most Informative Heads Enhance Performance While Amplifying Bias
Avinash Madasu, Vasudev Lal, Phillip Howard
摘要
CLIP is one of the most popular foundation models and is heavily used for many visionlanguage tasks, yet little is known about its inner workings. As CLIP is increasingly deployed in real-world applications, it is becoming even more critical to understand its limitations and embedded social biases to mitigate potentially harmful downstream consequences. However, the question of what internal mechanisms drive both the impressive capabilities as well as problematic shortcomings of CLIP has largely remained unanswered. To bridge this gap, we study the conceptual consistency of text descriptions for attention heads in CLIPlike models. Specifically, we propose Concept Consistency Score (CCS), a novel interpretability metric that measures how consistently individual attention heads in CLIP models align with specific concepts. Our soft-pruning experiments reveal that high CCS heads are critical for preserving model performance, as pruning them leads to a significantly larger performance drop than pruning random or low CCS heads. Notably, we find that high CCS heads capture essential concepts and play a key role in out-ofdomain detection, concept-specific reasoning, and video-language understanding. Moreover, we prove that high CCS heads learn spurious correlations which amplify social biases. These results position CCS as a powerful interpretability metric exposing the paradox of performance and social biases in CLIP models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 被引用 629 次
- Interpreting CLIP's Image Representation via Text-Based DecompositionYossi Gandelsman, Alexei A. Efros, Jacob SteinhardtICLR 2024 · 被引用 179 次
- CLIPood: Generalizing CLIP to Out-of-DistributionsYang Shu, Xingzhuo Guo, Jialong Wu, Ximei Wang 等ICML 2023 · 被引用 122 次
- Rosetta Neurons: Mining the Common Units in a Model ZooAmil Dravid, Yossi Gandelsman, Alexei A. Efros, Assaf ShocherICCV 2023 · 被引用 46 次
- VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language TransformersEstelle Aflalo, Meng Du, Shao-Yen Tseng, Yongfei Liu 等CVPR 2022 · 被引用 34 次
相关 Paper
- Joint Vision-Language Social Bias Removal for CLIPHaoyu Zhang, Yangyang Guo, Mohan S. KankanhalliCVPR 2025
- Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance ApproachAishwarya Agarwal, Srikrishna Karanam, Vineet GandhiCVPR 2026 · 被引用 2 次
- From Weights to Concepts: Data-Free Interpretability of CLIP via Singular Vector DecompositionFrancesco Gentile, Nicola DallAsen, Francesco Tonini, Massimiliano Mancini 等CVPR 2026
- Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)Usha Bhalla, Alex Oesterling, Suraj Srinivas, Flávio P. Calmon 等NeurIPS 2024 · 被引用 146 次
- WRING Out The Bias: A Rotation-Based Alternative To Projection DebiasingWalter Gerych, Cassandra Parent, Quinn Perian, Rafiya Javed 等ICLR 2026
