Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
Jeonghyeon Kim, Sangheum Hwang
摘要
Prior research on out-of-distribution detection (OoDD) has primarily focused on single-modality models. Recently, with the advent of large-scale pretrained vision-language models such as CLIP, OoDD methods utilizing such multimodal representations through zero-shot and prompt learning strategies have emerged. However, these methods typically involve either freezing the pretrained weights or only partially tuning them, which can be suboptimal for downstream datasets. In this paper, we highlight that multimodal fine-tuning (MMFT) can achieve notable OoDD performance. Despite some recent works demonstrating the impact of fine-tuning methods for OoDD, there remains significant potential for performance improvement. We investigate the limitation of naïve fine-tuning methods, examining why they fail to fully leverage the pretrained knowledge. Our empirical analysis suggests that this issue could stem from the modality gap within in-distribution (ID) embeddings. To address this, we propose a training objective that enhances cross-modal alignment by regularizing the distances between image and text embeddings of ID data. This adjustment helps in better utilizing pretrained textual information by aligning similar semantics from different modalities (i.e., text and image) more closely in the hyperspherical representation space. We theoretically demonstrate that the proposed regularization corresponds to the maximum likelihood estimation of an energy-based model on a hypersphere. Utilizing ImageNet-1k OoD benchmark datasets, we show that our method, combined with posthoc OoDD approaches leveraging pretrained knowledge (e.g., NegLabel), significantly outperforms existing methods, achieving state-of-the-art OoDD performance and leading ID accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and ReasoningWenjie Zhu, Yabin Zhang, Xin Jin, Wenjun Zeng 等CVPR 2026 · 被引用 12 次
- T-REGS: Minimum Spanning Tree Regularization for Self-Supervised LearningJulie Mordacq, David Loiseaux, Vicky Kalogeiton, Steve OudotNeurIPS 2025 · 被引用 2 次
- Activation Matters: Test-time Activated Negative Labels for OOD Detection with Vision-Language ModelsYabin Zhang, Maya Varma, Yunhe Gao, Jean-Benoit Delbrouck 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 被引用 2,213 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma 等ICLR 2022 · 被引用 911 次
相关 Paper
- Amend to Alignment: Decoupled Prompt Tuning for Mitigating Spurious Correlation in Vision-Language ModelsJie Zhang, Xiaosong Ma, Song Guo, Peng Li 等ICML 2024 · 被引用 10 次
- LoCoOp: Few-Shot Out-of-Distribution Detection via Prompt LearningAtsuyuki Miyai, Qing Yu, Go Irie, Kiyoharu AizawaNeurIPS 2023 · 被引用 174 次
- Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image RecognitionYicheng Liu, Jie Wen, Chengliang Liu, Xiaozhao Fang 等ICML 2024 · 被引用 7 次
- A Novel Fine-Tuned CLIP-OOD Detection Method with Double Loss Constraint Through Optimal Transport Semantic AlignmentHengyang Lu, Xin Guo, Shuai Feng, Wenyu Jiang 等AAAI 2026
- Mitigate the Gap: Improving Cross-Modal Alignment in CLIPSedigheh Eslami, Gerard de MeloICLR 2025 · 被引用 1 次
