LICO: Explainable Models with Language-Image COnsistency
Yiming Lei, Zilong Li, Yangyang Li, Junping Zhang, Hongming Shan
Abstract
Interpreting the decisions of deep learning models has been actively studied since the explosion of deep neural networks. One of the most convincing interpretation approaches is salience-based visual interpretation, such as Grad-CAM, where the generation of attention maps depends merely on categorical labels. Although existing interpretation methods can provide explainable decision clues, they often yield partial correspondence between image and saliency maps due to the limited discriminative information from one-hot labels. This paper develops a Language-Image COnsistency model for explainable image classification, termed LICO, by correlating learnable linguistic prompts with corresponding visual features in a coarse-to-fine manner. Specifically, we first establish a coarse global manifold structure alignment by minimizing the distance between the distributions of image and language features. We then achieve fine-grained saliency maps by applying optimal transport (OT) theory to assign local feature maps with class-specific prompts. Extensive experimental results on eight benchmark datasets demonstrate that the proposed LICO achieves a significant improvement in generating more explainable attention maps in conjunction with existing interpretation methods such as Grad-CAM. Remarkably, LICO improves the classification performance of existing models without introducing any computational overhead during inference. Source code is made available at https://github.com/ymLeiFDU/LICO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Evidential Copula Concept Embedding ModelsYanjie Qiu, Xiaodong Yue, Xuhui Fan, Yufei Chen et al.ICML 2026 · 9 citations
- Denoising Diffusion Path: Attribution Noise Reduction with An Auxiliary Diffusion ModelYiming Lei, Zilong Li, Junping Zhang, Hongming ShanNeurIPS 2024 · 9 citations
- Narrowing Information Bottleneck Theory for Multimodal Image-Text Representations InterpretabilityZhiyu Zhu, Zhibo Jin, Jiayu Zhang, Nan Yang et al.ICLR 2025
- REPEAT: Improving Uncertainty Estimation in Representation Learning ExplainabilityKristoffer K. Wickstrøm, Thea Brüsch, Michael C. Kampffmeyer, Robert JenssenAAAI 2025
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo LabelingBowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu et al.NeurIPS 2021 · 1,389 citations
- Graph Optimal Transport for Cross-Domain AlignmentLiqun Chen, Zhe Gan, Yu Cheng, Linjie Li et al.ICML 2020 · 193 citations
Related papers
- Weakly Supervised Referring Image Segmentation with Intra-Chunk and Inter-Chunk ConsistencyJungbeom Lee, Sungjin Lee, Jinseok Nam, Seunghak Yu et al.ICCV 2023 · 28 citations
- Explainable Models with Consistent InterpretationsVipin Pillai, Hamed PirsiavashAAAI 2021 · 46 citations
- Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment RewardShizhan Gong, Minda Hu, Qiyuan Zhang, Chen Ma et al.CVPR 2026 · 1 citation
- Consistent Explanations by Contrastive LearningVipin Pillai, Soroush Abbasi Koohpayegani, Ashley Ouligian, Dennis Fong et al.CVPR 2022 · 15 citations
- Improved Visual Grounding through Self-Consistent ExplanationsRuozhen He, Paola Cascante-Bonilla, Ziyan Yang, Alexander C. Berg et al.CVPR 2024
