Align2Concept: Language Guided Interpretable Image Recognition by Visual Prototype and Textual Concept Alignment
Jiaqi Wang, Pichao Wang, Yi Feng, Huafeng Liu, Chang Gao, Liping Jing
摘要
Most works of interpretable neural networks strive for learning the semantics concepts merely from single modal information such as images. However, humans usually learn semantic concepts from multiple modalities and the semantics is encoded by the brain from fused multi-modal information. Inspired by cognitive science and vision-language learning, we propose a Prototype-Concept Alignment Network (ProCoNet) for learning visual prototypes under the guidance of textual concepts. In the ProCoNet, we have designed a visual encoder to decompose the input image into regional features of prototypes, while also developing a prompt generation strategy that incorporates in-context learning to prompt large language models to generate textual concepts. To align visual prototypes with textual concepts, we leverage the multimodal space provided by the pre-trained CLIP as a bridge. Specifically, the regional features from the vision space and the cropped regions of prototypes encoded by CLIP reside on different but semantically highly correlated manifolds, i.e. follow a multi-manifold distribution. We transform the multi-manifold distribution alignment problem into optimizing the projection matrix by Cayley transform on the Stiefel manifold. Through the learned projection matrix, visual prototypes can be projected into the multimodal space to align with semantically similar textual concept features encoded by CLIP. We conducted two case studies on the CUB-200-2011 and Oxford Flower dataset. Our experiments show that the ProCoNet provides higher accuracy and better interpretability compared to the single-modality interpretable model. Furthermore, ProCoNet offers a level of interpretability not previously available in other interpretable methods.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- PRISM: Prototype-based Reasoning with Inter-modal Semantic Mining for Interpretable Image RecognitionAnni Yu, Yu-Bin YangCVPR 2026
- Prototype-Grounded Concept Models for Verifiable Concept AlignmentStefano Colamonaco, David Debot, Pietro Barbiero, Giuseppe MarraICML 2026
相关 Paper
- Concept-Guided Prompt Learning for Generalization in Vision-Language ModelsYi Zhang, Ce Zhang, Ke Yu, Yushun Tang 等AAAI 2024 · 被引用 37 次
- ADAPT: Vision-Language Navigation with Modality-Aligned Action PromptsBingqian Lin, Yi Zhu, Zicong Chen, Xiwen Liang 等CVPR 2022 · 被引用 45 次
- Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIPZhongxing Xu, Feilong Tang, Zhe Chen, Yingxue Su 等AAAI 2025 · 被引用 23 次
- Multimodal Prompt Alignment for Facial Expression RecognitionFuyan Ma, Yiran He, Bin Sun, Shutao LiICCV 2025 · 被引用 5 次
- V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept TokenizerHangzhou He, Lei Zhu, Xinliang Zhang, Shuang Zeng 等AAAI 2025 · 被引用 11 次
