Delving into Spectral Clustering with Vision-Language Representations
Bo Peng, Yuanwei Hu, Bo Liu, Ling Chen, Jie Lu, Zhen Fang
摘要
Spectral clustering is known as a powerful technique in unsupervised data analysis. The vast majority of approaches to spectral clustering are driven by a single modality, leaving the rich information in multi-modal representations untapped. Inspired by the recent success of vision-language pre-training, this paper enriches the landscape of spectral clustering from a single-modal to a multi-modal regime. Particularly, we propose Neural Tangent Kernel Spectral Clustering that leverages cross-modal alignment in pre-trained vision-language models. By anchoring the neural tangent kernel with positive nouns, i.e., those semantically close to the images of interest, we arrive at formulating the affinity between images as a coupling of their visual proximity and semantic overlap. We show that this formulation amplifies within-cluster connections while suppressing spurious ones across clusters, hence encouraging block-diagonal structures. In addition, we present a regularized affinity diffusion mechanism that adaptively ensembles affinity matrices induced by different prompts. Extensive experiments on 16 benchmarks---including classical, large-scale, fine-grained and domain-shifted datasets---manifest that our method consistently outperforms the state-of-the-art by a large margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Respecting Modality Gap in Post-hoc Out-of-distribution Detection with Pre-trained Vision-Language ModelsYuanwei Hu, Bo Peng, Yadan Luo, zhen fang 等ICML 2026
- MAGIC: Multi-Granularity Language-Informed Image ClusteringXiaohan Zhang, Chao Zhang, Chunlin Chen, Huaxiong LiICML 2026
- A Close Look at Negative Label Guided Out-of-distribution Detection in Pre-trained Vision-Language ModelsBo Peng, Jie Lu, zhen fang, Guangquan ZhangICML 2026
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
相关 Paper
- Semantic-Augmented Image Clustering via Adaptive Multi-Modal CollaborationXiaohan Zhang, Chao Zhang, Deng Xu, Hong Yu 等AAAI 2026
- Learning Neural Eigenfunctions for Unsupervised Semantic SegmentationZhijie Deng, Yucen LuoICCV 2023 · 被引用 7 次
- Multi-modal Alignment using Representation CodebookJiali Duan, Liqun Chen, Son Tran, Jinyu Yang 等CVPR 2022 · 被引用 56 次
- Task-Aware Clustering for Prompting Vision-Language ModelsFusheng Hao, Fengxiang He, Fuxiang Wu, Tichao Wang 等CVPR 2025
- Contrastive Learning is Spectral Clustering on Similarity GraphZhiquan Tan, Yifan Zhang, Jingqin Yang, Yang YuanICLR 2024 · 被引用 34 次
