Customized Multiple Clustering via Multi-Modal Subspace Proxy Learning
Jiawei Yao, Qi Qian, Juhua Hu
Abstract
Multiple clustering aims to discover various latent structures of data from different aspects. Deep multiple clustering methods have achieved remarkable performance by exploiting complex patterns and relationships in data. However, existing works struggle to flexibly adapt to diverse user-specific needs in data grouping, which may require manual understanding of each clustering. To address these limitations, we introduce Multi-Sub, a novel end-to-end multiple clustering approach that incorporates a multi-modal subspace proxy learning framework in this work. Utilizing the synergistic capabilities of CLIP and GPT-4, Multi-Sub aligns textual prompts expressing user preferences with their corresponding visual representations. This is achieved by automatically generating proxy words from large language models that act as subspace bases, thus allowing for the customized representation of data in terms specific to the user's interests. Our method consistently outperforms existing baselines across a broad set of datasets in visual multiple clustering tasks. Our code is available at https://github.com/Alexander-Yao/Multi-Sub.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e8cdcff-f747-4b47-99bd-1d8426fbcfc6Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language ModelsManli Shu, Weili Nie, De-An Huang, Zhiding Yu et al.NeurIPS 2022 · 603 citations
- SoftTriple Loss: Deep Metric Learning Without Triplet SamplingQi Qian, Lei Shang, Baigui Sun, Juhua Hu et al.ICCV 2019 · 419 citations
Related papers
- Multi-Modal Proxy Learning Towards Personalized Visual Multiple ClusteringJiawei Yao, Qi Qian, Juhua HuCVPR 2024 · 19 citations
- ESMC: MLLM-Based Embedding Selection for Explainable Multiple ClusteringXinyue Wang, Yuheng Jia, Hui Liu, Junhui HouAAAI 2026 · 1 citation
- MLLM Enriched Explainable Multiple ClusteringShan Zhang, Liangrui Ren, Qiaoyu Tan, Carlotta Domeniconi et al.AAAI 2026
- Multi-modal Dynamic Proxy Learning for Personalized Multiple ClusteringJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li et al.AAAI 2026
- Align2Concept: Language Guided Interpretable Image Recognition by Visual Prototype and Textual Concept AlignmentJiaqi Wang, Pichao Wang, Yi Feng, Huafeng Liu et al.ACM MM 2024 · 1 citation
