Multi-Modal Proxy Learning Towards Personalized Visual Multiple Clustering
Jiawei Yao, Qi Qian, Juhua Hu
摘要
Multiple clustering has gained significant attention in recent years due to its potential to reveal multiple hidden structures of data from different perspectives. The advent of deep multiple clustering techniques has notably advanced the performance by uncovering complex patterns and relationships within large datasets. However, a major challenge arises as users often do not need all the clusterings that algorithms generate, and figuring out the one needed requires a substantial understanding of each clustering result. Traditionally, aligning a user's brief keyword of interest with the corresponding vision components was challenging, but the emergence of multi-modal and large language models (LLMs) has begun to bridge this gap. In response, given unlabeled target visual data, we propose Multi-MaP, a novel method employing a multi-modal proxy learning process. It leverages CLIP encoders to extract coherent text and image embeddings, with GPT-4 integrating users' interests to formulate effective textual contexts. Moreover, reference word constraint and concept-level constraint are designed to learn the optimal text proxy according to the user's interest. Multi-MaP not only adeptly captures a user's interest via a keyword but also facilitates identifying relevant clusterings. Our extensive experiments show that Multi-MaP consistently outperforms state-of-the-art methods in all benchmark multi-clustering vision tasks. Our code is available at https://github.com/Alexander-Yao/Multi-MaP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Customized Multiple Clustering via Multi-Modal Subspace Proxy LearningJiawei Yao, Qi Qian, Juhua HuNeurIPS 2024 · 被引用 17 次
- Conditional Representation Learning for Customized TasksHonglin Liu, Chao Sun, Peng Hu, Yunfan Li 等NeurIPS 2025 · 被引用 6 次
- FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion ModelZiyu Yao, Xuxin Cheng, Zhiqi HuangACM MM 2024 · 被引用 5 次
- ESMC: MLLM-Based Embedding Selection for Explainable Multiple ClusteringXinyue Wang, Yuheng Jia, Hui Liu, Junhui HouAAAI 2026 · 被引用 1 次
- Open Ad-hoc Categorization with Contextualized Feature LearningZilin Wang, Sangwoo Mo, Stella X. Yu, Sima Behpour 等CVPR 2025
它引用的顶会 Paper8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language ModelsManli Shu, Weili Nie, De-An Huang, Zhiding Yu 等NeurIPS 2022 · 被引用 603 次
- SoftTriple Loss: Deep Metric Learning Without Triplet SamplingQi Qian, Lei Shang, Baigui Sun, Juhua Hu 等ICCV 2019 · 被引用 419 次
- Multi-View Multiple Clusterings Using Deep Matrix FactorizationShaowei Wei, Jun Wang, Guoxian Yu, Carlotta Domeniconi 等AAAI 2020 · 被引用 87 次
相关 Paper
- MLLM Enriched Explainable Multiple ClusteringShan Zhang, Liangrui Ren, Qiaoyu Tan, Carlotta Domeniconi 等AAAI 2026
- Multi-modal Dynamic Proxy Learning for Personalized Multiple ClusteringJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li 等AAAI 2026
- Align2Concept: Language Guided Interpretable Image Recognition by Visual Prototype and Textual Concept AlignmentJiaqi Wang, Pichao Wang, Yi Feng, Huafeng Liu 等ACM MM 2024 · 被引用 1 次
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMsShengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma 等CVPR 2024 · 被引用 111 次
- Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image RecognitionYicheng Liu, Jie Wen, Chengliang Liu, Xiaozhao Fang 等ICML 2024 · 被引用 7 次
