Collaborative Consortium of Foundation Models for Open-World Few-Shot Learning
Shuai Shao, Yu Bai, Yan Wang, Baodi Liu, Bin Liu
Abstract
Open-World Few-Shot Learning (OFSL) is a crucial research field dedicated to accurately identifying target samples in scenarios where data is limited and labels are unreliable. This research holds significant practical implications and is highly relevant to real-world applications. Recently, the advancements in foundation models like CLIP and DINO have showcased their robust representation capabilities even in resource-constrained settings with scarce data. This realization has brought about a transformative shift in focus, moving away from "building models from scratch" towards "effectively harnessing the potential of foundation models to extract pertinent prior knowledge suitable for OFSL and utilizing it sensibly". Motivated by this perspective, we introduce the Collaborative Consortium of Foundation Models (CO3), which leverages CLIP, DINO, GPT-3, and DALL-E to collectively address the OFSL problem. CO3 comprises four key blocks: (1) the Label Correction Block (LC-Block) corrects unreliable labels, (2) the Data Augmentation Block (DA-Block) enhances available data, (3) the Feature Extraction Block (FE-Block) extracts multi-modal features, and (4) the Text-guided Fusion Adapter (TeFu-Adapter) integrates multiple features while mitigating the impact of noisy labels through semantic constraints. Only the adapter's parameters are adjustable, while the others remain frozen. Through collaboration among these foundation models, CO3 effectively unlocks their potential and unifies their capabilities to achieve state-of-the-art performance on multiple benchmark datasets. https://github.com/The-Shuai/CO3 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be038044-5f89-4a86-9327-e27424f3666aCited by top-tier papers4
- DePT: Decoupled Prompt TuningJi Zhang, Shihan Wu, Lianli Gao, Heng Tao Shen et al.CVPR 2024 · 36 citations
- FAST: A Dual-tier Few-Shot Learning Paradigm for Whole Slide Image ClassificationKexue Fu, Xiaoyuan Luo, Linhao Qu, Shuo Wang et al.NeurIPS 2024 · 12 citations
- Towards Effective Foundation Model Adaptation for Extreme Cross-Domain Few-Shot LearningFei Zhou, Peng Wang, Lei Zhang, Wei Wei et al.ICCV 2025 · 2 citations
- Rank-guided Diffusion for Noise Few-Shot LearningZelei Wu, Kun Zhou, xulun ye, Yifan Mei et al.ICML 2026
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Meta-Learning to Detect Rare ObjectsYu-Xiong Wang, Deva Ramanan, Martial HebertICCV 2019 · 339 citations
Related papers
- DeIL: Direct-and-Inverse CLIP for Open-World Few-Shot LearningShuai Shao, Yu Bai, Yan Wang, Baodi Liu et al.CVPR 2024
- Prompt, Generate, Then Cache: Cascade of Foundation Models Makes Strong Few-Shot LearnersRenrui Zhang, Xiangfei Hu, Bohao Li, Siyuan Huang et al.CVPR 2023
- OCRT: Boosting Foundation Models in the Open World with Object-Concept-Relation TriadLuyao Tang, Yuxuan Yuan, Chaoqi Chen, Zeyu Zhang et al.CVPR 2025
- Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene UnderstandingJinlong Li, Cristiano Saltori, Fabio Poiesi, Nicu SebeCVPR 2025
- Unknown Text Learning for Clip-Based Few-Shot Open-Set RecognitionRui Ma, Qilong Wang, Bing Cao, Qinghua Hu et al.ICCV 2025 · 1 citation
