Collaborative Consortium of Foundation Models for Open-World Few-Shot Learning
Shuai Shao, Yu Bai, Yan Wang, Baodi Liu, Bin Liu
摘要
Open-World Few-Shot Learning (OFSL) is a crucial research field dedicated to accurately identifying target samples in scenarios where data is limited and labels are unreliable. This research holds significant practical implications and is highly relevant to real-world applications. Recently, the advancements in foundation models like CLIP and DINO have showcased their robust representation capabilities even in resource-constrained settings with scarce data. This realization has brought about a transformative shift in focus, moving away from "building models from scratch" towards "effectively harnessing the potential of foundation models to extract pertinent prior knowledge suitable for OFSL and utilizing it sensibly". Motivated by this perspective, we introduce the Collaborative Consortium of Foundation Models (CO3), which leverages CLIP, DINO, GPT-3, and DALL-E to collectively address the OFSL problem. CO3 comprises four key blocks: (1) the Label Correction Block (LC-Block) corrects unreliable labels, (2) the Data Augmentation Block (DA-Block) enhances available data, (3) the Feature Extraction Block (FE-Block) extracts multi-modal features, and (4) the Text-guided Fusion Adapter (TeFu-Adapter) integrates multiple features while mitigating the impact of noisy labels through semantic constraints. Only the adapter's parameters are adjustable, while the others remain frozen. Through collaboration among these foundation models, CO3 effectively unlocks their potential and unifies their capabilities to achieve state-of-the-art performance on multiple benchmark datasets. https://github.com/The-Shuai/CO3 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DePT: Decoupled Prompt TuningJi Zhang, Shihan Wu, Lianli Gao, Heng Tao Shen 等CVPR 2024 · 被引用 36 次
- FAST: A Dual-tier Few-Shot Learning Paradigm for Whole Slide Image ClassificationKexue Fu, Xiaoyuan Luo, Linhao Qu, Shuo Wang 等NeurIPS 2024 · 被引用 12 次
- Towards Effective Foundation Model Adaptation for Extreme Cross-Domain Few-Shot LearningFei Zhou, Peng Wang, Lei Zhang, Wei Wei 等ICCV 2025 · 被引用 2 次
- Rank-guided Diffusion for Noise Few-Shot LearningZelei Wu, Kun Zhou, xulun ye, Yifan Mei 等ICML 2026
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Meta-Learning to Detect Rare ObjectsYu-Xiong Wang, Deva Ramanan, Martial HebertICCV 2019 · 被引用 339 次
相关 Paper
- DeIL: Direct-and-Inverse CLIP for Open-World Few-Shot LearningShuai Shao, Yu Bai, Yan Wang, Baodi Liu 等CVPR 2024
- Prompt, Generate, Then Cache: Cascade of Foundation Models Makes Strong Few-Shot LearnersRenrui Zhang, Xiangfei Hu, Bohao Li, Siyuan Huang 等CVPR 2023
- OCRT: Boosting Foundation Models in the Open World with Object-Concept-Relation TriadLuyao Tang, Yuxuan Yuan, Chaoqi Chen, Zeyu Zhang 等CVPR 2025
- Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene UnderstandingJinlong Li, Cristiano Saltori, Fabio Poiesi, Nicu SebeCVPR 2025
- Unknown Text Learning for Clip-Based Few-Shot Open-Set RecognitionRui Ma, Qilong Wang, Bing Cao, Qinghua Hu 等ICCV 2025 · 被引用 1 次
