Text as Any-Modality for Zero-Shot Classification by Consistent Prompt Tuning
Xiangyu Wu, Feng Yu, Yang Yang, Jianfeng Lu
Abstract
The integration of prompt tuning with multimodal learning has shown significant generalization abilities for various downstream tasks. Despite advancements, existing methods heavily depend on massive modality-specific labeled data (e.g., video, audio, and image), or are customized for a single modality. In this study, we present Text as Any-Modality by Consistent Prompt Tuning (TaAM-CPT), a scalable approach for constructing a general representation model toward unlimited modalities using solely text data. TaAM-CPT comprises modality prompt pools, text construction, and modality-aligned text encoders from pre-trained models, which allows for extending new modalities by simply adding prompt pools and modality-aligned text encoders. To harmonize the learning across different modalities, TaAM-CPT designs intra- and inter-modal learning objectives, which can capture category details within modalities while maintaining semantic consistency across different modalities. Benefiting from its scalable architecture and pre-trained models, TaAM-CPT can be seamlessly extended to accommodate unlimited modalities. Remarkably, without any modality-specific labeled data, TaAM-CPT achieves leading results on diverse datasets spanning various modalities, including video classification, image classification, and audio classification. The code is available at https://github.com/Jinx630/TaAM-CPT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 150fcb4e-9713-4e70-b111-b8c7177aa383Cited by top-tier papers3
- Dual Latent Memory for Visual Multi-agent SystemXinlei Yu, Chengming Xu, Zhangquan Chen, Bo Yin et al.ICML 2026 · 5 citations
- Multimodal Classification via Total Correlation MaximizationFeng Yu, Xiangyu Wu, Yang Yang, Jianfeng LuICLR 2026 · 4 citations
- Adaptive Debiasing Tsallis Entropy for Test-Time AdaptationXiangyu Wu, Dongming Jiang, Feng Yu, Yueying Tian et al.ICLR 2026 · 2 citations
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng et al.ICCV 2019 · 1,018 citations
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li et al.ICLR 2024 · 467 citations
Related papers
- Texts as Images in Prompt Tuning for Multi-Label Image RecognitionZixian Guo, Bowen Dong, Zhilong Ji, Jinfeng Bai et al.CVPR 2023
- AnyGPT: Unified Multimodal LLM with Discrete Sequence ModelingJun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou et al.ACL 2024
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang et al.AAAI 2024 · 54 citations
- LAMM: Label Alignment for Multi-Modal Prompt LearningJingsheng Gao, Jiacheng Ruan, Suncheng Xiang, Zefang Yu et al.AAAI 2024 · 33 citations
- Amend to Alignment: Decoupled Prompt Tuning for Mitigating Spurious Correlation in Vision-Language ModelsJie Zhang, Xiaosong Ma, Song Guo, Peng Li et al.ICML 2024 · 10 citations
