CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features
Po-han Li, Sandeep P. Chinchali, Ufuk Topcu
Abstract
Multimodal encoders like CLIP excel in tasks such as zero-shot image classification and cross-modal retrieval. However, they require excessive training data. We propose canonical similarity analysis (CSA), which uses two unimodal encoders to replicate multimodal encoders using limited data. CSA maps unimodal features into a multimodal space, using a new similarity score to retain only the multimodal information. CSA only involves the inference of unimodal encoders and a cubiccomplexity matrix decomposition, eliminating the need for extensive GPU-based model training. Experiments show that CSA outperforms CLIP while requiring 50, 000× fewer multimodal data pairs to bridge the modalities given pre-trained unimodal encoders on ImageNet classification and misinformative news caption detection. CSA surpasses the state-of-the-art method to map unimodal features to multimodal features. We also demonstrate the ability of CSA with modalities beyond image and text, paving the way for future modality pairs with limited paired multimodal data but abundant unpaired unimodal data, such as lidar and text. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb59eb72-8ec1-4af8-900c-3cc10e25184dCited by top-tier papers5
- With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide YouFabian Gröger, Shuo Wen, Huyen Le, Maria BrbicNeurIPS 2025 · 16 citations
- Learning Shared Representations from Unpaired DataAmitai Yacobi, Nir Ben-Ari, Ronen Talmon, Uri ShahamNeurIPS 2025 · 3 citations
- UNIALIGN: Scaling Multimodal Alignment within One Unified ModelBo Zhou, Liulei Li, Yujia Wang, Huafeng Liu et al.CVPR 2025
- Supervised Classification Heads as Semantic Prototypes: Unlocking Vision-Language Alignment via Weight RecyclingDavid Méndez, Roberto Confalonieri, Natalia Díaz-RodríguezICML 2026
- A Cross Modal Knowledge Distillation & Data Augmentation Recipe for Improving Transcriptomics Representations through Morphological FeaturesIhab Bendidi, Yassir El Mesbahi, Alisandra Kaye Denton, Karush Suri et al.ICML 2025
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Learning Relationships between Text, Audio, and Video via Deep Canonical Correlation for Multimodal Language AnalysisZhongkai Sun, Prathusha Kameswara Sarma, William A. Sethares, Yingyu LiangAAAI 2020 · 419 citations
- Rethinking Visual Geo-localization for Large-Scale ApplicationsGabriele Moreno Berton, Carlo Masone, Barbara CaputoCVPR 2022 · 235 citations
Related papers
- Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot LearningTianjiao Jiang, Zhen Zhang, Yuhang Liu, Javen Qinfeng ShiICCV 2025 · 3 citations
- Harnessing Frozen Unimodal Encoders for Flexible Multimodal AlignmentMayug Maniparambil, Raiymbek Akshulakov, Yasser Abdelaziz Dahou Djilali, Sanath Narayan et al.CVPR 2025
- ASIF: Coupled Data Turns Unimodal Models to Multimodal without TrainingAntonio Norelli, Marco Fumero, Valentino Maiorca, Luca Moschella et al.NeurIPS 2023 · 57 citations
- Do Vision and Language Encoders Represent the World Similarly?Mayug Maniparambil, Raiymbek Akshulakov, Yasser Abdelaziz Dahou Djilali, Mohamed El Amine Seddik et al.CVPR 2024 · 3 citations
- Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality InversionMarco Mistretta, Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini et al.ICLR 2025
