THE MORE, THE MERRIER: CONTRASTIVE FUSION FOR HIGHER-ORDER MULTIMODAL ALIGNMENT
Stefanos Koutoupis, Michaela Areti Zervou, Konstantinos Kontras, Maarten De Vos, Panagiotis Tsakalides, Grigorios Tsagkatakis
Abstract
Learning joint representations across multiple modalities remains a central challenge in multimodal machine learning. Prevailing approaches predominantly operate in pairwise settings, aligning two modalities at a time. While some recent methods aim to capture higher-order interactions among multiple modalities, they often overlook or insufficiently preserve pairwise relationships, limiting their effectiveness on single-modality tasks. In this work, we introduce Contrastive Fusion (ConFu), a framework that jointly embeds both individual modalities and their fused combinations into a unified representation space, where modalities and their fused counterparts are aligned. ConFu extends traditional pairwise contrastive objectives with an additional fused-modality contrastive term, encouraging the joint embedding of modality pairs with a third modality. This formulation enables ConFu to capture higher-order dependencies, such as XOR-like relationships, that cannot be recovered through pairwise alignment alone, while still maintaining strong pairwise correspondence. We evaluate ConFu on synthetic and real-world multimodal benchmarks, assessing its ability to exploit cross-modal complementarity, capture higher-order dependencies, and scale with increasing multimodal complexity. Across these settings, ConFu demonstrates competitive performance on retrieval and classification tasks, while supporting unified one-to-one and two-to-one retrieval within a single contrastive framework. We release our code and dataset at https://github.com/estafons/confu.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46ccde70-d74d-4832-9005-ec787bb06b62Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentBin Zhu, Bin Lin, Munan Ning, Yang Yan et al.ICLR 2024 · 403 citations
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang et al.CVPR 2022 · 264 citations
Related papers
- Contrasting with Symile: Simple Model-Agnostic Representation Learning for Unlimited ModalitiesAdriel Saporta, Aahlad Manas Puli, Mark Goldstein, Rajesh RanganathNeurIPS 2024 · 28 citations
- InfMasking: Unleashing Synergistic Information by Contrastive Multimodal InteractionsLiangjian Wen, Qun Dai, Jianzhuang Liu, Jiangtao Zheng et al.NeurIPS 2025 · 10 citations
- Contrastive Multimodal Fusion with TupleInfoNCEYunze Liu, Qingnan Fan, Shanghang Zhang, Hao Dong et al.ICCV 2021 · 84 citations
- When Kernels Multiply, Clusters Unify: Fusing Embeddings with the Kronecker ProductYouqi Wu, Jingwei Zhang, Farzan FarniaNeurIPS 2025 · 7 citations
- Easy2Hard: From Partially to Fully Unmatched Modalities as Negative Samples in Contrastive LearningZhicheng Yang, Yichen Liu, Chang Ge, Xiaopeng JiangCVPR 2026
