Contrasting Multiple Representations with the Multi-Marginal Matching Gap
Zoe Piran, Michal Klein, James Thornton, Marco Cuturi
Abstract
Learning meaningful representations of complex objects that can be seen through multiple () views or modalities is a core task in machine learning. Existing methods use losses originally intended for paired views, and extend them to views, either by instantiating loss-pairs, or by using reduced embeddings, following a one vs. average-of-rest strategy. We propose the multi-marginal matching gap (M3G), a loss that borrows tools from multi-marginal optimal transport (MM-OT) theory to simultaneously incorporate all views. Given a batch of points, each seen as a -tuple of views subsequently transformed into embeddings, our loss contrasts the cost of matching these ground-truth -tuples with the MM-OT polymatching cost, which seeks optimally arranged -tuples chosen within these vectors. While the exponential complexity ) of the MM-OT problem may seem daunting, we show in experiments that a suitable generalization of the Sinkhorn algorithm for that problem can scale to, e.g., views using mini-batches of size . Our experiments demonstrate improved performance over multiview extensions of pairwise losses, for both self-supervised and multimodal tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Inverse Optimal Transport for Efficient Adaptation of Vision-Language ModelsShupeng Qiu, Chuan-Xian RenAAAI 2026 · 1 citation
- DisPOSE: Projected Polystochastic Diffusion for Self-Supervised Multi-View 3D Human Pose EstimationTony Danjun Wang, Tolga Birdal, Nassir Navab, Lennart BastianICML 2026
- Transitive Representation Learning Enhances Histopathology AnnotationMoritz Schaefer, Zoe Piran, Nils Philipp Walter, Animesh Awasthi et al.ICML 2026
- Disentangled Representation Learning with the Gromov-Monge GapThéo Uscidda, Luca Eyring, Karsten Roth, Fabian J. Theis et al.ICLR 2025
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Batch Greenkhorn Algorithm for Entropic-Regularized Multimarginal Optimal Transport: Linear Rate of Convergence and Iteration ComplexityVladimir R. Kostic, Saverio Salzo, Massimiliano PontilICML 2022 · 3 citations
- Poly-View Contrastive LearningAmitis Shidani, R. Devon Hjelm, Jason Ramapuram, Russell Webb et al.ICLR 2024 · 9 citations
- Efficient Discrete Multi Marginal Optimal Transport RegularizationRonak Mehta, Jeffery Kline, Vishnu Suresh Lokhande, Glenn Fung et al.ICLR 2023
- OTKGE: Multi-modal Knowledge Graph Embeddings via Optimal TransportZongsheng Cao, Qianqian Xu, Zhiyong Yang, Yuan He et al.NeurIPS 2022 · 117 citations
- Representational Alignment Across Model Layers and Brain Regions with Multi-Level Optimal TransportShaan Shah, Meenakshi KhoslaICLR 2026 · 4 citations
