Contrasting Multiple Representations with the Multi-Marginal Matching Gap
Zoe Piran, Michal Klein, James Thornton, Marco Cuturi
摘要
Learning meaningful representations of complex objects that can be seen through multiple () views or modalities is a core task in machine learning. Existing methods use losses originally intended for paired views, and extend them to views, either by instantiating loss-pairs, or by using reduced embeddings, following a one vs. average-of-rest strategy. We propose the multi-marginal matching gap (M3G), a loss that borrows tools from multi-marginal optimal transport (MM-OT) theory to simultaneously incorporate all views. Given a batch of points, each seen as a -tuple of views subsequently transformed into embeddings, our loss contrasts the cost of matching these ground-truth -tuples with the MM-OT polymatching cost, which seeks optimally arranged -tuples chosen within these vectors. While the exponential complexity ) of the MM-OT problem may seem daunting, we show in experiments that a suitable generalization of the Sinkhorn algorithm for that problem can scale to, e.g., views using mini-batches of size . Our experiments demonstrate improved performance over multiview extensions of pairwise losses, for both self-supervised and multimodal tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Inverse Optimal Transport for Efficient Adaptation of Vision-Language ModelsShupeng Qiu, Chuan-Xian RenAAAI 2026 · 被引用 1 次
- DisPOSE: Projected Polystochastic Diffusion for Self-Supervised Multi-View 3D Human Pose EstimationTony Danjun Wang, Tolga Birdal, Nassir Navab, Lennart BastianICML 2026
- Transitive Representation Learning Enhances Histopathology AnnotationMoritz Schaefer, Zoe Piran, Nils Philipp Walter, Animesh Awasthi 等ICML 2026
- Disentangled Representation Learning with the Gromov-Monge GapThéo Uscidda, Luca Eyring, Karsten Roth, Fabian J. Theis 等ICLR 2025
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Batch Greenkhorn Algorithm for Entropic-Regularized Multimarginal Optimal Transport: Linear Rate of Convergence and Iteration ComplexityVladimir R. Kostic, Saverio Salzo, Massimiliano PontilICML 2022 · 被引用 3 次
- Poly-View Contrastive LearningAmitis Shidani, R. Devon Hjelm, Jason Ramapuram, Russell Webb 等ICLR 2024 · 被引用 9 次
- Efficient Discrete Multi Marginal Optimal Transport RegularizationRonak Mehta, Jeffery Kline, Vishnu Suresh Lokhande, Glenn Fung 等ICLR 2023
- OTKGE: Multi-modal Knowledge Graph Embeddings via Optimal TransportZongsheng Cao, Qianqian Xu, Zhiyong Yang, Yuan He 等NeurIPS 2022 · 被引用 117 次
- Representational Alignment Across Model Layers and Brain Regions with Multi-Level Optimal TransportShaan Shah, Meenakshi KhoslaICLR 2026 · 被引用 4 次
