Lune

ICML2024顶会

Contrasting Multiple Representations with the Multi-Marginal Matching Gap

Zoe Piran, Michal Klein, James Thornton, Marco Cuturi

2024年份
9被引次数
4顶会引用

摘要

Learning meaningful representations of complex objects that can be seen through multiple (k≥3k\geq 3) views or modalities is a core task in machine learning. Existing methods use losses originally intended for paired views, and extend them to kk views, either by instantiating 12k(k−1)\tfrac12k(k-1) loss-pairs, or by using reduced embeddings, following a one vs. average-of-rest strategy. We propose the multi-marginal matching gap (M3G), a loss that borrows tools from multi-marginal optimal transport (MM-OT) theory to simultaneously incorporate all kk views. Given a batch of nn points, each seen as a kk-tuple of views subsequently transformed into kk embeddings, our loss contrasts the cost of matching these nn ground-truth kk-tuples with the MM-OT polymatching cost, which seeks nn optimally arranged kk-tuples chosen within these n×kn\times k vectors. While the exponential complexity O(nkO(n^k) of the MM-OT problem may seem daunting, we show in experiments that a suitable generalization of the Sinkhorn algorithm for that problem can scale to, e.g., k=3∼6k=3\sim 6 views using mini-batches of size 64 ∼12864~\sim128. Our experiments demonstrate improved performance over multiview extensions of pairwise losses, for both self-supervised and multimodal tasks.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper4

问问它们各自怎么用它

它引用的顶会 Paper21

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖