A Trainable Optimal Transport Embedding for Feature Aggregation and its Relationship to Attention
Grégoire Mialon, Dexiong Chen, Alexandre d'Aspremont, Julien Mairal
摘要
We address the problem of learning on sets of features, motivated by the need of performing pooling operations in long biological sequences of varying sizes, with long-range dependencies, and possibly few labeled data. To address this challenging task, we introduce a parametrized representation of fixed size, which embeds and then aggregates elements from a given input set according to the optimal transport plan between the set and a trainable reference. Our approach scales to large datasets and allows end-to-end training of the reference, while also providing a simple unsupervised learning mechanism with small computational cost. Our aggregation technique admits two useful interpretations: it may be seen as a mechanism related to attention layers in neural networks, or it may be seen as a scalable surrogate of a classical optimal transport-based kernel. We experimentally demonstrate the effectiveness of our approach on biological sequences, achieving state-of-the-art results for protein fold recognition and detection of chromatin profiles tasks, and, as a proof of concept, we show promising results for processing natural language sequences. We provide an open-source implementation of our embedding that can be used alone or as a module in larger learning models at https://github.com/claying/OTK.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Wasserstein Embedding for Graph LearningSoheil Kolouri, Navid NaderiAlizadeh, Gustavo K. Rohde, Heiko HoffmannICLR 2021 · 被引用 99 次
- Multimodal Prototyping for cancer survival predictionAndrew H. Song, Richard J. Chen, Guillaume Jaume, Anurag J. Vaidya 等ICML 2024 · 被引用 53 次
- Morphological Prototyping for Unsupervised Slide Representation Learning in Computational PathologyAndrew H. Song, Richard J. Chen, Tong Ding, Drew F. K. Williamson 等CVPR 2024 · 被引用 51 次
- Keep It SimPool: Who Said Supervised Transformers Suffer from Attention Deficit?Bill Psomas, Ioannis Kakogeorgiou, Konstantinos Karantzalos, Yannis AvrithisICCV 2023 · 被引用 18 次
- Generalized Sum Pooling for Metric LearningYeti Ziya Gürbüz, Ozan Sener, A. Aydin AlatanICCV 2023 · 被引用 10 次
它引用的顶会 Paper3
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 被引用 629 次
- Hard-Coded Gaussian Attention for Neural Machine TranslationWeiqiu You, Simeng Sun, Mohit IyyerACL 2020 · 被引用 55 次
相关 Paper
- Differentiable Expectation-Maximization for Set Representation LearningMinyoung KimICLR 2022 · 被引用 18 次
- Local Aggregation for Unsupervised Learning of Visual EmbeddingsChengxu Zhuang, Alex Lin Zhai, Daniel YaminsICCV 2019 · 被引用 462 次
- Pooling by Sliced-Wasserstein EmbeddingNavid NaderiAlizadeh, Joseph F. Comer, Reed W. Andrews, Heiko Hoffmann 等NeurIPS 2021 · 被引用 38 次
- SeTformer Is What You Need for Vision and LanguagePourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger, Michael FelsbergAAAI 2024 · 被引用 8 次
- PoNet: Pooling Network for Efficient Token Mixing in Long SequencesChao-Hong Tan, Qian Chen, Wen Wang, Qinglin Zhang 等ICLR 2022 · 被引用 15 次
