Integrating Language Guidance into Vision-based Deep Metric Learning
Karsten Roth, Oriol Vinyals, Zeynep Akata
摘要
Deep Metric Learning (DML) proposes to learn metric spaces which encode semantic similarities as embedding space distances. These spaces should be transferable to classes beyond those seen during training. Commonly, DML methods task networks to solve contrastive ranking tasks defined over binary class assignments. However, such approaches ignore higher-level semantic relations between the actual classes. This causes learned embedding spaces to encode incomplete semantic context and misrepresent the semantic relation between classes, impacting the generalizability of the learned metric space. To tackle this issue, we propose a language guidance objective for visual similarity learning. Leveraging language embeddings of expert- and pseudo-classnames, we contextualize and realign visual representation spaces corresponding to meaningful language semantics for better semantic consistency. Extensive experiments and ablations provide a strong motivation for our proposed approach and show language guidance offering significant, model-agnostic improvements for DML, achieving competitive and state-of-the-art results on all benchmarks. Code available at github.com/ExplainableML/LanguageGuidance-for_DML.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Waffling around for Performance: Visual Classification with Random Words and Broad ConceptsKarsten Roth, Jae-Myung Kim, A. Sophia Koepke, Oriol Vinyals 等ICCV 2023 · 被引用 124 次
- Vision-by-Language for Training-Free Compositional Image RetrievalShyamgopal Karthik, Karsten Roth, Massimiliano Mancini, Zeynep AkataICLR 2024 · 被引用 120 次
- HiTeA: Hierarchical Temporal-Aware Video-Language Pre-trainingQinghao Ye, Guohai Xu, Ming Yan, Haiyang Xu 等ICCV 2023 · 被引用 102 次
- Introducing Language Guidance in Prompt-based Continual LearningMuhammad Gul Zain Ali Khan, Muhammad Ferjad Naeem, Luc Van Gool, Didier Stricker 等ICCV 2023 · 被引用 71 次
- Towards Universal Image Embeddings: A Large-Scale Dataset and Challenge for Generic Image RepresentationsNikolaos-Antonios Ypsilantis, Kaifeng Chen, Bingyi Cao, Mário Lipovský 等ICCV 2023 · 被引用 31 次
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
相关 Paper
- Revisiting Training Strategies and Generalization Performance in Deep Metric LearningKarsten Roth, Timo Milbich, Samarth Sinha, Prateek Gupta 等ICML 2020 · 被引用 187 次
- Simultaneous Similarity-based Self-Distillation for Deep Metric LearningKarsten Roth, Timo Milbich, Björn Ommer, Joseph Paul Cohen 等ICML 2021 · 被引用 16 次
- Deep Metric Learning with Graph ConsistencyBinghui Chen, Pengyu Li, Zhaoyi Yan, Biao Wang 等AAAI 2021 · 被引用 7 次
- Towards Interpretable Deep Metric Learning with Structural MatchingWenliang Zhao, Yongming Rao, Ziyi Wang, Jiwen Lu 等ICCV 2021 · 被引用 52 次
- Deep Relational Metric LearningWenzhao Zheng, Borui Zhang, Jiwen Lu, Jie ZhouICCV 2021 · 被引用 53 次
