Visual Grounding in Video for Unsupervised Word Translation
Gunnar A. Sigurdsson, Jean-Baptiste Alayrac, Aida Nematzadeh, Lucas Smaira, Mateusz Malinowski, João Carreira, Phil Blunsom, Andrew Zisserman
摘要
There are thousands of actively spoken languages on Earth, but a single visual world. Grounding in this visual world has the potential to bridge the gap between all these languages. Our goal is to use visual grounding to improve unsupervised word mapping between languages. The key idea is to establish a common visual representation between two languages by learning embeddings from unpaired instructional videos narrated in the native language. Given this shared embedding we demonstrate that (i) we can map words between the languages, particularly the 'visual' words; (ii) that the shared embedding provides a good initialization for existing unsupervised text-based word translation techniques, forming the basis for our proposed hybrid visual-text mapping algorithm, MUVE; and (iii) our approach achieves superior performance by addressing the shortcomings of text-based methods -it is more robust, handles datasets with less commonality, and is applicable to low-resource languages. We apply these methods to translate words from English to French, Korean, and Japanese -all without any parallel corpora and simply by watching many videos of people speaking while doing things.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Visual Pivoting for (Unsupervised) Entity AlignmentFangyu Liu, Muhao Chen, Dan Roth, Nigel CollierAAAI 2021 · 被引用 159 次
- Broaden Your Views for Self-Supervised Video LearningAdrià Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang 等ICCV 2021 · 被引用 139 次
- Revisiting Weakly Supervised Pre-Training of Visual Perception ModelsMannat Singh, Laura Gustafson, Aaron Adcock, Vinicius de Freitas Reis 等CVPR 2022 · 被引用 88 次
- VALHALLA: Visual Hallucination for Machine TranslationYi Li, Rameswar Panda, Yoon Kim, Chun-Fu Richard Chen 等CVPR 2022 · 被引用 31 次
- CLIPTrans: Transferring Visual Knowledge with Pre-trained Models for Multimodal Machine TranslationDevaansh Gupta, Siddhant Kharbanda, Jiawei Zhou, Wanhua Li 等ICCV 2023 · 被引用 28 次
它引用的顶会 Paper5
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li 等ICCV 2019 · 被引用 688 次
- Language-Conditioned Graph Networks for Relational ReasoningRonghang Hu, Anna Rohrbach, Trevor Darrell, Kate SaenkoICCV 2019 · 被引用 183 次
- Learning to Follow Directions in Street ViewKarl Moritz Hermann, Mateusz Malinowski, Piotr Mirowski, Andras Banki-Horvath 等AAAI 2020 · 被引用 78 次
- End-to-End Learning of Visual Representations From Uncurated Instructional VideosAntoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev 等CVPR 2020
相关 Paper
- Globetrotter: Connecting Languages by Connecting ImagesDídac Surís, Dave Epstein, Carl VondrickCVPR 2022 · 被引用 7 次
- MULE: Multimodal Universal Language EmbeddingDonghyun Kim, Kuniaki Saito, Kate Saenko, Stan Sclaroff 等AAAI 2020 · 被引用 45 次
- CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge TransferYabing Wang, Fan Wang, Jianfeng Dong, Hao LuoAAAI 2024 · 被引用 20 次
- Improving Low-Resource Languages in Pre-Trained Multilingual Language ModelsViktor Hangya, Hossain Shaikh Saadi, Alexander FraserEMNLP 2022 · 被引用 17 次
- Unsupervised Multimodal Neural Machine Translation with Pseudo Visual PivotingPo-Yao Huang, Junjie Hu, Xiaojun Chang, Alexander G. HauptmannACL 2020 · 被引用 43 次
