Visual Grounding in Video for Unsupervised Word Translation
Gunnar A. Sigurdsson, Jean-Baptiste Alayrac, Aida Nematzadeh, Lucas Smaira, Mateusz Malinowski, João Carreira, Phil Blunsom, Andrew Zisserman
Abstract
There are thousands of actively spoken languages on Earth, but a single visual world. Grounding in this visual world has the potential to bridge the gap between all these languages. Our goal is to use visual grounding to improve unsupervised word mapping between languages. The key idea is to establish a common visual representation between two languages by learning embeddings from unpaired instructional videos narrated in the native language. Given this shared embedding we demonstrate that (i) we can map words between the languages, particularly the 'visual' words; (ii) that the shared embedding provides a good initialization for existing unsupervised text-based word translation techniques, forming the basis for our proposed hybrid visual-text mapping algorithm, MUVE; and (iii) our approach achieves superior performance by addressing the shortcomings of text-based methods -it is more robust, handles datasets with less commonality, and is applicable to low-resource languages. We apply these methods to translate words from English to French, Korean, and Japanese -all without any parallel corpora and simply by watching many videos of people speaking while doing things.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- Visual Pivoting for (Unsupervised) Entity AlignmentFangyu Liu, Muhao Chen, Dan Roth, Nigel CollierAAAI 2021 · 159 citations
- Broaden Your Views for Self-Supervised Video LearningAdrià Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang et al.ICCV 2021 · 139 citations
- Revisiting Weakly Supervised Pre-Training of Visual Perception ModelsMannat Singh, Laura Gustafson, Aaron Adcock, Vinicius de Freitas Reis et al.CVPR 2022 · 88 citations
- VALHALLA: Visual Hallucination for Machine TranslationYi Li, Rameswar Panda, Yoon Kim, Chun-Fu Richard Chen et al.CVPR 2022 · 31 citations
- CLIPTrans: Transferring Visual Knowledge with Pre-trained Models for Multimodal Machine TranslationDevaansh Gupta, Siddhant Kharbanda, Jiawei Zhou, Wanhua Li et al.ICCV 2023 · 28 citations
Builds on5
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- Language-Conditioned Graph Networks for Relational ReasoningRonghang Hu, Anna Rohrbach, Trevor Darrell, Kate SaenkoICCV 2019 · 183 citations
- Learning to Follow Directions in Street ViewKarl Moritz Hermann, Mateusz Malinowski, Piotr Mirowski, Andras Banki-Horvath et al.AAAI 2020 · 78 citations
- End-to-End Learning of Visual Representations From Uncurated Instructional VideosAntoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev et al.CVPR 2020
Related papers
- Globetrotter: Connecting Languages by Connecting ImagesDídac Surís, Dave Epstein, Carl VondrickCVPR 2022 · 7 citations
- MULE: Multimodal Universal Language EmbeddingDonghyun Kim, Kuniaki Saito, Kate Saenko, Stan Sclaroff et al.AAAI 2020 · 45 citations
- CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge TransferYabing Wang, Fan Wang, Jianfeng Dong, Hao LuoAAAI 2024 · 20 citations
- Improving Low-Resource Languages in Pre-Trained Multilingual Language ModelsViktor Hangya, Hossain Shaikh Saadi, Alexander FraserEMNLP 2022 · 17 citations
- Unsupervised Multimodal Neural Machine Translation with Pseudo Visual PivotingPo-Yao Huang, Junjie Hu, Xiaojun Chang, Alexander G. HauptmannACL 2020 · 43 citations
