Language Features Matter: Effective Language Representations for Vision-Language Tasks
Andrea Burns, Reuben Tan, Kate Saenko, Stan Sclaroff, Bryan A. Plummer
摘要
Shouldn't language and vision features be treated equally in vision-language (VL) tasks? Many VL approaches treat the language component as an afterthought, using simple language models that are either built upon fixed word embeddings trained on text-only data or are learned from scratch. We conclude that language features deserve more attention, which has been informed by experiments which compare different word embeddings, language models, and embedding augmentation steps on five common VL tasks: image-sentence retrieval, image captioning, visual question answering, phrase grounding, and text-to-clip retrieval. Our experiments provide some striking results; an average embedding language model outperforms a LSTM on retrieval-style tasks; state-of-the-art representations such as BERT perform relatively poorly on vision-language tasks. From this comprehensive set of experiments we can propose a set of best practices for incorporating the language component of vision-language tasks. To further elevate language features, we also show that knowledge in vision-language problems can be transferred across tasks to gain performance with multi-task training. This multi-task training is applied to a new Graph Oriented Vision-Language Embedding (GrOVLE), which we adapt from Word2Vec using WordNet and an original visual-language graph built from Visual Genome, providing a ready-to-use vision-language embedding: http://ai.bu.edu/grovle.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- TeachText: CrossModal Generalized Distillation for Text-Video RetrievalIoana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu, Hailin Jin 等ICCV 2021 · 被引用 147 次
- Adaptive Cross-Modal Embeddings for Image-Text AlignmentJonatas Wehrmann, Camila Kolling, Rodrigo C. BarrosAAAI 2020 · 被引用 86 次
- MULE: Multimodal Universal Language EmbeddingDonghyun Kim, Kuniaki Saito, Kate Saenko, Stan Sclaroff 等AAAI 2020 · 被引用 45 次
- Telling the What while Pointing to the Where: Multimodal Queries for Image RetrievalSoravit Changpinyo, Jordi Pont-Tuset, Vittorio Ferrari, Radu SoricutICCV 2021 · 被引用 30 次
- Sign Language Video Retrieval with Free-Form Textual QueriesAmanda Cardoso Duarte, Samuel Albanie, Xavier Giró-i-Nieto, Gül VarolCVPR 2022 · 被引用 27 次
相关 Paper
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- Bridging Vision and Language Spaces with Assignment PredictionJungin Park, Jiyoung Lee, Kwanghoon SohnICLR 2024 · 被引用 15 次
- 12-in-1: Multi-Task Vision and Language Representation LearningJiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh 等CVPR 2020
- ViCo: Word Embeddings From Visual Co-OccurrencesTanmay Gupta, Alexander G. Schwing, Derek HoiemICCV 2019 · 被引用 26 次
- HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language EmbeddingChenxin Tao, Shiqian Su, Xizhou Zhu, Chenyu Zhang 等CVPR 2025
