ViCo: Word Embeddings From Visual Co-Occurrences
Tanmay Gupta, Alexander G. Schwing, Derek Hoiem
摘要
We propose to learn word embeddings from visual co-occurrences. Two words co-occur visually if both words apply to the same image or image region. Specifically, we extract four types of visual co-occurrences between object and attribute words from large-scale, textually-annotated visual databases like VisualGenome and ImageNet. We then train a multi-task log-bilinear model that compactly encodes word ``meanings'' represented by each co-occurrence type into a single visual word-vector. Through unsupervised clustering, supervised partitioning, and a zero-shot-like generalization analysis we show that our word embeddings complement text-only embeddings like GloVe by better representing similarities and differences between visual concepts that are difficult to obtain from text corpora alone. We further evaluate our embeddings on five downstream applications, four of which are vision-language tasks. Augmenting GloVe with our embeddings yields gains on all tasks. We also find that random embeddings perform comparably to learned embeddings on all supervised vision-language tasks, contrary to conventional wisdom.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- LayoutTransformer: Layout Generation and Completion with Self-attentionKamal Gupta, Justin Lazarow, Alessandro Achille, Larry Davis 等ICCV 2021 · 被引用 184 次
- MULE: Multimodal Universal Language EmbeddingDonghyun Kim, Kuniaki Saito, Kate Saenko, Stan Sclaroff 等AAAI 2020 · 被引用 45 次
- Explainable Semantic Space by Grounding Language to Vision with Cross-Modal Contrastive LearningYizhen Zhang, Minkyu Choi, Kuan Han, Zhongming LiuNeurIPS 2021 · 被引用 20 次
- Temporal Slowness in Central Vision Drives Semantic Object LearningTimothy Schaumlöffel, Arthur Aubret, Gemma Roig, Jochen TrieschICLR 2026
- PIGLeT: Language Grounding Through Neuro-Symbolic Interaction in a 3D WorldRowan Zellers, Ari Holtzman, Matthew E. Peters, Roozbeh Mottaghi 等ACL 2021
相关 Paper
- Language Features Matter: Effective Language Representations for Vision-Language TasksAndrea Burns, Reuben Tan, Kate Saenko, Stan Sclaroff 等ICCV 2019 · 被引用 28 次
- VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot LearningWenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele 等CVPR 2022 · 被引用 61 次
- On the Emergence of Linear Analogies in Word EmbeddingsDaniel J. Korchinski, Dhruva Karkada, Yasaman Bahri, Matthieu WyartNeurIPS 2025 · 被引用 10 次
- Semantic and Expressive Variations in Image Captions Across LanguagesAndre Ye, Sebastin Santy, Jena D. Hwang, Amy X. Zhang 等CVPR 2025
- Pixel Aligned Language ModelsJiarui Xu, Xingyi Zhou, Shen Yan, Xiuye Gu 等CVPR 2024 · 被引用 6 次
