Weakly-Supervised Generation and Grounding of Visual Descriptions with Conditional Generative Models
Effrosyni Mavroudi, René Vidal
摘要
Given weak supervision from image- or video-caption pairs, we address the problem of grounding (localizing) each object word of a ground-truth or generated sentence describing a visual input. Recent weakly-supervised approaches leverage region proposals and ground words based on the region attention coefficients of captioning models. To predict each next word in the sentence they attend over regions using a summary of the previous words as a query, and then ground the word by selecting the most attended regions. However, this leads to sub-optimal grounding, since attention coefficients are computed without taking into account the word that needs to be localized. To address this shortcoming, we propose a novel Grounded Visual Description Conditional Variational Autoencoder (GVD-CVAE) and leverage its latent variables for grounding. In particular, we introduce a discrete random variable that models each word-to-region alignment, and learn its approximate posterior distribution given the full sentence. Experiments on challenging image and video datasets (Flickr30k Entities, YouCook2, ActivityNet Entities) validate the effectiveness of our conditional generative model, showing that it can substantially outperform soft-attention-based baselines in grounding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Helping Hands: An Object-Aware Ego-Centric Video Recognition ModelChuhan Zhang, Ankush Gupta, Andrew ZissermanICCV 2023 · 被引用 39 次
- Weakly Supervised Referring Image Segmentation with Intra-Chunk and Inter-Chunk ConsistencyJungbeom Lee, Sungjin Lee, Jinseok Nam, Seunghak Yu 等ICCV 2023 · 被引用 28 次
- Learning to Segment Referred Objects from Narrated Egocentric VideosYuhan Shen, Huiyu Wang, Xitong Yang, Matt Feiszli 等CVPR 2024 · 被引用 1 次
- MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language ModelsYifan Xu, Chao Zhang, Ruifei Ma, Fei Gao 等CVPR 2026
它引用的顶会 Paper11
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou 等ICCV 2021 · 被引用 468 次
- Align2Ground: Weakly Supervised Phrase Grounding Guided by Image-Caption AlignmentSamyak Datta, Karan Sikka, Anirban Roy, Karuna Ahuja 等ICCV 2019 · 被引用 113 次
- Sequential Latent Spaces for Modeling the Intention During Diverse Image CaptioningJyoti Aneja, Harsh Agrawal, Dhruv Batra, Alexander G. SchwingICCV 2019 · 被引用 71 次
- Prophet Attention: Predicting Attention with Future AttentionFenglin Liu, Xuancheng Ren, Xian Wu, Shen Ge 等NeurIPS 2020 · 被引用 52 次
- Weakly-Supervised Video Object Grounding by Exploring Spatio-Temporal ContextsXun Yang, Xueliang Liu, Meng Jian, Xinjian Gao 等ACM MM 2020 · 被引用 47 次
相关 Paper
- Distributed Attention for Grounded Image CaptioningNenglun Chen, Xingjia Pan, Runnan Chen, Lei Yang 等ACM MM 2021 · 被引用 19 次
- Relational Graph Learning for Grounded Video Description GenerationWenqiao Zhang, Xin Eric Wang, Siliang Tang, Haizhou Shi 等ACM MM 2020 · 被引用 25 次
- AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual GroundingYidan Wang, Chenyi Zhuang, Wutao Liu, Pan Gao 等ACM MM 2025 · 被引用 2 次
- Pixel Aligned Language ModelsJiarui Xu, Xingyi Zhou, Shen Yan, Xiuye Gu 等CVPR 2024 · 被引用 6 次
- Similarity Maps for Self-Training Weakly-Supervised Phrase GroundingTal Shaharabany, Lior WolfCVPR 2023
