Grounded Video Situation Recognition
Zeeshan Khan, C. V. Jawahar, Makarand Tapaswi
摘要
Dense video understanding requires answering several questions such as who is doing what to whom, with what, how, why, and where. Recently, Video Situation Recognition (VidSitu) is framed as a task for structured prediction of multiple events, their relationships, and actions and various verb-role pairs attached to descriptive entities. This task poses several challenges in identifying, disambiguating, and co-referencing entities across multiple verb-role pairs, but also faces some challenges of evaluation. In this work, we propose the addition of spatiotemporal grounding as an essential component of the structured prediction task in a weakly supervised setting, and present a novel three stage Transformer model, VideoWhisperer, that is empowered to make joint predictions. In stage one, we learn contextualised embeddings for video features in parallel with key objects that appear in the video clips to enable fine-grained spatio-temporal reasoning. The second stage sees verb-role queries attend and pool information from object embeddings, localising answers to questions posed about the action. The final stage generates these answers as captions to describe each verb-role pair present in the video. Our model operates on a group of events (clips) simultaneously and predicts verbs, verb-role pairs, their nouns, and their grounding on-the-fly. When evaluated on a grounding-augmented version of the VidSitu dataset, we observe a large improvement in entity captioning accuracy, as well as the ability to localize verb-roles without grounding annotations at training time. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Constructing Holistic Spatio-Temporal Scene Graph for Video Semantic Role LabelingYu Zhao, Hao Fei, Yixin Cao, Bobo Li 等ACM MM 2023 · 被引用 31 次
- Chirality in Action: Time-Aware Video Representation Learning by Latent StraighteningPiyush Bagad, Andrew ZissermanNeurIPS 2025 · 被引用 14 次
- VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in VideosBaoyu Liang, Qile Su, Shoutai Zhu, Yuchen Liang 等AAAI 2025 · 被引用 5 次
- HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video UnderstandingTrong-Thuan Nguyen, Pha A. Nguyen, Khoa LuuCVPR 2024 · 被引用 5 次
- MICap: A Unified Model for Identity-Aware Movie DescriptionsHaran Raajesh, Naveen Reddy Desanur, Zeeshan Khan, Makarand TapaswiCVPR 2024 · 被引用 4 次
它引用的顶会 Paper10
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding 等ICCV 2019 · 被引用 709 次
- TubeDETR: Spatio-Temporal Video Grounding with TransformersAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等CVPR 2022 · 被引用 87 次
- Rethinking the Two-Stage Framework for Grounded Situation RecognitionMeng Wei, Long Chen, Wei Ji, Xiaoyu Yue 等AAAI 2022 · 被引用 38 次
相关 Paper
- Weakly Supervised Video Representation Learning with Unaligned Text for Sequential VideosSixun Dong, Huazhang Hu, Dongze Lian, Weixin Luo 等CVPR 2023
- GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention RefinementZhi-Qi Cheng, Qi Dai, Siyao Li, Teruko Mitamura 等ACM MM 2022 · 被引用 40 次
- Comprehensive Visual Grounding for Video DescriptionWenhui Jiang, Yibo Cheng, Linxin Liu, Yuming Fang 等AAAI 2024 · 被引用 5 次
- Dense Video Object Captioning from Disjoint SupervisionXingyi Zhou, Anurag Arnab, Chen Sun, Cordelia SchmidICLR 2025
- STVGBert: A Visual-linguistic Transformer based Framework for Spatio-temporal Video GroundingRui Su, Qian Yu, Dong XuICCV 2021 · 被引用 75 次
