Visual Semantic Role Labeling for Video Understanding
Arka Sadhu, Tanmay Gupta, Mark Yatskar, Ram Nevatia, Aniruddha Kembhavi
Abstract
Verb: deflect (block, avoid) Arg0 (deflector) woman with shield Arg1 (thing deflected) boulder Scene city park Verb: talk (speak) Arg0 (talker) woman with shield Arg2 (hearer) man with trident ArgM (manner) urgently Scene city park Verb: leap (physically leap) Arg0 (jumper) man with trident Arg1 (obstacle) over stairs ArgM (direction) towards shirtless man ArgM (goal) to attack shirtless man Scene city park Verb: punch (to hit) Arg0 (agent) shirtless man Arg1 (entity punched) man with trident ArgM (direction) far into distance Scene city park Verb: punch (to hit) Arg0 (agent) shirtless man Arg1 (entity punched) woman with shield ArgM (direction) down the stairs Scene city park 2 Seconds
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- Knowing Where to Focus: Event-aware Transformer for Video GroundingJinhyun Jang, Jungin Park, Jin Kim, Hyeongjun Kwon et al.ICCV 2023 · 103 citations
- Verbs in Action: Improving verb understanding in video-language modelsLiliane Momeni, Mathilde Caron, Arsha Nagrani, Andrew Zisserman et al.ICCV 2023 · 93 citations
- ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language ModelsIlker Kesen, Andrea Pedrotti, Mustafa Dogan, Michele Cafagna et al.ICLR 2024 · 25 citations
- Hierarchical Self-supervised Representation Learning for Movie UnderstandingFanyi Xiao, Kaustav Kundu, Joseph Tighe, Davide ModoloCVPR 2022 · 21 citations
- Grounded Video Situation RecognitionZeeshan Khan, C. V. Jawahar, Makarand TapaswiNeurIPS 2022 · 19 citations
Builds on12
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingLinjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan et al.EMNLP 2020 · 387 citations
Related papers
- MoMask: Generative Masked Modeling of 3D Human MotionsChuan Guo, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang et al.CVPR 2024
- Dream3D: Zero-Shot Text-to-3D Synthesis Using 3D Shape Prior and Text-to-Image Diffusion ModelsJiale Xu, Xintao Wang, Weihao Cheng, Yan-Pei Cao et al.CVPR 2023
- DualVector: Unsupervised Vector Font Synthesis with Dual-Part RepresentationYing-Tian Liu, Zhifei Zhang, Yuan-Chen Guo, Matthew Fisher et al.CVPR 2023
- DreamDistribution: Learning Prompt Distribution for Diverse In-distribution GenerationBrian Nlong Zhao, Yuhang Xiao, Jiashu Xu, Xinyang Jiang et al.ICLR 2025
- WorldSimBench: Towards Video Generation Models as World SimulatorsYiran Qin, Zhelun Shi, Jiwen Yu, Xijun Wang et al.ICML 2025
