Visual Semantic Role Labeling for Video Understanding
Arka Sadhu, Tanmay Gupta, Mark Yatskar, Ram Nevatia, Aniruddha Kembhavi
摘要
Verb: deflect (block, avoid) Arg0 (deflector) woman with shield Arg1 (thing deflected) boulder Scene city park Verb: talk (speak) Arg0 (talker) woman with shield Arg2 (hearer) man with trident ArgM (manner) urgently Scene city park Verb: leap (physically leap) Arg0 (jumper) man with trident Arg1 (obstacle) over stairs ArgM (direction) towards shirtless man ArgM (goal) to attack shirtless man Scene city park Verb: punch (to hit) Arg0 (agent) shirtless man Arg1 (entity punched) man with trident ArgM (direction) far into distance Scene city park Verb: punch (to hit) Arg0 (agent) shirtless man Arg1 (entity punched) woman with shield ArgM (direction) down the stairs Scene city park 2 Seconds
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Knowing Where to Focus: Event-aware Transformer for Video GroundingJinhyun Jang, Jungin Park, Jin Kim, Hyeongjun Kwon 等ICCV 2023 · 被引用 103 次
- Verbs in Action: Improving verb understanding in video-language modelsLiliane Momeni, Mathilde Caron, Arsha Nagrani, Andrew Zisserman 等ICCV 2023 · 被引用 93 次
- ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language ModelsIlker Kesen, Andrea Pedrotti, Mustafa Dogan, Michele Cafagna 等ICLR 2024 · 被引用 25 次
- Hierarchical Self-supervised Representation Learning for Movie UnderstandingFanyi Xiao, Kaustav Kundu, Joseph Tighe, Davide ModoloCVPR 2022 · 被引用 21 次
- Grounded Video Situation RecognitionZeeshan Khan, C. V. Jawahar, Makarand TapaswiNeurIPS 2022 · 被引用 19 次
它引用的顶会 Paper12
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding 等ICCV 2019 · 被引用 709 次
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li 等ICCV 2019 · 被引用 688 次
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingLinjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan 等EMNLP 2020 · 被引用 387 次
相关 Paper
- MoMask: Generative Masked Modeling of 3D Human MotionsChuan Guo, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang 等CVPR 2024
- Dream3D: Zero-Shot Text-to-3D Synthesis Using 3D Shape Prior and Text-to-Image Diffusion ModelsJiale Xu, Xintao Wang, Weihao Cheng, Yan-Pei Cao 等CVPR 2023
- DualVector: Unsupervised Vector Font Synthesis with Dual-Part RepresentationYing-Tian Liu, Zhifei Zhang, Yuan-Chen Guo, Matthew Fisher 等CVPR 2023
- DreamDistribution: Learning Prompt Distribution for Diverse In-distribution GenerationBrian Nlong Zhao, Yuhang Xiao, Jiashu Xu, Xinyang Jiang 等ICLR 2025
- WorldSimBench: Towards Video Generation Models as World SimulatorsYiran Qin, Zhelun Shi, Jiwen Yu, Xijun Wang 等ICML 2025
