ATM: Action Temporality Modeling for Video Question Answering
Junwen Chen, Jie Zhu, Yu Kong
摘要
Despite significant progress in video question answering (VideoQA), existing methods fall short of questions that require causal/temporal reasoning across frames. This can be attributed to imprecise motion representations. We introduce Action Temporality Modeling (ATM) for temporality reasoning via three-fold uniqueness: (1) rethinking the optical flow and realizing that optical flow is effective in capturing the long horizon temporality reasoning; (2) training the visual-text embedding by contrastive learning in an action-centric manner, leading to better action representations in both vision and text modalities; and (3) preventing the model from answering the question given the shuffled video in the fine-tuning stage, to avoid spurious correlation between appearance and motion and hence ensure faithful temporality reasoning. In the experiments, we show that ATM outperforms existing approaches in terms of the accuracy on multiple VideoQAs and exhibits better true temporality reasoning ability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Unleashing the Power of Chain-of-Prediction for Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Girish Chandar Ganesan, Xiaoming LiuCVPR 2026 · 被引用 13 次
- FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human RecognitionJie Zhu, Xiao Guo, Yiyang Su, Anil K. Jain 等CVPR 2026 · 被引用 7 次
- Towards Intrinsic-Aware Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Xiaoming LiuCVPR 2026 · 被引用 5 次
- A Quality-Guided Mixture of Score-Fusion Experts Framework for Human RecognitionJie Zhu, Yiyang Su, Minchul Kim, Anil K. Jain 等ICCV 2025 · 被引用 3 次
- H-MoRe: Learning Human-centric Motion Representation for Action AnalysisZhanbo Huang, Xiaoming Liu, Yu KongCVPR 2025
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
相关 Paper
- Three Stream Based Multi-level Event Contrastive Learning for Text-Video Event ExtractionJiaqi Li, Chuanyi Zhang, Miaozeng Du, Dehai Min 等EMNLP 2023 · 被引用 1 次
- Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal CluesYan Zhang, Gangyan Zeng, Huawen Shen, Daiqing Wu 等AAAI 2025 · 被引用 1 次
- Gather and Trace: Rethinking Video TextVQA from an Instance-oriented PerspectiveYan Zhang, Gangyan Zeng, Daiqing Wu, Huawen Shen 等ACM MM 2025 · 被引用 2 次
- Map the Flow: Revealing Hidden Pathways of Information in VideoLLMsMinji Kim, Taekyung Kim, Bohyung HanICLR 2026 · 被引用 8 次
- Structured Video-Language Modeling with Temporal Grouping and Spatial GroundingYuanhao Xiong, Long Zhao, Boqing Gong, Ming-Hsuan Yang 等ICLR 2024
