Action Detail Matters: Refining Video Recognition with Local Action Queries
Mengmeng Wang, Zeyi Huang, Xiangjie Kong, Guojiang Shen, Guang Dai, Jingdong Wang, Yong Liu
摘要
Video action recognition involves interpreting both global context and specific details to accurately identify actions. While previous models are effective at capturing spatiotemporal features, they often lack a focused representation of key action details. To address this, we introduce Fo-cusVideo, a framework designed for refining video action recognition through integrated global and local feature learning. Inspired by human visual cognition theory, our approach balances the focus on both broad contextual changes and action-specific details, minimizing the influence of irrelevant background noise. We first employ learnable action queries to selectively emphasize action-relevant regions without requiring region-specific labels. Next, these queries are learned by a local action streaming branch that enables progressive query propagation. Moreover, we introduce a parameter-free feature interaction mechanism for effective multi-scale interaction between global and local features with minimal additional overhead. Extensive experiments demonstrate that FocusVideo achieves stateof-the-art performance across multiple action recognition datasets, validating its effectiveness and robustness in handling action-relevant details.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- FineTec: Fine-Grained Action Recognition Under Temporal Corruption via Skeleton Decomposition and Sequence CompletionDian Shao, Mingfei Shi, Like LiuAAAI 2026
- VidPrism: Heterogeneous Mixture of Experts for Image-to-Video TransferRui Lin, Chuanming Wang, Huadong MaCVPR 2026
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
相关 Paper
- PRVQL: Progressive Knowledge-Guided Refinement for Robust Egocentric Visual Query LocalizationBing Fan, Yunhe Feng, Yapeng Tian, James Chenhao Liang 等ICCV 2025 · 被引用 1 次
- Object-Centric Framework for Video Moment RetrievalZongyao Li, Yongkang Wong, Satoshi Yamazaki, Jianquan Liu 等AAAI 2026
- Few-Shot Common Action Localization via Cross-Attentional Fusion of Context and Temporal DynamicsJuntae Lee, Mihir Jain, Sungrack YunICCV 2023 · 被引用 5 次
- TubeR: Tubelet Transformer for Video Action DetectionJiaojiao Zhao, Yanyi Zhang, Xinyu Li, Hao Chen 等CVPR 2022 · 被引用 77 次
- Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal InteractionMingda Jia, Weiliang Meng, Zenghuang Fu, Yiheng Li 等AAAI 2026 · 被引用 1 次
