Exploiting Frequency Dynamics for Enhanced Multimodal Event-Based Action Recognition
Meiqi Cao, Xiangbo Shu, Xin Jiang, Rui Yan, Yazhou Yao, Jinhui Tang
摘要
While event cameras excel in capturing microsecond temporal dynamics, they suffer from sparse spatial representations compared to traditional RGB data. Thus, multimodal event-based action recognition approaches aim to synergize complementary strengths by independently extracting and integrating paired RGB-Event features. However, this paradigm inevitably introduces additional data acquisition costs, while eroding the inherent privacy advantages of event-based sensing. Drawing inspiration from event-to-image reconstruction, texture-enriched visual representation directly reconstructed from asynchronous event streams is a promising solution. In response, we propose an Enhanced Multimodal Perceptual (EMP) framework that hierarchically explores multimodal cues (e.g., edges and textures) from raw event streams through two synergistic innovations spanning representation to feature levels. Specifically, we introduce Cross-Modal Frequency Enhancer (CFE) that leverages complementary frequency characteristics between reconstructed frames and stacked frames to refine event representations. Furthermore, to achieve unified feature encoding across modalities, we develop High-Frequency Guided Selector (HGS) for semantic consistency token selection guided by dynamic edge features while suppressing redundant multimodal information interference adaptively. Extensive experiments on four benchmark datasets demonstrate the superior effectiveness of our proposed framework. The code is available at https://github.com/caomq123/EMP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Fine-Grained Image Retrieval via Dual-Vision AdaptationXin Jiang, Meiqi Cao, Hao Tang, Fei Shen 等AAAI 2026 · 被引用 1 次
- Seeing Motion Through Polarity for Event-based Action RecognitionMeiqi Cao, Jiachao Zhang, Xin Jiang, Rui Yan 等CVPR 2026
- DiT-Distill: Open-Set Fine-Grained Retrieval via Generative Curriculum KnowledgeXin Jiang, Hao Tang, Meiqi Cao, Junyao Gao 等CVPR 2026
- Spatiotemporal-Untrammelled Mixture of Experts for Multi-Person Motion PredictionZheng Yin, Chengjian Li, Xiangbo Shu, Meiqi Cao 等AAAI 2026
它引用的顶会 Paper23
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo 等NeurIPS 2021 · 被引用 2,148 次
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu 等NeurIPS 2021 · 被引用 1,343 次
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski 等AAAI 2022 · 被引用 529 次
相关 Paper
- CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training FrameworkWentao Wu, Xiao Wang, Chenglong Li, Bo Jiang 等ACM MM 2025 · 被引用 2 次
- Frame-Event Alignment and Fusion Network for High Frame Rate TrackingJiqing Zhang, Yuanchen Wang, Wenxi Liu, Meng Li 等CVPR 2023
- ExACT: Language-Guided Conceptual Reasoning and Uncertainty Estimation for Event-Based Action Recognition and MoreJiazhou Zhou, Xu Zheng, Yuanhuiyi Lyu, Lin WangCVPR 2024 · 被引用 15 次
- Frequency-Aware Event-Based Video Deblurring for Real-World Motion BlurTaewoo Kim, Hoonhee Cho, Kuk-Jin YoonCVPR 2024
- Complementing Event Streams and RGB Frames for Hand Mesh ReconstructionJianping Jiang, Xinyu Zhou, Bingxuan Wang, Xiaoming Deng 等CVPR 2024
