Exploiting Frequency Dynamics for Enhanced Multimodal Event-Based Action Recognition
Meiqi Cao, Xiangbo Shu, Xin Jiang, Rui Yan, Yazhou Yao, Jinhui Tang
Abstract
While event cameras excel in capturing microsecond temporal dynamics, they suffer from sparse spatial representations compared to traditional RGB data. Thus, multimodal event-based action recognition approaches aim to synergize complementary strengths by independently extracting and integrating paired RGB-Event features. However, this paradigm inevitably introduces additional data acquisition costs, while eroding the inherent privacy advantages of event-based sensing. Drawing inspiration from event-to-image reconstruction, texture-enriched visual representation directly reconstructed from asynchronous event streams is a promising solution. In response, we propose an Enhanced Multimodal Perceptual (EMP) framework that hierarchically explores multimodal cues (e.g., edges and textures) from raw event streams through two synergistic innovations spanning representation to feature levels. Specifically, we introduce Cross-Modal Frequency Enhancer (CFE) that leverages complementary frequency characteristics between reconstructed frames and stacked frames to refine event representations. Furthermore, to achieve unified feature encoding across modalities, we develop High-Frequency Guided Selector (HGS) for semantic consistency token selection guided by dynamic edge features while suppressing redundant multimodal information interference adaptively. Extensive experiments on four benchmark datasets demonstrate the superior effectiveness of our proposed framework. The code is available at https://github.com/caomq123/EMP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f02694a-aec4-4cba-b134-bde37e9cecfcCited by top-tier papers4
- Fine-Grained Image Retrieval via Dual-Vision AdaptationXin Jiang, Meiqi Cao, Hao Tang, Fei Shen et al.AAAI 2026 · 1 citation
- Seeing Motion Through Polarity for Event-based Action RecognitionMeiqi Cao, Jiachao Zhang, Xin Jiang, Rui Yan et al.CVPR 2026
- DiT-Distill: Open-Set Fine-Grained Retrieval via Generative Curriculum KnowledgeXin Jiang, Hao Tang, Meiqi Cao, Junyao Gao et al.CVPR 2026
- Spatiotemporal-Untrammelled Mixture of Experts for Multi-Person Motion PredictionZheng Yin, Chengjian Li, Xiangbo Shu, Meiqi Cao et al.AAAI 2026
Builds on23
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo et al.NeurIPS 2021 · 2,148 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
Related papers
- CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training FrameworkWentao Wu, Xiao Wang, Chenglong Li, Bo Jiang et al.ACM MM 2025 · 2 citations
- Frame-Event Alignment and Fusion Network for High Frame Rate TrackingJiqing Zhang, Yuanchen Wang, Wenxi Liu, Meng Li et al.CVPR 2023
- ExACT: Language-Guided Conceptual Reasoning and Uncertainty Estimation for Event-Based Action Recognition and MoreJiazhou Zhou, Xu Zheng, Yuanhuiyi Lyu, Lin WangCVPR 2024 · 15 citations
- Frequency-Aware Event-Based Video Deblurring for Real-World Motion BlurTaewoo Kim, Hoonhee Cho, Kuk-Jin YoonCVPR 2024
- Complementing Event Streams and RGB Frames for Hand Mesh ReconstructionJianping Jiang, Xinyu Zhou, Bingxuan Wang, Xiaoming Deng et al.CVPR 2024
