Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning
Wenrui Li, Xi-Le Zhao, Zhengyu Ma, Xingtao Wang, Xiaopeng Fan, Yonghong Tian
Abstract
Audio-visual zero-shot learning (ZSL) has attracted board attention, as it could classify video data from classes that are not observed during training. However, most of the existing methods are restricted to background scene bias and fewer motion details by employing a single-stream network to process scenes and motion information as a unified entity. In this paper, we address this challenge by proposing a novel dual-stream architecture Motion-Decoupled Spiking Transformer (MDFT) to explicitly decouple the contextual semantic information and highly sparsity dynamic motion information. Specifically, The Recurrent Joint Learning Unit (RJLU) could extract contextual semantic information effectively and understand the environment in which actions occur by capturing joint knowledge between different modalities. By converting RGB images to events, our approach effectively captures motion information while mitigating the influence of background scene biases, leading to more accurate classification results. We utilize the inherent strengths of Spiking Neural Networks (SNNs) to process highly sparsity event data efficiently. Additionally, we introduce a Discrepancy Analysis Block (DAB) to model the audio motion features. To enhance the efficiency of SNNs in extracting dynamic temporal and motion information, we dynamically adjust the threshold of Leaky Integrate-and-Fire (LIF) neurons based on the statistical cues of global motion and contextual semantic information. Our experiments demonstrate the effectiveness of MDFT, which consistently outperforms state-of-the-art methods across mainstream benchmarks. Moreover, we find that motion information serves as a powerful regularization for video networks, where using it improves the accuracy of HM and ZSL by 19.1% and 38.4%, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers5
- OTIAS: OcTree Implicit Adaptive Sampling for Multispectral and Hyperspectral Image FusionShangqi Deng, Jun Ma, Liang-Jian Deng, Ping WeiAAAI 2025 · 10 citations
- A Unified Understanding of Adversarial Vulnerability Regarding Unimodal Models and Vision-Language Pre-training ModelsHaonan Zheng, Xinyang Deng, Wen Jiang, Wenrui LiACM MM 2024 · 4 citations
- Toward a Stable, Fair, and Comprehensive Evaluation of Object Hallucination in Large Vision-Language ModelsHongliang Wei, Xingtao Wang, Xianqi Zhang, Xiaopeng Fan et al.NeurIPS 2024 · 4 citations
- SAMPLE: Semantic Alignment through Temporal-Adaptive Multimodal Prompt Learning for Event-Based Open-Vocabulary Action RecognitionJing Wang, Rui Zhao, Ruiqin Xiong, Xingtao Wang et al.ICCV 2025 · 2 citations
- Mind marginal non-crack regions: Clustering-inspired representation learning for crack segmentationZhuangzhuang Chen, Zhuonan Lai, Jie Chen, Jianqiang LiCVPR 2024
Related papers
- SpikingVTG: A Spiking Detection Transformer for Video Temporal GroundingMalyaban Bal, Brian Matejek, Susmit Jha, Adam D. CobbNeurIPS 2025 · 1 citation
- Isomer: Isomerous Transformer for Zero-shot Video Object SegmentationYichen Yuan, Yifan Wang, Lijun Wang, Xiaoqi Zhao et al.ICCV 2023 · 16 citations
- Dual Conditioned Motion Diffusion for Pose-Based Video Anomaly DetectionHongsong Wang, Andi Xu, Pinle Ding, Jie GuiAAAI 2025 · 8 citations
- DFCNet: Dual-Factor Compensatory Clustering Network for Modality-Imbalanced Generalized Zero-Shot LearningXiangyu Shan, Heng Song, Junwu ZhuACM MM 2025
- UCF-Crime-DVS: A Novel Event-Based Dataset for Video Anomaly Detection with Spiking Neural NetworksYuanbin Qian, Shuhan Ye, Chong Wang, Xiaojie Cai et al.AAAI 2025 · 18 citations
