Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning
Wenrui Li, Xi-Le Zhao, Zhengyu Ma, Xingtao Wang, Xiaopeng Fan, Yonghong Tian
摘要
Audio-visual zero-shot learning (ZSL) has attracted board attention, as it could classify video data from classes that are not observed during training. However, most of the existing methods are restricted to background scene bias and fewer motion details by employing a single-stream network to process scenes and motion information as a unified entity. In this paper, we address this challenge by proposing a novel dual-stream architecture Motion-Decoupled Spiking Transformer (MDFT) to explicitly decouple the contextual semantic information and highly sparsity dynamic motion information. Specifically, The Recurrent Joint Learning Unit (RJLU) could extract contextual semantic information effectively and understand the environment in which actions occur by capturing joint knowledge between different modalities. By converting RGB images to events, our approach effectively captures motion information while mitigating the influence of background scene biases, leading to more accurate classification results. We utilize the inherent strengths of Spiking Neural Networks (SNNs) to process highly sparsity event data efficiently. Additionally, we introduce a Discrepancy Analysis Block (DAB) to model the audio motion features. To enhance the efficiency of SNNs in extracting dynamic temporal and motion information, we dynamically adjust the threshold of Leaky Integrate-and-Fire (LIF) neurons based on the statistical cues of global motion and contextual semantic information. Our experiments demonstrate the effectiveness of MDFT, which consistently outperforms state-of-the-art methods across mainstream benchmarks. Moreover, we find that motion information serves as a powerful regularization for video networks, where using it improves the accuracy of HM and ZSL by 19.1% and 38.4%, respectively.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- OTIAS: OcTree Implicit Adaptive Sampling for Multispectral and Hyperspectral Image FusionShangqi Deng, Jun Ma, Liang-Jian Deng, Ping WeiAAAI 2025 · 被引用 10 次
- A Unified Understanding of Adversarial Vulnerability Regarding Unimodal Models and Vision-Language Pre-training ModelsHaonan Zheng, Xinyang Deng, Wen Jiang, Wenrui LiACM MM 2024 · 被引用 4 次
- Toward a Stable, Fair, and Comprehensive Evaluation of Object Hallucination in Large Vision-Language ModelsHongliang Wei, Xingtao Wang, Xianqi Zhang, Xiaopeng Fan 等NeurIPS 2024 · 被引用 4 次
- SAMPLE: Semantic Alignment through Temporal-Adaptive Multimodal Prompt Learning for Event-Based Open-Vocabulary Action RecognitionJing Wang, Rui Zhao, Ruiqin Xiong, Xingtao Wang 等ICCV 2025 · 被引用 2 次
- Mind marginal non-crack regions: Clustering-inspired representation learning for crack segmentationZhuangzhuang Chen, Zhuonan Lai, Jie Chen, Jianqiang LiCVPR 2024
相关 Paper
- SpikingVTG: A Spiking Detection Transformer for Video Temporal GroundingMalyaban Bal, Brian Matejek, Susmit Jha, Adam D. CobbNeurIPS 2025 · 被引用 1 次
- Isomer: Isomerous Transformer for Zero-shot Video Object SegmentationYichen Yuan, Yifan Wang, Lijun Wang, Xiaoqi Zhao 等ICCV 2023 · 被引用 16 次
- Dual Conditioned Motion Diffusion for Pose-Based Video Anomaly DetectionHongsong Wang, Andi Xu, Pinle Ding, Jie GuiAAAI 2025 · 被引用 8 次
- DFCNet: Dual-Factor Compensatory Clustering Network for Modality-Imbalanced Generalized Zero-Shot LearningXiangyu Shan, Heng Song, Junwu ZhuACM MM 2025
- UCF-Crime-DVS: A Novel Event-Based Dataset for Video Anomaly Detection with Spiking Neural NetworksYuanbin Qian, Shuhan Ye, Chong Wang, Xiaojie Cai 等AAAI 2025 · 被引用 18 次
