SVL: Empowering Spiking Neural Networks for Efficient 3D Open-World Understanding
Xuerui Qiu, Shaowei Gu, Peixi Wu, JiaKui Hu, Yaozhi Wen, Yuqi Pan, Xinhao Luo, Bo XU, Guoqi Li
摘要
Spiking Neural Networks (SNNs) offer an energy--efficient route to 3D spatio--temporal perception, yet they lag behind Artificial Neural Networks (ANNs) due to weak pretraining and heavy inference stacks, limiting generalization and multimodal reasoning (e.g., zero--shot 3D classification and open--world QA). We present a universal Spike--based Vision--Language pretraining framework (SVL) that equips SNNs with open--world 3D understanding while preserving end--to--end spike efficiency. SVL comprises two core components: (i) Multi--scale Triple Alignment (MTA), a label--free triplet contrastive objective aligning 3D, image, and text; and (ii) Re--parameterizable Vision--Language Integration (Rep--VLI), which converts offline text embeddings into lightweight weights for text--encoder--free inference. Moreover, we present the first fully spike--driven point Transformer, Spike-driven PointFormer, whose 3D spike--driven self--attention (3D-SDSA) reduces interactions to sparse additions, enabling faster, more efficient training. Extensive experiments show that SVL attains strong zero--shot 3D classification (85.4% top--1) and consistently outperforms prior SNNs on downstream tasks (e.g., +6.1% 3D cls, +2.1% DVS actions, +1.1% detection, +2.1% segmentation) while enabling open--world 3D question answering, sometimes outperforming ANNs. To the best of our knowledge, SVL represents the first scalable, generalizable, and hardware-friendly paradigm for 3D open-world understanding, effectively bridging the gap between SNNs and ANNs in complex open-world understanding tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou 等AAAI 2021 · 被引用 1,128 次
- Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World DataMikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen 等ICCV 2019 · 被引用 1,003 次
相关 Paper
- SpikingResformer: Bridging ResNet and Vision Transformer in Spiking Neural NetworksXinyu Shi, Zecheng Hao, Zhaofei YuCVPR 2024 · 被引用 53 次
- UniGS: Unified Language-Image-3D Pretraining with Gaussian SplattingHaoyuan Li, Yanpeng Zhou, Tao Tang, Jifei Song 等ICLR 2025
- QKFormer: Hierarchical Spiking Transformer using Q-K AttentionChenlin Zhou, Han Zhang, Zhaokun Zhou, Liutao Yu 等NeurIPS 2024 · 被引用 126 次
- Spikformer: When Spiking Neural Network Meets TransformerZhaokun Zhou, Yuesheng Zhu, Chao He, Yaowei Wang 等ICLR 2023 · 被引用 103 次
- ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D UnderstandingLe Xue, Mingfei Gao, Chen Xing, Roberto Martín-Martín 等CVPR 2023
