Proxy Target: Bridging the Gap Between Discrete Spiking Neural Networks and Continuous Control
Zijie Xu, Tong Bu, Zecheng Hao, Jianhao Ding, Zhaofei Yu
Abstract
Spiking Neural Networks (SNNs) offer low-latency and energy-efficient decision making on neuromorphic hardware, making them attractive for Reinforcement Learning (RL) in resource-constrained edge devices. However, most RL algorithms for continuous control are designed for Artificial Neural Networks (ANNs), particularly the target network soft update mechanism, which conflicts with the discrete and non-differentiable dynamics of spiking neurons. We show that this mismatch destabilizes SNN training and degrades performance. To bridge the gap between discrete SNNs and continuous-control algorithms, we propose a novel proxy target framework. The proxy network introduces continuous and differentiable dynamics that enable smooth target updates, stabilizing the learning process. Since the proxy operates only during training, the deployed SNN remains fully energy-efficient with no additional inference overhead. Extensive experiments on continuous control benchmarks demonstrate that our framework consistently improves stability and achieves up to higher performance across various spiking neuron models. Notably, to the best of our knowledge, this is the first approach that enables SNNs with simple Leaky Integrate and Fire (LIF) neurons to surpass their ANN counterparts in continuous control. This work highlights the importance of SNN-tailored RL algorithms and paves the way for neuromorphic agents that combine high performance with low power consumption. Code is available at https://github.com/xuzijie32/Proxy-Target.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd953402-0d08-48a1-9355-7a44d516f9f8Cited by top-tier papers3
- CaRe-BN: Precise Moving Statistics for Stabilizing Spiking Neural Networks in Reinforcement LearningZijie Xu, Xinyu Shi, Yiting Dong, Zihan Huang et al.ICLR 2026 · 4 citations
- Error Amplification Limits ANN-to-SNN Conversion in Continuous ControlZijie Xu, Zihan Huang, Yiting Dong, Kang Chen et al.ICML 2026 · 2 citations
- SpikeVLA: Vision-Language-Action Models with Spiking Neural NetworksRuiqi Song, Dujun Nie, Siyu Teng, Baiyong Ding et al.ICML 2026 · 1 citation
Builds on9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Incorporating Learnable Membrane Time Constant to Enhance Learning of Spiking Neural NetworksWei Fang, Zhaofei Yu, Yanqi Chen, Timothée Masquelier et al.ICCV 2021 · 731 citations
- Optimal ANN-SNN Conversion for High-accuracy and Ultra-low-latency Spiking Neural NetworksTong Bu, Wei Fang, Jianhao Ding, Penglin Dai et al.ICLR 2022 · 272 citations
- Optimized Potential Initialization for Low-Latency Spiking Neural NetworksTong Bu, Jianhao Ding, Zhaofei Yu, Tiejun HuangAAAI 2022 · 112 citations
- Strategy and Benchmark for Converting Deep Q-Networks to Event-Driven Spiking Neural NetworksWeihao Tan, Devdhar Patel, Robert KozmaAAAI 2021 · 54 citations
Related papers
- SpikeDyn: A Framework for Energy-Efficient Spiking Neural Networks with Continual and Unsupervised Learning Capabilities in Dynamic EnvironmentsRachmad Vidya Wicaksana Putra, Muhammad ShafiqueDAC 2021 · 4 citations
- Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike RepresentationQingyan Meng, Mingqing Xiao, Shen Yan, Yisen Wang et al.CVPR 2022 · 114 citations
- Differentiable Spike: Rethinking Gradient-Descent for Training Spiking Neural NetworksYuhang Li, Yufei Guo, Shanghang Zhang, Shikuang Deng et al.NeurIPS 2021 · 288 citations
- SpikeCLR: Self-Supervised Contrastive Learning for Visual Representations with Spiking Neural NetworksChengwei Zhou, Gourav DattaICML 2026
- Temporal-Coded Deep Spiking Neural Network with Easy Training and Robust PerformanceShibo Zhou, Xiaohua Li, Ying Chen, Sanjeev Tannirkulam Chandrasekaran et al.AAAI 2021 · 114 citations
