Spike-HTR: Spiking Neural Transformer for Handwritten Text Recognition
Xiubo Liang, Jinxing Han, Yuke Li, Haoqi Zhu, Yu Zhao, Hongzhi Wang
Abstract
Handwritten Text Recognition (HTR) is computationally imbalanced in two ways: most image pixels are background, and many width-axis sequence positions are blank-dominated. This creates a mismatch for Spiking Neural Networks (SNNs): handwriting is observed as a static image, whereas spiking computation unfolds over timesteps. We propose Spike-HTR, a hybrid spiking recognizer that controls both the number of spiking steps and the number of width positions processed by the deep sequence mixer. To make a static image suitable for short-horizon spiking inference, InkCoder converts it into a coarse-to-fine input stream, where early steps cover broad stroke regions and later steps emphasize sharper stroke details. To reduce sequence computation, a CTC-guided length reducer keeps likely character or uncertain positions and compresses long blank-dominated stretches before deep mixing. With , Spike-HTR trains only on target data, decodes without language models or lexicons, and reaches validation/test CERs of 3.5/5.4, 2.3/2.5, and 4.2/3.9 on IAM, LAM, and READ2016. Codes are available at https://github.com/QomolangmaH/SpikeHTR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0efa54f6-c8fe-47ba-99c4-21e112fe9282Builds on10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Incorporating Learnable Membrane Time Constant to Enhance Learning of Spiking Neural NetworksWei Fang, Zhaofei Yu, Yanqi Chen, Timothée Masquelier et al.ICCV 2021 · 731 citations
- Going Deeper With Directly-Trained Larger Spiking Neural NetworksHanle Zheng, Yujie Wu, Lei Deng, Yifan Hu et al.AAAI 2021 · 694 citations
- Spike-driven TransformerMan Yao, Jiakui Hu, Zhaokun Zhou, Li Yuan et al.NeurIPS 2023 · 368 citations
Related papers
- Learning to Generate Stylized Handwritten Text via a Unified Representation of Style, Content, and NoiseHonglie Wang, Yan-Ming Zhang, Wangzi Yao, Fei Yin et al.ICLR 2026
- SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command RecognitionJiaqi Wang, Liutao Yu, Xiongri Shen, Sihang Guo et al.AAAI 2026 · 1 citation
- Complex Dynamic Neurons Improved Spiking Transformer Network for Efficient Automatic Speech RecognitionQingyu Wang, Tielin Zhang, Minglun Han, Yi Wang et al.AAAI 2023 · 40 citations
- Sketchformer: Transformer-Based Representation for Sketched StructureLeo Sampaio Ferraz Ribeiro, Tu Bui, John P. Collomosse, Moacir PontiCVPR 2020
- OrigamiNet: Weakly-Supervised, Segmentation-Free, One-Step, Full Page Text Recognition by learning to unfoldMohamed Yousef, Tom E. BishopCVPR 2020
