Zipformer: A faster and better encoder for automatic speech recognition
Zengwei Yao, Liyong Guo, Xiaoyu Yang, Wei Kang, Fangjun Kuang, Yifan Yang, Zengrui Jin, Long Lin, Daniel Povey
摘要
The Conformer has become the most popular encoder model for automatic speech recognition (ASR). It adds convolution modules to a transformer to learn both local and global dependencies. In this work we describe a faster, more memory-efficient, and better-performing transformer, called Zipformer. Modeling changes include: 1) a U-Net-like encoder structure where middle stacks operate at lower frame rates; 2) reorganized block structure with more modules, within which we re-use attention weights for efficiency; 3) a modified form of LayerNorm called BiasNorm allows us to retain some length information; 4) new activation functions SwooshR and SwooshL work better than Swish. We also propose a new optimizer, called ScaledAdam, which scales the update by each tensor's current scale to keep the relative change about the same, and also explictly learns the parameter scale. It achieves faster convergence and better performance than Adam. Extensive experiments on LibriSpeech, Aishell-1, and WenetSpeech datasets demonstrate the effectiveness of our proposed Zipformer over other state-of-the-art ASR models. Our code is publicly available at https://github.com/k2-fsa/icefall.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- EXP-Bench: Can AI Conduct AI Research Experiments?Patrick Tser Jern Kon, Qiuyi Ding, Jiachen Liu, Xinyi Zhu 等ICLR 2026 · 被引用 35 次
- Speech Recognition Meets Large Language Model: Benchmarking, Models, and ExplorationZiyang Ma, Guanrou Yang, Yifan Yang, Zhifu Gao 等AAAI 2025 · 被引用 17 次
- TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal GroundingMingyue Huo, Yiwen Shao, Yuheng ZhangACL 2026 · 被引用 12 次
- SPEAR: A Unified SSL Framework for Learning Speech and Audio RepresentationsXiaoyu Yang, Yifan Yang, Zengrui Jin, Ziyun Cui 等ICML 2026 · 被引用 12 次
- ZIPA: A family of efficient models for multilingual phone recognitionJian Zhu, Farhan Samir, Eleanor Chodroff, David R. MortensenACL 2025 · 被引用 10 次
它引用的顶会 Paper3
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Branchformer: Parallel MLP-Attention Architectures to Capture Local and Global Context for Speech Recognition and UnderstandingYifan Peng, Siddharth Dalmia, Ian R. Lane, Shinji WatanabeICML 2022 · 被引用 203 次
- Squeezeformer: An Efficient Transformer for Automatic Speech RecognitionSehoon Kim, Amir Gholami, Albert E. Shaw, Nicholas Lee 等NeurIPS 2022 · 被引用 152 次
相关 Paper
- Dual-mode ASR: Unify and Improve Streaming ASR with Full-context ModelingJiahui Yu, Wei Han, Anmol Gulati, Chung-Cheng Chiu 等ICLR 2021 · 被引用 80 次
- DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted AveragingMatteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin JaggiNeurIPS 2024 · 被引用 60 次
- Brainformers: Trading Simplicity for EfficiencyYanqi Zhou, Nan Du, Yanping Huang, Daiyi Peng 等ICML 2023 · 被引用 38 次
- LiteASR: Efficient Automatic Speech Recognition with Low-Rank ApproximationKeisuke Kamahori, Jungo Kasai, Noriyuki Kojima, Baris KasikciEMNLP 2025 · 被引用 1 次
- Deconstructing What Makes a Good Optimizer for Autoregressive Language ModelsRosie Zhao, Depen Morwani, David Brandfonbrener, Nikhil Vyas 等ICLR 2025
