Branchformer: Parallel MLP-Attention Architectures to Capture Local and Global Context for Speech Recognition and Understanding
Yifan Peng, Siddharth Dalmia, Ian R. Lane, Shinji Watanabe
Abstract
Conformer has proven to be effective in many speech processing tasks. It combines the benefits of extracting local dependencies using convolutions and global dependencies using self-attention. Inspired by this, we propose a more flexible, interpretable and customizable encoder alternative, Branchformer, with parallel branches for modeling various ranged dependencies in end-to-end speech processing. In each encoder layer, one branch employs self-attention or its variant to capture long-range dependencies, while the other branch utilizes an MLP module with convolutional gating (cgMLP) to extract local relationships. We conduct experiments on several speech recognition and spoken language understanding benchmarks. Results show that our model outperforms both Transformer and cgMLP. It also matches with or outperforms state-of-the-art results achieved by Conformer. Furthermore, we show various strategies to reduce computation thanks to the two-branch architecture, including the ability to have variable inference complexity in a single trained model. The weights learned for merging branches indicate how local and global dependencies are utilized in different layers, which benefits model designing. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd5706bd-e42a-40f3-af78-aafaf959a382Cited by top-tier papers15
- Zipformer: A faster and better encoder for automatic speech recognitionZengwei Yao, Liyong Guo, Xiaoyu Yang, Wei Kang et al.ICLR 2024 · 155 citations
- How Much Temporal Long-Term Context is Needed for Action Segmentation?Emad Bahrami Rad, Gianpiero Francesca, Juergen GallICCV 2023 · 54 citations
- Efficient Multi-View Graph Clustering with Local and Global Structure PreservationYi Wen, Suyuan Liu, Xinhang Wan, Siwei Wang et al.ACM MM 2023 · 39 citations
- Graph Convolutions Enrich the Self-Attention in Transformers!Jeongwhan Choi, Hyowon Wi, Jayoung Kim, Yehjin Shin et al.NeurIPS 2024 · 24 citations
- Towards Robust Speech Representation Learning for Thousands of LanguagesWilliam Chen, Wangyou Zhang, Yifan Peng, Xinjian Li et al.EMNLP 2024 · 19 citations
Builds on6
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- Pay Attention to MLPsHanxiao Liu, Zihang Dai, David R. So, Quoc V. LeNeurIPS 2021 · 912 citations
- SLURP: A Spoken Language Understanding Resource PackageEmanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, Verena RieserEMNLP 2020 · 129 citations
Related papers
- Squeezeformer: An Efficient Transformer for Automatic Speech RecognitionSehoon Kim, Amir Gholami, Albert E. Shaw, Nicholas Lee et al.NeurIPS 2022 · 152 citations
- Long-Short Decision Transformer: Bridging Global and Local Dependencies for Generalized Decision-MakingJincheng Wang, Penny Karanasou, Pengyuan Wei, Elia Gatti et al.ICLR 2025
- Brainformers: Trading Simplicity for EfficiencyYanqi Zhou, Nan Du, Yanping Huang, Daiyi Peng et al.ICML 2023 · 38 citations
- Separate and Reconstruct: Asymmetric Encoder-Decoder for Speech SeparationUi-Hyeop Shin, Sangyoun Lee, Taehan Kim, Hyung-Min ParkNeurIPS 2024 · 46 citations
- An efficient encoder-decoder architecture with top-down attention for speech separationKai Li, Runxuan Yang, Xiaolin HuICLR 2023 · 16 citations
