TRAMBA: A Hybrid Transformer and Mamba Architecture for Practical Audio and Bone Conduction Speech Super Resolution and Enhancement on Mobile and Wearable Platforms
Yueyuan Sui, Minghui Zhao, Junxi Xia, Xiaofan Jiang, Stephen Xia
Abstract
We propose TRAMBA, a hybrid transformer and Mamba architecture for acoustic and bone conduction speech enhancement, suitable for mobile and wearable platforms. Bone conduction speech enhancement has been impractical to adopt in mobile and wearable platforms for several reasons: (i) data collection is labor-intensive, resulting in scarcity; (ii) there exists a performance gap between state-of-art models with memory footprints of hundreds of MBs and methods better suited for resource-constrained systems. To adapt TRAMBA to vibration-based sensing modalities, we pre-train TRAMBA with audio speech datasets that are widely available. Then, users fine-tune with a small amount of bone conduction data. TRAMBA outperforms state-of-art GANs by up to 7.3% in Perceptual Evaluation of Speech Quality (PESQ) and 1.8% in Short-Time Objective Intelligibility (STOI), with an order of magnitude smaller memory footprint and an inference speed up of up to 465 times. We integrate TRAMBA into real systems and show that TRAMBA (i) improves battery life of wearables by up to 160% by requiring less data sampling and transmission; (ii) generates higher quality voice in noisy environments than over-the-air speech; (iii) requires a memory footprint of less than 20.0 MB.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6338203b-4164-4048-b5e1-b86f05632879Cited by top-tier papers2
- CIS-BWE: Chaos-Informed Speech Bandwidth ExtensionTarikul Islam Tamiti, Tonmoy Das, Nursadul Mamun, Anomadarshi BaruaACL 2026
- SIGMA-ASL: Sensor-Integrated Multimodal Dataset for Sign Language RecognitionXiaofang Xiao, Guangchao Li, Guangrong Zhao, Qi Lin et al.UbiComp 2026
Builds on8
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab et al.NeurIPS 2021 · 1,280 citations
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song et al.ICLR 2021 · 122 citations
- MuteIt: Jaw Motion Based Unvoiced Command Recognition Using EarableTanmay Srivastava, Prerna Khanna, Shijia Pan, Phuc Nguyen et al.UbiComp 2022 · 52 citations
Related papers
- BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech EnhancementCunhang Fan, Enrui Liu, Andong Li, Jianhua Tao et al.AAAI 2025
- Jamba: Hybrid Transformer-Mamba Language ModelsBarak Lenz, Opher Lieber, Alan Arazi, Amir Bergman et al.ICLR 2025
- Hybrid Neural Networks for On-Device Directional HearingAnran Wang, Maruchi Kim, Hao Zhang, Shyamnath GollakotaAAAI 2022 · 18 citations
- NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech InteractionJun Rekimoto, Yu Nishimura, Bojian YangCHI 2026 · 1 citation
- mmMUSE: An mmWave-based Motion-resilient Universal Speech Enhancement SystemLingyu Wang, Kai Wang, Dequan Wang, You Zuo et al.UbiComp 2026 · 2 citations
