TRAMBA: A Hybrid Transformer and Mamba Architecture for Practical Audio and Bone Conduction Speech Super Resolution and Enhancement on Mobile and Wearable Platforms
Yueyuan Sui, Minghui Zhao, Junxi Xia, Xiaofan Jiang, Stephen Xia
摘要
We propose TRAMBA, a hybrid transformer and Mamba architecture for acoustic and bone conduction speech enhancement, suitable for mobile and wearable platforms. Bone conduction speech enhancement has been impractical to adopt in mobile and wearable platforms for several reasons: (i) data collection is labor-intensive, resulting in scarcity; (ii) there exists a performance gap between state-of-art models with memory footprints of hundreds of MBs and methods better suited for resource-constrained systems. To adapt TRAMBA to vibration-based sensing modalities, we pre-train TRAMBA with audio speech datasets that are widely available. Then, users fine-tune with a small amount of bone conduction data. TRAMBA outperforms state-of-art GANs by up to 7.3% in Perceptual Evaluation of Speech Quality (PESQ) and 1.8% in Short-Time Objective Intelligibility (STOI), with an order of magnitude smaller memory footprint and an inference speed up of up to 465 times. We integrate TRAMBA into real systems and show that TRAMBA (i) improves battery life of wearables by up to 160% by requiring less data sampling and transmission; (ii) generates higher quality voice in noisy environments than over-the-air speech; (iii) requires a memory footprint of less than 20.0 MB.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CIS-BWE: Chaos-Informed Speech Bandwidth ExtensionTarikul Islam Tamiti, Tonmoy Das, Nursadul Mamun, Anomadarshi BaruaACL 2026
- SIGMA-ASL: Sensor-Integrated Multimodal Dataset for Sign Language RecognitionXiaofang Xiao, Guangchao Li, Guangrong Zhao, Qi Lin 等UbiComp 2026
它引用的顶会 Paper8
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab 等NeurIPS 2021 · 被引用 1,280 次
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song 等ICLR 2021 · 被引用 122 次
- MuteIt: Jaw Motion Based Unvoiced Command Recognition Using EarableTanmay Srivastava, Prerna Khanna, Shijia Pan, Phuc Nguyen 等UbiComp 2022 · 被引用 52 次
相关 Paper
- BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech EnhancementCunhang Fan, Enrui Liu, Andong Li, Jianhua Tao 等AAAI 2025
- Jamba: Hybrid Transformer-Mamba Language ModelsBarak Lenz, Opher Lieber, Alan Arazi, Amir Bergman 等ICLR 2025
- Hybrid Neural Networks for On-Device Directional HearingAnran Wang, Maruchi Kim, Hao Zhang, Shyamnath GollakotaAAAI 2022 · 被引用 18 次
- NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech InteractionJun Rekimoto, Yu Nishimura, Bojian YangCHI 2026 · 被引用 1 次
- mmMUSE: An mmWave-based Motion-resilient Universal Speech Enhancement SystemLingyu Wang, Kai Wang, Dequan Wang, You Zuo 等UbiComp 2026 · 被引用 2 次
