UALM: Unified Audio Language Model for Understanding, Generation and Reasoning
Jinchuan Tian, Sang-gil Lee, Zhifeng Kong, Sreyan Ghosh, Arushi Goel, Chao-Han Huck Yang, Wenliang Dai, Zihan Liu, Hanrong Ye, Shinji Watanabe, Mohammad Shoeybi, Bryan Catanzaro
Abstract
Recent advances in the audio language modeling (ALM) domain tackle audio understanding and text-to-audio generation as separate tasks. Very few studies attempt to unify these tasks -an essential step toward advanced multimodal reasoning. This paper introduces Unified Audio Language Model (UALM), which aims to unify audio understanding, text-to-audio generation, and multimodal reasoning in a single model. To achieve this goal, we first present UALM-Gen, a text-to-audio language model that directly predicts audio tokens and is comparable to state-of-the-art diffusion-based models. We then demonstrate, using proper data blending, training recipes, and inference techniques, that our single UALM model matches the quality of state-of-the-art specialized models in audio understanding, text-to-audio generation, and text reasoning. Furthermore, we present UALM-Reason, a multimodal reasoning model that utilizes both text and audio in the intermediate thinking steps to facilitate complex generation tasks. To our knowledge, this is the first demonstration in audio research of cross-modal generative reasoning, with its effectiveness confirmed by subjective evaluations. * Equal Contribution; Alphabetic Order. CMU 1 . NVIDIA 2 . UMD 3 . †: Work done during an internship at NVIDIA. ‡: Technical Lead. Preliminary work. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe527f22-0d86-4f24-98be-af7bb4325b53Cited by top-tier papers2
- AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion ForcingWilliam Chen, Prem Seetharaman, Rithesh Kumar, Oriol Nieto et al.ICML 2026 · 7 citations
- Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and EditingZeyue Tian, Binxin Yang, Zhaoyang Liu, Jiexuan Zhang et al.SIGGRAPH 2026
Builds on24
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar et al.NeurIPS 2023 · 910 citations
Related papers
- Audio Entailment: Assessing Deductive Reasoning for Audio UnderstandingSoham Deshmukh, Shuo Han, Hazim T. Bukhari, Benjamin Elizalde et al.AAAI 2025 · 23 citations
- Echo: Towards Advanced Audio Comprehension via Audio-Interleaved ReasoningDaiqing Wu, Xuan Zhang, Dongbao Yang, Jiashu Yao et al.ICLR 2026 · 6 citations
- U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generationxiang deng, Feng Gao, Yong Zhang, Youxin Pang et al.CVPR 2026 · 2 citations
- Audio-Reasoner: Improving Reasoning Capability in Large Audio Language ModelsZhifei Xie, Mingbao Lin, Zihang Liu, Pengcheng Wu et al.EMNLP 2025 · 5 citations
- UniAudio 1.5: Large Language Model-Driven Audio Codec is A Few-Shot Audio Task LearnerDongchao Yang, Haohan Guo, Yuanyuan Wang, Rongjie Huang et al.NeurIPS 2024 · 55 citations
