Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
Juncheng Wang, Chao Xu, Cheng Yu, Zhe Hu, Haoyu Xie, Guoqi Yu, Lei Shang, Shujun Wang
Abstract
While language models (LMs) paired with residual vector quantization (RVQ) tokenizers have shown promise in text-to-audio (T2A) generation, they still lag behind diffusion-based models by a non-trivial margin. We identify a critical dilemma underpinning this gap: incorporating more RVQ layers improves audio reconstruction fidelity but exceeds the generation capacity of conventional LMs. To address this, we first analyze RVQ dynamics and uncover two key limitations: 1) orthogonality of features across RVQ layers hinders effective LMs training, and 2) descending semantic richness in tokens from deeper RVQ layers exacerbates exposure bias during autoregressive decoding. Based on these insights, we propose Siren, a novel LM-based framework that employs multiple isolated transformers with causal conditioning and anti-causal alignment via reinforcement learning. Extensive experiments demonstrate that Siren outperforms both existing LM-based and diffusion-based T2A systems, achieving state-of-the-art results. By bridging the representational strengths of LMs with the fidelity demands of audio synthesis, our approach repositions LMs as competitive contenders against diffusion models in T2A tasks. Moreover, by aligning audio representations with linguistic structures, Siren facilitates a promising pathway toward unified multimodal generation frameworks. The code is released at https://github.com/wjc2830/Siren.git .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bbc82d8d-e984-487f-bcc4-a3e28294add3Cited by top-tier papers2
- Decentralized Attention Fails Centralized Signals: Rethinking Transformers for Medical Time SeriesGuoqi Yu, Juncheng Wang, Chen Yang, Jing Qin et al.ICLR 2026 · 6 citations
- OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video GenerationDonghao Zhou, Guisheng Liu, Hao Yang, Jiatong Li et al.ICML 2026 · 3 citations
Builds on34
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Hierarchical Semantic-Acoustic Modeling via Semi-Discrete Residual Representations for Expressive End-to-End Speech SynthesisYixuan Zhou, Guoyang Zeng, Xin Liu, Xiang Li et al.ICLR 2026
- Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token PredictionShu-Wen Yang, Byeonggeun Kim, Kuan-Po Huang, Qingming Tang et al.ICML 2025
- Closing the Modality Reasoning Gap for Speech Large Language ModelsChaoren Wang, Heng Lu, Xueyao Zhang, Shujie Liu et al.ACL 2026 · 10 citations
- Scaling Transformers for End-to-End Discrete Audio TokenizationYitian Gong, Kuangwei Chen, Zhaoye Fei, Xiaogui Yang et al.ICML 2026
- Token-Efficient Long-Term Interest Sketching and Internalized Reasoning for LLM-based RecommendationZhihao Ding, Jinming Li, Shuai Mu, Jieming ShiICLR 2026
