MoonCast: High-Quality Zero-Shot Podcast Generation
Zeqian Ju, Dongchao Yang, Kai Shen, Yichong Leng, Zhengtao Wang, Songxiang Liu, Xinyu Zhou, Tao Qin, Xiangyang Li, Jianwei Yu, Xu Tan
Abstract
Recent advances in text-to-speech synthesis have achieved notable success in generating high-quality short utterances for individual speakers. However, these systems still face challenges when extending their capabilities to long, multi-speaker, and spontaneous dialogues, typical of real-world scenarios such as podcasts. These limitations arise from two primary challenges: 1) long speech: podcasts typically span several minutes, exceeding the upper limit of most existing work; 2) spontaneity: podcasts are marked by their spontaneous, oral nature, which sharply contrasts with formal, written contexts; existing works often fall short in capturing this spontaneity. In this paper, we propose MoonCast, a solution for high-quality zero-shot podcast generation, aiming to synthesize spontaneous podcast-style speech from text-only sources (e.g., stories, technical reports, news in TXT, PDF, or Web URL formats) using the voices of unseen speakers. To enable long audio generation, we employ a language model with parameter, data, and context scaling to process sequences in an innovative format designed for modeling entire multi-speaker, multi-turn speech interactions. To enhance spontaneity, we observe that ASR transcripts capture spontaneous speech details (e.g., filler words indicating hesitations, and specific punctuation and spaces reflecting breathing pauses), suggesting that these transcripts can serve as a partial indicator of speech spontaneity. Building upon this assumption, we utilize a script generation module to generate scripts incorporating these spontaneous elements. Experiments show MoonCast outperforms baselines, with notable improvements in contextual coherence and spontaneity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f7ac23d-f83a-4f44-b770-d2a05729139fCited by top-tier papers6
- CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow MatchingLeying Zhang, Yao Qian, Xiaofei Wang, Manthan Thakker et al.NeurIPS 2025 · 16 citations
- From Natural Alignment to Conditional Controllability in Multimodal DialogueZeyu Jin, Songtao Zhou, Haoyu Wang, Minghao Tian et al.ICLR 2026 · 2 citations
- TellWhisper: Tell Whisper Who Speaks WhenYifan Hu, Peiji Yang, Zhisheng Wang, Yicheng Zhong et al.ACL 2026 · 1 citation
- VibeVoice: Expressive Podcast Generation with Next-Token DiffusionZhiliang Peng, Jianwei Yu, Wenhui Wang, Yaoyao Chang et al.ICLR 2026
- PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script GenerationChenning Xu, Mao Zheng, Mingyu Zheng, Mingyang SongACL 2026
Builds on11
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar et al.NeurIPS 2023 · 910 citations
- Vector-quantized Image Modeling with Improved VQGANJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang et al.ICLR 2022 · 753 citations
Related papers
- Towards Abstractive Grounded Summarization of Podcast TranscriptsKaiqiang Song, Chen Li, Xiaoyang Wang, Dong Yu et al.ACL 2022 · 11 citations
- Long-Form Speech Generation with Spoken Language ModelsSe Jin Park, Julian Salazar, Aren Jansen, Keisuke Kinoshita et al.ICML 2025
- Beyond Transcripts: A Renewed Perspective on Audio ChapteringFabian Retkowski, Maike Züfle, Thai-Binh Nguyen, Jan Niehues et al.ACL 2026 · 2 citations
- FillerSpeech: Towards Human-Like Text-to-Speech Synthesis with Filler Insertion and Filler Style ControlSeung-Bin Kim, Junhyeok Cha, Hyung-Seok Oh, Heejin Choi et al.EMNLP 2025
- A Vector Quantized Approach for Text to Speech Synthesis on Real-World Spontaneous SpeechLi-Wei Chen, Shinji Watanabe, Alexander RudnickyAAAI 2023 · 49 citations
