PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation
Chenning Xu, Mao Zheng, Mingyu Zheng, Mingyang Song
Abstract
Podcast script generation requires LLMs to synthesize structured, context-grounded dialogue from diverse inputs, yet systematic evaluation resources for this task remain limited. To bridge this gap, we introduce PodBench, a benchmark comprising 800 samples with inputs up to 21K tokens and complex multi-speaker instructions. We propose a multifaceted evaluation framework that integrates quantitative constraints with LLM-based quality assessment. Extensive experiments reveal that while proprietary models generally excel, open-source models equipped with explicit reasoning demonstrate superior robustness in handling long contexts and multi-speaker coordination compared to standard baselines. However, our analysis uncovers a persistent divergence where high instruction following does not guarantee high content substance. PodBench offers a reproducible testbed to address these challenges in long-form, audio-centric generation. The code and data are publicly accessible here. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf3ad157-0c51-44d4-9058-a24a29fba2ebBuilds on5
- Reverse-Engineered Reasoning for Open-Ended GenerationHaozhe Wang, Haoran Que, Qixin Xu, Minghao Liu et al.ICLR 2026 · 38 citations
- MoonCast: High-Quality Zero-Shot Podcast GenerationZeqian Ju, Dongchao Yang, Kai Shen, Yichong Leng et al.NeurIPS 2025 · 32 citations
- Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement LearningXuanyu Lei, Chenliang Li, Yuning Wu, Kaiming Liu et al.ACL 2026 · 8 citations
- LongWriter-V: Enabling Ultra-Long and High-Fidelity Generation in Vision-Language ModelsShangqing Tu, Yucheng Wang, Daniel Zhang-Li, Yushi Bai et al.ACM MM 2025
- LongWriter: Unleashing 10, 000+ Word Generation from Long Context LLMsYushi Bai, Jiajie Zhang, Xin Lv, Linzhi Zheng et al.ICLR 2025
Related papers
- LongGenBench: Benchmarking Long-Form Generation in Long Context LLMsYuhao Wu, Ming Shan Hee, Zhiqiang Hu, Roy Ka-Wei LeeICLR 2025
- ReFF: Reinforcing Format Faithfulness in Language Models Across Varied TasksJiashu Yao, Heyan Huang, Zeming Liu, Haoyu Wen et al.AAAI 2025 · 1 citation
- FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language ModelsYuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong et al.ACL 2024 · 10 citations
- ınftyBench: Extending Long Context Evaluation Beyond 100K TokensXinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu et al.ACL 2024
- CFBench: A Comprehensive Constraints-Following Benchmark for LLMsTao Zhang, Chenglin Zhu, Yanjun Shen, Wenjing Luo et al.ACL 2025 · 53 citations
