LLark: A Multimodal Instruction-Following Language Model for Music
Joshua Patrick Gardner, Simon Durand, Daniel Stoller, Rachel M. Bittner
Abstract
Music has a unique and complex structure which is challenging for both expert humans and existing AI systems to understand, and presents unique challenges relative to other forms of audio. We present LLark, an instruction-tuned multimodal model for music understanding. We detail our process for dataset creation, which involves augmenting the annotations of diverse open-source music datasets and converting them to a unified instruction-tuning format. We propose a multimodal architecture for LLark, integrating a pretrained generative model for music with a pretrained language model. In evaluations on three types of tasks (music understanding, captioning, reasoning), we show that LLark matches or outperforms existing baselines in music understanding, and that humans show a high degree of agreement with its responses in captioning and reasoning tasks. LLark is trained entirely from open-source music data and models, and we make our training code available along with the release of this paper. Additional results and audio examples are at https://bit.ly/llark, and our source code is available at https://github.com/spotify-research/llark .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Music Flamingo: Scaling Music Understanding in Audio Language ModelsSreyan Ghosh, Arushi Goel, Lasha Koroshinadze, Sang-gil Lee et al.ICLR 2026 · 33 citations
- LLM2Fx-Tools: Tool Calling for Music Post-ProductionSeungHeon Doh, Junghyun Koo, Marco A. Martínez-Ramírez, Woosung Choi et al.ICLR 2026 · 10 citations
- NatureLM-audio: an Audio-Language Foundation Model for BioacousticsDavid Robinson, Marius Miron, Masato Hagiwara, Olivier PietquinICLR 2025
- AnyGPT: Unified Multimodal LLM with Discrete Sequence ModelingJun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou et al.ACL 2024
- DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction TuningZhuoyuan Mao, Mengjie Zhao, Qiyu Wu, Hiromi Wakaki et al.EMNLP 2025
Builds on2
Related papers
- Listen, Think, and UnderstandYuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky et al.ICLR 2024 · 247 citations
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language ModelsSreyan Ghosh, Arushi Goel, Jaehyeon Kim, Sonal Kumar et al.NeurIPS 2025 · 299 citations
- WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music ReasoningGagan Mundada, Yash Vishe, Amit Namburi, Xin Xu et al.EMNLP 2025 · 1 citation
- Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning AbilitiesSreyan Ghosh, Zhifeng Kong, Sonal Kumar, S. Sakshi et al.ICML 2025
- Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical ScoresCongren Dai, Yue Yang, Krinos Li, Huichi Zhou et al.ACL 2026 · 4 citations
