LLark: A Multimodal Instruction-Following Language Model for Music
Joshua Patrick Gardner, Simon Durand, Daniel Stoller, Rachel M. Bittner
摘要
Music has a unique and complex structure which is challenging for both expert humans and existing AI systems to understand, and presents unique challenges relative to other forms of audio. We present LLark, an instruction-tuned multimodal model for music understanding. We detail our process for dataset creation, which involves augmenting the annotations of diverse open-source music datasets and converting them to a unified instruction-tuning format. We propose a multimodal architecture for LLark, integrating a pretrained generative model for music with a pretrained language model. In evaluations on three types of tasks (music understanding, captioning, reasoning), we show that LLark matches or outperforms existing baselines in music understanding, and that humans show a high degree of agreement with its responses in captioning and reasoning tasks. LLark is trained entirely from open-source music data and models, and we make our training code available along with the release of this paper. Additional results and audio examples are at https://bit.ly/llark, and our source code is available at https://github.com/spotify-research/llark .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Music Flamingo: Scaling Music Understanding in Audio Language ModelsSreyan Ghosh, Arushi Goel, Lasha Koroshinadze, Sang-gil Lee 等ICLR 2026 · 被引用 33 次
- LLM2Fx-Tools: Tool Calling for Music Post-ProductionSeungHeon Doh, Junghyun Koo, Marco A. Martínez-Ramírez, Woosung Choi 等ICLR 2026 · 被引用 10 次
- NatureLM-audio: an Audio-Language Foundation Model for BioacousticsDavid Robinson, Marius Miron, Masato Hagiwara, Olivier PietquinICLR 2025
- AnyGPT: Unified Multimodal LLM with Discrete Sequence ModelingJun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou 等ACL 2024
- DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction TuningZhuoyuan Mao, Mengjie Zhao, Qiyu Wu, Hiromi Wakaki 等EMNLP 2025
它引用的顶会 Paper2
相关 Paper
- Listen, Think, and UnderstandYuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky 等ICLR 2024 · 被引用 247 次
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language ModelsSreyan Ghosh, Arushi Goel, Jaehyeon Kim, Sonal Kumar 等NeurIPS 2025 · 被引用 299 次
- WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music ReasoningGagan Mundada, Yash Vishe, Amit Namburi, Xin Xu 等EMNLP 2025 · 被引用 1 次
- Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning AbilitiesSreyan Ghosh, Zhifeng Kong, Sonal Kumar, S. Sakshi 等ICML 2025
- Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical ScoresCongren Dai, Yue Yang, Krinos Li, Huichi Zhou 等ACL 2026 · 被引用 4 次
