DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning
Zhuoyuan Mao, Mengjie Zhao, Qiyu Wu, Hiromi Wakaki, Yuki Mitsufuji
Abstract
Recent advancements in music large language models (LLMs) have significantly improved music understanding tasks, which involve the model's ability to analyze and interpret various musical elements. These improvements primarily focused on integrating both music and text inputs. However, the potential of incorporating additional modalities such as images, videos and textual music features to enhance music understanding remains unexplored. To bridge this gap, we propose DeepResonance, a multimodal music understanding LLM fine-tuned via multiway instruction tuning with multi-way aligned music, text, image, and video data. To this end, we construct Music4way-MI2T, Music4way-MV2T, and Music4way-Any2T, three 4-way training and evaluation datasets designed to enable DeepResonance to integrate both visual and textual music feature content. We also introduce multi-sampled ImageBind embeddings and a pre-LLM fusion Transformer to enhance modality fusion prior to input into text LLMs, tailoring for multi-way instruction tuning. Our model achieves state-of-the-art performances across six music understanding tasks, highlighting the benefits of the auxiliary modalities and the structural superiority of DeepResonance. We open-source the codes, models and datasets we constructed: https: //github.com/sony/DeepResonance .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
Related papers
- LLark: A Multimodal Instruction-Following Language Model for MusicJoshua Patrick Gardner, Simon Durand, Daniel Stoller, Rachel M. BittnerICML 2024 · 34 citations
- DeepAlign: Mitigating Modality Conflict through Modality-Specific AlignmentShuo Li, Bingchen Miao, Wendong Bu, Juncheng Li et al.CVPR 2026
- Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth FusionJiuhai Chen, Jianwei Yang, Haiping Wu, Dianqi Li et al.CVPR 2025
- Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical ScoresCongren Dai, Yue Yang, Krinos Li, Huichi Zhou et al.ACL 2026 · 4 citations
- AnyGPT: Unified Multimodal LLM with Discrete Sequence ModelingJun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou et al.ACL 2024
