Harmonic Canvas: Inversion-Free Editing for Visually-Guided Music Style Transfer
Yue Lei, Siqi Yang, Ting Zhong, Fan Zhou
Abstract
Music style transfer (MST) aims to reinterpret existing musical pieces in new stylistic forms while maintaining their melodic coherence. Conventional approaches conditioned on text or audio overlook the profoundly multimodal character of musical style. Visual ambience -reflected in color, lighting, and composition -encodes affective attributes that parallel timbre, rhythm, and harmony, which, however, remain underexplored in MST context. We introduce a flow-based, inversion-free framework for multimodal music style transfer that unifies textual and visual guidance. Our approach tackles two challenges: (1) capturing crossmodal semantics beyond language through a dual-encoder fusion module that merges CLIP-and ViT-derived embeddings, and (2) preserving melodic identity using a differentiable normalized chroma constraint that regulates pitchclass consistency along the generative flow. We reorganize and extend the MeLBench and MusicCaps collections into a genre-structured multimodal dataset to support style-aware analysis. Quantitative and perceptual evaluations demonstrate that our approach achieves superior control, structural fidelity, and cross-modal expressiveness, underscoring the role of visual perception in music generation. Our demo is available online.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 24acc87f-e73e-43c1-9c4f-4038005783ffBuilds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song et al.ICLR 2022 · 2,128 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
Related papers
- MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech SynthesisWenhao Guan, Yishuang Li, Tao Li, Hukai Huang et al.AAAI 2024 · 25 citations
- Music Style Transfer with Time-Varying Inversion of Diffusion ModelsSifei Li, Yuxin Zhang, Fan Tang, Chongyang Ma et al.AAAI 2024 · 20 citations
- MusFlow: Multimodal Music Generation via Conditional Flow MatchingJiahao Song, Yuzhao WangACM MM 2025 · 3 citations
- Crossing You in Style: Cross-modal Style Transfer from Music to Visual ArtsCheng-Che Lee, Wan-Yi Lin, Yen-Ting Shih, Pei-Yi (Patricia) Kuo et al.ACM MM 2020 · 16 citations
- MELFuSION: Synthesizing Music from Image and Language Cues Using Diffusion ModelsSanjoy Chowdhury, Sayan Nag, K. J. Joseph, Balaji Vasan Srinivasan et al.CVPR 2024
