Harmonic Canvas: Inversion-Free Editing for Visually-Guided Music Style Transfer
Yue Lei, Siqi Yang, Ting Zhong, Fan Zhou
摘要
Music style transfer (MST) aims to reinterpret existing musical pieces in new stylistic forms while maintaining their melodic coherence. Conventional approaches conditioned on text or audio overlook the profoundly multimodal character of musical style. Visual ambience -reflected in color, lighting, and composition -encodes affective attributes that parallel timbre, rhythm, and harmony, which, however, remain underexplored in MST context. We introduce a flow-based, inversion-free framework for multimodal music style transfer that unifies textual and visual guidance. Our approach tackles two challenges: (1) capturing crossmodal semantics beyond language through a dual-encoder fusion module that merges CLIP-and ViT-derived embeddings, and (2) preserving melodic identity using a differentiable normalized chroma constraint that regulates pitchclass consistency along the generative flow. We reorganize and extend the MeLBench and MusicCaps collections into a genre-structured multimodal dataset to support style-aware analysis. Quantitative and perceptual evaluations demonstrate that our approach achieves superior control, structural fidelity, and cross-modal expressiveness, underscoring the role of visual perception in music generation. Our demo is available online.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song 等ICLR 2022 · 被引用 2,128 次
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang 等NeurIPS 2021 · 被引用 1,553 次
相关 Paper
- MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech SynthesisWenhao Guan, Yishuang Li, Tao Li, Hukai Huang 等AAAI 2024 · 被引用 25 次
- Music Style Transfer with Time-Varying Inversion of Diffusion ModelsSifei Li, Yuxin Zhang, Fan Tang, Chongyang Ma 等AAAI 2024 · 被引用 20 次
- MusFlow: Multimodal Music Generation via Conditional Flow MatchingJiahao Song, Yuzhao WangACM MM 2025 · 被引用 3 次
- Crossing You in Style: Cross-modal Style Transfer from Music to Visual ArtsCheng-Che Lee, Wan-Yi Lin, Yen-Ting Shih, Pei-Yi (Patricia) Kuo 等ACM MM 2020 · 被引用 16 次
- MELFuSION: Synthesizing Music from Image and Language Cues Using Diffusion ModelsSanjoy Chowdhury, Sayan Nag, K. J. Joseph, Balaji Vasan Srinivasan 等CVPR 2024
