Crossing You in Style: Cross-modal Style Transfer from Music to Visual Arts
Cheng-Che Lee, Wan-Yi Lin, Yen-Ting Shih, Pei-Yi (Patricia) Kuo, Li Su
Abstract
Music-to-visual style transfer is a challenging yet important cross-modal learning problem in the practice of creativity. Its major difference from the traditional image style transfer problem is that the style information is provided by music rather than images. Assuming that musical features can be properly mapped to visual contents through semantic links between the two domains, we solve the music-to-visual style transfer problem in two steps: music visualization and style transfer. The music visualization network utilizes an encoder-generator architecture with a conditional generative adversarial network to generate image-based music representations from music data. This network is integrated with an image style transfer method to accomplish the style transfer process. Experiments are conducted on WikiArt-IMSLP, a newly compiled dataset including Western music recordings and paintings listed by decades. By utilizing such a label to learn the semantic connection between paintings and music, we demonstrate that the proposed framework can generate diverse image style representations from a music piece, and these representations can unveil certain art forms of the same era. Subjective testing results also emphasize the role of the era label in improving the perceptual quality on the compatibility between music and visual content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59839da0-ce54-44fc-afed-da81ba629773Cited by top-tier papers4
- Sound-Guided Semantic Image ManipulationSeung Hyun Lee, Wonseok Roh, Wonmin Byeon, Sang Ho Yoon et al.CVPR 2022 · 43 citations
- Generating Realistic Images from In-the-wild SoundsTaegyeong Lee, Jeonghun Kang, Hyeonyu Kim, Taehwan KimICCV 2023 · 11 citations
- Audio-sync Video Instance Editing with Granularity-Aware Mask RefinerHaojie Zheng, Shuchen Weng, Jingqi Liu, Siqi Yang et al.CVPR 2026 · 8 citations
- Music2Palette: Emotion-aligned Color Palette Generation via Cross-Modal Representation LearningJiayun Hu, Yueyi He, Tianyi Liang, Changbo Wang et al.ACM MM 2025 · 2 citations
Related papers
- Harmonic Canvas: Inversion-Free Editing for Visually-Guided Music Style TransferYue Lei, Siqi Yang, Ting Zhong, Fan ZhouCVPR 2026
- MusFlow: Multimodal Music Generation via Conditional Flow MatchingJiahao Song, Yuzhao WangACM MM 2025 · 3 citations
- Context-aware Image-to-Music Generation via Bridging Modalities through Musical CaptionsShilin Liu, Kyohei Kamikawa, Keisuke Maeda, Takahiro Ogawa et al.ACM MM 2025
- Stochastic Interpolants for Revealing Stylistic Flows Across the History of ArtPingchuan Ma, Ming Gui, Johannes Schusterbauer, Xiaopei Yang et al.ICCV 2025
- Diverse Image Style Transfer via Invertible Cross-Space MappingHaibo Chen, Lei Zhao, Huiming Zhang, Zhizhong Wang et al.ICCV 2021 · 42 citations
