Amuse: Human-AI Collaborative Songwriting with Multimodal Inspirations
Yewon Kim, Sung-Ju Lee, Chris Donahue
Abstract
Songwriting is often driven by multimodal inspirations, such as imagery, narratives, or existing music, yet songwriters remain unsupported by current music AI systems in incorporating these multimodal inputs into their creative processes. We introduce Amuse, a songwriting assistant that transforms multimodal (image, text, or audio) inputs into chord progressions that can be seamlessly incorporated into songwriters' creative process. A key feature of Amuse is its novel method for generating coherent chords that are relevant to music keywords in the absence of datasets with paired examples of multimodal inputs and chords. Specifically, we propose a method that leverages multimodal LLMs to convert multimodal inputs into noisy chord suggestions and uses a unimodal chord model to filter the suggestions. A user study with songwriters shows that Amuse effectively supports transforming multimodal ideas into coherent musical suggestions, enhancing users' agency and creativity throughout the songwriting process.
• Human-centered computing → Interactive systems and tools; • Applied computing → Sound and music computing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d57c1e9b-b5f7-4bdb-be92-40a13dcc3578Cited by top-tier papers10
- Partnering with Generative AI: Experimental Evaluation of Model-Led and Human-Led Interaction in Human-AI Co-CreationSebastian Maier, Manuel Schneider, Stefan FeuerriegelCHI 2026 · 5 citations
- PaperTok: Exploring the Use of Generative AI for Creating Short-form Videos for Research CommunicationMeziah Ruby Cristobal, Hyeonjeong Byeon, Tze-Yu Chen, Ruoxi Shang et al.CHI 2026 · 2 citations
- SoundStager: Interactive Design of Story-Driven GenAI Soundscapes for VideoSuhyeon Yoo, Adolfo Hernandez Santisteban, Prem Seetharaman, Justin Salamon et al.CHI 2026 · 2 citations
- Co-Ideation Across Time: Revitalizing Legacy Design Sketchnotes with Conversational AI Agents to Foster Intergenerational CollaborationYuqing Lucy Li, Quincy Kuang, Xiao Xiao, Jean-Baptiste Labrune et al.CHI 2026 · 2 citations
- MoSound: An Interactive Tool for Generative Sound Design in Motion GraphicsJialin Huang, Prem Seetharaman, Timothy Richard Langlois, Li-Yi Wei et al.CHI 2026 · 2 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model PromptsTongshuang Wu, Michael Terry, Carrie Jun CaiCHI 2022 · 465 citations
Related papers
- Exploring the Potential of Music Generative AI for Music-Making by Deaf and Hard of Hearing PeopleYoujin Choi, JaeYoung Moon, Jinyoung Yoo, Jin-Hyuk HongCHI 2025 · 15 citations
- MusFlow: Multimodal Music Generation via Conditional Flow MatchingJiahao Song, Yuzhao WangACM MM 2025 · 3 citations
- Context-aware Image-to-Music Generation via Bridging Modalities through Musical CaptionsShilin Liu, Kyohei Kamikawa, Keisuke Maeda, Takahiro Ogawa et al.ACM MM 2025
- Designing a Generative AI-Assisted Music Psychotherapy Tool for Deaf and Hard-of-Hearing IndividualsYoujin Choi, JaeYoung Moon, Jinyoung Yoo, Jennifer G. Kim et al.CHI 2026 · 2 citations
- SongDriver: Real-time Music Accompaniment Generation without Logical Latency nor Exposure BiasZihao Wang, Kejun Zhang, Yuxing Wang, Chen Zhang et al.ACM MM 2022 · 12 citations
