Amuse: Human-AI Collaborative Songwriting with Multimodal Inspirations
Yewon Kim, Sung-Ju Lee, Chris Donahue
摘要
Songwriting is often driven by multimodal inspirations, such as imagery, narratives, or existing music, yet songwriters remain unsupported by current music AI systems in incorporating these multimodal inputs into their creative processes. We introduce Amuse, a songwriting assistant that transforms multimodal (image, text, or audio) inputs into chord progressions that can be seamlessly incorporated into songwriters' creative process. A key feature of Amuse is its novel method for generating coherent chords that are relevant to music keywords in the absence of datasets with paired examples of multimodal inputs and chords. Specifically, we propose a method that leverages multimodal LLMs to convert multimodal inputs into noisy chord suggestions and uses a unimodal chord model to filter the suggestions. A user study with songwriters shows that Amuse effectively supports transforming multimodal ideas into coherent musical suggestions, enhancing users' agency and creativity throughout the songwriting process.
• Human-centered computing → Interactive systems and tools; • Applied computing → Sound and music computing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Partnering with Generative AI: Experimental Evaluation of Model-Led and Human-Led Interaction in Human-AI Co-CreationSebastian Maier, Manuel Schneider, Stefan FeuerriegelCHI 2026 · 被引用 5 次
- PaperTok: Exploring the Use of Generative AI for Creating Short-form Videos for Research CommunicationMeziah Ruby Cristobal, Hyeonjeong Byeon, Tze-Yu Chen, Ruoxi Shang 等CHI 2026 · 被引用 2 次
- SoundStager: Interactive Design of Story-Driven GenAI Soundscapes for VideoSuhyeon Yoo, Adolfo Hernandez Santisteban, Prem Seetharaman, Justin Salamon 等CHI 2026 · 被引用 2 次
- Co-Ideation Across Time: Revitalizing Legacy Design Sketchnotes with Conversational AI Agents to Foster Intergenerational CollaborationYuqing Lucy Li, Quincy Kuang, Xiao Xiao, Jean-Baptiste Labrune 等CHI 2026 · 被引用 2 次
- MoSound: An Interactive Tool for Generative Sound Design in Motion GraphicsJialin Huang, Prem Seetharaman, Timothy Richard Langlois, Li-Yi Wei 等CHI 2026 · 被引用 2 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez 等NeurIPS 2023 · 被引用 843 次
- AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model PromptsTongshuang Wu, Michael Terry, Carrie Jun CaiCHI 2022 · 被引用 465 次
相关 Paper
- Exploring the Potential of Music Generative AI for Music-Making by Deaf and Hard of Hearing PeopleYoujin Choi, JaeYoung Moon, Jinyoung Yoo, Jin-Hyuk HongCHI 2025 · 被引用 15 次
- MusFlow: Multimodal Music Generation via Conditional Flow MatchingJiahao Song, Yuzhao WangACM MM 2025 · 被引用 3 次
- Context-aware Image-to-Music Generation via Bridging Modalities through Musical CaptionsShilin Liu, Kyohei Kamikawa, Keisuke Maeda, Takahiro Ogawa 等ACM MM 2025
- Designing a Generative AI-Assisted Music Psychotherapy Tool for Deaf and Hard-of-Hearing IndividualsYoujin Choi, JaeYoung Moon, Jinyoung Yoo, Jennifer G. Kim 等CHI 2026 · 被引用 2 次
- SongDriver: Real-time Music Accompaniment Generation without Logical Latency nor Exposure BiasZihao Wang, Kejun Zhang, Yuxing Wang, Chen Zhang 等ACM MM 2022 · 被引用 12 次
