Steering Autoregressive Music Generation with Recursive Feature Machines
Daniel Zhao, Daniel Beaglehole, Julian J. McAuley, Taylor Berg-Kirkpatrick, Zachary Novack
摘要
Controllable music generation remains a significant challenge, with existing methods often requiring model retraining or introducing audible artifacts. We introduce MusicRFM, a framework that adapts Recursive Feature Machines (RFMs) (Radhakrishnan et al., 2023) to enable fine-grained, interpretable control over frozen, pre-trained music models by directly steering their internal activations. RFMs analyze a model's internal gradients to produce interpretable "concept directions", or specific axes in the activation space that correspond to musical attributes like notes or chords. We first train lightweight RFM probes to discover these directions within MUSICGEN's hidden states; then, during inference, we inject them back into the model to guide the generation process in real-time without per-step optimization. We present advanced mechanisms for this control, including dynamic, time-varying schedules and methods for the simultaneous enforcement of multiple musical properties. Our method successfully navigates the trade-off between control and generation quality: we can increase the accuracy of generating a target musical note from 0.23 to 0.82, while text prompt adherence remains within approximately 0.02 of the unsteered baseline, demonstrating effective control with minimal impact on prompt fidelity. We release code 1 to encourage further exploration on RFMs in the music domain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez 等NeurIPS 2023 · 被引用 843 次
- DITTO: Diffusion Inference-Time T-Optimization for Music GenerationZachary Novack, Julian J. McAuley, Taylor Berg-Kirkpatrick, Nicholas J. BryanICML 2024 · 被引用 81 次
- Rethinking FID: Towards a Better Evaluation Metric for Image GenerationSadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner 等CVPR 2024
- Steering Llama 2 via Contrastive Activation AdditionNina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong 等ACL 2024
相关 Paper
- General and Efficient Steering of Unconditional Diffusion ModelsQingsong Wang, Misha Belkin, Yusu WangICML 2026
- Symbolic Music Generation with Non-Differentiable Rule Guided DiffusionYujia Huang, Adishree Ghatare, Yuanzhe Liu, Ziniu Hu 等ICML 2024 · 被引用 48 次
- JEN-1 DreamStyler: Customized Musical Concept Learning via Pivotal Parameters TuningBoyu Chen, Peike Li, Yao Yao, Alex WangAAAI 2025 · 被引用 3 次
- RSA-Control: A Pragmatics-Grounded Lightweight Controllable Text Generation FrameworkYifan Wang, Vera DembergEMNLP 2024 · 被引用 2 次
- MusicRL: Aligning Music Generation to Human PreferencesGeoffrey Cideron, Sertan Girgin, Mauro Verzetti, Damien Vincent 等ICML 2024 · 被引用 41 次
