Steering Autoregressive Music Generation with Recursive Feature Machines
Daniel Zhao, Daniel Beaglehole, Julian J. McAuley, Taylor Berg-Kirkpatrick, Zachary Novack
Abstract
Controllable music generation remains a significant challenge, with existing methods often requiring model retraining or introducing audible artifacts. We introduce MusicRFM, a framework that adapts Recursive Feature Machines (RFMs) (Radhakrishnan et al., 2023) to enable fine-grained, interpretable control over frozen, pre-trained music models by directly steering their internal activations. RFMs analyze a model's internal gradients to produce interpretable "concept directions", or specific axes in the activation space that correspond to musical attributes like notes or chords. We first train lightweight RFM probes to discover these directions within MUSICGEN's hidden states; then, during inference, we inject them back into the model to guide the generation process in real-time without per-step optimization. We present advanced mechanisms for this control, including dynamic, time-varying schedules and methods for the simultaneous enforcement of multiple musical properties. Our method successfully navigates the trade-off between control and generation quality: we can increase the accuracy of generating a target musical note from 0.23 to 0.82, while text prompt adherence remains within approximately 0.02 of the unsteered baseline, demonstrating effective control with minimal impact on prompt fidelity. We release code 1 to encourage further exploration on RFMs in the music domain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2621a56-2b6e-4ae3-adcd-f94e8d3698efBuilds on4
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- DITTO: Diffusion Inference-Time T-Optimization for Music GenerationZachary Novack, Julian J. McAuley, Taylor Berg-Kirkpatrick, Nicholas J. BryanICML 2024 · 81 citations
- Rethinking FID: Towards a Better Evaluation Metric for Image GenerationSadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner et al.CVPR 2024
- Steering Llama 2 via Contrastive Activation AdditionNina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong et al.ACL 2024
Related papers
- General and Efficient Steering of Unconditional Diffusion ModelsQingsong Wang, Misha Belkin, Yusu WangICML 2026
- Symbolic Music Generation with Non-Differentiable Rule Guided DiffusionYujia Huang, Adishree Ghatare, Yuanzhe Liu, Ziniu Hu et al.ICML 2024 · 48 citations
- JEN-1 DreamStyler: Customized Musical Concept Learning via Pivotal Parameters TuningBoyu Chen, Peike Li, Yao Yao, Alex WangAAAI 2025 · 3 citations
- RSA-Control: A Pragmatics-Grounded Lightweight Controllable Text Generation FrameworkYifan Wang, Vera DembergEMNLP 2024 · 2 citations
- MusicRL: Aligning Music Generation to Human PreferencesGeoffrey Cideron, Sertan Girgin, Mauro Verzetti, Damien Vincent et al.ICML 2024 · 41 citations
