Lune

ICLR2026Top-tier venue

Parameter-Efficient Reinforcement Learning using Prefix Optimization

Itamar Rocha Filho, Rosie Zhao, Sham M. Kakade, Eran Malach, Samy Jelassi

2026Year

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) is a leading approach for tuning language models on mathematical reasoning tasks. However, it remains unclear whether RLVR's gains stem from genuine reasoning improvements or simply from steering the model toward answer formats that already appear in the reference distribution. Inspired by recent evidence (Zhao et al., 2025; Yue et al., 2025) , we study this question by optimizing only the first k tokens (e.g. k = 32) of each solution, generating the remainder of the response from the reference model. We study two methods for prefix optimization, using a naive algorithm that clusters prefixes and selects the best prefix (Prefix Clustering), and a method that optimizes the prefix by finetuning a lightweight adapter model with RL (Prefix-RL). We show that tuning only the first k tokens can significantly improve the accuracy on math, suggesting that at least some of the gains from RL are due to upweighting a preferable solution strategy. Our results suggest that simple prefix optimization methods can provide an efficient alternative to RL, delivering substantial improvements across different models and benchmarks for a tiny fraction of the compute required for standard RL, and that these gains are robust across prefix lengths and random seeds.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext a2bfc830-c52b-4294-ae16-31fbbf3ad273

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines