Lune

ICML2026Top-tier venue

Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection

Jianghao Wu, Jianfei Cai, Weiqiang Wang, Jin Ye, Daniel F Schmidt, Yasmeen George

2026Year

Abstract

Reinforcement learning with verifiable rewards (RLVR) can yield large reasoning gains from very few training instances, yet its strong sensitivity to which instances are used makes data selection a central bottleneck. Most existing selection pipelines rely on training-time optimization signals and/or require access to verifiable rewards or ground-truth answers over large candidate pools, which is costly and often infeasible in specialized domains. We study RLVR data selection in a setting where selection must be performed before any RL training and without labels or reward evaluation on the full pool. We propose SHIFT, a one-shot, training-free selector based solely on inference-time hidden-state dynamics. For each candidate instance, SHIFT runs a single deterministic reasoning rollout and computes a reasoning-induced representation shift (RIRS) as the start-to-end hidden-state delta. SHIFT uses the RIRS magnitude as a lightweight proxy for instance utility and enforces coverage via a qualityweighted farthest-first CoreSet procedure in an RIRS-augmented feature space, producing compact subsets that scale to large unlabeled pools. Across mathematical reasoning and medical QA benchmarks under ultra-low budgets, SHIFT consistently outperforms training-free diversity and difficulty/uncertainty baselines, improving both in-domain accuracy and transfer to harder evaluation settings. Ablations show that RIRS-based coverage and quality-weighting contribute complementary gains, and analyses indicate that RIRS is not explained by simple input/output length statistics. Code is available at https: //github.com/JianghaoWu/SHIFT . Figure 1. RLVR is highly example-sensitive (a), while prior selection methods rely on training-time and/or ground-truth signals (b,c). We enable one-shot, training-free and label-free selection via reasoning-induced representation shifts (d).

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ddff99c0-ed14-4047-b399-cb191024f231

Builds on17

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines