Leveraging Sparsity for Sample-Efficient Preference Learning: A Theoretical Perspective
Yunzhen Yao, Lie He, Michael Gastpar
Abstract
This paper considers the sample-efficiency of preference learning, which models and predicts human choices based on comparative judgments. The minimax optimal estimation error rate Θ(d/n) in classical estimation theory requires that the number of samples n scales linearly with the dimensionality of the feature space d. However, the high dimensionality of the feature space and the high cost of collecting human-annotated data challenge the efficiency of traditional estimation methods. To remedy this, we leverage sparsity in the preference model and establish sharp error rates. We show that under the sparse random utility model, where the parameter of the reward function is k-sparse, the minimax optimal rate can be reduced to Θ(k/n log(d/k)). Furthermore, we analyze the ℓ 1 -regularized estimator and show that it achieves near-optimal rate under mild assumptions on the Gram matrix. Experiments on synthetic data and LLM alignment data validate our theoretical findings, showing that sparsityaware methods significantly reduce sample complexity and improve prediction accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Learning Parametric Distributions from Samples and PreferencesMarc Jourdan, Gizem Yüce, Nicolas FlammarionICML 2025
- Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPORuizhe Shi, Minhak Song, Runlong Zhou, Zihan Zhang et al.ICML 2026
Builds on18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li et al.ICML 2024 · 569 citations
Related papers
- SparseRM: A Lightweight Preference Modeling with Sparse AutoencoderDengcan Liu, Jiahao Li, Zheren Fu, Yi Tu et al.AAAI 2026
- Spread Preference Annotation: Direct Preference Judgment for Efficient LLM AlignmentDongyoung Kim, Kimin Lee, Jinwoo Shin, Jaehyung KimICLR 2025
- The Sign Estimator: Preference Modeling for LLM Alignment under HeterogeneityAli Aouad, Aymane El Gadarri, Vivek FariasICML 2026 · 2 citations
- Comparing Few to Rank Many: Active Human Preference Learning Using Randomized Frank-Wolfe MethodKiran Koshy Thekumparampil, Gaurush Hiranandani, Kousha Kalantari, Shoham Sabach et al.ICML 2025
- Constrain Alignment with Sparse AutoencodersQingyu Yin, Chak Tou Leong, Hongbo Zhang, Minjun Zhu et al.ICML 2025
