Toward Understanding Privileged Features Distillation in Learning-to-Rank
Shuo Yang, Sujay Sanghavi, Holakou Rahmanian, Jan Bakus, S. V. N. Vishwanathan
Abstract
In learning-to-rank problems, a privileged feature is one that is available during model training, but not available at test time. Such features naturally arise in merchandised recommendation systems; for instance, "user clicked this item" as a feature is predictive of "user purchased this item" in the offline data, but is clearly not available during online serving. Another source of privileged features is those that are too expensive to compute online but feasible to be added offline. Privileged features distillation (PFD) refers to a natural idea: train a "teacher" model using all features (including privileged ones) and then use it to train a "student" model that does not use the privileged features. In this paper, we first study PFD empirically on three public ranking datasets and an industrial-scale ranking problem derived from Amazon's logs. We show that PFD outperforms several baselines (no-distillation, pretraining-finetuning, self-distillation, and generalized distillation) on all these datasets. Next, we analyze why and when PFD performs well via both empirical ablation studies and theoretical analysis for linear models. Both investigations uncover an interesting non-monotone behavior: as the predictive power of a privileged feature increases, the performance of the resulting student model initially increases but then decreases. We show the reason for the later decreasing performance is that a very predictive privileged teacher produces predictions with high variance, which lead to high variance student estimates and inferior testing performance. * This work was done while Shuo Yang was interning at Amazon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75ead04e-19c2-41b5-970b-e6a8a5d19c98Cited by top-tier papers7
- When does Privileged information Explain Away Label Noise?Guillermo Ortiz-Jiménez, Mark Collier, Anant Nawalgaria, Alexander Nicholas D'Amour et al.ICML 2023 · 16 citations
- Conformal Prediction with Corrupted Labels: Uncertain Imputation and Robust Re-weightingShai Feldman, Stephen Bates, Yaniv RomanoICLR 2026 · 5 citations
- KINDLE: Knowledge-Guided Distillation for Prior-Free Gene Regulatory Network InferenceRui Peng, Yuchen Lu, Qichen Sun, Yuxing Lu et al.NeurIPS 2025 · 1 citation
- Improving Target Sound Extraction via Disentangled Codec Representations with Privileged Knowledge DistillationDail Kim, Joon-Hyuk ChangNeurIPS 2025 · 1 citation
- BLEND: Behavior-guided Neural Population Dynamics Modeling via Privileged Knowledge DistillationZhengrui Guo, Fangxu Zhou, Wei Wu, Qichen Sun et al.ICLR 2025
Builds on3
- Are Neural Rankers still Outperformed by Gradient Boosted Decision Trees?Zhen Qin, Le Yan, Honglei Zhuang, Yi Tay et al.ICLR 2021 · 41 citations
- Privileged Graph Distillation for Cold Start RecommendationShuai Wang, Kun Zhang, Le Wu, Haiping Ma et al.SIGIR 2021 · 32 citations
- Transfer and Marginalize: Explaining Away Label Noise with Privileged InformationMark Collier, Rodolphe Jenatton, Effrosyni Kokiopoulou, Jesse BerentICML 2022 · 19 citations
Related papers
- Privileged Knowledge State Distillation for Reinforcement Learning-based Educational Path RecommendationQingyao Li, Wei Xia, Li'ang Yin, Jiarui Jin et al.KDD 2024 · 6 citations
- Bidirectional Distillation for Top-K Recommender SystemWonbin Kweon, SeongKu Kang, Hwanjo YuWWW 2021 · 58 citations
- Explicit Intent-Enhanced Knowledge Distillation for Trip RecommendationShuliang Wang, Xiaoting Leng, Sijie Ruan, Dingqi Yang et al.AAAI 2026
- Exploring Feature-based Knowledge Distillation for Recommender System: A Frequency PerspectiveZhangchi Zhu, Wei ZhangKDD 2025 · 1 citation
- Privileged Information Distillation for Language ModelsEmiliano Penaloza, Dheeraj Vattikonda, Nicolas Gontier, Alexandre Lacoste et al.ICML 2026 · 61 citations
