Toward Understanding Privileged Features Distillation in Learning-to-Rank
Shuo Yang, Sujay Sanghavi, Holakou Rahmanian, Jan Bakus, S. V. N. Vishwanathan
摘要
In learning-to-rank problems, a privileged feature is one that is available during model training, but not available at test time. Such features naturally arise in merchandised recommendation systems; for instance, "user clicked this item" as a feature is predictive of "user purchased this item" in the offline data, but is clearly not available during online serving. Another source of privileged features is those that are too expensive to compute online but feasible to be added offline. Privileged features distillation (PFD) refers to a natural idea: train a "teacher" model using all features (including privileged ones) and then use it to train a "student" model that does not use the privileged features. In this paper, we first study PFD empirically on three public ranking datasets and an industrial-scale ranking problem derived from Amazon's logs. We show that PFD outperforms several baselines (no-distillation, pretraining-finetuning, self-distillation, and generalized distillation) on all these datasets. Next, we analyze why and when PFD performs well via both empirical ablation studies and theoretical analysis for linear models. Both investigations uncover an interesting non-monotone behavior: as the predictive power of a privileged feature increases, the performance of the resulting student model initially increases but then decreases. We show the reason for the later decreasing performance is that a very predictive privileged teacher produces predictions with high variance, which lead to high variance student estimates and inferior testing performance. * This work was done while Shuo Yang was interning at Amazon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- When does Privileged information Explain Away Label Noise?Guillermo Ortiz-Jiménez, Mark Collier, Anant Nawalgaria, Alexander Nicholas D'Amour 等ICML 2023 · 被引用 16 次
- Conformal Prediction with Corrupted Labels: Uncertain Imputation and Robust Re-weightingShai Feldman, Stephen Bates, Yaniv RomanoICLR 2026 · 被引用 5 次
- KINDLE: Knowledge-Guided Distillation for Prior-Free Gene Regulatory Network InferenceRui Peng, Yuchen Lu, Qichen Sun, Yuxing Lu 等NeurIPS 2025 · 被引用 1 次
- Improving Target Sound Extraction via Disentangled Codec Representations with Privileged Knowledge DistillationDail Kim, Joon-Hyuk ChangNeurIPS 2025 · 被引用 1 次
- BLEND: Behavior-guided Neural Population Dynamics Modeling via Privileged Knowledge DistillationZhengrui Guo, Fangxu Zhou, Wei Wu, Qichen Sun 等ICLR 2025
它引用的顶会 Paper3
- Are Neural Rankers still Outperformed by Gradient Boosted Decision Trees?Zhen Qin, Le Yan, Honglei Zhuang, Yi Tay 等ICLR 2021 · 被引用 41 次
- Privileged Graph Distillation for Cold Start RecommendationShuai Wang, Kun Zhang, Le Wu, Haiping Ma 等SIGIR 2021 · 被引用 32 次
- Transfer and Marginalize: Explaining Away Label Noise with Privileged InformationMark Collier, Rodolphe Jenatton, Effrosyni Kokiopoulou, Jesse BerentICML 2022 · 被引用 19 次
相关 Paper
- Privileged Knowledge State Distillation for Reinforcement Learning-based Educational Path RecommendationQingyao Li, Wei Xia, Li'ang Yin, Jiarui Jin 等KDD 2024 · 被引用 6 次
- Bidirectional Distillation for Top-K Recommender SystemWonbin Kweon, SeongKu Kang, Hwanjo YuWWW 2021 · 被引用 58 次
- Explicit Intent-Enhanced Knowledge Distillation for Trip RecommendationShuliang Wang, Xiaoting Leng, Sijie Ruan, Dingqi Yang 等AAAI 2026
- Exploring Feature-based Knowledge Distillation for Recommender System: A Frequency PerspectiveZhangchi Zhu, Wei ZhangKDD 2025 · 被引用 1 次
- Privileged Information Distillation for Language ModelsEmiliano Penaloza, Dheeraj Vattikonda, Nicolas Gontier, Alexandre Lacoste 等ICML 2026 · 被引用 61 次
