Towards End-to-End Alignment of User Satisfaction via Questionnaire in Video Recommendation
Na Li, Jiaqi Yu, Minzhi Xie, Tiantian He, Xiaoxiao Xu, Zixiu Wang, Lantao Hu, Yongqi Liu, Han Li, Kaiqiao Zhan, Kun Gai
摘要
Short-video recommender systems typically optimize ranking models using dense user behavioral signals, such as clicks and watch time. However, these signals are only indirect proxies of user satisfaction and often suffer from noise and bias. Recently, explicit satisfaction feedback collected through questionnaires has emerged as a high-quality direct alignment supervision, but is extremely sparse and easily overwhelmed by abundant behavioral data, making it difficult to incorporate into online recommendation models. To address these challenges, we propose a novel framework which is towards End-to-End Alignment of user Satisfaction via Questionnaire, named EASQ, to enable real-time alignment of ranking models with true user satisfaction. Specifically, we first construct an independent parameter pathway for sparse questionnaire signals by combining a multi-task architecture and a lightweight LoRA module. The multi-task design separates sparse satisfaction supervision from dense behavioral signals, preventing the former from being overwhelmed. The LoRA module pre-inject these preferences in a parameter-isolated manner, ensuring stability in the backbone while optimizing user satisfaction. Furthermore, we employ a DPO-based optimization objective tailored for online learning, which aligns the main model outputs with sparse satisfaction signals in real time. This design enables end-to-end online learning, allowing the model to continuously adapt to new questionnaire feedback while maintaining the stability and effectiveness of the backbone. Extensive offline experiments and large-scale online A/B tests demonstrate that EASQ consistently improves user satisfaction metrics across multiple scenarios. EASQ has been successfully deployed in a production short-video recommendation system, delivering significant and stable business gains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Causal Intervention for Leveraging Popularity Bias in RecommendationYang Zhang, Fuli Feng, Xiangnan He, Tianxin Wei 等SIGIR 2021 · 被引用 431 次
- On Softmax Direct Preference Optimization for RecommendationYuxin Chen, Junfei Tan, An Zhang, Zhengyi Yang 等NeurIPS 2024 · 被引用 126 次
相关 Paper
- MTRec: Learning to Align with User Preferences via Mental Reward ModelsMengchen Zhao, Yifan Gao, Yaqing Hou, Xiangyang Li 等NeurIPS 2025 · 被引用 1 次
- Multimodal-aware Multi-intention Learning for RecommendationWei Yang, Qingchen YangACM MM 2024 · 被引用 4 次
- Relative Advantage Debiasing for Watch-Time Prediction in Short-Video RecommendationEmily Liu, Kuan Han, Minfeng Zhan, Bocheng Zhao 等AAAI 2026
- CoPL: Collaborative Preference Learning for Personalizing LLMsYoungbin Choi, Seunghyuk Cho, Minjong Lee, MoonJeong Park 等EMNLP 2025
- Generative Regression Based Watch Time Prediction for Short-Video RecommendationHongxu Ma, Kai Tian, Tao Zhang, Xuefeng Zhang 等WWW 2026 · 被引用 6 次
