Generative Regression Based Watch Time Prediction for Short-Video Recommendation
Hongxu Ma, Kai Tian, Tao Zhang, Xuefeng Zhang, Han Zhou, Chenghou Jin, Chunjie Chen, Han Li, Jihong Guan, Shuigeng Zhou
Abstract
Watch time prediction (WTP) has emerged as a pivotal task in short video recommendation systems, designed to quantify user engagement through continuous interaction modeling. Predicting users' watch times on videos often encounters fundamental challenges, including wide value ranges and imbalanced data distributions, which can lead to significant estimation bias when directly applying regression techniques. Recent studies have attempted to address these issues by converting the continuous watch time estimation into an ordinal regression task. While these methods demonstrate partial effectiveness, they exhibit notable limitations: (1) The discretization process frequently relies on bucket partitioning, inherently reducing prediction flexibility and accuracy. (2) The interdependencies among different partition intervals remain underutilized, missing opportunities for effective error correction. Inspired by language modeling paradigms, we propose a novel Generative Regression (GR) framework that reformulates WTP as a sequence generation task. Our approach employs structural discretization to enable nearly lossless value reconstruction while maintaining prediction flexibility. Through carefully designed vocabulary construction and label encoding schemes, each watch time is bijectively mapped to a token sequence. To mitigate the training-inference discrepancy caused by teacher-forcing, we introduce a curriculum learning with embedding mixup strategy that gradually transitions from guided to free-generation modes. We test our models extensively on two public datasets, a large-scale offline industrial dataset, and an online A/B test on Kuaishou App with over 400 million daily active users (DAU) and GR consistently outperforms existing state-of-the-art approaches significantly. Our code is available at https://github.com/snailma0229/GR.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32d8298d-9901-48e7-81ae-509286e8f76aCited by top-tier papers9
- HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language ModelsZhaolu Kang, Junhao Gong, Jiaxu Yan, Wanke Xia et al.ICLR 2026 · 24 citations
- MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic LearningHongxu Ma, Guanshuo Wang, Fufu Yu, Qiong Jia et al.ACM MM 2025 · 9 citations
- Fine-grained Zero-Shot Object DetectionHongxu Ma, Chenbo Zhang, Lu Zhang, Jiaogen Zhou et al.ACM MM 2025 · 4 citations
- RL-ScanIQA: Reinforcement-Learned Scanpaths for Blind 360deg Image Quality AssessmentYujia Wang, Yuyan Li, Jiuming Liu, Fang-Lue Zhang et al.CVPR 2026 · 3 citations
- DiffoR: A Unified Continuous Generative Framework for Universal Ordinal RegressionHongxu Ma, Lin Wang, Chenghou Jin, Han Zhou et al.KDD 2026 · 1 citation
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- DVR: Micro-Video Recommendation Optimizing Watch-Time-Gain under Duration BiasYu Zheng, Chen Gao, Jingtao Ding, Lingling Yi et al.ACM MM 2022 · 28 citations
- Counteracting Duration Bias in Video Recommendation via Counterfactual Watch TimeHaiyuan Zhao, Guohao Cai, Jieming Zhu, Zhenhua Dong et al.KDD 2024 · 9 citations
- TeaForN: Teacher-Forcing with N-gramsSebastian Goodman, Nan Ding, Radu SoricutEMNLP 2020 · 1 citation
- CREAD: A Classification-Restoration Framework with Error Adaptive Discretization for Watch Time Prediction in Video Recommender SystemsJie Sun, Zhaoying Ding, Xiaoshuang Chen, Qi Chen et al.AAAI 2024
Related papers
- FlowTime: Towards Continuous Generative Watch Time Prediction via Flow-based Personalized PriorsHongxu Ma, Han Zhou, Chenghou Jin, Jie Zhang et al.KDD 2026 · 1 citation
- An Action-Aware Generative Sequence Modeling for Short Video RecommendationWenhao Li, Zihan Lin, Zhengxiao Guo, Jie Zhou et al.SIGIR 2026
- Calibrating Video Watch-time Predictions with Credible Prototype AlignmentChao Cui, Shisong Tang, Fan Li, Jiechao Gao et al.ICML 2025
- Dancing with Shackles, Meet the Challenge of Industrial Adaptive Streaming via Offline Reinforcement LearningLianchen Jia, Chao Zhou, Tianchi Huang, Chaoyang Li et al.INFOCOM 2024 · 2 citations
- Relative Advantage Debiasing for Watch-Time Prediction in Short-Video RecommendationEmily Liu, Kuan Han, Minfeng Zhan, Bocheng Zhao et al.AAAI 2026
