UNO! UNified Offline Training Paradigm for Learning Path Recommendation
Linzhi Peng, Wentao Zhu, Ke Cheng, Heng Chang, Junchen Ye, Bowen Du, Weifeng Lv
摘要
With the wide adoption of online education platforms, adaptive learning systems have become increasingly important. Learning Path Recommendation (LPR) aims to dynamically adjust learning content to optimize learning efficiency based on individual student needs. However, current LPR methods suffer from sparse reward for precise assessment and only focus on anonymous sessions that overlook more personalized and effective paths. To address these challenges, we propose UNO, UNified Offline Training Paradigm for Learning Path Recommendation. This approach introduces an offline training paradigm in reinforcement learning based LPR to provide dense process rewards by a personalized advantage based on a reward model, which can estimate the students' internal knowledge levels on the learning targets. Additionally, we propose UniLPR model, a personalized recommendation system that unifies modeling the implicit relationships between students' long-term accumulation and evolving requirements for questions, and refines through Group Relative Policy Optimization(GRPO). Finally, we design learning tasks that encompass historical reviewing, recent learning, and long-term exploratory learning to simulate the comprehensive and diverse needs of students. Our UNO achieves state-of-the-art performance across all tasks, demonstrating its effectiveness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative RecommendationsJiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang 等ICML 2024 · 被引用 200 次
- Set-to-Sequence Ranking-Based Concept-Aware Learning Path RecommendationXianyu Chen, Jian Shen, Wei Xia, Jiarui Jin 等AAAI 2023 · 被引用 25 次
- TimeSGN: Scalable and Effective Temporal Graph Neural NetworkYuanyuan Xu, Wenjie Zhang, Ying Zhang, Maria E. Orlowska 等ICDE 2024 · 被引用 15 次
相关 Paper
- GenAL: Generative Agent for Adaptive LearningRui Lv, Qi Liu, Weibo Gao, Haotian Zhang 等AAAI 2025 · 被引用 7 次
- GraphRAG-Induced Dual Knowledge Structure Graphs for Personalized Learning Path RecommendationXinghe Cheng, Zihan Zhang, Jiapu Wang, Liangda Fang 等AAAI 2026 · 被引用 1 次
- Rethinking Reinforcement Learning for Recommendation: A Prompt PerspectiveXin Xin, Tiago Pimentel, Alexandros Karatzoglou, Pengjie Ren 等SIGIR 2022 · 被引用 48 次
- Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy OptimizationKun Lei, Zhengmao He, Chenhao Lu, Kaizhe Hu 等ICLR 2024 · 被引用 31 次
- Item-Difficulty-Aware Learning Path Recommendation: From a Real Walking PerspectiveHaotian Zhang, Shuanghong Shen, Bihan Xu, Zhenya Huang 等KDD 2024 · 被引用 3 次
