UNO! UNified Offline Training Paradigm for Learning Path Recommendation
Linzhi Peng, Wentao Zhu, Ke Cheng, Heng Chang, Junchen Ye, Bowen Du, Weifeng Lv
Abstract
With the wide adoption of online education platforms, adaptive learning systems have become increasingly important. Learning Path Recommendation (LPR) aims to dynamically adjust learning content to optimize learning efficiency based on individual student needs. However, current LPR methods suffer from sparse reward for precise assessment and only focus on anonymous sessions that overlook more personalized and effective paths. To address these challenges, we propose UNO, UNified Offline Training Paradigm for Learning Path Recommendation. This approach introduces an offline training paradigm in reinforcement learning based LPR to provide dense process rewards by a personalized advantage based on a reward model, which can estimate the students' internal knowledge levels on the learning targets. Additionally, we propose UniLPR model, a personalized recommendation system that unifies modeling the implicit relationships between students' long-term accumulation and evolving requirements for questions, and refines through Group Relative Policy Optimization(GRPO). Finally, we design learning tasks that encompass historical reviewing, recent learning, and long-term exploratory learning to simulate the comprehensive and diverse needs of students. Our UNO achieves state-of-the-art performance across all tasks, demonstrating its effectiveness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f05e320-5434-4f10-8455-4fae915446ecBuilds on10
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative RecommendationsJiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang et al.ICML 2024 · 200 citations
- Set-to-Sequence Ranking-Based Concept-Aware Learning Path RecommendationXianyu Chen, Jian Shen, Wei Xia, Jiarui Jin et al.AAAI 2023 · 25 citations
- TimeSGN: Scalable and Effective Temporal Graph Neural NetworkYuanyuan Xu, Wenjie Zhang, Ying Zhang, Maria E. Orlowska et al.ICDE 2024 · 15 citations
Related papers
- GenAL: Generative Agent for Adaptive LearningRui Lv, Qi Liu, Weibo Gao, Haotian Zhang et al.AAAI 2025 · 7 citations
- GraphRAG-Induced Dual Knowledge Structure Graphs for Personalized Learning Path RecommendationXinghe Cheng, Zihan Zhang, Jiapu Wang, Liangda Fang et al.AAAI 2026 · 1 citation
- Rethinking Reinforcement Learning for Recommendation: A Prompt PerspectiveXin Xin, Tiago Pimentel, Alexandros Karatzoglou, Pengjie Ren et al.SIGIR 2022 · 48 citations
- Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy OptimizationKun Lei, Zhengmao He, Chenhao Lu, Kaizhe Hu et al.ICLR 2024 · 31 citations
- Item-Difficulty-Aware Learning Path Recommendation: From a Real Walking PerspectiveHaotian Zhang, Shuanghong Shen, Bihan Xu, Zhenya Huang et al.KDD 2024 · 3 citations
