Beyond Pedagogical Principles: Multi-Horizon Preference Optimization for Efficient Socratic Tutoring
Xin Shi, Chao Zhang, Yifan Zhu, Xueqiao Zhang, Yawei Luo
Abstract
The development of LLM-based tutor agents faces challenges in simultaneously ensuring adherence to pedagogical principles and achieving optimal pedagogical effectiveness, particularly in dynamic, multi-turn interactions. Existing methods are often constrained by static data or sparse reward signals in online settings. To address this gap, we propose M ulti-H orizon P reference O ptimization ( MHPO ), a novel framework that iteratively refines tutor agents using a multi-horizon reward function within a dynamic teacher-student simulation environment. Specifically, this reward function is designed to capture both turn-level pedagogical quality and trajectory-level pedagogical effectiveness, which is estimated via Monte Carlo rollouts. We further investigate two distinct strategies to aggregate these rewards for policy optimization. Our experiments demonstrate that MHPO significantly enhances base model performance, achieving a superior balance be-tween principles and effectiveness compared to various baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- ORPO: Monolithic Preference Optimization without Reference ModelJiwoo Hong, Noah Lee, James ThorneEMNLP 2024 · 71 citations
- Automatic Generation of Socratic Subquestions for Teaching Math Word ProblemsKumar Shridhar, Jakub Macina, Mennatallah El-Assady, Tanmay Sinha et al.EMNLP 2022 · 31 citations
Related papers
- From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement LearningDavid Dinucu-Jianu, Jakub Macina, Nico Daheim, Ido Hakimi et al.EMNLP 2025
- Implicit Turn-Wise Policy Optimization for Proactive User-LLM InteractionHaoyu Wang, Yuxin Chen, Liang Luo, Buyun Zhang et al.ICML 2026 · 3 citations
- Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy OptimizationYifeng Ding, Hung Le, Songyang Han, Kangrui Ruan et al.ACL 2026 · 5 citations
- InfoPO: Information-Driven Policy Optimization for User-Centric AgentsFanqi Kong, Jiayi Zhang, Mingyi Deng, Chenglin Wu et al.ICML 2026
- Planning-Guided Tutoring with Assessment-Driven Memory for Pedagogical LLM TutorsZechen Li, Qiannan Zhu, Mei Wang, Jia Li et al.ACL 2026
