Direct Post-Training Preference Alignment for Multi-Agent Motion Generation Model Using Implicit Feedback from Pre-training Demonstrations
Thomas Tian, Kratarth Goel
Abstract
Recent advancements in Large Language Models (LLMs) have revolutionized motion generation models in embodied applications such as autonomous driving and robotic manipulation. While LLM-type auto-regressive motion generation models benefit from training scalability, there remains a discrepancy between their token prediction objectives and human preferences. As a result, models pre-trained solely with token-prediction objectives often generate behaviors that deviate from what humans would prefer, making post-training preference alignment crucial for producing human-preferred motions. Unfortunately, post-training alignment requires extensive preference rankings of motions generated by the pre-trained model, which are costly and time-consuming to annotate-especially in multi-agent motion generation settings. Recently, there has been growing interest in leveraging expert demonstrations previously used during pre-training to scalably generate preference data for post-training alignment. However, these methods often adopt an adversarial assumption, treating all pretrained model-generated samples as unpreferred examples and relying solely on pre-training expert demonstrations to construct preferred examples. This adversarial approach overlooks the valuable signal provided by preference rankings among the model's own generations, ultimately reducing alignment effectiveness and potentially leading to misaligned behaviors. In this work, instead of treating all generated samples as equally bad, we propose a principled approach that leverages implicit preferences encoded in pre-training expert demonstrations to construct preference rankings among the pre-trained model's generations, offering more nuanced preference alignment guidance with zero human cost. We apply our approach to large-scale traffic simulation (more than 100 agents) and demonstrate its effectiveness in improving the realism of pre-trained model's generated behaviors, making a lightweight 1M motion generation model comparable to state-of-the-art large imitation-based models by relying solely on implicit feedback from pre-training demonstrations, without requiring additional post-training human preference annotations or incurring high computational costs. Furthermore, we provide an in-depth analysis of preference data scaling laws and their effects on over-optimization, offering valuable insights for future studies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- RLFTSim: Realistic and Controllable Multi-Agent Traffic Simulation via Reinforcement Learning Fine-TuningEhsan Ahmadi, Hunter Schofield, Behzad Khamidehi, Fazel Arasteh et al.CVPR 2026 · 2 citations
- U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generationxiang deng, Feng Gao, Yong Zhang, Youxin Pang et al.CVPR 2026 · 2 citations
Builds on17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard et al.ICML 2024 · 598 citations
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language ModelsZixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji et al.ICML 2024 · 527 citations
- MotionLM: Multi-Agent Motion Forecasting as Language ModelingAri Seff, Brian Cera, Dian Chen, Mason Ng et al.ICCV 2023 · 186 citations
Related papers
- Self-Guided Alignment: Adaptive Preference Sensing for Multi-Objective GenerationNing Wang, Zhanyang Liu, Taotao Zhou, Xinrui Zhang et al.ACL 2026
- Understanding the Learning Dynamics of Alignment with Human FeedbackShawn Im, Yixuan LiICML 2024 · 18 citations
- Learning Preference Model for LLMs via Automatic Preference Data GenerationShijia Huang, Jianqiao Zhao, Yanyang Li, Liwei WangEMNLP 2023 · 3 citations
- Implicit Preference Alignment for Human Image AnimationYuanzhi Wang, Xuhua Ren, Jiaxiang Cheng, bing ma et al.ICML 2026
- Alignment-Aware DecodingFrédéric Berdoz, Luca Lanzendörfer, René Caky, Roger WattenhoferICML 2026 · 1 citation
