Learning Calibratable Policies using Programmatic Style-Consistency
Eric Zhan, Albert Tseng, Yisong Yue, Adith Swaminathan, Matthew J. Hausknecht
摘要
We study the problem of controllable generation of long-term sequential behaviors, where the goal is to calibrate to multiple behavior styles simultaneously. In contrast to the well-studied areas of controllable generation of images, text, and speech, there are two questions that pose significant challenges when generating long-term behaviors: how should we specify the factors of variation to control, and how can we ensure that the generated behavior faithfully demonstrates combinatorially many styles? We leverage programmatic labeling functions to specify controllable styles, and derive a formal notion of styleconsistency as a learning objective, which can then be solved using conventional policy learning approaches. We evaluate our framework using demonstrations from professional basketball players and agents in the MuJoCo physics environment, and show that existing approaches that do not explicitly enforce style-consistency fail to generate diverse behaviors whereas our learned policies can be calibrated for up to 4 5 (1024) distinct style combinations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Uni[MASK]: Unified Inference in Sequential Decision ProblemsMicah Carroll, Orr Paradise, Jessy Lin, Raluca Georgescu 等NeurIPS 2022 · 被引用 29 次
- Animal behavioral analysis and neural encoding with transformer-based self-supervised pretrainingYanchen Wang, Han Yu, Ari Blau, Yizi Zhang 等ICLR 2026 · 被引用 8 次
- Semi-Supervised Generative Models for Multiagent TrajectoriesDennis Fassmeyer, Pascal Fassmeyer, Ulf BrefeldNeurIPS 2022 · 被引用 7 次
- Synthesizing Trajectory Queries from ExamplesStephen Mell, Favyen Bastani, Steve Zdancewic, Osbert BastaniCAV 2023 · 被引用 5 次
- Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation LearningHanlin Yang, Jian Yao, Weiming Liu, Qing Wang 等ICLR 2025
它引用的顶会 Paper2
相关 Paper
- MultiAct: Long-Term 3D Human Motion Generation from Multiple Action LabelsTaeryung Lee, Gyeongsik Moon, Kyoung Mu LeeAAAI 2023 · 被引用 62 次
- Control strategies for physically simulated characters performing two-player competitive sportsJungdam Won, Deepak Gopinath, Jessica K. HodginsSIGGRAPH 2021 · 被引用 73 次
- UniPhys: Unified Planner and Controller with Diffusion for Flexible Physics-Based Character ControlYan Wu, Korrawe Karunratanakul, Zhengyi Luo, Siyu TangICCV 2025 · 被引用 3 次
- Auto-Regressive Diffusion for Generating 3D Human-Object InteractionsZichen Geng, Zeeshan Hayder, Wei Liu, Ajmal Saeed MianAAAI 2025 · 被引用 8 次
- Composite Motion Learning with Task ControlPei Xu, Xiumin Shang, Victor B. Zordan, Ioannis KaramouzasSIGGRAPH 2023 · 被引用 27 次
