Learning Calibratable Policies using Programmatic Style-Consistency
Eric Zhan, Albert Tseng, Yisong Yue, Adith Swaminathan, Matthew J. Hausknecht
Abstract
We study the problem of controllable generation of long-term sequential behaviors, where the goal is to calibrate to multiple behavior styles simultaneously. In contrast to the well-studied areas of controllable generation of images, text, and speech, there are two questions that pose significant challenges when generating long-term behaviors: how should we specify the factors of variation to control, and how can we ensure that the generated behavior faithfully demonstrates combinatorially many styles? We leverage programmatic labeling functions to specify controllable styles, and derive a formal notion of styleconsistency as a learning objective, which can then be solved using conventional policy learning approaches. We evaluate our framework using demonstrations from professional basketball players and agents in the MuJoCo physics environment, and show that existing approaches that do not explicitly enforce style-consistency fail to generate diverse behaviors whereas our learned policies can be calibrated for up to 4 5 (1024) distinct style combinations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 367af859-fc47-4f6f-9e3f-e1ebabd17fc0Cited by top-tier papers9
- Uni[MASK]: Unified Inference in Sequential Decision ProblemsMicah Carroll, Orr Paradise, Jessy Lin, Raluca Georgescu et al.NeurIPS 2022 · 29 citations
- Animal behavioral analysis and neural encoding with transformer-based self-supervised pretrainingYanchen Wang, Han Yu, Ari Blau, Yizi Zhang et al.ICLR 2026 · 8 citations
- Semi-Supervised Generative Models for Multiagent TrajectoriesDennis Fassmeyer, Pascal Fassmeyer, Ulf BrefeldNeurIPS 2022 · 7 citations
- Synthesizing Trajectory Queries from ExamplesStephen Mell, Favyen Bastani, Steve Zdancewic, Osbert BastaniCAV 2023 · 5 citations
- Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation LearningHanlin Yang, Jian Yao, Weiming Liu, Qing Wang et al.ICLR 2025
Builds on2
Related papers
- MultiAct: Long-Term 3D Human Motion Generation from Multiple Action LabelsTaeryung Lee, Gyeongsik Moon, Kyoung Mu LeeAAAI 2023 · 62 citations
- Control strategies for physically simulated characters performing two-player competitive sportsJungdam Won, Deepak Gopinath, Jessica K. HodginsSIGGRAPH 2021 · 73 citations
- UniPhys: Unified Planner and Controller with Diffusion for Flexible Physics-Based Character ControlYan Wu, Korrawe Karunratanakul, Zhengyi Luo, Siyu TangICCV 2025 · 3 citations
- Auto-Regressive Diffusion for Generating 3D Human-Object InteractionsZichen Geng, Zeeshan Hayder, Wei Liu, Ajmal Saeed MianAAAI 2025 · 8 citations
- Composite Motion Learning with Task ControlPei Xu, Xiumin Shang, Victor B. Zordan, Ioannis KaramouzasSIGGRAPH 2023 · 27 citations
