Autonomous Functional Play with Correspondence-Driven Trajectory Warping
William Liang, Sam Wang, Hung-Ju Wang, Osbert Bastani, Yecheng Jason Ma, Dinesh Jayaraman
Abstract
The ability to conduct and learn from interaction and experience is a central challenge in robotics, offering a scalable alternative to labor-intensive human demonstrations. However, realizing such "play" requires (1) a policy robust to diverse, potentially out-of-distribution environment states, and (2) a procedure that continuously produces useful robot experience. To address these challenges, we introduce Tether, a method for autonomous functional play involving structured, task-directed interactions. First, we design a novel open-loop policy that warps actions from a small set of source demonstrations (≤ 10) by anchoring them to semantic keypoint correspondences in the target scene. We show that this design is extremely data-efficient and robust even under significant spatial and semantic variations. Second, we deploy this policy for autonomous functional play in the real world via a continuous cycle of task selection, execution, evaluation, and improvement, guided by the visual understanding capabilities of vision-language models. This procedure generates diverse, high-quality datasets with minimal human intervention. In a household-like multi-object setup, our method is the first to perform many hours of autonomous multi-task play in the real world starting from only a handful of demonstrations. This produces a stream of data that consistently improves the performance of closed-loop imitation policies over time, ultimately yielding over 1000 expert-level trajectories and training policies competitive with those learned from human-collected demonstrations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 120cb0be-df68-46b4-9f3d-f77cda0fc526Builds on8
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- LIV: Language-Image Representations and Rewards for Robotic ControlYecheng Jason Ma, Vikash Kumar, Amy Zhang, Osbert Bastani et al.ICML 2023 · 212 citations
- GenSim: Generating Robotic Simulation Tasks via Large Language ModelsLirui Wang, Yiyang Ling, Zhecheng Yuan, Mohit Shridhar et al.ICLR 2024 · 143 citations
- Autonomous Reinforcement Learning via Subgoal CurriculaArchit Sharma, Abhishek Gupta, Sergey Levine, Karol Hausman et al.NeurIPS 2021 · 41 citations
- AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic ManipulationJiafei Duan, Wilbert Pumacay, Nishanth Kumar, Yi Ru Wang et al.ICLR 2025 · 4 citations
Related papers
- UniJEPA: Enhancing Robot Policy via Unified Continuous and Discrete Representation LearningJianke Zhang, Yucheng Hu, Yanjiang Guo, Xiaoyu Chen et al.ICML 2026
- STRAP: Robot Sub-Trajectory Retrieval for Augmented Policy LearningMarius Memmel, Jacob Berg, Bingqing Chen, Abhishek Gupta et al.ICLR 2025
- Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAsJunhao Shi, Siyin Wang, Xiaopeng Yu, Li Ji et al.ICML 2026
- AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Affordance CorrespondenceJiawei Zhang, Kaizhe Hu, Yingqian Huang, Yuanchen Ju et al.CVPR 2026
- KISA: A Unified Keyframe Identifier and Skill Annotator for Long-Horizon Robotics DemonstrationsLongxin Kou, Fei Ni, Yan Zheng, Jinyi Liu et al.ICML 2024 · 5 citations
