Generating Language Corrections for Teaching Physical Control Tasks
Megha Srivastava, Noah D. Goodman, Dorsa Sadigh
Abstract
AI assistance continues to help advance applications in education, from language learning to intelligent tutoring systems, yet current methods for providing students feedback are still quite limited. Most automatic feedback systems either provide binary correctness feedback, which may not help a student understand how to improve, or require hand-coding feedback templates, which may not generalize to new domains. This can be particularly challenging for physical control tasks, where the rich diversity in student behavior and specialized domains make it challenging to leverage general-purpose assistive tools for providing feedback. We design and build CORGI, a model trained to generate language corrections for physical control tasks, such as learning to ride a bike. CORGI takes in as input a pair of student and expert trajectories, and then generates natural language corrections to help the student improve. We collect and train CORGI over data from three diverse physical control tasks (drawing, steering, and joint movement). Through both automatic and human evaluations, we show that CORGI can (i) generate valid feedback for novel student trajectories, (ii) outperform baselines on domains with novel control dynamics, and (iii) improve student learning in an interactive drawing task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
- X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text RetrievalYiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan et al.ACM MM 2022 · 314 citations
- Mapping Language Models to Grounded Conceptual SpacesRoma Patel, Ellie PavlickICLR 2022 · 197 citations
Related papers
- Teaching Models to Improve on TapeLiat Bezalel, Eyal Orgad, Amir GlobersonAAAI 2025
- ExpertAF: Expert Actionable Feedback from VideoKumar Ashutosh, Tushar Nagarajan, Georgios Pavlakos, Kris Kitani et al.CVPR 2025
- Generative Muscle Stimulation: Providing Users with Physical Assistance by Constraining Multimodal-AI with Embodied KnowledgeYun Ho, Romain Nith, Peili Jiang, Steven He et al.CHI 2026 · 1 citation
- CigTime: Corrective Instruction Generation Through Inverse Motion EditingQihang Fang, Chengcheng Tang, Bugra Tekin, Yanchao YangNeurIPS 2024 · 3 citations
- CueTip: An Interactive and Explainable Physics-aware Pool AssistantSean Memery, Kevin Denamganaï, Jiaxin Zhang, Zehai Tu et al.SIGGRAPH 2025
