Look Ma, No Hands! Agent-Environment Factorization of Egocentric Videos
Matthew Chang, Aditya Prakash, Saurabh Gupta
Abstract
The analysis and use of egocentric videos for robotic tasks is made challenging by occlusion due to the hand and the visual mismatch between the human hand and a robot end-effector. In this sense, the human hand presents a nuisance. However, often hands also provide a valuable signal, e.g. the hand pose may suggest what kind of object is being held. In this work, we propose to extract a factored representation of the scene that separates the agent (human hand) and the environment. This alleviates both occlusion and mismatch while preserving the signal, thereby easing the design of models for downstream robotics tasks. At the heart of this factorization is our proposed Video Inpainting via Diffusion Model (VIDM) that leverages both a prior on real-world images (through a large-scale pre-trained diffusion model) and the appearance of the object in earlier frames of the video (through attention). Our experiments demonstrate the effectiveness of VIDM at improving inpainting quality on egocentric videos and the power of our factored representation for numerous tasks: object detection, 3D reconstruction of manipulated objects, and learning of reward functions, policies, and affordances from videos. Project website: https://matthewchang.github.io/vidm . Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5e5fdcdf-7a95-43bf-b723-4508146a6f01Cited by top-tier papers4
- Zero-Shot Robotic Manipulation via 3D Gaussian Splatting-Enhanced Multimodal Retrieval-Augmented GenerationZilong Xie, Jingyu Gong, Xin Tan, Zhizhong Zhang et al.AAAI 2026
- VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic ManipulationHanzhi Chen, Boyang Sun, Anran Zhang, Marc Pollefeys et al.CVPR 2025
- 2HandedAfforder: Learning Precise Actionable Bimanual Affordances from Human VideosMarvin Heidinger, Snehal Jauhri, Vignesh Prasad, Georgia ChalvatzakiICCV 2025
- How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday InteractionsAditya Prakash, Benjamin Lundell, Dmitry Andreychuk, David Forsyth et al.CVPR 2025
Builds on39
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
Related papers
- Diffusion-Guided Reconstruction of Everyday Hand-Object Interaction ClipsYufei Ye, Poorvi Hebbar, Abhinav Gupta, Shubham TulsianiICCV 2023 · 80 citations
- Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-TrainingHaoran He, Chenjia Bai, Ling Pan, Weinan Zhang et al.NeurIPS 2024 · 38 citations
- EgoControl: Controllable Egocentric Video Generation via 3D Full-Body PosesEnrico Pallotta, Sina Mokhtarzadeh Azar, Lars Doorenbos, Serdar Ozsoy et al.CVPR 2026 · 7 citations
- TACO: Taming Diffusion for In-the-Wild Video Amodal CompletionRuijie Lu, Yixin Chen, Yu Liu, Jiaxiang Tang et al.ICCV 2025 · 3 citations
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsYucheng Hu, Yanjiang Guo, Pengchao Wang, Xiaoyu Chen et al.ICML 2025
