Dexterous World Models
Byungjun Kim, Taeksoo Kim, Junyoung Lee, Hanbyul Joo
Abstract
Recent progress in 3D reconstruction has made it easy to create realistic digital twins from everyday environments. However, current digital twins remain largely static-limited to navigation and view synthesis without embodied interactivity. To bridge this gap, we introduce Dexterous World Model (DWM), a scene-action-conditioned video diffusion framework that models how dexterous human actions induce dynamic changes in static 3D scenes. Given a static 3D scene rendering and an egocentric hand motion sequence, DWM generates temporally coherent videos depicting plausible human-scene interactions. Our approach conditions video generation on (1) static scene renderings following a specified camera trajectory to ensure spatial consistency, and (2) egocentric hand mesh renderings that encode both geometry and motion cues in the egocentric view to model action-conditioned dynamics directly. To train DWM, we * Equal contribution construct a hybrid interaction video dataset: synthetic egocentric interactions provide fully aligned supervision for joint locomotion-manipulation learning, while fixed-camera real-world videos contribute diverse and realistic object dynamics. Experiments demonstrate that DWM enables realistic, physically plausible interactions, such as grasping, opening, or moving objects, while maintaining camera and scene consistency. This framework establishes the first step toward video diffusion-based interactive digital twins, enabling embodied simulation from egocentric actions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22b08418-871b-40d2-a071-edf741a4e61cCited by top-tier papers1
Ask how each one uses itBuilds on38
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- AGILE: Hand-object Interaction Reconstruction from Video via Agentic GenerationJin-Chuan Shi, Binhong Ye, Tao Liu, Xiaoyang Liu et al.SIGGRAPH 2026
- Diffusion-Guided Reconstruction of Everyday Hand-Object Interaction ClipsYufei Ye, Poorvi Hebbar, Abhinav Gupta, Shubham TulsianiICCV 2023 · 80 citations
- AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D DiffusionHongjie Li, Heng Yu, Jiaman Li, Hong-Xing Yu et al.CVPR 2026 · 2 citations
- SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction ScenariosLingwei Dang, Ruizhi Shao, Hongwen Zhang, Wei Min et al.NeurIPS 2025 · 12 citations
- DreamDojo: A Real-Time Robot World Model from Large-Scale Human VideosShenyuan Gao, William Liang, Kaiyuan Zheng, Ayaan Malik et al.ICML 2026 · 96 citations
