RealTwin: Concept Graph Representation and Grounding Framework for Reality-Preserving Digital Twin Reconstruction
Zisu Li, Ruohao Li, Jiawei Li, Chao Liu, Junyi Zhu, Daniela Rus, Chen Liang, Mingming Fan
Abstract
Reconstructing realistic digital twins has become crucial as advances in mixed reality, metaverse, and robotics demand more accurate simulations for the physical world. Despite technical progress, building high-fidelity digital twins from a systematic and human-centered perspective remains underexplored. Drawing from the human processing model, we decompose human-centric reality into perception, motion, and cognition, and define a reality-preserving digital twin (RPDT) as a reconstruction integrating these dimensions. We present RealTwin, an attribute-graph-based representation and inference framework for RPDT. Leveraging the grounding capabilities of Multimodal Large Language Models (MLLMs), RealTwin chains AI tools to construct attribute graphs that faithfully encode real-world properties. We validate RealTwin through both technical evaluation, showing promising success in graph parsing and attribute inference, and a user study, assessing its applicability across diverse user groups. Enlightened by RealTwin, we discuss critical issues, including ecology, interaction space, and real-world adoption, for future end-to-end, fine-grained, and scalable digital twin reconstruction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on40
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
- Text2Tex: Text-driven Texture Synthesis via Diffusion ModelsDave Zhenyu Chen, Yawar Siddiqui, Hsin-Ying Lee, Sergey Tulyakov et al.ICCV 2023 · 262 citations
- VR-GS: A Physical Dynamics-Aware Interactive Gaussian Splatting System in Virtual RealityYing Jiang, Chang Yu, Tianyi Xie, Xuan Li et al.SIGGRAPH 2024 · 153 citations
- Holistic++ Scene Understanding: Single-View 3D Holistic Scene Parsing and Human Pose Estimation With Human-Object Interaction and Physical CommonsenseYixin Chen, Siyuan Huang, Tao Yuan, Yixin Zhu et al.ICCV 2019 · 130 citations
- RigNet: neural rigging for articulated charactersZhan Xu, Yang Zhou, Evangelos Kalogerakis, Chris Landreth et al.SIGGRAPH 2020 · 127 citations
Related papers
- ArtLLM: Generating Articulated Assets via 3D LLMPenghao Wang, Siyuan Xie, Hongyu Yan, Xianghui Yang et al.CVPR 2026 · 7 citations
- URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language ModelZhe Li, Xiang Bai, Jieyu Zhang, Zhuangzhe Wu et al.NeurIPS 2025 · 24 citations
- Continuously Updating Digital Twins using Large Language ModelsHarry Amad, Nicolás Astorga, Mihaela van der SchaarICML 2025
- Online Reasoning Video Segmentation with Just-in-Time Digital TwinsYiqing Shen, Bohan Liu, Chenjia Li, Lalithkumar Seenivasan et al.ICCV 2025 · 7 citations
- Ditto: Building Digital Twins of Articulated Objects from InteractionZhenyu Jiang, Cheng-Chun Hsu, Yuke ZhuCVPR 2022 · 77 citations
