Physics-guided Shape-from-Template: Monocular Video Perception through Neural Surrogate Models
David Stotko, Nils Wandel, Reinhard Klein
Abstract
3D reconstruction of dynamic scenes is a long-standing problem in computer graphics and increasingly difficult the less information is available. Shape-from-Template (SfT) methods aim to reconstruct a template-based geometry from RGB images or video sequences, often leveraging just a single monocular camera without depth information, such as regular smartphone recordings. Unfortunately, existing reconstruction methods are either unphysical and noisy or slow in optimization. To solve this problem, we propose a novel SfT reconstruction algorithm for cloth using a pre-trained neural surrogate model that is fast to evaluate, stable, and produces smooth reconstructions due to a regularizing physics simulation. Differentiable rendering of the simulated mesh enables pixel-wise comparisons between the reconstruction and a target video sequence that can be used for a gradient-based optimization procedure to extract not only shape information but also physical parameters such as stretching, shearing, or bending stiffness of the cloth. This allows to retain a precise, stable, and smooth reconstructed geometry while reducing the runtime by a factor of 400–500 compared to ϕ-SfT, a state-of-the-art physics-based SfT approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- SAFT: Shape and Appearance of Fabrics from Template via Differentiable Physical Simulations from Monocular VideoDavid Stotko, Reinhard KleinICCV 2025 · 1 citation
- Thin-Shell-SfT: Fine-Grained Monocular Non-rigid 3D Surface Tracking with Neural Deformation FieldsNavami Kairanda, Marc Habermann, Shanthika Naik, Christian Theobalt et al.CVPR 2025
- Image-Guided Shape-From-Template Using Mesh Inextensibility ConstraintsThuy Tran, Ruochen Chen, Shaifali ParasharICCV 2025
- Physics from Video: Identifiability of Time-Invariant Second-Order ODEs under Minimal Trajectory ConditionsYuanyuan Wang, Wenjie Wang, Kun Zhang, Mingming GongICML 2026
- Metamizer: A Versatile Neural Optimizer for Fast and Accurate Physics SimulationsNils Wandel, Stefan Schulz, Reinhard KleinICLR 2025
Builds on13
- Learning Mesh-Based Simulation with Graph NetworksTobias Pfaff, Meire Fortunato, Alvaro Sanchez-Gonzalez, Peter W. BattagliaICLR 2021 · 1,175 citations
- DiffTaichi: Differentiable Programming for Physical SimulationYuanming Hu, Luke Anderson, Tzu-Mao Li, Qi Sun et al.ICLR 2020 · 479 citations
- PhysDiff: Physics-Guided Human Motion Diffusion ModelYe Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat et al.ICCV 2023 · 414 citations
- InterDiff: Generating 3D Human-Object Interactions with Physics-Informed DiffusionSirui Xu, Zhengyuan Li, Yu-Xiong Wang, Liang-Yan GuiICCV 2023 · 201 citations
- Scalable Differentiable Physics for Learning and ControlYi-Ling Qiao, Junbang Liang, Vladlen Koltun, Ming C. LinICML 2020 · 133 citations
Related papers
- -SfT: Shape-from-Template with a Physics-Based Deformation ModelNavami Kairanda, Edith Tretschk, Mohamed A. Elgharib, Christian Theobalt et al.CVPR 2022 · 20 citations
- SelfRecon: Self Reconstruction Your Digital Avatar from Monocular VideoBoyi Jiang, Yang Hong, Hujun Bao, Juyong ZhangCVPR 2022 · 142 citations
- NGD: Neural Gradient Based Deformation for Monocular Garment ReconstructionSoham Dasgupta, Shanthika Naik, Preet Savalia, Sujay Kumar Ingle et al.ICCV 2025 · 2 citations
- NSF: Neural Surface Fields for Human Modeling from Monocular DepthYuxuan Xue, Bharat Lal Bhatnagar, Riccardo Marin, Nikolaos Sarafianos et al.ICCV 2023 · 26 citations
- REC-MV: REconstructing 3D Dynamic Cloth from Monocular VideosLingteng Qiu, Guanying Chen, Jiapeng Zhou, Mutian Xu et al.CVPR 2023
