SAFT: Shape and Appearance of Fabrics from Template via Differentiable Physical Simulations from Monocular Video
David Stotko, Reinhard Klein
Abstract
The reconstruction of three-dimensional dynamic scenes is a well-established yet challenging task within the domain of computer vision. In this paper, we propose a novel approach that combines the domains of 3D geometry reconstruction and appearance estimation for physically based rendering and present a system that is able to perform both tasks for fabrics, utilizing only a single monocular RGB video sequence as input. In order to obtain realistic and high-quality deformations and renderings, a physical simulation of the cloth geometry and differentiable rendering are employed. In this paper, we introduce two novel regularization terms for the 3D reconstruction task that improve the plausibility of the reconstruction by addressing the depth ambiguity problem in monocular video. In comparison with the most recent methods in the field, we have reduced the error in the 3D reconstruction by a factor of 2.64 while requiring a medium runtime of 30 min per scene. Furthermore, the optimized motion achieves sufficient quality to perform an appearance estimation of the deforming object, recovering sharp details from this single monocular RGB video.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c496e5cc-23dd-4dac-b42b-4b0fc507a345Builds on25
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Learning Mesh-Based Simulation with Graph NetworksTobias Pfaff, Meire Fortunato, Alvaro Sanchez-Gonzalez, Peter W. BattagliaICLR 2021 · 1,175 citations
- DiffTaichi: Differentiable Programming for Physical SimulationYuanming Hu, Luke Anderson, Tzu-Mao Li, Qi Sun et al.ICLR 2020 · 479 citations
- PhysDiff: Physics-Guided Human Motion Diffusion ModelYe Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat et al.ICCV 2023 · 414 citations
- Extracting Triangular 3D Models, Materials, and Lighting From ImagesJacob Munkberg, Wenzheng Chen, Jon Hasselgren, Alex Evans et al.CVPR 2022 · 306 citations
Related papers
- Physics-guided Shape-from-Template: Monocular Video Perception through Neural Surrogate ModelsDavid Stotko, Nils Wandel, Reinhard KleinCVPR 2024 · 6 citations
- MonoCloth: Reconstruction and Animation of Cloth-Decoupled Human Avatars from Monocular VideosDaisheng Jin, Ying HeAAAI 2026 · 1 citation
- Learning Motion-Dependent Appearance for High-Fidelity Rendering of Dynamic Humans from a Single CameraJae Shin Yoon, Duygu Ceylan, Tuanfeng Y. Wang, Jingwan Lu et al.CVPR 2022 · 10 citations
- -SfT: Shape-from-Template with a Physics-Based Deformation ModelNavami Kairanda, Edith Tretschk, Mohamed A. Elgharib, Christian Theobalt et al.CVPR 2022 · 20 citations
- NGD: Neural Gradient Based Deformation for Monocular Garment ReconstructionSoham Dasgupta, Shanthika Naik, Preet Savalia, Sujay Kumar Ingle et al.ICCV 2025 · 2 citations
