ShowMak3r: Compositional TV Show Reconstruction
Sangmin Kim, Seunguk Do, Jaesik Park
摘要
Reconstructing dynamic radiance fields from video clips is challenging, especially when entertainment videos like TV shows are given. Many challenges make the reconstruction difficult due to actors occluding with each other and having diverse facial expressions, cluttered stages, and small baseline views or sudden shot changes. Reconstruction becomes even more challenging when dealing with general monocular web videos, which present an even greater degree of unpredictability and complexity compared to controlled environments. To address these issues, we present ShowMak3r++, a unified reconstruction pipeline that targets both controlled settings like TV shows and uncontrolled scenarios like web videos. Our pipeline allows the editing of scenes like how video clips are made in a production control room after the reconstruction is done. In our pipeline, we propose a spatio-temporal positioning module that locates actors on the stage by using depth prior while maintaining 2D image alignment and natural 3D motions. ShotMatcher module then tracks the actors under shot changes. Finally, a face-fitting network dynamically recovers the actors' expressions. Experiments on Sitcoms3D and CMU Panoptic datasets show that our pipeline can reassemble TV show scenes with new cameras at different timestamps. We also demonstrate that our method can successfully reconstruct challenging web videos, including dynamic action clips, dance videos, and movie clips. Furthermore, we demonstrate that our pipeline enables interesting applications such as synthetic shot-making, actor relocation, insertion, deletion, and pose manipulation. Project
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- 4D Human-Scene Reconstruction from Low-Overlap CapturesMinhyuk Hwang, Sangmin Kim, Seunguk Do, Daneul Kim 等SIGGRAPH 2026
- Surface-Aware Feed-Forward Quadratic Gaussian for Frame Interpolation with Large MotionZaoming Yan, Yaomin Huang, Pengcheng Lei, Qizhou Chen 等NeurIPS 2025
它引用的顶会 Paper55
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao 等NeurIPS 2023 · 被引用 1,498 次
相关 Paper
- DeepFaceFlow: In-the-Wild Dense 3D Facial Motion EstimationMohammad Rami Koujan, Anastasios Roussos, Stefanos ZafeiriouCVPR 2020
- Robust Dynamic Radiance FieldsYu-Lun Liu, Chen Gao, Andreas Meuleman, Hung-Yu Tseng 等CVPR 2023
- Facial hair tracking for high fidelity performance captureSebastian Winberg, Gaspard Zoss, Prashanth Chandran, Paulo F. U. Gotardo 等SIGGRAPH 2022 · 被引用 11 次
- KeyTr: Keypoint Transporter for 3D Reconstruction of Deformable Objects in VideosDavid Novotný, Ignacio Rocco, Samarth Sinha, Alexandre Carlier 等CVPR 2022 · 被引用 11 次
- Augmenting TV Shows via Uncalibrated Camera Small Motion Tracking in Dynamic SceneYizhen Lao, Jie Yang, Xinying Wang, Jianxin Lin 等ACM MM 2021 · 被引用 1 次
