ShowMak3r: Compositional TV Show Reconstruction
Sangmin Kim, Seunguk Do, Jaesik Park
Abstract
Reconstructing dynamic radiance fields from video clips is challenging, especially when entertainment videos like TV shows are given. Many challenges make the reconstruction difficult due to actors occluding with each other and having diverse facial expressions, cluttered stages, and small baseline views or sudden shot changes. Reconstruction becomes even more challenging when dealing with general monocular web videos, which present an even greater degree of unpredictability and complexity compared to controlled environments. To address these issues, we present ShowMak3r++, a unified reconstruction pipeline that targets both controlled settings like TV shows and uncontrolled scenarios like web videos. Our pipeline allows the editing of scenes like how video clips are made in a production control room after the reconstruction is done. In our pipeline, we propose a spatio-temporal positioning module that locates actors on the stage by using depth prior while maintaining 2D image alignment and natural 3D motions. ShotMatcher module then tracks the actors under shot changes. Finally, a face-fitting network dynamically recovers the actors' expressions. Experiments on Sitcoms3D and CMU Panoptic datasets show that our pipeline can reassemble TV show scenes with new cameras at different timestamps. We also demonstrate that our method can successfully reconstruct challenging web videos, including dynamic action clips, dance videos, and movie clips. Furthermore, we demonstrate that our pipeline enables interesting applications such as synthetic shot-making, actor relocation, insertion, deletion, and pose manipulation. Project
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d8d0be0-2552-4f96-accd-2bb2bc50785cCited by top-tier papers2
- 4D Human-Scene Reconstruction from Low-Overlap CapturesMinhyuk Hwang, Sangmin Kim, Seunguk Do, Daneul Kim et al.SIGGRAPH 2026
- Surface-Aware Feed-Forward Quadratic Gaussian for Frame Interpolation with Large MotionZaoming Yan, Yaomin Huang, Pengcheng Lei, Qizhou Chen et al.NeurIPS 2025
Builds on55
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
Related papers
- DeepFaceFlow: In-the-Wild Dense 3D Facial Motion EstimationMohammad Rami Koujan, Anastasios Roussos, Stefanos ZafeiriouCVPR 2020
- Robust Dynamic Radiance FieldsYu-Lun Liu, Chen Gao, Andreas Meuleman, Hung-Yu Tseng et al.CVPR 2023
- Facial hair tracking for high fidelity performance captureSebastian Winberg, Gaspard Zoss, Prashanth Chandran, Paulo F. U. Gotardo et al.SIGGRAPH 2022 · 11 citations
- KeyTr: Keypoint Transporter for 3D Reconstruction of Deformable Objects in VideosDavid Novotný, Ignacio Rocco, Samarth Sinha, Alexandre Carlier et al.CVPR 2022 · 11 citations
- Augmenting TV Shows via Uncalibrated Camera Small Motion Tracking in Dynamic SceneYizhen Lao, Jie Yang, Xinying Wang, Jianxin Lin et al.ACM MM 2021 · 1 citation
