Generating Humanless Environment Walkthroughs from Egocentric Walking Tour Videos
Yujin Ham, Junho Kim, Vivek Boominathan, Guha Balakrishnan
Abstract
Egocentric ``walking tour'' videos provide a rich source of image data to develop rich and diverse visual models of environments around the world. However, the significant presence of humans in frames of these videos due to crowds and eye-level camera perspectives mitigates their usefulness in environment modeling applications. We focus on addressing this challenge by developing a generative algorithm that can realistically remove (i.e., inpaint) humans and their associated shadow effects from walking tour videos. Key to our approach is the construction of a rich semi-synthetic dataset of video clip pairs to train this generative model. Each pair in the dataset consists of an environment-only background clip, and a composite clip of walking humans with simulated shadows overlaid on the background. We randomly sourced both foreground and background components from real egocentric walking videos around the world to maintain visual diversity. We then used this dataset to fine-tune the state-of-the-art Casper video diffusion model for object and effects inpainting, and demonstrate that the resulting model performs far better than Casper both qualitatively and quantitatively at removing humans from walking tour clips with significant human presence and complex backgrounds. Finally, we show that the resulting generated clips can be used to build successful 3D Gaussian Splatting models of urban locations which was otherwise not possible from the original clips.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- ProPainter: Improving Propagation and Transformer for Video InpaintingShangchen Zhou, Chongyi Li, Kelvin C. K. Chan, Chen Change LoyICCV 2023 · 205 citations
Related papers
- AutoRemover: Automatic Object Removal for Autonomous Driving VideosRong Zhang, Wei Li, Peng Wang, Chenye Guan et al.AAAI 2020 · 16 citations
- 3D StreetUnveiler with Semantic-aware 2DGS - a simple baselineJingwei Xu, Yikai Wang, Yiqun Zhao, Yanwei Fu et al.ICLR 2025
- HouseTour: A Virtual Real Estate A(I)gentAta Çelen, Marc Pollefeys, Dániel Baráth, Iro ArmeniICCV 2025 · 1 citation
- EffectErase: Joint Video Object Removal and Insertion for High-Quality Effect ErasingYANG FU, Yike Zheng, Ziyun Dai, Henghui DingCVPR 2026 · 14 citations
- DeGauss: Dynamic-Static Decomposition with Gaussian Splatting for Distractor-Free 3D ReconstructionRui Wang, Quentin Lohmeyer, Mirko Meboldt, Siyu TangICCV 2025 · 12 citations
