VoluMe - Authentic 3D Video Calls from Live Gaussian Splat Prediction
Martin de La Gorce, Charlie Hewitt, Tibor Takács, Robert Gerdisch, Zafiirah Hosenie, Givi Meishvili, Marek Kowalski, Thomas J. Cashman, Antonio Criminisi
Abstract
Virtual 3D meetings offer the potential to enhance copresence, increase engagement and thus improve effectiveness of remote meetings compared to standard 2D video calls. However, representing people in 3D meetings remains a challenge; existing solutions achieve high quality by using complex hardware, making use of fixed appearance via enrolment, or by inverting a pre-trained generative model. These approaches lead to constraints that are unwelcome and ill-fitting for videoconferencing applications. We present the first method to predict 3D Gaussian reconstructions in real time from a single 2D webcam feed, where the 3D representation is not only live and realistic, but also authentic to the input video. By conditioning the 3D representation on each video frame independently, our reconstruction faithfully recreates the input video from the captured viewpoint (a property we call authenticity), while generalizing realistically to novel viewpoints. Additionally, we introduce a stability loss to obtain reconstructions that are temporally stable on video sequences. We show that our method delivers state-of-the-art accuracy in visual quality and stability metrics compared to existing methods, and demonstrate our approach in live one-to-one 3D meetings using only a standard 2D camera and display. This demonstrates that our approach can allow anyone to communicate volumetrically, via a method for 3D videoconferencing that is not only highly accessible, but also realistic and authentic.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 827e471b-16f5-4c9b-8863-03fac446bcc2Cited by top-tier papers1
Ask how each one uses itBuilds on21
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano et al.CVPR 2022 · 984 citations
- Mixture of volumetric primitives for efficient neural renderingStephen Lombardi, Tomas Simon, Gabriel Schwartz, Michael Zollhöfer et al.SIGGRAPH 2021 · 240 citations
- GaussianAvatars: Photorealistic Head Avatars with Rigged 3D GaussiansShenhan Qian, Tobias Kirschstein, Liam Schoneveld, Davide Davoli et al.CVPR 2024 · 175 citations
Related papers
- The eyes have it: an integrated eye and face model for photorealistic facial animationGabriel Schwartz, Shih-En Wei, Te-Li Wang, Stephen Lombardi et al.SIGGRAPH 2020 · 54 citations
- Coherent 3D Portrait Video Reconstruction via Triplane FusionShengze Wang, Xueting Li, Chao Liu, Matthew A. Chan et al.CVPR 2025
- GASP: Gaussian Avatars with Synthetic PriorsJack R. Saunders, Charlie Hewitt, Yanan Jian, Marek Kowalski et al.CVPR 2025
- VRGaussianAvatar: Integrating 3D Gaussian Avatars into VRHail Song, Boram Yoon, Seokhwan Yang, Seoyoung Kang et al.IEEE VR 2026 · 2 citations
- GaussianAvatar: Towards Realistic Human Avatar Modeling from a Single Video via Animatable 3D GaussiansLiangxiao Hu, Hongwen Zhang, Yuxiang Zhang, Boyao Zhou et al.CVPR 2024
