Scaling VR Video Conferencing
Mallesham Dasari, Edward Lu, Michael W. Farb, Nuno Pereira, Ivan Liang, Anthony Rowe
Abstract
Virtual Reality (VR) telepresence platforms are being challenged to support live performances, sporting events, and conferences with thousands of users across seamless virtual worlds. Current systems have struggled to meet these demands which has led to high-profile performance events with groups of users isolated in parallel sessions. The core difference in scaling VR environments compared to classic 2D video content delivery comes from the dynamic peer-to-peer spatial dependence on communication. Users have many pair-wise interactions that grow and shrink as they explore spaces. In this paper, we discuss the challenges of VR scaling and present an architecture that supports hundreds of users with spatial audio and video in a single virtual environment. We leverage the property of spatial locality with two key optimizations: (1) a Quality of Service (QoS) scheme to prioritize audio and video traffic based on users' locality, and (2) a resource manager that allocates client connections across multiple servers based on user proximity within the virtual world. Through real-world deployments and extensive evaluations under real and simulated environments, we demonstrate the scalability of our platform while showing improved QoS compared with existing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- MiniMates: Miniature Avatars for AR Remote Meetings within Limited Physical SpacesAkihiro Kiuchi, Jonathan Wieland, Takeo Igarashi, David LindlbauerCHI 2025 · 5 citations
- Roaming Free in the VR World with MP2Yifei Xu, Xumiao Zhang, Yuning Chen, Pan Hu et al.USENIX ATC 2025 · 2 citations
- eXpressSFU: Toward Super-Scalable Video Conferencing with SmartNICsTuan Tran, S. M. H. Hosseini, Seyeon Kim, Kyunghan Lee et al.NSDI 2026
Builds on3
- Immersive light field video with a layered mesh representationMichael Broxton, John Flynn, Ryan S. Overbeck, Daniel Erickson et al.SIGGRAPH 2020 · 271 citations
- XNect: real-time multi-person 3D motion capture with a single RGB cameraDushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller, Weipeng Xu et al.SIGGRAPH 2020 · 267 citations
- Pixel Codec AvatarsShugao Ma, Tomas Simon, Jason M. Saragih, Dawei Wang et al.CVPR 2021
Related papers
- MuV2: Scaling up Multi-user Mobile Volumetric Video Streaming via Content Hybridization and SharingYu Liu, Puqi Zhou, Zejun Zhang, Anlan Zhang et al.MobiCom 2024 · 16 citations
- Addressing Scalability for Real-time Multiuser Holo-portation: Introducing and Assessing a Multipoint Control Unit (MCU) for Volumetric VideoSergi Fernández, Mario Montagud, David Rincón, Juame Moragues et al.ACM MM 2023 · 14 citations
- Evaluating the Impact of Tiled User-Adaptive Real-Time Point Cloud Streaming on VR Remote CommunicationShishir Subramanyam, Irene Viola, Jack Jansen, Evangelos Alexiou et al.ACM MM 2022 · 17 citations
- Evaluating the Effect of Binaural Auralization on Audiovisual Plausibility and Communication Behavior in Virtual RealityFelix Immohr, Gareth Rendle, Anton Lammert, Annika Neidhardt et al.IEEE VR 2024 · 9 citations
- Influence of Audiovisual Realism on Communication Behaviour in Group-to-Group TelepresenceGareth Rendle, Felix Immohr, Christian Kehling, Anton Lammert et al.IEEE VR 2025 · 4 citations
