LITE: Loss-resilient Immersive Telepresence with Multi-modal Semantics
Ruizhi Cheng, Harshvardhan Takawale, Nan Wu, Nirupam Roy, Sennur Ulukus, Matteo Varvello, Eugene Chai, Bo Han
Abstract
Immersive telepresence has the potential to transform real-time communication through highly interactive and engaging experiences. Despite recent advances in reducing communication and computation costs, existing systems largely overlook packet loss, which can severely degrade the quality of experience (QoE). Recovering lost immersive content is considerably more challenging than in 2D video due to the complexity of dense 3D representations. Recovery must be both accurate and timely while minimizing the communication and computation overhead it incurs. To address these challenges, we present LITE, the first loss-resilient immersive telepresence system. LITE incorporates three key design principles: (1) leveraging semantic communication to transmit compact motion and audio semantics, which can be reconstructed into the remote user's immersive representation and voice, enabling fast semantic-level recovery and remaining robust to congestion-control-induced rate reductions under loss; (2) fusing audio and motion semantics via a lightweight multimodal model to achieve accurate, real-time recovery of motion semantics; and (3) encoding audio semantics from multiple past frames into succinct neural redundancy to enable robust recovery. We prototype LITE using a well-known parametric facial motion representation and extensively evaluate its performance across diverse networks. Our results demonstrate that LITE improves QoE by up to 109% compared with existing schemes, while sustaining real-time streaming at 30 frames per second and preserving high visual fidelity (structural similarity index measure above 0.9, where 1 indicates perfect similarity).
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- FarfetchFusion: Towards Fully Mobile Live 3D Telepresence PlatformKyungjin Lee, Juheon Yi, Youngki LeeMobiCom 2023 · 30 citations
- RIFTCast: A Template-Free End-to-End Multi-View Live Telepresence Framework and BenchmarkDomenic Zingsheim, Markus Plack, Hannah Dröge, Janelle Pfeifer et al.ACM MM 2025 · 2 citations
- HeadsetOff: Enabling Photorealistic Video Conferencing on Economical VR HeadsetsYili Jin, Xize Duan, Fangxin Wang, Xue LiuACM MM 2024 · 5 citations
- Evaluating the Impact of Tiled User-Adaptive Real-Time Point Cloud Streaming on VR Remote CommunicationShishir Subramanyam, Irene Viola, Jack Jansen, Evangelos Alexiou et al.ACM MM 2022 · 17 citations
- GauMVC: Generative Decoupled Gaussian Representation for Human-centric Multi-view Video CompressionRuoke Yan, Mingjia Yang, Xinfeng Zhang, Haocheng Tang et al.CVPR 2026
