LITE: Loss-resilient Immersive Telepresence with Multi-modal Semantics
Ruizhi Cheng, Harshvardhan Takawale, Nan Wu, Nirupam Roy, Sennur Ulukus, Matteo Varvello, Eugene Chai, Bo Han
摘要
Immersive telepresence has the potential to transform real-time communication through highly interactive and engaging experiences. Despite recent advances in reducing communication and computation costs, existing systems largely overlook packet loss, which can severely degrade the quality of experience (QoE). Recovering lost immersive content is considerably more challenging than in 2D video due to the complexity of dense 3D representations. Recovery must be both accurate and timely while minimizing the communication and computation overhead it incurs. To address these challenges, we present LITE, the first loss-resilient immersive telepresence system. LITE incorporates three key design principles: (1) leveraging semantic communication to transmit compact motion and audio semantics, which can be reconstructed into the remote user's immersive representation and voice, enabling fast semantic-level recovery and remaining robust to congestion-control-induced rate reductions under loss; (2) fusing audio and motion semantics via a lightweight multimodal model to achieve accurate, real-time recovery of motion semantics; and (3) encoding audio semantics from multiple past frames into succinct neural redundancy to enable robust recovery. We prototype LITE using a well-known parametric facial motion representation and extensively evaluate its performance across diverse networks. Our results demonstrate that LITE improves QoE by up to 109% compared with existing schemes, while sustaining real-time streaming at 30 frames per second and preserving high visual fidelity (structural similarity index measure above 0.9, where 1 indicates perfect similarity).
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- FarfetchFusion: Towards Fully Mobile Live 3D Telepresence PlatformKyungjin Lee, Juheon Yi, Youngki LeeMobiCom 2023 · 被引用 30 次
- RIFTCast: A Template-Free End-to-End Multi-View Live Telepresence Framework and BenchmarkDomenic Zingsheim, Markus Plack, Hannah Dröge, Janelle Pfeifer 等ACM MM 2025 · 被引用 2 次
- HeadsetOff: Enabling Photorealistic Video Conferencing on Economical VR HeadsetsYili Jin, Xize Duan, Fangxin Wang, Xue LiuACM MM 2024 · 被引用 5 次
- Evaluating the Impact of Tiled User-Adaptive Real-Time Point Cloud Streaming on VR Remote CommunicationShishir Subramanyam, Irene Viola, Jack Jansen, Evangelos Alexiou 等ACM MM 2022 · 被引用 17 次
- GauMVC: Generative Decoupled Gaussian Representation for Human-centric Multi-view Video CompressionRuoke Yan, Mingjia Yang, Xinfeng Zhang, Haocheng Tang 等CVPR 2026
