CaV3: Cache-assisted Viewport Adaptive Volumetric Video Streaming
Junhua Liu, Boxiang Zhu, Fangxin Wang, Yili Jin, Wenyi Zhang, Zihan Xu, Shuguang Cui
Abstract
Volumetric video (VV) recently emerges as a new form of video application providing a photorealistic immersive 3D viewing experience with 6 degree-of-freedom (DoF), which empowers many applications such as VR, AR, and Metaverse. A key problem therein is how to stream the enormous size VV through the network with limited bandwidth. Existing works mostly focused on predicting the viewport for a tiling-based adaptive VV streaming, which however only has quite a limited effect on resource saving. We argue that the content repeatability in the viewport can be further leveraged, and for the first time, propose a client-side cache-assisted strategy that aims to buffer the repeatedly appearing VV tiles in the near future so as to reduce the redundant VV content transmission. The key challenges exist in three aspects, including (1) feature extraction and mining in 6 DoF VV context, (2) accurate long-term viewing pattern estimation and (3) optimal caching scheduling with limited capacity. In this paper, we propose CaV3, an integrated cache-assisted viewport adaptive VV streaming framework to address the challenges. CaV3 employs a Long-short term Sequential prediction model (LSTSP) that achieves accurate short-term, mid-term and long-term viewing pattern prediction with a multi-modal fusion model by capturing the viewer's behavior inertia, current attention, and subjective intention. Besides, CaV3 also contains a contextual MAB-based caching adaptation algorithm (CCA) to fully utilize the viewing pattern and solve the optimal caching problem with a proved upper bound regret. Compared to existing VV datasets only containing single or co-located objects, we for the first time collect a comprehensive dataset with sufficient practical unbounded 360° scenes. The extensive evaluation of the dataset confirms the superiority of CaV3, which outperforms the SOTA algorithm by 15.6%-43% in viewport prediction and 13%-40% in system utility.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9cfcf2fa-1023-472c-a59b-017387259cc7Cited by top-tier papers4
- Understanding User Behavior in Volumetric Video Watching: Dataset, Analysis and PredictionKaiyuan Hu, Haowen Yang, Yili Jin, Junhua Liu et al.ACM MM 2023 · 37 citations
- FSVFG: Towards Immersive Full-Scene Volumetric Video Streaming with Adaptive Feature GridDaheng Yin, Jianxin Shi, Miao Zhang, Zhaowu Huang et al.ACM MM 2024 · 7 citations
- HeadsetOff: Enabling Photorealistic Video Conferencing on Economical VR HeadsetsYili Jin, Xize Duan, Fangxin Wang, Xue LiuACM MM 2024 · 5 citations
- CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian SplattingDaheng Yin, Yili Jin, Jianxin Shi, Isaac Ding et al.SIGGRAPH 2026
Related papers
- Personalized 360-Degree Video Streaming: A Meta-Learning ApproachYiyun Lu, Yifei Zhu, Zhi WangACM MM 2022 · 28 citations
- LiveObj: Object Semantics-based Viewport Prediction for Live Mobile Virtual Reality StreamingXianglong Feng, Zeyang Bao, Sheng WeiIEEE VR 2021 · 33 citations
- ViVo: visibility-aware mobile volumetric video streamingBo Han, Yu Liu, Feng QianMobiCom 2020 · 183 citations
- STAR-VP: Improving Long-term Viewport Prediction in 360° Videos via Space-aligned and Time-varying FusionBaoqi Gao, Daoxu Sheng, Lei Zhang, Qi Qi et al.ACM MM 2024 · 9 citations
- Towards Viewport-dependent 6DoF 360 Video Tiled Streaming for Virtual Reality SystemsJongBeom Jeong, Soonbin Lee, Il-Woong Ryu, Tuan Thanh Le et al.ACM MM 2020 · 27 citations
