ImViD: Immersive Volumetric Videos for Enhanced VR Engagement
Zhengxian Yang, Shi Pan, Shengqi Wang, Haoxiang Wang, Li Lin, Guanjun Li, Zhengqi Wen, Borong Lin, Jianhua Tao, Tao Yu
Abstract
User engagement is greatly enhanced by fully immersive multi-modal experiences that combine visual and auditory stimuli. Consequently, the next frontier in VR/AR technologies lies in immersive volumetric videos with complete scene capture, large 6-DoF interaction space, multimodal feedback, and high resolution & frame-rate contents. To stimulate the reconstruction of immersive volumetric videos, we introduce ImViD, a multi-view, multi-modal dataset featuring complete space-oriented data capture and various indoor/outdoor scenarios. Our capture rig supports multi-view video-audio capture while on the move, a capability absent in existing datasets, significantly enhancing the completeness, flexibility, and efficiency of data capture.The captured multi-view videos (with synchronized audios) are in 5K resolution at 60FPS, lasting from 1-5 minutes, and include rich foreground-background elements, and complex dynamics. We benchmark existing methods using our dataset and establish a base pipeline for constructing immersive volumetric videos from multi-view audiovisual inputs for 6-DoF multi-modal immersive VR experiences. The benchmark and the reconstruction and interaction results demonstrate the effectiveness of our dataset and baseline method, which we believe will stimulate future research on immersive volumetric video production. Project Page: https://yzxqh.github.io/ImViD/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on33
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian SplattingZeyu Yang, Hongye Yang, Zijie Pan, Li ZhangICLR 2024 · 529 citations
- 4D Gaussian Splatting for Real-Time Dynamic Scene RenderingGuanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie et al.CVPR 2024 · 513 citations
- Neural 3D Video Synthesis from Multi-view VideoTianye Li, Mira Slavcheva, Michael Zollhöfer, Simon Green et al.CVPR 2022 · 324 citations
- Immersive light field video with a layered mesh representationMichael Broxton, John Flynn, Ryan S. Overbeck, Daniel Erickson et al.SIGGRAPH 2020 · 271 citations
Related papers
- SceneHub4D: A Dataset and Evaluation Framework for 6-DoF 4D VR ScenesJaehong Kim, Tao Jin, Mallesham Dasari, Srinivasan Seshan et al.IEEE VR 2026 · 1 citation
- From an Image to a Scene: Learning to Imagine the World from a Million 360° VideosMatthew Wallingford, Anand Bhattad, Aditya Kusupati, Vivek Ramanujan et al.NeurIPS 2024 · 2 citations
- SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human GenerationYouliang Zhang, Zhaoyang Li, Duomin Wang, jiahe zhang et al.ICLR 2026 · 30 citations
- Replay: Multi-modal Multi-view Acted Videos for Casual HolographyRoman Shapovalov, Yanir Kleiman, Ignacio Rocco, David Novotný et al.ICCV 2023 · 11 citations
- OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World ModelingYang Zhou, Yifan Wang, Jianjun Zhou, Wenzheng Chang et al.ICLR 2026 · 58 citations
