Tile Classification Based Viewport Prediction with Multi-modal Fusion Transformer
Zhihao Zhang, Yiwei Chen, Weizhan Zhang, Caixia Yan, Qinghua Zheng, Qi Wang, Wangdu Chen
Abstract
Viewport prediction is a crucial aspect of tile-based 360° video streaming system. However, existing trajectory based methods lack of robustness, also oversimplify the process of information construction and fusion between different modality inputs, leading to the error accumulation problem. In this paper, we propose a tile classification based viewport prediction method with Multi-modal Fusion Transformer, namely MFTR. Specifically, MFTR utilizes transformer-based networks to extract the long-range dependencies within each modality, then mine intra- and inter-modality relations to capture the combined impact of user historical inputs and video contents on future viewport selection. In addition, MFTR categorizes future tiles into two categories: user interested or not, and selects future viewport as the region that contains most user interested tiles. Comparing with predicting head trajectories, choosing future viewport based on tile's binary classification results exhibits better robustness and interpretability. To evaluate our proposed MFTR, we conduct extensive experiments on two widely used PVS-HM and Xu-Gaze dataset. MFTR shows superior performance over state-of-the-art methods in terms of average prediction accuracy and overlap ratio, also presents competitive computation efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b7da52b-e0fe-4514-9947-18c6c426c014Cited by top-tier papers5
- Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model OutputsYiwei Chen, Soumyadeep Pal, Yimeng Zhang, Qing Qu et al.ICLR 2026 · 15 citations
- Unleashing the Power of Chain-of-Prediction for Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Girish Chandar Ganesan, Xiaoming LiuCVPR 2026 · 13 citations
- FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human RecognitionJie Zhu, Xiao Guo, Yiyang Su, Anil K. Jain et al.CVPR 2026 · 7 citations
- Towards Intrinsic-Aware Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Xiaoming LiuCVPR 2026 · 5 citations
- Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine UnlearningYiwei Chen, Yuguang Yao, Yihua Zhang, Bingquan Shen et al.ICLR 2026
Builds on6
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
Related papers
- STAR-VP: Improving Long-term Viewport Prediction in 360° Videos via Space-aligned and Time-varying FusionBaoqi Gao, Daoxu Sheng, Lei Zhang, Qi Qi et al.ACM MM 2024 · 9 citations
- A Spherical Convolution Approach for Learning Long Term Viewport Prediction in 360 Immersive VideoChenglei Wu, Ruixiao Zhang, Zhi Wang, Lifeng SunAAAI 2020 · 48 citations
- Personalized 360-Degree Video Streaming: A Meta-Learning ApproachYiyun Lu, Yifei Zhu, Zhi WangACM MM 2022 · 28 citations
- CaV3: Cache-assisted Viewport Adaptive Volumetric Video StreamingJunhua Liu, Boxiang Zhu, Fangxin Wang, Yili Jin et al.IEEE VR 2023 · 40 citations
- Towards Viewport-dependent 6DoF 360 Video Tiled Streaming for Virtual Reality SystemsJongBeom Jeong, Soonbin Lee, Il-Woong Ryu, Tuan Thanh Le et al.ACM MM 2020 · 27 citations
