ScanTD: 360° Scanpath Prediction based on Time-Series Diffusion
Yujia Wang, Fang-Lue Zhang, Neil A. Dodgson
Abstract
Scanpath generation in 360° images aims to model the realistic trajectories of gaze points that viewers follow when exploring panoramic environments. Existing methods for scanpath genera- tion suffer from various limitations, including a lack of global atten-tion to panoramic environments, insufficient diversity in generated scanpaths, and inadequate consideration of the temporal sequence of gaze points. To address these challenges, we propose a novel approach, named ScanTD, which employs a conditional Diffusion Model-based method to generate multiple scanpaths. Notably, a transformer-based time-series (TTS) module with a novel attention mechanism is integrated into ScanTD to capture the temporal de- pendency of gaze points effectively. Additionally, ScanTD utilizes a Vision Transformer-based method for image feature extraction, en- abling better learning of scene semantic information. Experimental results demonstrate that our approach outperforms state-of-the-art methods across three datasets. We further demonstrate its general- izability by applying it to the 360° saliency detection task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Target Scanpath-Guided 360-Degree Image EnhancementYujia Wang, Fang-Lue Zhang, Neil A. DodgsonAAAI 2025 · 21 citations
- RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video GenerationTianyi Yan, Wencheng Han, Xia Zhou, Xueyang Zhang et al.NeurIPS 2025 · 9 citations
- CompTrack: Information Bottleneck-Guided Low-Rank Dynamic Token Compression for Point Cloud TrackingSifan Zhou, Yichao Cao, Jiahao Nie, Yuqian Fu et al.AAAI 2026 · 9 citations
- MoleBridge: Synthetic Space Projecting with Discrete Markov BridgesRongchao Zhang, Yu Huang, Yongzhi Cao, Hanpin WangNeurIPS 2025 · 8 citations
- DiffEye: Diffusion-Based Continuous Eye-Tracking Data Generation Conditioned on Natural ImagesOzgur Kara, Harris Nisar, James M. RehgNeurIPS 2025 · 7 citations
Builds on19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- ScanDMM: A Deep Markov Model of Scanpath Prediction for 360° ImagesXiangjie Sui, Yuming Fang, Hanwei Zhu, Shiqi Wang et al.CVPR 2023
- Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath PredictionGiuseppe Cartella, Vittorio Cuculo, Alessandro D'Amelio, Marcella Cornia et al.ICCV 2025 · 3 citations
- ScanGAN360: A Generative Model of Realistic Scanpaths for 360° ImagesDaniel Martin, Ana Serrano, Alexander W. Bergman, Gordon Wetzstein et al.IEEE VR 2022 · 71 citations
- SalDiff-DTM: A Novel Dual-Temporal Modulated Diffusion Model for Omnidirectional Images Scanpath PredictionXiaohui Kong, Qian Liu, Dandan Zhu, Kaiwei Zhang et al.AAAI 2026
- SpFormer: Spatio-Temporal Modeling for Scanpaths with TransformerWenqi Zhong, Linzhi Yu, Chen Xia, Junwei Han et al.AAAI 2024 · 7 citations
