ScanTD: 360° Scanpath Prediction based on Time-Series Diffusion
Yujia Wang, Fang-Lue Zhang, Neil A. Dodgson
摘要
Scanpath generation in 360° images aims to model the realistic trajectories of gaze points that viewers follow when exploring panoramic environments. Existing methods for scanpath genera- tion suffer from various limitations, including a lack of global atten-tion to panoramic environments, insufficient diversity in generated scanpaths, and inadequate consideration of the temporal sequence of gaze points. To address these challenges, we propose a novel approach, named ScanTD, which employs a conditional Diffusion Model-based method to generate multiple scanpaths. Notably, a transformer-based time-series (TTS) module with a novel attention mechanism is integrated into ScanTD to capture the temporal de- pendency of gaze points effectively. Additionally, ScanTD utilizes a Vision Transformer-based method for image feature extraction, en- abling better learning of scene semantic information. Experimental results demonstrate that our approach outperforms state-of-the-art methods across three datasets. We further demonstrate its general- izability by applying it to the 360° saliency detection task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Target Scanpath-Guided 360-Degree Image EnhancementYujia Wang, Fang-Lue Zhang, Neil A. DodgsonAAAI 2025 · 被引用 21 次
- RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video GenerationTianyi Yan, Wencheng Han, Xia Zhou, Xueyang Zhang 等NeurIPS 2025 · 被引用 9 次
- CompTrack: Information Bottleneck-Guided Low-Rank Dynamic Token Compression for Point Cloud TrackingSifan Zhou, Yichao Cao, Jiahao Nie, Yuqian Fu 等AAAI 2026 · 被引用 9 次
- MoleBridge: Synthetic Space Projecting with Discrete Markov BridgesRongchao Zhang, Yu Huang, Yongzhi Cao, Hanpin WangNeurIPS 2025 · 被引用 8 次
- DiffEye: Diffusion-Based Continuous Eye-Tracking Data Generation Conditioned on Natural ImagesOzgur Kara, Harris Nisar, James M. RehgNeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- ScanDMM: A Deep Markov Model of Scanpath Prediction for 360° ImagesXiangjie Sui, Yuming Fang, Hanwei Zhu, Shiqi Wang 等CVPR 2023
- Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath PredictionGiuseppe Cartella, Vittorio Cuculo, Alessandro D'Amelio, Marcella Cornia 等ICCV 2025 · 被引用 3 次
- ScanGAN360: A Generative Model of Realistic Scanpaths for 360° ImagesDaniel Martin, Ana Serrano, Alexander W. Bergman, Gordon Wetzstein 等IEEE VR 2022 · 被引用 71 次
- SalDiff-DTM: A Novel Dual-Temporal Modulated Diffusion Model for Omnidirectional Images Scanpath PredictionXiaohui Kong, Qian Liu, Dandan Zhu, Kaiwei Zhang 等AAAI 2026
- SpFormer: Spatio-Temporal Modeling for Scanpaths with TransformerWenqi Zhong, Linzhi Yu, Chen Xia, Junwei Han 等AAAI 2024 · 被引用 7 次
