SalDiff-DTM: A Novel Dual-Temporal Modulated Diffusion Model for Omnidirectional Images Scanpath Prediction
Xiaohui Kong, Qian Liu, Dandan Zhu, Kaiwei Zhang, Xiongkuo Min
摘要
Scanpath prediction in omnidirectional images (ODIs) serves as a critical component for optimizing foveated rendering efficiency and enhancing interactive quality in virtual reality systems. However, existing scanpath prediction methods for ODIs still suffer from fundamental limitations: (1) inadequate modeling and capturing of long-range temporal dependencies in fixation regions, and (2) suboptimal integration of spatial and temporal visual features, ultimately compromising prediction performance. To address these limitations, we propose a novel Dual-Temporal Modulated Diffusion model for Omnidirectional Images Scanpath Prediction, named SalDiff-DTM model, to effectively generate realistic human eye viewing trajectories. Specifically, to effectively model spatial relationships, we propose a novel Dual-Graph Convolutional Network (Dual-GCN) module that simultaneously captures semantic-level and image-level correlations. By integrating both local spatial details and global contextual information across the internal temporal dimension, this module achieves comprehensive and robust modeling of spatial relationships. To further enhance the modeling of temporal dependencies inherent in diverse fixation patterns, we introduce TABiMamba (Temporal-Aware BiLSTM-Mamba), a dedicated module that synergistically combines the contextual sensitivity of BiLSTM with the long-range sequence modeling capabilities of Mamba. This design facilitates deep information flow and context-aware sequential reasoning, thereby enabling high-fidelity capture of intricate temporal correlations. Inspired by the progressive refinement mechanism of diffusion models in various generative tasks, we propose a saliency-guided diffusion module that formulates the prediction problem as a conditional generative process, iteratively yielding accurate and perceptually plausible scanpaths. Extensive experiments demonstrate that SalDiff-DTM significantly outperforms state-of-the-art models, paving the way for future advancements in eye-tracking technologies and cognitive modeling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
- Palette: Image-to-Image Diffusion ModelsChitwan Saharia, William Chan, Huiwen Chang, Chris A. Lee 等SIGGRAPH 2022 · 被引用 1,638 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- CSDI: Conditional Score-based Diffusion Models for Probabilistic Time Series ImputationYusuke Tashiro, Jiaming Song, Yang Song, Stefano ErmonNeurIPS 2021 · 被引用 1,245 次
- Autoregressive Denoising Diffusion Models for Multivariate Probabilistic Time Series ForecastingKashif Rasul, Calvin Seward, Ingmar Schuster, Roland VollgrafICML 2021 · 被引用 500 次
相关 Paper
- ScanTD: 360° Scanpath Prediction based on Time-Series DiffusionYujia Wang, Fang-Lue Zhang, Neil A. DodgsonACM MM 2024 · 被引用 11 次
- ScanDMM: A Deep Markov Model of Scanpath Prediction for 360° ImagesXiangjie Sui, Yuming Fang, Hanwei Zhu, Shiqi Wang 等CVPR 2023
- DiffEye: Diffusion-Based Continuous Eye-Tracking Data Generation Conditioned on Natural ImagesOzgur Kara, Harris Nisar, James M. RehgNeurIPS 2025 · 被引用 7 次
- OSMamba: Omnidirectional Spectral Mamba with Dual-Domain Prior Generator for Exposure CorrectionGehui Li, Bin Chen, Chen Zhao, Lei Zhang 等CVPR 2025
- Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath PredictionGiuseppe Cartella, Vittorio Cuculo, Alessandro D'Amelio, Marcella Cornia 等ICCV 2025 · 被引用 3 次
