PRE-MAP: Personalized Reinforced Eye-tracking Multimodal LLM for High-Resolution Multi-Attribute Point Prediction
Hanbing Wu, Ping Jiang, Anyang Su, Chenxu Zhao, Tianyu Fu, Minghui Wu, Beiping Tan, Huiying Li
Abstract
Visual selective attention, driven by individual preferences, regulates human prioritization of visual stimuli by bridging subjective cognitive mechanisms with objective visual elements, thereby steering the semantic interpretation and hierarchical processing of dynamic visual scenes. However, existing models and datasets predominantly neglect the influence of subjective cognitive diversity on fixation behavior. Conventional saliency prediction models, typically employing segmentation approaches, rely on low-resolution imagery to generate saliency heatmaps, subsequently upscaled to native resolutions, which limiting their capacity to capture personalized attention patterns. Furthermore, MLLMs are constrained by factors such as hallucinations, making it very costly to strictly adhere to the expected format in tasks involving multiple point predictions, and achieving precise point positioning is challenging. To address these limitations, we present Subjective Personalized Attention for Ad vertisement Videos, namely SPA-ADV, a large-scale multimodal dataset capturing gaze behaviors from over 4,500 participants varying in age and gender with 486 videos. Furthermore, we propose PRE-MAP, a novel eye-tracking saliency model that characterizes Personalized visual disparities through Reinforcement learning-optimized Eye-tracking, built upon MLLMs and guided by Multi-Attribute user profiles to predict Points. To ensure MLLMs produce prediction points that are both format-correct and spatially accurate, we introduce Consistency Group Relative Policy Optimization (C-GRPO), inspired by the variability in eye movement points and Multi-Attribute profiles. Extensive experiments on SPA-ADV and other benchmarks demonstrate the effectiveness of our approach. The code and dataset are available at https://github.com/mininglamp-MLLM/PRE-MAP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b172e843-6993-4531-852b-de7e2bf54e4bBuilds on5
- DeepGaze IIE: Calibrated prediction in and out-of-domain for state-of-the-art saliency modelingAkis Linardos, Matthias Kümmerer, Ori Press, Matthias BethgeICCV 2021 · 98 citations
- Does text attract attention on e-commerce images: A novel saliency prediction dataset and methodLai Jiang, Yifei Li, Shengxi Li, Mai Xu et al.CVPR 2022 · 21 citations
- Hypergraph Multi-modal Large Language Model: Exploiting EEG and Eye-tracking Modalities to Evaluate Heterogeneous Responses for Video UnderstandingMinghui Wu, Chenxu Zhao, Anyang Su, Donglin Di et al.ACM MM 2024 · 8 citations
- TempSAL - Uncovering Temporal Information for Deep Saliency PredictionBahar Aydemir, Ludo Hoffstetter, Tong Zhang, Mathieu Salzmann et al.CVPR 2023
- Beyond Average: Individualized Visual Scanpath PredictionXianyu Chen, Ming Jiang, Qi ZhaoCVPR 2024
Related papers
- Predicting Goal-Directed Human Attention Using Inverse Reinforcement LearningZhibo Yang, Lihan Huang, Yupei Chen, Zijun Wei et al.CVPR 2020
- Predicting Human Scanpaths in Visual Question AnsweringXianyu Chen, Ming Jiang, Qi ZhaoCVPR 2021
- Learning from Unique Perspectives: User-aware Saliency ModelingShi Chen, Nachiappan Valliappan, Shaolei Shen, Xinyu Ye et al.CVPR 2023
- DiffEye: Diffusion-Based Continuous Eye-Tracking Data Generation Conditioned on Natural ImagesOzgur Kara, Harris Nisar, James M. RehgNeurIPS 2025 · 7 citations
- Guiding Cross-Modal Representations with MLLM Priors via Preference AlignmentPengfei Zhao, Rongbo Luan, Wei Zhang, Peng Wu et al.NeurIPS 2025 · 3 citations
