Modeling Saliency Dataset Bias
Matthias Kümmerer, Harneet Singh Khanuja, Matthias Bethge
Abstract
Recent advances in image-based saliency prediction are approaching gold standard performance levels on existing benchmarks. Despite this success, we show that predicting fixations across multiple saliency datasets remains challenging due to dataset bias. We find a significant performance drop (around 40%) when models trained on one dataset are applied to another. Surprisingly, increasing dataset diversity does not resolve this inter-dataset gap, with close to 60% attributed to dataset-specific biases. To address this remaining generalization gap, we propose a novel architecture extending a mostly dataset-agnostic encoder-decoder structure with fewer than 20 dataset-specific parameters that govern interpretable mechanisms such as multi-scale structure, center bias, and fixation spread. Adapting only these parameters to new data accounts for more than 75% of the generalization gap, with a large fraction of the improvement achieved with as few as 50 samples. Our model sets a new state-of-the-art on all three datasets of the MIT/Tuebingen Saliency Benchmark (MIT300, CAT2000, and COCO-Freeview), even when purely generalizing from unrelated datasets, but with a substantial boost when adapting to the respective training datasets. The model also provides valuable insights into spatial saliency properties, revealing complex multi-scale effects that combine both absolute and relative sizes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- What Moves the Eyes: Doubling Mechanistic Model Performance Using Deep Networks to Discover and Test Cognitive HypothesesFederico D'Agostino, Lisa Schwetlick, Matthias Bethge, Matthias KümmererNeurIPS 2025 · 4 citations
- Attend to Anything: Foundation Model for Unified Human Attention ModelingWenzhuo Zhao, Ronghao Xian, Keren Fu, Qijun ZhaoICML 2026
Builds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Improving robustness against common corruptions by covariate shift adaptationSteffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann et al.NeurIPS 2020 · 688 citations
- DeepGaze IIE: Calibrated prediction in and out-of-domain for state-of-the-art saliency modelingAkis Linardos, Matthias Kümmerer, Ori Press, Matthias BethgeICCV 2021 · 98 citations
- When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It?Mert Yüksekgönül, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky et al.ICLR 2023 · 37 citations
Related papers
- When Random Saliency Looks Trained: Architectural Center Bias in CNN InterpretabilityKeying Kuang, Iain Carmichael, Elizabeth PurdomICML 2026
- Mesh Saliency: An Independent Perceptual Measure or a Derivative of Image Saliency?Ran Song, Wei Zhang, Yitian Zhao, Yonghuai Liu et al.CVPR 2021
- Mitigating Gender Bias in Captioning SystemsRuixiang Tang, Mengnan Du, Yuening Li, Zirui Liu et al.WWW 2021 · 77 citations
- DiffSal: Joint Audio and Video Learning for Diffusion Saliency PredictionJunwen Xiong, Peng Zhang, Tao You, Chuanyue Li et al.CVPR 2024 · 12 citations
- Revisiting Saliency Metrics: Farthest-Neighbor Area Under CurveSen Jia, Neil D. B. BruceCVPR 2020
