Rethinking 360° Image Visual Attention Modelling with Unsupervised Learning
Yasser Abdelaziz Dahou Djilali, Tarun Krishna, Kevin McGuinness, Noel E. O'Connor
Abstract
Despite the success of self-supervised representation learning on planar data, to date it has not been studied on 360°images. In this paper, we extend recent advances in contrastive learning to learn latent representations that are sufficiently invariant to be highly effective for spherical saliency prediction as a downstream task. We argue that omni-directional images are particularly suited to such an approach due to the geometry of the data domain. To verify this hypothesis, we design an unsupervised framework that effectively maximizes the mutual information between the different views from both the equator and the poles. We show that the decoder is able to learn good quality saliency distributions from the encoder embeddings. Our model compares favorably with fully-supervised learning methods on the Salient360!, VR-EyeTracking and Sitzman datasets. This performance is achieved using an encoder that is trained in a completely unsupervised way and a relatively lightweight supervised decoder (3.8 × fewer parameters in the case of the ResNet50 encoder). We believe that this combination of supervised and unsupervised learning is an important step toward flexible formulations of human visual attention. The results can be reproduced on GitHub * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b81079b-45b5-449a-b099-d71d36dc4e03Builds on6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Automatic Shortcut Removal for Self-Supervised Representation LearningMatthias Minderer, Olivier Bachem, Neil Houlsby, Michael TschannenICML 2020 · 78 citations
- Self-Supervised Learning of Pretext-Invariant RepresentationsIshan Misra, Laurens van der MaatenCVPR 2020
Related papers
- CoCoNets: Continuous Contrastive 3D Scene RepresentationsShamit Lal, Mihir Prabhudesai, Ishita Mediratta, Adam W. Harley et al.CVPR 2021
- SalGCN: Saliency Prediction for 360-Degree Images Based on Spherical Graph Convolutional NetworksHaoran Lv, Qin Yang, Chenglin Li, Wenrui Dai et al.ACM MM 2020 · 24 citations
- EquiAV: Leveraging Equivariance for Audio-Visual Contrastive LearningJongsuk Kim, Hyeongkeun Lee, Kyeongha Rho, Junmo Kim et al.ICML 2024 · 15 citations
- SalBiNet360: Saliency Prediction on 360° Images with Local-Global Bifurcated Deep NetworkDongwen Chen, Chunmei Qing, Xiangmin Xu, Huansheng ZhuIEEE VR 2020 · 3 citations
- Spherical Pseudo-Cylindrical Representation for Omnidirectional Image Super-resolutionQing Cai, Mu Li, Dongwei Ren, Jun Lyu et al.AAAI 2024 · 11 citations
