Time to augment self-supervised visual representation learning
Arthur Aubret, Markus Roland Ernst, Céline Teulière, Jochen Triesch
Abstract
Biological vision systems are unparalleled in their ability to learn visual representations without supervision. In machine learning, self-supervised learning (SSL) has led to major advances in forming object representations in an unsupervised fashion. Such systems learn representations invariant to augmentation operations over images, like cropping or flipping. In contrast, biological vision systems exploit the temporal structure of the visual experience during natural interactions with objects. This gives access to "augmentations" not commonly used in SSL, like watching the same object from multiple viewpoints or against different backgrounds. Here, we systematically investigate and compare the potential benefits of such time-based augmentations during natural interactions for learning object categories. Our results show that time-based augmentations achieve large performance gains over state-of-the-art image augmentations. Specifically, our analyses reveal that: 1) 3-D object manipulations drastically improve the learning of object categories; 2) viewing objects against changing backgrounds is important for learning to discard background-related information from the latent representation. Overall, we conclude that time-based augmentations during natural interactions with objects can substantially improve self-supervised learning, narrowing the gap between artificial and biological vision systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e40cd96-f225-4ea8-9b7e-16c9e1e9ba73Cited by top-tier papers8
- Learning Efficient Coding of Natural Images with Maximum Manifold Capacity RepresentationsThomas E. Yerxa, Yilun Kuang, Eero P. Simoncelli, SueYeon ChungNeurIPS 2023 · 44 citations
- Time Does Tell: Self-Supervised Time-Tuning of Dense Image RepresentationsMohammadreza Salehi, Efstratios Gavves, Cees G. M. Snoek, Yuki M. AsanoICCV 2023 · 34 citations
- Curriculum Learning With Infant Egocentric VideosSaber Sheybani, Himanshu Hansaria, Justin Wood, Linda B. Smith et al.NeurIPS 2023 · 26 citations
- Are Vision Transformers More Data Hungry Than Newborn Visual Systems?Lalit Pandey, Samantha M. W. Wood, Justin N. WoodNeurIPS 2023 · 24 citations
- Towards Principled Representation Learning from Videos for Reinforcement LearningDipendra Misra, Akanksha Saran, Tengyang Xie, Alex Lamb et al.ICLR 2024 · 9 citations
Builds on17
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li et al.AAAI 2020 · 4,134 citations
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan et al.NeurIPS 2020 · 1,631 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
Related papers
- Learning Without Augmenting: Unsupervised Time Series Representation Learning via Frame ProjectionsBerken Utku Demirel, Christian HolzNeurIPS 2025 · 1 citation
- Time-Equivariant Contrastive Video Representation LearningSimon Jenni, Hailin JinICCV 2021 · 64 citations
- Contrastive-Equivariant Self-Supervised Learning Improves Alignment with Primate Visual Area ITThomas E. Yerxa, Jenelle Feather, Eero P. Simoncelli, SueYeon ChungNeurIPS 2024 · 12 citations
- Demystifying Contrastive Self-Supervised Learning: Invariances, Augmentations and Dataset BiasesSenthil Purushwalkam, Abhinav GuptaNeurIPS 2020 · 240 citations
- Learning predictable and robust neural representations by straightening image sequencesXueyan Niu, Cristina Savin, Eero P. SimoncelliNeurIPS 2024 · 13 citations
