Self-supervised learning through the eyes of a child
A. Emin Orhan, Vaibhav V. Gupta, Brenden M. Lake
Abstract
Within months of birth, children develop meaningful expectations about the world around them. How much of this early knowledge can be explained through generic learning mechanisms applied to sensory data, and how much of it requires more substantive innate inductive biases? Addressing this fundamental question in its full generality is currently infeasible, but we can hope to make real progress in more narrowly defined domains, such as the development of high-level visual categories, thanks to improvements in data collecting technology and recent progress in deep learning. In this paper, our goal is precisely to achieve such progress by utilizing modern self-supervised deep learning methods and a recent longitudinal, egocentric video dataset recorded from the perspective of three young children (Sullivan et al., 2020). Our results demonstrate the emergence of powerful, high-level visual representations from developmentally realistic natural videos using generic self-supervised learning objectives.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec381cbb-5a9b-4d93-aa6a-3f24928189b2Cited by top-tier papers24
- Hard Negative Mixing for Contrastive LearningYannis Kalantidis, Mert Bülent Sariyildiz, Noé Pion, Philippe Weinzaepfel et al.NeurIPS 2020 · 805 citations
- Partial success in closing the gap between human and machine visionRobert Geirhos, Kantharaju Narayanappa, Benjamin Mitzkus, Tizian Thieringer et al.NeurIPS 2021 · 304 citations
- Grounded Language Learning Fast and SlowFelix Hill, Olivier Tieleman, Tamara von Glehn, Nathaniel Wong et al.ICLR 2021 · 85 citations
- Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled videoShashanka Venkataramanan, Mamshad Nayeem Rizve, João Carreira, Yuki M. Asano et al.ICLR 2024 · 42 citations
- Self-Supervised Representation Learning from Flow EquivarianceYuwen Xiong, Mengye Ren, Wenyuan Zeng, Raquel Urtasun WaabiICCV 2021 · 32 citations
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Selectivity considered harmful: evaluating the causal impact of class selectivity in DNNsMatthew L. Leavitt, Ari S. MorcosICLR 2021 · 34 citations
- Unsupervised Learning From Video With Deep Neural EmbeddingsChengxu Zhuang, Tianwei She, Alex Andonian, Max Sobol Mark et al.CVPR 2020
- Momentum Contrast for Unsupervised Visual Representation LearningKaiming He, Haoqi Fan, Yuxin Wu, Saining Xie et al.CVPR 2020
Related papers
- Curriculum Learning With Infant Egocentric VideosSaber Sheybani, Himanshu Hansaria, Justin Wood, Linda B. Smith et al.NeurIPS 2023 · 26 citations
- Learning to See Through a Baby’s Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and MachinesYusen Cai, Qing Lin, BHARGAVA SATYA NUNNA, Mengmi ZhangCVPR 2026 · 4 citations
- Self-Supervised Object Detection from Egocentric VideosPeri Akiva, Jing Huang, Kevin J. Liang, Rama Kovvuri et al.ICCV 2023 · 9 citations
- Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric VideoYuting Tan, Xilong Cheng, Yunxiao Qin, Zhengnan Li et al.CVPR 2026 · 1 citation
- Demystifying Contrastive Self-Supervised Learning: Invariances, Augmentations and Dataset BiasesSenthil Purushwalkam, Abhinav GuptaNeurIPS 2020 · 240 citations
