Discovering Hidden Visual Concepts Beyond Linguistic Input in Infant Learning
Xueyi Ke, Satoshi Tsutsui, Yayun Zhang, Bihan Wen
Abstract
recognize objects beyond the model's original vocabulary. Furthermore, we compare the differences in representation between infant models and those in modern computer vision models, such as CLIP and ImageNet pre-trained model. Ultimately, our work bridges cognitive science and computer vision by analyzing the internal representations of a computational model trained on an infant visual and linguistic inputs. Our code is available at https://github.com/ Kexueyi/discover_infant_vis .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02f8e77d-7a2b-474b-8000-cc261f163c48Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 229 citations
- Natural Language Descriptions of Deep Visual FeaturesEvan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili et al.ICLR 2022 · 160 citations
- Self-supervised learning through the eyes of a childA. Emin Orhan, Vaibhav V. Gupta, Brenden M. LakeNeurIPS 2020 · 119 citations
Related papers
- BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation ModelsShengao Wang, Wenqi Wang, Zecheng Wang, Max Whitton et al.CVPR 2026 · 4 citations
- BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant LearningShengao Wang, Arjun Chandra, Aoming Liu, Venkatesh Saligrama et al.ICCV 2025 · 8 citations
- BabyVision: Visual Reasoning Beyond LanguageLiang Chen, Weichu Xie, Liang Yiyan, Hongfeng He et al.ICML 2026 · 25 citations
- Is a Caption Worth a Thousand Images? A Study on Representation LearningShibani Santurkar, Yann Dubois, Rohan Taori, Percy Liang et al.ICLR 2023 · 9 citations
- What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable InsightsXin Wen, Bingchen Zhao, Yilun Chen, Jiangmiao Pang et al.NeurIPS 2024 · 19 citations
