Internet Explorer: Targeted Representation Learning on the Open Web
Alexander Cong Li, Ellis Langham Brown, Alexei A. Efros, Deepak Pathak
摘要
Modern vision models typically rely on finetuning general-purpose models pre-trained on large, static datasets. These general-purpose models only capture the knowledge within their pretraining datasets, which are tiny, out-of-date snapshots of the Internet-where billions of images are uploaded each day. We suggest an alternate approach: rather than hoping our static datasets transfer to our desired tasks after large-scale pretraining, we propose dynamically utilizing the Internet to quickly train a small-scale model that does extremely well on the task at hand. Our approach, called Internet Explorer, explores the web in a self-supervised manner to progressively find relevant examples that improve performance on a desired target dataset. It cycles between searching for images on the Internet with text queries, self-supervised training on downloaded images, determining which images were useful, and prioritizing what to search for next. We evaluate Internet Explorer across several datasets and show that it outperforms or matches CLIP oracle performance by using just a single GPU desktop to actively query the Internet for 30-40 hours. Results, visualizations, videos, and code on our website: internet-explorer-ssl.github.io/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo 等NeurIPS 2024 · 被引用 1,004 次
- The Neglected Tails in Vision-Language ModelsShubham Parashar, Zhiqiu Lin, Tian Liu, Xiangjue Dong 等CVPR 2024 · 被引用 15 次
- Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific ModelsRaviteja Vemulapalli, Hadi Pouransari, Fartash Faghri, Sachin Mehta 等ICML 2024 · 被引用 15 次
- GenDataAgent: On-the-fly Dataset Augmentation with Synthetic DataZhiteng Li, Lele Chen, Jerone T. A. Andrews, Yunhao Ba 等ICLR 2025
- Few-Shot Recognition via Stage-Wise Retrieval-Augmented FinetuningTian Liu, Huixin Zhang, Shubham Parashar, Shu KongCVPR 2025
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun 等ICML 2021 · 被引用 2,942 次
相关 Paper
- Visual-Language Navigation Pretraining via Prompt-based Environmental Self-explorationXiwen Liang, Fengda Zhu, Lingling Li, Hang Xu 等ACL 2022 · 被引用 37 次
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal 等ICLR 2022 · 被引用 503 次
- Classification Done Right for Vision-Language Pre-TrainingZilong Huang, Qinghao Ye, Bingyi Kang, Jiashi Feng 等NeurIPS 2024 · 被引用 13 次
- Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote AlignmentUtkarsh Mall, Cheng Perng Phoo, Meilin Kelsey Liu, Carl Vondrick 等ICLR 2024 · 被引用 90 次
- Masked Unsupervised Self-training for Label-free Image ClassificationJunnan Li, Silvio Savarese, Steven C. H. HoiICLR 2023 · 被引用 7 次
