Lune

NeurIPS2023顶会

StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation Learners

Yonglong Tian, Lijie Fan, Phillip Isola, Huiwen Chang, Dilip Krishnan

2023年份
251被引次数
64顶会引用

摘要

We investigate the potential of learning visual representations using synthetic images generated by text-to-image models. This is a natural question in the light of the excellent performance of such models in generating high-quality images. We consider specifically the Stable Diffusion, one of the leading open source text-toimage models. We show that (1) when the generative model is configured with proper classifier-free guidance scale, training self-supervised methods on synthetic images can match or beat the real image counterpart; (2) by treating the multiple images generated from the same text prompt as positives for each other, we develop a multi-positive contrastive learning method, which we call StableRep. With solely synthetic images, the representations learned by StableRep surpass the performance of representations learned by SimCLR and CLIP using the same set of text prompts and corresponding real images, on large scale datasets. When we further add language supervision, StableRep trained with 20M synthetic images achieves better accuracy than CLIP trained with 50M real images. Generative Models Stable Diffusion (SD) Data Engine Embedding Real data (A) Traditional Representation Learning (B) Representation Learning with Synthetic Data Synthetic Data Real data Embedding Synthetic data Encoder Encoder Figure 1: Left: traditional visual representation learning relies on a dataset of real images to train an image embedding function. Right: we view generative models as datasets that allow us to sample images from the data distribution. In our study, we leverage text-to-image models (Stable Diffusion [61]) and treat multiple images synthesized from the same prompt as positives for contrastive representation learning.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 607866d0-2182-4a1d-beb3-e47cfbc0337c

引用它的顶会 Paper64

问问它们各自怎么用它

它引用的顶会 Paper29

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖