Lune

ICML2023顶会

Architecture-Agnostic Masked Image Modeling - From ViT back to CNN

Siyuan Li, Di Wu, Fang Wu, Zelin Zang, Stan Z. Li

2023年份
60被引次数
15顶会引用

摘要

Masked image modeling, an emerging self-supervised pre-training method, has shown impressive success across numerous downstream vision tasks with Vision transformers. Its underlying idea is simple: a portion of the input image is masked out and then reconstructed via a pre-text task. However, the working principle behind MIM is not well explained, and previous studies insist that MIM primarily works for the Transformer family but is incompatible with CNNs. In this work, we observe that MIM essentially teaches the model to learn better middle-order interactions among patches for more generalized feature extraction. We then propose an Architecture-Agnostic Masked Image Modeling framework (A2^2MIM), which is compatible with both Transformers and CNNs in a unified way. Extensive experiments on popular benchmarks show that A2^2MIM learns better representations without explicit design and endows the backbone model with the stronger capability to transfer to various downstream tasks.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 6cfd3a7d-9a64-4e3b-ac40-d6175bbb4107

引用它的顶会 Paper15

问问它们各自怎么用它

它引用的顶会 Paper38

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖