Lune

CVPR2023Top-tier venue

OmniMAE: Single Model Masked Pretraining on Images and Videos

Rohit Girdhar, Alaaeldin El-Nouby, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, Ishan Misra

2023Year
19Top-tier citations

Abstract

Figure 1 . OmniMAE is a single model for images and videos that is trained using masked autoencoding [40] . We use a plain Vision Transformer [24] architecture but with spatio-temporal patches as input. At training, we 'patchify' the visual input (images or videos), and feed the encoder only a subset of the patches. The decoder reconstructs the pixels for the missing patches using the encoder's output. The encoder-decoder model is trained using a pixel reconstruction loss. After training, our single plain Transformer encoder performs competitively compared to specialized architectures on downstream image and video recognition tasks.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext b6c595af-f38b-43e4-a067-da0f4dcd471f

Cited by top-tier papers19

Ask how each one uses it

Builds on47

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines