Does Visual Pretraining Help End-to-End Reasoning?
Chen Sun, Calvin Luo, Xingyi Zhou, Anurag Arnab, Cordelia Schmid
Abstract
We aim to investigate whether end-to-end learning of visual reasoning can be achieved with general-purpose neural networks, with the help of visual pretraining. A positive result would refute the common belief that explicit visual abstraction (e.g. object detection) is essential for compositional generalization on visual reasoning, and confirm the feasibility of a neural network "generalist" to solve visual recognition and reasoning tasks. We propose a simple and general self-supervised framework which "compresses" each video frame into a small set of tokens with a transformer network, and reconstructs the remaining frames based on the compressed temporal context. To minimize the reconstruction loss, the network must learn a compact representation for each image, as well as capture temporal dynamics and object permanence from temporal context. We perform evaluation on two visual reasoning benchmarks, CATER and ACRE. We observe that pretraining is essential to achieve compositional generalization for end-to-end visual reasoning. Our proposed framework outperforms traditional supervised pretraining, including image classification and explicit object detection, by large margins.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9dcdcf4f-50ac-4145-b2ef-bbabba7b2a3aCited by top-tier papers1
Ask how each one uses itBuilds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- Generative Video Transformer: Can Objects be the Words?Yi-Fu Wu, Jaesik Yoon, Sungjin AhnICML 2021 · 37 citations
- Hopper: Multi-hop Transformer for Spatiotemporal ReasoningHonglu Zhou, Asim Kadav, Farley Lai, Alexandru Niculescu-Mizil et al.ICLR 2021 · 19 citations
- CATER: A diagnostic dataset for Compositional Actions & TEmporal ReasoningRohit Girdhar, Deva RamananICLR 2020 · 198 citations
- An Empirical Study of Autoregressive Pre-Training from VideosJathushan Rajasegaran, Ilija Radosavovic, Rahul Ravishankar, Yossi Gandelsman et al.ICCV 2025
- Attention over Learned Object Embeddings Enables Complex Visual ReasoningDavid Ding, Felix Hill, Adam Santoro, Malcolm Reynolds et al.NeurIPS 2021 · 87 citations
