Towards In-context Scene Understanding
Ivana Balazevic, David Steiner, Nikhil Parthasarathy, Relja Arandjelovic, Olivier J. Hénaff
摘要
In-context learningthe ability to configure a model's behavior with different promptshas revolutionized the field of natural language processing, alleviating the need for task-specific models and paving the way for generalist models capable of assisting with any query. Computer vision, in contrast, has largely stayed in the former regime: specialized decoders and finetuning protocols are generally required to perform dense tasks such as semantic segmentation and depth estimation. In this work we explore a simple mechanism for in-context learning of such scene understanding tasks: nearest neighbor retrieval from a prompt of annotated features. We propose a new pretraining protocolleveraging attention within and across imageswhich yields representations particularly useful in this regime. The resulting Hummingbird model, suitably prompted, performs various scene understanding tasks without modification while approaching the performance of specialists that have been finetuned for each task. Moreover, Hummingbird can be configured to perform new tasks much more efficiently than finetuned models, raising the possibility of scene understanding in the interactive assistant regime.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- Locality-Attending Vision TransformerSina Hajimiri, Farzad Beizaee, Fereshteh Shakeri, Christian Desrosiers 等ICLR 2026 · 被引用 427 次
- Explore In-Context Learning for 3D Point Cloud UnderstandingZhongbin Fang, Xiangtai Li, Xia Li, Joachim M. Buhmann 等NeurIPS 2023 · 被引用 47 次
- Time Does Tell: Self-Supervised Time-Tuning of Dense Image RepresentationsMohammadreza Salehi, Efstratios Gavves, Cees G. M. Snoek, Yuki M. AsanoICCV 2023 · 被引用 34 次
- Context-Aware Meta-LearningChristopher Fifty, Dennis Duan, Ronald G. Junkins, Ehsan Amid 等ICLR 2024 · 被引用 28 次
- Franca: Nested Matryoshka Clustering for Scalable Visual Representation LearningShashanka Venkataramanan, Valentinos Pariza, Mohammadreza Salehi, Lukas Knobel 等CVPR 2026 · 被引用 26 次
它引用的顶会 Paper33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
相关 Paper
- What Makes Good Examples for Visual In-Context Learning?Yuanhan Zhang, Kaiyang Zhou, Ziwei LiuNeurIPS 2023 · 被引用 219 次
- Skeleton-in-Context: Unified Skeleton Sequence Modeling with In-Context LearningXinshun Wang, Zhongbin Fang, Xia Li, Xiangtai Li 等CVPR 2024 · 被引用 12 次
- Towards More Unified In-Context Visual UnderstandingDianmo Sheng, Dongdong Chen, Zhentao Tan, Qiankun Liu 等CVPR 2024 · 被引用 9 次
- Stable Diffusion Models Are Secretly Good at Visual In-Context LearningTrevine Oorloff, Vishwanath Sindagi, Wele Gedara Chaminda Bandara, Ali Shafahi 等ICCV 2025 · 被引用 11 次
- Show and Segment: Universal Medical Image Segmentation via In-Context LearningYunhe Gao, Di Liu, Zhuowei Li, Yunsheng Li 等CVPR 2025
