LocCa: Visual Pretraining with Location-aware Captioners
Bo Wan, Michael Tschannen, Yongqin Xian, Filip Pavetic, Ibrahim M. Alabdulmohsin, Xiao Wang, André Susano Pinto, Andreas Steiner, Lucas Beyer, Xiaohua Zhai
摘要
Image captioning has been shown as an effective pretraining method similar to contrastive pretraining. However, the incorporation of location-aware information into visual pretraining remains an area with limited research. In this paper, we propose a simple visual pretraining method with location-aware captioners (LocCa). LocCa uses a simple image captioner task interface, to teach a model to read out rich information, i.e. bounding box coordinates, and captions, conditioned on the image pixel input. Thanks to the multitask capabilities of an encoder-decoder architecture, we show that an image captioner can easily handle multiple tasks during pretraining. Our experiments demonstrate that LocCa outperforms standard captioners significantly on localization downstream tasks while maintaining comparable performance on holistic tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Cambrian-S: Towards Spatial Supersensing in VideoShusheng Yang, Jihan Yang, Pinzhi Huang, Ellis Brown 等ICLR 2026 · 被引用 139 次
- Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence LearningApoorv Vyas, Heng-Jui Chang, Cheng-Fu Yang, Po-Yao Huang 等CVPR 2026 · 被引用 23 次
- No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language ModelsAngéline Pouget, Lucas Beyer, Emanuele Bugliarello, Xiao Wang 等NeurIPS 2024 · 被引用 17 次
- TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text AlignmentBingyi Cao, Koert Chen, Kevis-Kokitsi Maninis, Kaifeng Chen 等CVPR 2026 · 被引用 14 次
- Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf InteractionsHubert Baniecki, Maximilian Muschalik, Fabian Fumagalli, Barbara Hammer 等NeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper44
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
相关 Paper
- LocTex: Learning Data-Efficient Visual Representations from Localized Textual SupervisionZhijian Liu, Simon Stent, Jie Li, John Gideon 等ICCV 2021 · 被引用 10 次
- Contrastive Localized Language-Image Pre-TrainingHong-You Chen, Zhengfeng Lai, Haotian Zhang, Xinze Wang 等ICML 2025
- CONICA: A Contrastive Image Captioning Framework with Robust Similarity LearningLin Deng, Yuzhong Zhong, Maoning Wang, Jianwei ZhangACM MM 2023 · 被引用 4 次
- CapOnImage: Context-driven Dense-Captioning on ImageYiqi Gao, Xinglin Hou, Yuanmeng Zhang, Tiezheng Ge 等EMNLP 2022 · 被引用 5 次
- Improving fine-grained understanding in image-text pre-trainingIoana Bica, Anastasija Ilic, Matthias Bauer, Goker Erdogan 等ICML 2024 · 被引用 53 次
