Pixel Aligned Language Models
Jiarui Xu, Xingyi Zhou, Shen Yan, Xiuye Gu, Anurag Arnab, Chen Sun, Xiaolong Wang, Cordelia Schmid
摘要
Large language models have achieved great success in recent years, so as their variants in vision. Existing vision-language models can describe images in natural languages, answer visual-related questions, or perform complex reasoning about the image. However, it is yet unclear how localization tasks, such as word grounding or referring localization, can be performed using large language models. In this work, we aim to develop a vision-language model that can take locations, for example, a set of points or boxes, as either inputs or outputs. When taking locations as inputs, the model performs location-conditioned captioning, which generates captions for the indicated object or region. When generating locations as outputs, our model regresses pixel coordinates for each output word generated by the language model, and thus performs dense word grounding. Our model is pre-trained on the Localized Narrative dataset, which contains pixel-word-aligned captioning from human attention. We show our model can be applied to various location-aware vision-language tasks, including referring localization, location-conditioned captioning, and dense object captioning, archiving state-of-the-art performance on RefCOCO and Visual Genome.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMsSoroush Nasiriany, Fei Xia, Wenhao Yu, Ted Xiao 等ICML 2024 · 被引用 212 次
- VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language TasksJiannan Wu, Muyan Zhong, Sen Xing, Zeqiang Lai 等NeurIPS 2024 · 被引用 179 次
- Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene UnderstandingYunze Man, Shuhong Zheng, Zhipeng Bao, Martial Hebert 等NeurIPS 2024 · 被引用 56 次
- ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language ModelsMingrui Wu, Xinyue Cai, Jiayi Ji, Jiale Li 等NeurIPS 2024 · 被引用 50 次
- SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial ReasoningYang Liu, Ming Ma, Xiaomin Yu, Pengxiang Ding 等NeurIPS 2025 · 被引用 46 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
相关 Paper
- Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual GroundingWenbo Chen, Zhen Xu, Ruotao Xu, Si Wu 等CVPR 2025
- Connecting the Dots: Training-Free Visual Grounding via Agentic ReasoningLiqin Luo, Guangyao Chen, Xiawu Zheng, Yongxing Dai 等AAAI 2026
- FlexCap: Describe Anything in Images in Controllable DetailDebidatta Dwibedi, Vidhi Jain, Jonathan Tompson, Andrew Zisserman 等NeurIPS 2024 · 被引用 11 次
- Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual GroundingSeil Kang, Jinyeong Kim, Junhyeok Kim, Seong Jae HwangCVPR 2025
- Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?Gregor Geigle, Radu Timofte, Goran GlavasEMNLP 2024 · 被引用 1 次
