Describe Anything: Detailed Localized Image and Video Captioning
Long Lian, Yifan Ding, Yunhao Ge, Sifei Liu, Hanzi Mao, Boyi Li, Marco Pavone, Ming-Yu Liu, Trevor Darrell, Adam Yala, Yin Cui
Abstract
Generating detailed and accurate descriptions for specific regions in images and videos remains a fundamental challenge for vision-language models. We introduce the Describe Anything Model (DAM), a model designed for detailed localized captioning (DLC). DAM preserves both local details and global context through two key innovations: a focal prompt, which ensures high-resolution encoding of targeted regions, and a localized vision backbone, which integrates precise localization with its broader context. To tackle the scarcity of high-quality DLC data, we propose a Semi-supervised learning (SSL)-based Data Pipeline (DLC-SDP). DLC-SDP starts with existing segmentation datasets and expands to unlabeled web images using SSL. We introduce DLC-Bench, a benchmark designed to evaluate DLC without relying on reference captions. DAM sets new state-of-the-art on 7 benchmarks spanning keyword-level, phrase-level, and detailed multi-sentence localized image and video captioning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers37
- Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and VideosWeifeng Lin, Xinyu Wei, Ruichuan An, Tianhe Ren et al.NeurIPS 2025 · 47 citations
- Describe Anything Anywhere At Any MomentNicolas Gorlo, Lukas Schmid, Luca CarloneCVPR 2026 · 27 citations
- ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMsXiyao Wang, Zhengyuan Yang, Chao Feng, Yuhang Zhou et al.NeurIPS 2025 · 27 citations
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMsHaochen Wang, Yuhao Wang, Tao Zhang, Yikang Zhou et al.ICLR 2026 · 18 citations
- Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D WorldYuzhi Huang, Kairun Wen, Rongxin Gao, Dongxuan Liu et al.CVPR 2026 · 15 citations
Builds on39
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
Related papers
- Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image CaptioningFan Lu, Wei Wu, Kecheng Zheng, Shuailei Ma et al.CVPR 2025
- CALVIN: Improved Contextual Video Captioning via Instruction TuningGowthami Somepalli, Arkabandhu Chowdhury, Jonas Geiping, Ronen Basri et al.NeurIPS 2024 · 4 citations
- CaReBench: A Fine-grained Benchmark for Video Captioning and RetrievalYifan Xu, Xinhao Li, Yichun Yang, Desen Meng et al.ICLR 2026 · 10 citations
- Segment and Caption AnythingXiaoke Huang, Jianfeng Wang, Yansong Tang, Zheng Zhang et al.CVPR 2024
- Fine-grained Audible Video DescriptionXuyang Shen, Dong Li, Jinxing Zhou, Zhen Qin et al.CVPR 2023
