Describe Anything: Detailed Localized Image and Video Captioning
Long Lian, Yifan Ding, Yunhao Ge, Sifei Liu, Hanzi Mao, Boyi Li, Marco Pavone, Ming-Yu Liu, Trevor Darrell, Adam Yala, Yin Cui
摘要
Generating detailed and accurate descriptions for specific regions in images and videos remains a fundamental challenge for vision-language models. We introduce the Describe Anything Model (DAM), a model designed for detailed localized captioning (DLC). DAM preserves both local details and global context through two key innovations: a focal prompt, which ensures high-resolution encoding of targeted regions, and a localized vision backbone, which integrates precise localization with its broader context. To tackle the scarcity of high-quality DLC data, we propose a Semi-supervised learning (SSL)-based Data Pipeline (DLC-SDP). DLC-SDP starts with existing segmentation datasets and expands to unlabeled web images using SSL. We introduce DLC-Bench, a benchmark designed to evaluate DLC without relying on reference captions. DAM sets new state-of-the-art on 7 benchmarks spanning keyword-level, phrase-level, and detailed multi-sentence localized image and video captioning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and VideosWeifeng Lin, Xinyu Wei, Ruichuan An, Tianhe Ren 等NeurIPS 2025 · 被引用 47 次
- Describe Anything Anywhere At Any MomentNicolas Gorlo, Lukas Schmid, Luca CarloneCVPR 2026 · 被引用 27 次
- ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMsXiyao Wang, Zhengyuan Yang, Chao Feng, Yuhang Zhou 等NeurIPS 2025 · 被引用 27 次
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMsHaochen Wang, Yuhao Wang, Tao Zhang, Yikang Zhou 等ICLR 2026 · 被引用 18 次
- Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D WorldYuzhi Huang, Kairun Wen, Rongxin Gao, Dongxuan Liu 等CVPR 2026 · 被引用 15 次
它引用的顶会 Paper39
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
相关 Paper
- Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image CaptioningFan Lu, Wei Wu, Kecheng Zheng, Shuailei Ma 等CVPR 2025
- CALVIN: Improved Contextual Video Captioning via Instruction TuningGowthami Somepalli, Arkabandhu Chowdhury, Jonas Geiping, Ronen Basri 等NeurIPS 2024 · 被引用 4 次
- CaReBench: A Fine-grained Benchmark for Video Captioning and RetrievalYifan Xu, Xinhao Li, Yichun Yang, Desen Meng 等ICLR 2026 · 被引用 10 次
- Segment and Caption AnythingXiaoke Huang, Jianfeng Wang, Yansong Tang, Zheng Zhang 等CVPR 2024
- Fine-grained Audible Video DescriptionXuyang Shen, Dong Li, Jinxing Zhou, Zhen Qin 等CVPR 2023
