Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
Fan Lu, Wei Wu, Kecheng Zheng, Shuailei Ma, Biao Gong, Jiawei Liu, Wei Zhai, Yang Cao, Yujun Shen, Zheng-Jun Zha
摘要
Generating detailed captions comprehending text-rich visual content in images has received growing attention for Large Vision-Language Models (LVLMs). However, few studies have developed benchmarks specifically tailored for detailed captions to measure their accuracy and comprehensiveness. In this paper, we introduce a detailed caption benchmark, termed as CompreCap, to evaluate the visual context from a directed scene graph view. Concretely, we first manually segment the image into semantically mean-This CVPR paper is the Open Access version, provided by the Computer Vision Foundation.
Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.
ingful regions (i.e., semantic segmentation mask) according to common-object vocabulary, while also distinguishing attributes of objects within all those regions. Then directional relation labels of these objects are annotated to compose a directed scene graph that can well encode rich compositional information of the image. Based on our directed scene graph, we develop a pipeline to assess the generated detailed captions from LVLMs on multiple levels, including the object-level coverage, the accuracy of attribute descriptions, the score of key relationships, etc. Experimental results on the CompreCap dataset confirm that our evaluation method aligns closely with human evaluation scores across LVLMs. We have released the code and the dataset here to support the community.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- CaptionQA: Is Your Caption as Useful as the Image Itself?Shijia Yang, Yunong Liu, Bohan Zhai, Ximeng Sun 等CVPR 2026 · 被引用 15 次
- Panoptic Captioning: An Equivalence Bridge for Image and TextKun-Yu Lin, Hongjun Wang, Weining Ren, Kai HanNeurIPS 2025 · 被引用 7 次
- FINER: MLLMs Hallucinate under Fine-grained Negative QueriesRui Xiao, Sanghwan Kim, Yongqin Xian, Zeynep Akata 等CVPR 2026 · 被引用 3 次
- RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual ReconstructionYuchi Wang, Yishuo Cai, Shuhuai Ren, Sihan Yang 等EMNLP 2025 · 被引用 1 次
- Multigranular Evaluation for Brain Visual DecodingWeihao Xia, Cengiz ÖztireliAAAI 2026 · 被引用 1 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
相关 Paper
- VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality EvaluationShi-Xue Zhang, Hongfa Wang, Duojun Huang, Xin Li 等AAAI 2026 · 被引用 5 次
- ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded SegmentationAli Athar, Xueqing Deng, Liang-Chieh ChenCVPR 2025
- IF-VidCap: Can Video Caption Models Follow Instructions?Shihao Li, Yuanxing Zhang, Jiangtao Wu, Zhide Lei 等ICLR 2026 · 被引用 7 次
- Describe Anything: Detailed Localized Image and Video CaptioningLong Lian, Yifan Ding, Yunhao Ge, Sifei Liu 等ICCV 2025 · 被引用 14 次
- VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image CaptionsKazuki Matsuda, Yuiga Wada, Shinnosuke Hirano, Seitaro Otsuki 等EMNLP 2025
