Top-Down Semantic Refinement for Image Captioning
Jusheng Zhang, Kaitong Cai, Jing Yang, Jian Wang, Chengpei Tang, Keze Wang
摘要
Large Vision-Language Models (VLMs) face an inherent contradiction in image captioning: their powerful single-step generation capabilities often lead to a myopic decision-making process. This makes it difficult to maintain global narrative coherence while capturing rich details, a limitation that is particularly pronounced in tasks that require multi-step and complex scene description. To overcome this fundamental challenge, we redefine image captioning as a goal-oriented hierarchical refinement planning problem, and further propose a novel framework, named Top-Down Semantic Refinement (TDSR), which models the generation process as a Markov Decision Process (MDP). However, planning within the vast state space of a VLM presents a significant computational hurdle. Our core contribution, therefore, is the design of a highly efficient Monte Carlo Tree Search (MCTS) algorithm tailored for VLMs. By incorporating a visual-guided parallel expansion and a lightweight value network, our TDSR reduces the call frequency to the expensive VLM by an order of magnitude without sacrificing planning quality. Furthermore, an adaptive early stopping mechanism dynamically matches computational overhead to the image's complexity. Extensive experiments on multiple benchmarks, including DetailCaps, COMPOSITIONCAP, and POPE, demonstrate that our TDSR, as a plug-and-play module, can significantly enhance the performance of existing VLMs (e.g., LLaVA-1.5, Qwen2.5-VL) by achieving state-of-the-art or highly competitive results in fine-grained description, compositional generalization, and hallucination suppression.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ContextFlow: Training-Free Video Object Editing via Adaptive Context EnrichmentYiyang Chen, Xuanhua He, Xiujun Ma, Jack MaAAAI 2026 · 被引用 17 次
- Cost-Effective Communication: An Auction-based Method for Language Agent InteractionYijia Fan, Jusheng Zhang, Kaitong Cai, Jing Yang 等AAAI 2026 · 被引用 15 次
- MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion TransformerPenghui Liu, Jiangshan Wang, Yutong Shen, Shanhui Mo 等AAAI 2026 · 被引用 2 次
- LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination CorrectionJusheng Zhang, Ningyuan Liu, Yijia Fan, Zihao Huang 等AAAI 2026
它引用的顶会 Paper25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM CaptioningAnkan Deria, Adinath Madhavrao Dukre, Feilong Tang, Sara Atito 等NeurIPS 2025 · 被引用 2 次
- Multi-Modal Hallucination Control by Visual Information GroundingAlessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary 等CVPR 2024
- Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree SearchLinhao Yu, Xingguang Ji, Yahui Liu, Fanheng Kong 等ACL 2025 · 被引用 2 次
- Scaling Inference-Time Search with Vision Value Model for Improved Visual ComprehensionXiyao Wang, Zhengyuan Yang, Linjie Li, Hongjin Lu 等ICCV 2025 · 被引用 1 次
- From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data CalibrationMingyang Song, Xiaoye Qu, Jiawei Zhou, Yu ChengCVPR 2025
