Image Captioning with Multimodal Guidance and Search Space Optimization
Yimou Guo, Yaochen Li, Jingze Liu, Jiahui Feng, Haoyi Lou, Zhimin Chen, Yuan Gao, Yuanqi Su
摘要
Image captioning bridges the gap between visual perception and natural language understanding by transforming image content into descriptive text. While existing methods have made significant progress in visual feature extraction, encoding, and cross-modal semantic alignment, challenges remain in terms of fine-grained feature representation, cross-modal alignment efficiency, and suboptimal search strategies. To address these issues, a multimodal-guided and search space-optimized image captioning model is proposed. In the visual encoding stage, we construct a hierarchical network that integrates regional and grid features through a geometry-constrained multi-layer feature aggregation mechanism, which enhances the model's capability to jointly capture global semantics and local details. In the decoding stage, we introduce a dynamic grouped beam width adjustment strategy to improve semantic path exploration. Additionally, a diversity-driven scoring function is designed to enforce intra-group diversity rewards and inter-group similarity penalties, encouraging the generation of more diverse captions. Finally, we incorporate a two-level pruning algorithm based on syntactic and spatial logic constraints to refine the search space from both hard and soft constraint perspectives, improving both the accuracy and diversity of generated captions. A 3% improvement in CIDEr is achieved by the proposed method over state-of-the-art (SOTA) models, as demonstrated by experiments on the COCO and Flickr30k datasets.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Image Captioning with Context-Aware Auxiliary GuidanceZeliang Song, Xiaofei Zhou, Zhendong Mao, Jianlong TanAAAI 2021 · 被引用 36 次
- HAAV: Hierarchical Aggregation of Augmented Views for Image CaptioningChia-Wen Kuo, Zsolt KiraCVPR 2023
- Improving Image Captioning through Visual and Semantic Mutual PromotionJing Zhang, Yingshuai Xie, Xiaoqiang LiuACM MM 2023 · 被引用 4 次
- Generating Diverse and Descriptive Image Captions Using Visual ParaphrasesLixin Liu, Jiajun Tang, Xiaojun Wan, Zongming GuoICCV 2019 · 被引用 48 次
- Comprehending and Ordering Semantics for Image CaptioningYehao Li, Yingwei Pan, Ting Yao, Tao MeiCVPR 2022 · 被引用 124 次
