Control Image Captioning Spatially and Temporally
Kun Yan, Lei Ji, Huaishao Luo, Ming Zhou, Nan Duan, Shuai Ma
Abstract
Generating image captions with user intention is an emerging need. The recently published Localized Narratives dataset takes mouse traces as another input to the image captioning task, which is an intuitive and efficient way for a user to control what to describe in the image. However, how to effectively employ traces to improve generation quality and controllability is still under exploration. This paper aims to solve this problem by proposing a novel model called LoopCAG, which connects Contrastive constraints and Attention Guidance in a Loop manner, engaged explicit spatial and temporal constraints to the generating process. Precisely, each generated sentence is temporally aligned to the corresponding trace sequence through a contrastive learning strategy. Besides, each generated text token is supervised to attend to the correct visual objects under heuristic spatial attention guidance. Comprehensive experimental results demonstrate that our LoopCAG model learns better correspondence among the three modalities(vision, language, and traces) and achieves SOTA performance on trace controlled image captioning task. Moreover, the controllability and explainability of LoopCAG are validated by analyzing spatial and temporal sensitivity during the generation process.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 094e3c1e-d99a-4a00-8b10-584621a09b3bCited by top-tier papers5
- Voila-A: Aligning Vision-Language Models with User's Gaze AttentionKun Yan, Zeyu Wang, Lei Ji, Yuntao Wang et al.NeurIPS 2024 · 43 citations
- G-VOILA: Gaze-Facilitated Information Querying in Daily ScenariosZeyu Wang, Yuanchun Shi, Yuntao Wang, Yuchen Yao et al.UbiComp 2024 · 23 citations
- CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal GroundingZhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao et al.ACL 2023 · 17 citations
- SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioningXu Zhang, Jin Yuan, Hanwang Zhang, Guojin Zhong et al.AAAI 2025 · 2 citations
- Personalized Image Descriptions from Attention SequencesRuoyu Xue, Hieu Le, Jingyi Xu, Sounak Mondal et al.CVPR 2026 · 2 citations
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- A Mutual Information Maximization Perspective of Language Representation LearningLingpeng Kong, Cyprien de Masson d'Autume, Lei Yu, Wang Ling et al.ICLR 2020 · 179 citations
Related papers
- Connecting What To Say With Where To Look by Modeling Human Attention TracesZihang Meng, Licheng Yu, Ning Zhang, Tamara L. Berg et al.CVPR 2021
- LocTex: Learning Data-Efficient Visual Representations from Localized Textual SupervisionZhijian Liu, Simon Stent, Jie Li, John Gideon et al.ICCV 2021 · 10 citations
- TGT: Text-Grounded Trajectories for Locally Controlled Video GenerationGuofeng Zhang, Angtian Wang, Jacob Zhiyuan Fang, Liming Jiang et al.CVPR 2026 · 5 citations
- Pixel Aligned Language ModelsJiarui Xu, Xingyi Zhou, Shen Yan, Xiuye Gu et al.CVPR 2024 · 6 citations
- LocCa: Visual Pretraining with Location-aware CaptionersBo Wan, Michael Tschannen, Yongqin Xian, Filip Pavetic et al.NeurIPS 2024 · 38 citations
