Cycle-Consistency Learning for Captioning and Grounding
Ning Wang, Jiajun Deng, Mingbo Jia
Abstract
We present that visual grounding and image captioning, which perform as two mutually inverse processes, can be bridged together for collaborative training by careful designs. By consolidating this idea, we introduce CyCo, a cyclic-consistent learning framework to ameliorate the independent training pipelines of visual grounding and image captioning. The proposed framework (1) allows the semi-weakly supervised training of visual grounding; (2) improves the performance of fully supervised visual grounding; (3) yields a general captioning model that can describe arbitrary image regions. Extensive experiments show that our fully supervised grounding model achieves state-of-the-art performance, and the semi-weakly supervised one also exhibits competitive performance compared to the fully supervised counterparts. Our image captioning model has the capability to freely describe image regions and meanwhile shows impressive performance on prevalent captioning benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b9179af-8ffe-4ebe-b966-8b2ab7dd1359Cited by top-tier papers8
- OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring ModelingLinhui Xiao, Xiaoshan Yang, Fang Peng, Yaowei Wang et al.NeurIPS 2024 · 45 citations
- Question-Answering Dense Video EventsHangyu Qin, Junbin Xiao, Angela YaoSIGIR 2025 · 5 citations
- Cycle-Consistent Learning for Joint Layout-to-Image Generation and Object DetectionXinhao Cai, Qiuxia Lai, Gensheng Pei, Xiangbo Shu et al.ICCV 2025 · 1 citation
- Temporal Calibrating and Distilling for Scene-Text Aware Text-Video RetrievalZhiqian Zhao, Liang Li, Lei Shen, Xichun Sheng et al.AAAI 2026 · 1 citation
- From Pixels to Logic: A Perception-Reasoning Decomposition Framework for Open-World Referring Expression ComprehensionLihong Huang, Sheng-hua Zhong, Zhi Zhang, Yan LiuAAAI 2026
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
Related papers
- 3DJCG: A Unified Framework for Joint Dense Captioning and Visual Grounding on 3D Point CloudsDaigang Cai, Lichen Zhao, Jing Zhang, Lu Sheng et al.CVPR 2022 · 80 citations
- UniT3D: A Unified Transformer for 3D Dense Captioning and Visual GroundingDave Zhenyu Chen, Ronghang Hu, Xinlei Chen, Matthias Nießner et al.ICCV 2023 · 82 citations
- AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual GroundingYidan Wang, Chenyi Zhuang, Wutao Liu, Pan Gao et al.ACM MM 2025 · 2 citations
- Improving Image Captioning with Better Use of CaptionZhan Shi, Xu Zhou, Xipeng Qiu, Xiaodan ZhuACL 2020 · 84 citations
- Distributed Attention for Grounded Image CaptioningNenglun Chen, Xingjia Pan, Runnan Chen, Lei Yang et al.ACM MM 2021 · 19 citations
