Crossing the Gap: Domain Generalization for Image Captioning
Yuchen Ren, Zhendong Mao, Shancheng Fang, Yan Lu, Tong He, Hao Du, Yongdong Zhang, Wanli Ouyang
Abstract
Existing image captioning methods are under the assumption that the training and testing data are from the same domain or that the data from the target domain (i.e., the domain that testing data lie in) are accessible. However, this assumption is invalid in real-world applications where the data from the target domain is inaccessible. In this paper, we introduce a new setting called Domain Generalization for Image Captioning (DGIC), where the data from the target domain is unseen in the learning process. We first construct a benchmark dataset for DGIC, which helps us to investigate models' domain generalization (DG) ability on unseen domains. With the support of the new benchmark, we further propose a new framework called language-guided semantic metric learning (LSML) for the DGIC setting. Experiments on multiple datasets demonstrate the challenge of the task and the effectiveness of our newly proposed benchmark and LSML framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00c0e985-b26b-442d-8a8b-71dcb5461bc1Cited by top-tier papers2
- Region-Aware Exposure Consistency Network for Mixed Exposure CorrectionJin Liu, Huiyuan Fu, Chuanming Wang, Huadong MaAAAI 2024 · 23 citations
- LG-VQ: Language-Guided Codebook LearningGuotao Liang, Baoquan Zhang, Yaowei Wang, Yunming Ye et al.NeurIPS 2024 · 14 citations
Builds on17
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 999 citations
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- Domain Randomization and Pyramid Consistency: Simulation-to-Real Generalization Without Accessing Target Domain DataXiangyu Yue, Yang Zhang, Sicheng Zhao, Alberto L. Sangiovanni-Vincentelli et al.ICCV 2019 · 462 citations
Related papers
- Open-Vocabulary Domain Generalization in Urban-Scene SegmentationDong Zhao, Qi Zang, Nan Pu, Wenjing Li et al.CVPR 2026 · 3 citations
- Rethinking the Reference-based Distinctive Image CaptioningYangjun Mao, Long Chen, Zhihong Jiang, Dong Zhang et al.ACM MM 2022 · 22 citations
- Domain Generalization in CLIP via Learning with Diverse Text PromptsChangsong Wen, Zelin Peng, Yu Huang, Xiaokang Yang et al.CVPR 2025
- UNISON: Unpaired Cross-Lingual Image CaptioningJiahui Gao, Yi Zhou, Philip L. H. Yu, Shafiq R. Joty et al.AAAI 2022 · 18 citations
- Generalized Semantic Segmentation by Self-Supervised Source Domain Projection and Multi-Level Contrastive LearningLiwei Yang, Xiang Gu, Jian SunAAAI 2023 · 25 citations
