Paraphrasing Is All You Need for Novel Object Captioning
Cheng-Fu Yang, Yao-Hung Hubert Tsai, Wan-Cyuan Fan, Russ Salakhutdinov, Louis-Philippe Morency, Frank Wang
Abstract
Novel object captioning (NOC) aims to describe images containing objects without observing their ground truth captions during training. Due to the absence of caption annotation, captioning models cannot be directly optimized via sequenceto-sequence training or CIDEr optimization. As a result, we present Paraphrasingto-Captioning (P2C), a two-stage learning framework for NOC, which would heuristically optimize the output captions via paraphrasing. With P2C, the captioning model first learns paraphrasing from a language model pre-trained on text-only corpus, allowing expansion of the word bank for improving linguistic fluency. To further enforce the output caption sufficiently describing the visual content of the input image, we perform self-paraphrasing for the captioning model with fidelity and adequacy objectives introduced. Since no ground truth captions are available for novel object images during training, our P2C leverages cross-modality (image-text) association modules to ensure the above caption characteristics can be properly preserved. In the experiments, we not only show that our P2C achieves state-of-the-art performances on nocaps and COCO Caption datasets, we also verify the effectiveness and flexibility of our learning framework by replacing language and cross-modality association models for NOC. Implementation details and code are available in the supplementary materials. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Zero-Shot Image Captioning with Multi-type Entity RepresentationsDelong Zeng, Ying Shen, Man Lin, Zihao Yi et al.AAAI 2025 · 3 citations
- RAPPER: Reinforced Rationale-Prompted Paradigm for Natural Language Explanation in Visual Question AnsweringKai-Po Chang, Chi-Pin Huang, Wei-Yuan Cheng, Fu-En Yang et al.ICLR 2024 · 3 citations
- Verbalized Representation Learning for Interpretable Few-Shot GeneralizationCheng-Fu Yang, Da Yin, Wenbo Hu, Heng Ji et al.ICCV 2025 · 1 citation
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- nocaps: novel object captioning at scaleHarsh Agrawal, Peter Anderson, Karan Desai, Yufei Wang et al.ICCV 2019 · 631 citations
Related papers
- VIVO: Visual Vocabulary Pre-Training for Novel Object CaptioningXiaowei Hu, Xi Yin, Kevin Lin, Lei Zhang et al.AAAI 2021 · 63 citations
- Generating Diverse and Descriptive Image Captions Using Visual ParaphrasesLixin Liu, Jiajun Tang, Xiaojun Wan, Zongming GuoICCV 2019 · 48 citations
- NOC-REK: Novel Object Captioning with Retrieved Vocabulary from External KnowledgeDuc Minh Vo, Hong Chen, Akihiro Sugimoto, Hideki NakayamaCVPR 2022 · 21 citations
- RCA-NOC: Relative Contrastive Alignment for Novel Object CaptioningJiashuo Fan, Yaoyuan Liang, Leyao Liu, Shao-Lun Huang et al.ICCV 2023 · 7 citations
- Noise-Aware Image Captioning with Progressively Exploring Mismatched WordsZhongtian Fu, Kefei Song, Luping Zhou, Yang YangAAAI 2024 · 36 citations
