Concadia: Towards Image-Based Text Generation with a Purpose
Elisa Kreiss, Fei Fang, Noah D. Goodman, Christopher Potts
Abstract
Current deep learning models often achieve excellent results on benchmark image-to-text datasets but fail to generate texts that are useful in practice. We argue that to close this gap, it is vital to distinguish descriptions from captions based on their distinct communicative roles. Descriptions focus on visual features and are meant to replace an image (often to increase accessibility), whereas captions appear alongside an image to supply additional information. To motivate this distinction and help people put it into practice, we introduce the publicly available Wikipedia-based dataset Concadia consisting of 96,918 images with corresponding English-language descriptions, captions, and surrounding context. Using insights from Concadia, models trained on it, and a preregistered human-subjects experiment with human- and model-generated texts, we characterize the commonalities and differences between descriptions and captions. In addition, we show that, for generating both descriptions and captions, it is useful to augment image-to-text models with representations of the textual context in which the image appeared.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- A fine-grained comparison of pragmatic language understanding in humans and language modelsJennifer Hu, Sammy Floyd, Olessia Jouravlev, Evelina Fedorenko et al.ACL 2023 · 45 citations
- Context Matters for Image Descriptions for Accessibility: Challenges for Referenceless Evaluation MetricsElisa Kreiss, Cynthia L. Bennett, Shayan Hooshmand, Eric Zelikman et al.EMNLP 2022 · 12 citations
- Alt-Text with Context: Improving Accessibility for Images on TwitterNikita Srivatsan, Sofía Samaniego, Omar Florez, Taylor Berg-KirkpatrickICLR 2024 · 9 citations
- ContextRef: Evaluating Referenceless Metrics for Image Description GenerationElisa Kreiss, Eric Zelikman, Christopher Potts, Nick HaberICLR 2024 · 6 citations
- Dealing with Semantic Underspecification in Multimodal NLPSandro PezzelleACL 2023 · 5 citations
Builds on5
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- "Person, Shoes, Tree. Is the Person Naked?" What People with Vision Impairments Want in Image DescriptionsAbigale Stangl, Meredith Ringel Morris, Danna GurariCHI 2020 · 136 citations
- Cross-modal Coherence Modeling for Caption GenerationMalihe Alikhani, Piyush Sharma, Shengjie Li, Radu Soricut et al.ACL 2020 · 42 citations
- Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual ConceptsSoravit Changpinyo, Piyush Sharma, Nan Ding, Radu SoricutCVPR 2021
- VinVL: Revisiting Visual Representations in Vision-Language ModelsPengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang et al.CVPR 2021
Related papers
- Widget Captioning: Generating Natural Language Description for Mobile User Interface ElementsYang Li, Gang Li, Luheng He, Jingjie Zheng et al.EMNLP 2020 · 46 citations
- CONICA: A Contrastive Image Captioning Framework with Robust Similarity LearningLin Deng, Yuzhong Zhong, Maoning Wang, Jianwei ZhangACM MM 2023 · 4 citations
- BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation ModelsZiheng Zhang, Xinyue Ma, Arpita Chowdhury, Elizabeth G Campolongo et al.ICLR 2026 · 3 citations
- ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic ArithmeticYoad Tewel, Yoav Shalev, Idan Schwartz, Lior WolfCVPR 2022 · 129 citations
- Diverse Image Captioning with Context-Object Split Latent SpacesShweta Mahajan, Stefan RothNeurIPS 2020 · 47 citations
