Impressions: Visual Semiotics and Aesthetic Impact Understanding
Julia Kruk, Caleb Ziems, Diyi Yang
Abstract
Is aesthetic impact different from beauty? Is visual salience a reflection of its capacity for effective communication? We present Impressions, a novel dataset through which to investigate the semiotics of images, and how specific visual features and design choices can elicit specific emotions, thoughts and beliefs. We posit that the impactfulness of an image extends beyond formal definitions of aesthetics, to its success as a communicative act, where style contributes as much to meaning formation as the subject matter. We also acknowledge that existing Image Captioning datasets are not designed to empower state-of-the-art architectures to model potential human impressions or interpretations of images. To fill this need, we design an annotation task heavily inspired by image analysis techniques in the Visual Arts to collect 1,440 image-caption pairs and 4,320 unique annotations exploring impact, pragmatic image description, impressions and aesthetic design choices. We show that existing multimodal image captioning and conditional generation models struggle to simulate plausible human responses to images. However, this dataset significantly improves their ability to model impressions and aesthetic evaluations of images through fine-tuning and few-shot adaptation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Similar Scenes Arouse Similar Emotions: Parallel Data Augmentation for Stylized Image CaptioningGuodun Li, Yuchen Zhai, Zehao Lin, Yin ZhangACM MM 2021 · 23 citations
Related papers
- ArtEmis: Affective Language for Visual ArtPanos Achlioptas, Maks Ovsjanikov, Kilichbek Haydarov, Mohamed Elhoseiny et al.CVPR 2021
- Can Machines Understand Composition? Dataset and Benchmark for Photographic Image Composition Embedding and UnderstandingZhaoran Zhao, Peng Lu, Anran Zhang, Peipei Li et al.CVPR 2025
- Aesthetically Relevant Image CaptioningZhipeng Zhong, Fei Zhou, Guoping QiuAAAI 2023 · 16 citations
- AI-generated Image Quality Assessment in Visual CommunicationYu Tian, Yixuan Li, Baoliang Chen, Hanwei Zhu et al.AAAI 2025 · 12 citations
- Concadia: Towards Image-Based Text Generation with a PurposeElisa Kreiss, Fei Fang, Noah D. Goodman, Christopher PottsEMNLP 2022 · 13 citations
