Explain Me the Painting: Multi-Topic Knowledgeable Art Description Generation
Zechen Bai, Yuta Nakashima, Noa Garcia
Abstract
Have you ever looked at a painting and wondered what is the story behind it? This work presents a framework to bring art closer to people by generating comprehensive descriptions of fine-art paintings. Generating informative descriptions for artworks, however, is extremely challenging, as it requires to 1) describe multiple aspects of the image such as its style, content, or composition, and 2) provide background and contextual knowledge about the artist, their influences, or the historical period. To address these challenges, we introduce a multi-topic and knowledgeable art description framework, which modules the generated sentences according to three artistic topics and, additionally, enhances each description with external knowledge. The framework is validated through an exhaustive analysis, both quantitative and qualitative, as well as a comparative human evaluation, demonstrating outstanding results in terms of both topic diversity and information veracity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40bdbaf4-8211-4406-b7aa-276faf5952abCited by top-tier papers10
- One Token to Seg Them All: Language Instructed Reasoning Segmentation in VideosZechen Bai, Tong He, Haiyang Mei, Pichao Wang et al.NeurIPS 2024 · 147 citations
- ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art UnderstandingShuai Wang, Ivona Najdenkoska, Hongyi Zhu, Stevan Rudinac et al.ACM MM 2025 · 6 citations
- Understanding Museum Exhibits using Vision-Language ReasoningAda-Astrid Balauca, Sanjana Garai, Stefan Balauca, Rasesh Udayakumar Shetty et al.ICCV 2025 · 3 citations
- VL-KGE: Vision-Language Models Meet Knowledge Graph EmbeddingsAthanasios Efthymiou, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg et al.WWW 2026 · 2 citations
- ImageSet2Text: Describing Sets of Images Through TextPiera Riccio, Francesco Galati, Kajetan Schweighofer, Noa Garcia et al.AAAI 2026 · 1 citation
Builds on1
Related papers
- Imagine, Reason and Write: Visual Storytelling with Graph Knowledge and Relational ReasoningChunpu Xu, Min Yang, Chengming Li, Ying Shen et al.AAAI 2021 · 39 citations
- GalleryGPT: Analyzing Paintings with Large Multimodal ModelsYi Bin, Wenhao Shi, Yujuan Ding, Zhiqiang Hu et al.ACM MM 2024 · 35 citations
- R-GAN: Exploring Human-like Way for Reasonable Text-to-Image Synthesis via Generative Adversarial NetworksYanyuan Qiao, Qi Chen, Chaorui Deng, Ning Ding et al.ACM MM 2021 · 18 citations
- MAGIC: Multimodal relAtional Graph adversarIal inferenCe for Diverse and Unpaired Text-Based Image CaptioningWenqiao Zhang, Haochen Shi, Jiannan Guo, Shengyu Zhang et al.AAAI 2022 · 52 citations
- Graph Neural Networks for Knowledge Enhanced Visual Representation of PaintingsAthanasios Efthymiou, Stevan Rudinac, Monika Kackovic, Marcel Worring et al.ACM MM 2021 · 25 citations
