ImageSet2Text: Describing Sets of Images Through Text
Piera Riccio, Francesco Galati, Kajetan Schweighofer, Noa Garcia, Nuria Oliver
Abstract
A collection of interior design images showcasing various interior spaces. The designs reflect a blend of modern and eclectic styles, capturing elegance through a sophisticated color palette. The images feature components such as light fixtures, a variety of textures, styled interior decorations, and a collection of furniture, all enhanced by a combination of natural light and soft artificial lighting. ImageSet2Text A collection of paintings depicting landscapes with various types of vegetation, specifically featuring trees. The landscapes are characterized by tranquil topography and include a background of mountains and/or rocky formations. These mountainous regions are portrayed during twilight and set within a scenic mountainous environment, showcasing various shades of blue, gray, and earthy tones. The paintings exhibit an expressive style that is romantic and impressionistic, utilizing a muted and cool color scheme. ImageSet2Text Figure 1: ImageSet2Text generates detailed and nuanced descriptions from large sets of images. We report exemplary generated descriptions for two groups of images (Sharma et al. 2018; Tan et al. 2019 ).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f5388770-c78b-4d25-afd3-7dc444f6e96eBuilds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
Related papers
- PreciseCam: Precise Camera Control for Text-to-Image GenerationEdurne Bernal-Berdun, Ana Serrano, Belén Masiá, Matheus Gadelha et al.CVPR 2025
- ArtAdapter: Text-to-Image Style Transfer using Multi-Level Style Encoder and Explicit AdaptationDar-Yen Chen, Hamish Tennent, Ching-Wen HsuCVPR 2024
- TextCraftor: Your Text Encoder can be Image Quality ControllerYanyu Li, Xian Liu, Anil Kag, Ju Hu et al.CVPR 2024
- Exploring Sparse MoE in GANs for Text-conditioned Image SynthesisJiapeng Zhu, Ceyuan Yang, Kecheng Zheng, Yinghao Xu et al.CVPR 2025
- StyleDrop: Text-to-Image Synthesis of Any StyleKihyuk Sohn, Lu Jiang, Jarred Barber, Kimin Lee et al.NeurIPS 2023 · 71 citations
