MetaCLUE: Towards Comprehensive Visual Metaphors Research
Arjun R. Akula, Brendan Driscoll, Pradyumna Narayana, Soravit Changpinyo, Zhiwei Jia, Suyash Damle, Garima Pruthi, Sugato Basu, Leonidas J. Guibas, William T. Freeman, Yuanzhen Li, Varun Jampani
摘要
Figure 1 . With MetaCLUE, we introduce several interesting tasks related to visual metaphors. We collect metaphor annotations (objects, abstract concepts, relationships and object boxes) for evaluating existing models on these tasks. Specifically we perform a comprehensive evaluation of vision and language models on four different tasks (Classification, Localization, Understanding, and gEneration). Comprehensive experiments in this work show that state-of-the-art techniques mostly focus on literal interpretation and perform poorly in understanding and generation of metaphor images.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Text-to-Image Generation for Abstract ConceptsJiayi Liao, Xu Chen, Qiang Fu, Lun Du 等AAAI 2024 · 被引用 25 次
- Creative Blends of Visual ConceptsZhida Sun, Zhenyao Zhang, Yue Zhang, Min Lu 等CHI 2025 · 被引用 11 次
- Impressions: Visual Semiotics and Aesthetic Impact UnderstandingJulia Kruk, Caleb Ziems, Diyi YangEMNLP 2023 · 被引用 5 次
- Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument UnderstandingJiwan Chung, Sungjae Lee, Minseo Kim, Seungju Han 等EMNLP 2024 · 被引用 4 次
- Beyond Input-Output: Rethinking Creativity through Design-by-Analogy in Human-AI CollaborationXuechen Li, Shuai Zhang, Nan Cao, Qing ChenCHI 2026 · 被引用 2 次
它引用的顶会 Paper8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
相关 Paper
- VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic PhenomenaLetitia Parcalabescu, Michele Cafagna, Lilitta Muradjan, Anette Frank 等ACL 2022 · 被引用 147 次
- MetaGPT: A Large Vision-Language Model for Meme Metaphor UnderstandingBo Xu, Chenyuan Wang, Xinyu Chen, Hongfei Lin 等AAAI 2026
- MultiMET: A Multimodal Dataset for Metaphor UnderstandingDongyu Zhang, Minghao Zhang, Heting Zhang, Liang Yang 等ACL 2021
- MemeCap: A Dataset for Captioning and Interpreting MemesEunjeong Hwang, Vered ShwartzEMNLP 2023 · 被引用 14 次
- GOME: Grounding-based Metaphor Binding With Conceptual Elaboration For Figurative Language IllustrationLinhao Zhang, Jintao Liu, Li Jin, Hao Wang 等EMNLP 2024 · 被引用 1 次
