DIVE: Towards Descriptive and Diverse Visual Commonsense Generation
Jun-Hyung Park, Hyuntae Park, Youjin Kang, Eojin Jeon, SangKeun Lee
摘要
Towards human-level visual understanding, visual commonsense generation has been introduced to generate commonsense inferences beyond images. However, current research on visual commonsense generation has overlooked an important human cognitive ability: generating descriptive and diverse inferences. In this work, we propose a novel visual commonsense generation framework, called DIVE, which aims to improve the descriptiveness and diversity of generated inferences. DIVE involves two methods, generic inference filtering and contrastive retrieval learning, which address the limitations of existing visual commonsense resources and training objectives. Experimental results verify that DIVE outperforms state-of-the-art models for visual commonsense generation in terms of both descriptiveness and diversity, while showing a superior quality in generating unique and novel inferences. Notably, DIVE achieves human-level descriptiveness and diversity on Visual Commonsense Graphs. Furthermore, human evaluations confirm that DIVE aligns closely with human judgments on descriptiveness and diversity
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- UNIFIED-IO: A Unified Model for Vision, Language, and Multi-modal TasksJiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi 等ICLR 2023 · 被引用 110 次
- Diverse and Informative Dialogue Generation with Context-Specific Commonsense Knowledge AwarenessSixing Wu, Ying Li, Dawei Zhang, Yang Zhou 等ACL 2020 · 被引用 104 次
相关 Paper
- Improving Commonsense in Vision-Language Models via Knowledge Graph RiddlesShuquan Ye, Yujia Xie, Dongdong Chen, Yichong Xu 等CVPR 2023
- Imagine, Reason and Write: Visual Storytelling with Graph Knowledge and Relational ReasoningChunpu Xu, Min Yang, Chengming Li, Ying Shen 等AAAI 2021 · 被引用 39 次
- DiscoSense: Commonsense Reasoning with Discourse ConnectivesPrajjwal Bhargava, Vincent NgEMNLP 2022 · 被引用 1 次
- FOCUS: Evaluating Pre-trained Vision-Language Models on Underspecification ReasoningKankan Zhou, Eason Lai, Kyriakos Mouratidis, Jing JiangACL 2025
- KGR4: Retrieval, Retrospect, Refine and Rethink for Commonsense GenerationXin Liu, Dayiheng Liu, Baosong Yang, Haibo Zhang 等AAAI 2022 · 被引用 9 次
