Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Commonsense Norms
Seungju Han, Junhyeok Kim, Jack Hessel, Liwei Jiang, Jiwan Chung, Yejin Son, Yejin Choi, Youngjae Yu
摘要
Commonsense norms are defeasible by context: reading books is usually great, but not when driving a car. While contexts can be explicitly described in language, in embodied scenarios, contexts are often provided visually. This type of visually grounded reasoning about defeasible commonsense norms is generally easy for humans, but (as we show) poses a challenge for machines, as it necessitates both visual understanding and reasoning about commonsense norms. We construct a new multimodal benchmark for studying visual-grounded commonsense norms: NORMLENS. NORMLENS consists of 10K human judgments accompanied by freeform explanations covering 2K multimodal situations, and serves as a probe to address two questions: (1) to what extent can models align with average human judgment? and (2) how well can models explain their predicted judgments? We find that state-of-theart model judgments and explanations are not well-aligned with human annotation. Additionally, we present a new approach to better align models with humans by distilling social commonsense knowledge from large language models. The data and code are released at https://seungjuhan.me/normlens .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument UnderstandingJiwan Chung, Sungjae Lee, Minseo Kim, Seungju Han 等EMNLP 2024 · 被引用 4 次
- CELLO: Causal Evaluation of Large Vision-Language ModelsMeiqi Chen, Bo Peng, Yan Zhang, Chaochao LuEMNLP 2024 · 被引用 4 次
- Pangea: A Fully Open Multilingual Multimodal LLM for 39 LanguagesXiang Yue, Yueqi Song, Akari Asai, Seungone Kim 等ICLR 2025
- Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable EventsAditya Chinchure, Sahithya Ravi, Raymond T. Ng, Vered Shwartz 等CVPR 2025
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami 等NeurIPS 2020 · 被引用 1,022 次
相关 Paper
- UI-Lens: Assessing General MLLMs' Potential to Automate UI Display Quality AssuranceWei Xiang, Yexinrui Wu, Xinli Chen, Xinran Li 等CVPR 2026
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMsChun-Hsiao Yeh, Chenyu Wang, Shengbang Tong, Ta Ying Cheng 等AAAI 2026 · 被引用 35 次
- LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language ModelsRuilin Yao, Bo Zhang, Jirui Huang, Xinwei Long 等ICLR 2026 · 被引用 8 次
- Expectation Alignment of Language Models for Real-World User ExpectationsMiaomiao Li, Yang Wang, Bin Liang, Shudong Liu 等ICML 2026
- Where Norms and References Collide: Evaluating LLMs on Normative ReasoningMitchell Abrams, Kaveh Eskandari Miandoab, Felix Gervits, Vasanth Sarathy 等AAAI 2026
