CAVE : Detecting and Explaining Commonsense Anomalies in Visual Environments
Rishika Bhagwatkar, Syrielle Montariol, Angelika Romanou, Beatriz Borges, Irina Rish, Antoine Bosselut
摘要
Humans can naturally identify, reason about, and explain anomalies in their environment. In computer vision, this long-standing challenge remains limited to industrial defects or unrealistic, synthetically generated anomalies, failing to capture the richness and unpredictability of real-world anomalies. In this work, we introduce CAVE, the first benchmark of real-world visual anomalies. CAVE supports three open-ended tasks: anomaly description, explanation, and justification; with fine-grained annotations for visual grounding and categorizing anomalies based on their visual manifestations, their complexity, severity, and commonness. These annotations draw inspiration from cognitive science research on how humans identify and resolve anomalies, providing a comprehensive framework for evaluating Vision-Language Models (VLMs) in detecting and understanding anomalies. We show that state-of-the-art VLMs struggle with visual anomaly perception and commonsense reasoning, even with advanced prompting strategies. By offering a realistic and cognitively grounded benchmark, CAVE serves as a valuable resource for advancing research in anomaly detection and commonsense reasoning in VLMs. We release the code and benchmark on our project webpage. 1 * Equal contribution. 1 https://smontariol.github.io/ cave-visual-anomalies/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 被引用 401 次
- Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language ModelsSiddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang 等ICML 2024 · 被引用 306 次
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh 等CVPR 2022 · 被引用 179 次
相关 Paper
- CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-WorldYating Yu, Congqi Cao, Zhaoying Wang, Weihua Meng 等AAAI 2026 · 被引用 1 次
- VALU: A Benchmark for Video Anomaly Temporal Localization and Understanding at Multiple Semantic LevelsYixiao He, Menghao Zhang, Haifeng Sun, Jing Wang 等ACL 2026
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMsBrigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh, Wamiq Reyaz Para 等CVPR 2026 · 被引用 3 次
- FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly UnderstandingJoão Alexandre Cardeira Pereira, Vasco Lopes, João Neves, David SemedoAAAI 2026
- Detecting Violations of Physical Common Sense in Images: A Challenge Dataset and Effective ModelWeibin Wu, Zitong Wang, Zhengjie Luo, Wenqing Chen 等ACM MM 2025
