Exploring Question Decomposition for Zero-Shot VQA
Zaid Khan, Vijay Kumar B. G, Samuel Schulter, Manmohan Chandraker, Yun Fu
摘要
Visual question answering (VQA) has traditionally been treated as a single-step task where each question receives the same amount of effort, unlike natural human question-answering strategies. We explore a question decomposition strategy for VQA to overcome this limitation. We probe the ability of recently developed large vision-language models to use human-written decompositions and produce their own decompositions of visual questions, finding they are capable of learning both tasks from demonstrations alone. However, we show that naive application of model-written decompositions can hurt performance. We introduce a model-driven selective decomposition approach for second-guessing predictions and correcting errors, and validate its effectiveness on eight VQA tasks across three domains, showing consistent improvements in accuracy, including improvements of > 20% on medical VQA datasets and boosting the zero-shot performance of BLIP-2 above chance on a VQA reformulation of the challenging Winoground task. Project Site: https://zaidkhan.me/decomposition-0shot-vqa/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DWIM: Towards Tool-Aware Visual Reasoning via Discrepancy-Aware Workflow Generation & Instruct-Masking TuningFucai Ke, Vijay Kumar B. G, Xingjian Leng, Zhixi Cai 等ICCV 2025 · 被引用 1 次
- Natural Language Inference Improves Compositionality in Vision-Language ModelsPaola Cascante-Bonilla, Yu Hou, Yang Trista Cao, Hal Daumé III 等ICLR 2025
- Knowledge Exchange with Confidence: Cost-Effective LLM Integration for Reliable and Efficient Visual Question AnsweringMahsa Mozaffari, Hitesh Sapkota, Xumin Liu, Qi YuICLR 2026
- Confidence-guided Refinement Reasoning for Zero-shot Question AnsweringYouwon Jang, Woo Suk Choi, Minjoon Jung, Minsu Lee 等EMNLP 2025
它引用的顶会 Paper29
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
相关 Paper
- DR-VQA: Decompose-then-Reconstruct for Visual Question Answering in BLV AssistanceBocheng Pan, Hailong Shi, Xingyu GaoACM MM 2025
- Looking Beyond the One: Operationalizing and Eliciting Visual Ambiguity in VLLMsYuchong Chen, Bowei Zou, Yuhan Chen, Yifan Fan 等ACL 2026
- Analyzing Modular Approaches for Visual Question DecompositionApoorv Khandelwal, Ellie Pavlick, Chen SunEMNLP 2023 · 被引用 2 次
- Learning by Correction: Efficient Tuning Task for Zero-Shot Generative Vision-Language ReasoningRongjie Li, Yu Wu, Xuming HeCVPR 2024
- Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMsKanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed, Michael S. Ryoo 等CVPR 2024 · 被引用 21 次
