DR-VQA: Decompose-then-Reconstruct for Visual Question Answering in BLV Assistance
Bocheng Pan, Hailong Shi, Xingyu Gao
Abstract
Visual impairment affects over 200 million individuals globally, creating significant challenges in daily visual tasks. While vision-language models offer transformative assistive potential, existing systems based on Multimodal Large Language Models (MLLMs) face a serious cross-contamination problem when processing real-world images captured by blind and low-vision (BLV) users: when jointly processing imperfect images and specific questions, current models are often misled by question assumptions rather than adhering to visual facts, generating hallucinations about objects not present in the image. We introduce DR-VQA (Decompose-then-Reconstruct Visual Question Answering), a novel framework that balances user intent with visual facts. Our approach prevents cross-contamination through structured reasoning. Our approach deliberately separates image processing from question analysis, ensuring model-generated descriptions are strictly based on image facts without being influenced by questions. Subsequently, through a structured decomposition mechanism, the system generates targeted sub-questions relevant to user intent, gradually aligning visual descriptions with user needs while minimizing question bias. During final synthesis, a memory-reset LLM reconstructs the reasoning chain with detailed information to generate responses that either provide evidence-supported conclusions or transparently acknowledge information limitations. Experimental evaluations demonstrate our framework's effectiveness in reducing hallucination risks while improving answer accuracy. By systematically balancing user intent with factual visual evidence, this work advances BLV-assistive technologies from probabilistic outputs to reliable visual assistance services.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 37c808dc-60bb-4fc9-8843-6c08e72bf6acRelated papers
- "This is My Fault", Really? Understanding Blind and Low-Vision People's Perception of Hallucination in Large Vision Language ModelsYilin Tang, Yuyang Fang, Tianle Wang, Lingyun Sun et al.UIST 2025 · 3 citations
- How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision PeopleRicardo E. Gonzalez Penuela, Crescentia Jung, Sharon Y. Lin, Ruiying Hu et al.CHI 2026 · 1 citation
- Right this way: Can VLMs Guide Us to See More to Answer Questions?Li Liu, Diji Yang, Sijia Zhong, Kalyana Suma Sree Tholeti et al.NeurIPS 2024 · 20 citations
- GenAssist: Making Image Generation AccessibleMina Huh, Yi-Hao Peng, Amy PavelUIST 2023 · 58 citations
- ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and WisdomJingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu et al.EMNLP 2025
