VL-DynaRefine: A Vision-Language Dynamic Refinement Approach for Visual Reasoning
Jing Ma, Haochen Sun, Zeyuan Zang, Fangxiang Feng, Caixia Yuan, Lei Ren, Huixing Jiang, Wei Chen, Xiaojie Wang
摘要
Visual reasoning is a key capability that significantly impacts the performance of multimodal tasks, such as compositional visual question answering and visual grounding. These tasks often require complex, multi-step reasoning processes. In recent years, several training-free methods for Vision-Language Models (VLMs) have emerged, with visual programming methods being proposed to enhance the capability of VLMs in visual reasoning tasks. While these methods have made some progress, they still face two primary challenges due to the lack of verification and refinement mechanisms for each action's output during the reasoning process: error accumulation and feedback delay, as well as insufficient utilization of multimodal contextual information. To address these challenges, we propose VL-DynaRefine, a training-free approach consisting of three modules: a planner, a verifier, and a refiner. The planner generates programmatic actions to solve the problem and executes each action in sequence, which is inspected by a verifier that reassesses the actions via confidence scores and determines whether refinement is necessary based on the evaluation results. In the refiner module, we incorporate a context-aware local refinement mechanism and a global refinement mechanism based on visual and action trajectories to reduce the impact of reasoning errors on the outcome. We evaluate our approach on multiple visual reasoning datasets, and the experimental results show that our method outperforms existing visual programming methods in both reasoning accuracy and efficiency, further validating its effectiveness in visual reasoning tasks.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- De-fine: Decomposing and Refining Visual Programs with Auto-FeedbackMinghe Gao, Juncheng Li, Hao Fei, Liang Pang 等ACM MM 2024 · 被引用 1 次
- Unveiling the Compositional Ability Gap in Vision-Language Reasoning ModelTianle Li, Jihai Zhang, Yongming Rao, Yu ChengNeurIPS 2025 · 被引用 17 次
- See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMsYongchang Zhang, Xianzheng Ma, Tianyi Liu, Guangquan Zhou 等CVPR 2026 · 被引用 2 次
- No Labels, No Problem: Training Visual Reasoners with Multimodal VerifiersDamiano Marsili, Georgia GkioxariICLR 2026 · 被引用 3 次
- DWIM: Towards Tool-Aware Visual Reasoning via Discrepancy-Aware Workflow Generation & Instruct-Masking TuningFucai Ke, Vijay Kumar B. G, Xingjian Leng, Zhixi Cai 等ICCV 2025 · 被引用 1 次
