LVLM-Aided Alignment of Task-Specific Vision Models
Alexander Koebler, Lukas Kuhn, Ingo Thon, Florian Buettner
Abstract
In high-stakes domains, small task-specific vision models are crucial due to their low computational requirements and the availability of numerous methods to explain their results. However, these explanations often reveal that the models do not align well with human domain knowledge, relying instead on spurious correlations. This might result in brittle behavior once deployed in the real-world. To address this issue, we introduce a novel and efficient method for aligning small task-specific vision models with human domain knowledge by leveraging the generalization capabilities of a Large Vision Language Model (LVLM). Our LVLM-Aided Visual Alignment (LVLM-VA) method provides a bidirectional interface that translates model behavior into natural language and maps human class-level specifications to image-level critiques, enabling effective interaction between domain experts and the model. Our method demonstrates substantial improvement in aligning model behavior with human specifications, as validated on both synthetic and real-world datasets. We show that it effectively reduces the model’s dependence on spurious features and on group-specific biases, without requiring fine-grained feedback.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan et al.ICML 2021 · 683 citations
- AnomalyGPT: Detecting Industrial Anomalies Using Large Vision-Language ModelsZhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen et al.AAAI 2024 · 312 citations
Related papers
- Visual Evidence Prompting Mitigates Hallucinations in Large Vision-Language ModelsWei Li, Zhen Huang, Houqiang Li, Le Lu et al.ACL 2025
- DRESS : Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language FeedbackYangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji et al.CVPR 2024
- Large-Small Model Synergy with Multimodal Fine-Grained Heuristics for Knowledge-Based Visual Question AnsweringZhongfan Sun, Kan Guo, Yongli Hu, Daxin Tian et al.ACM MM 2025
- ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative DecodingJialiang Kang, Han Shu, Wenshuo Li, Yingjie Zhai et al.NeurIPS 2025 · 24 citations
- Testing the Limits of Fine-Tuning for Improving Visual Cognition in Vision Language ModelsLuca M. Schulze Buschoff, Konstantinos Voudouris, Elif Akata, Matthias Bethge et al.ICML 2025
