PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models
Jenny Schmalfuss, Nadine Chang, Vibashan VS, Maying Shen, Andrés Bruhn, José M. Álvarez
Abstract
Figure 1. PARC prompt sensitivity analysis framework overview. Given a collection of VLMs and datasets, PARC identifies which prompt variations these VLMs are most sensitive to, and which VLMs are most agnostic to prompt variations [green]. To achieve this, PARC first applies systematic prompt variations [orange] to the language and vision components of the datasets, then evaluates the VLM performance on these varied datasets with multiple established scores and a novel reliability score [blue], and finally calibrates [red]
those scores to make them directly comparable across the diverse input datasets as well as PARC's prompt variations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 167478da-c0b9-4e9d-98ba-cae0725235e3Cited by top-tier papers2
- RobustSpring: Benchmarking Robustness to Image Corruptions for Optical Flow, Scene Flow and StereoVictor Oei, Jenny Schmalfuss, Lukas Mehl, Madlen Bartsch et al.ICLR 2026 · 9 citations
- vMFCoOp: Towards Equilibrium on a Unified Hyperspherical Manifold for Prompting Biomedical VLMsMinye Shao, Sihan Guo, Xinrun Li, Xingyu Miao et al.AAAI 2026
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Prompt-Robust Vision-Language Models via Meta-FinetuningHaohui Liang, Runlin Huang, Yingjun Du, Yujia Hu et al.ICLR 2026
- Understanding the Prompt SensitivityYang Liu, Chenhui ChuACL 2026 · 192 citations
- Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language ModelsKyle Cox, Jiawei Xu, Yikun Han, Rong Xu et al.AAAI 2025 · 6 citations
- Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual IllusionsXiaoxiao Sun, Mingyang Li, Kun Yuan, Min Woo Sun et al.CVPR 2026 · 8 citations
- Quantifying Memorization Advantage in Code LLMsAlberick Euraste Djire, Abdoul Kader Kaboré, Jordan Samhi, Earl Barr et al.ICSE 2026
