LOREAL: Mitigating Low-Resolution Challenges in Vision-Language Models with Attribute-driven Prompt Self-Distillation
Xucong Wang, Pengkun Wang, Zhe Zhao, Liheng Yu, Rui Mao, Yang Wang
Abstract
Prompt Learning (PL) has emerged as a parameter-efficient technique for adapting Vision-Language Models (VLMs) to downstream tasks. However, almost all existing PL methods are primarily designed and evaluated on wellcurated datasets, overlooking a critical post-deployment phenomenon, i.e., the intrinsic connection between input resolution and storage-memory consumption. Specifically, to satisfy the stringent storage-memory constraints on edge devices, models are often limited to low-resolution inputs (e.g., ≤ 224×224 for CLIP-ViT/B-16) and generate fewer tokens (with the position embedding resized), which poses a unique challenge in performance robustness. To tackle this issue, we propose LOREAL, an efficient prompt selfdistillation framework that learns resolution-robust representations by excavating attribute semantics. At the heart of LOREAL is a dual-student architecture, i.e., two student models fed with inputs at different resolutions synergistically learn from each other. Building upon this, we contextualize the students' prompt with resolution-robust attributes queried from the LLM, then leverage cross-modality meta-nets to generate attribute semantics. These metanets are bridged between the different encoders of two students, wherein we introduce Low-Level Distillation (LLD) and High-Level Distillation (HLD) to facilitate the learning of more cross-resolution representations. Extensive experiments show that LOREAL significantly improves VLMs' performance and robustness under varied resolution settings, underscoring significant practical utilities. Project
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on48
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Enhancing Pre-trained ViTs for Downstream Task Adaptation: A Locality-Aware Prompt Learning MethodShaokun Wang, Yifan Yu, Yuhang He, Yihong GongACM MM 2024 · 1 citation
- Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIPChen Huang, Skyler Seto, Samira Abnar, David Grangier et al.NeurIPS 2024 · 8 citations
- Fed-Duet: Dual Expert-Orchestrated Framework for Continual Federated Vision-Language LearningTao Guo, Junwei Chen, Laizhong CuiICLR 2026
- ProLoG: Hybrid Prompt and LoRA Based Adaptation of Vision-Language Models for OOD GeneralizationJungwuk Park, Dong-Jun Han, Jaekyun MoonAAAI 2026
- Hierarchical Cross-Modal Prompt Learning for Vision-Language ModelsHao Zheng, Shunzhi Yang, Zhuoxin He, Jinfeng Yang et al.ICCV 2025 · 5 citations
