HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
Runhui Huang, Xinpeng Ding, Chunwei Wang, Jianhua Han, Yulong Liu, Hengshuang Zhao, Hang Xu, Lu Hou, Wei Zhang, Xiaodan Liang
摘要
High-resolution image inputs allow Large Vision-Language Models (LVLMs) to capture finer visual details, improving comprehension. However, the increased training and computational costs associated with such inputs pose significant challenges. A common approach to mitigate these costs involves slicing the input into uniform patches using sliding windows, each aligned with the vision encoder's input size. While efficient, this method fragments the input, disrupting the continuity of context, which negatively impacts cross-patch perception tasks. To address these limitations, we propose HiRes-LLaVA, a novel framework designed to efficiently process high-resolution inputs of any size without altering the original contextual and geometric information. HiRes-LLaVA introduces two key components: (i) a SliceRestore Adapter (SRA) that reconstructs sliced patches into their original form, enabling efficient extraction of both global and local features through down-up-sampling and convolutional layers, and (ii) a Self-Mining Sampler (SMS) that compresses visual tokens based on internal relationships, preserving original context and positional information while reducing training overhead. To assess the ability of handling context fragmentation, we construct a new benchmark, EntityGrid-QA, consisting of edge-related tasks. Extensive experiments demonstrate the superiority of HiRes-LLaVA on both existing public benchmarks and EntityGrid-QA. For example, with SRA, our method achieves a performance improvement of ∼ 12% over state-of-the-art LVLMs in addressing fragmentation issues. Additionally, our SMS outperforms other visual token downsamplers, while offering high data efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & SparsificationWei Li, Renshan Zhang, Rui Shao, Jie He 等NeurIPS 2025 · 被引用 87 次
- DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and PlanningZhe Liu, Runhui Huang, Rui Yang, Siming Yan 等CVPR 2026 · 被引用 15 次
- Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token InteractionsLin Chen, zhaoxiaoke, Kun Ding, Weiwei Feng 等ICML 2026 · 被引用 4 次
- FALCON: Resolving Visual Redundancy and Fragmentation in High-Resolution Multimodal Large Language Models via Visual RegistersRenshan Zhang, Rui Shao, Gongwei Chen, Miao Zhang 等ICCV 2025 · 被引用 2 次
- When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token PruningJunwei Luo, Yingying Zhang, Xue Yang, Kang Wu 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language ModelsGen Luo, Yiyi Zhou, Yuxin Zhang, Xiawu Zheng 等ICLR 2025
- HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language ModelsKazi Hasan Ibn Arif, JinYi Yoon, Dimitrios S. Nikolopoulos, Hans Vandierendonck 等AAAI 2025 · 被引用 49 次
- iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal ModelsLianyu Hu, Liqing Gao, Fanhua Shang, Liang Wan 等ICLR 2026 · 被引用 9 次
- SPLIT-VLM: Salience-Guided Partitioning towards Local Coverage for Importance-Aware Token Dropping in Vision-Language ModelsSeungil Lee, Gilha lee, Hyun KimICML 2026
- GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K ResolutionFengxiang Wang, Mingshuo Chen, Yueying Li, Di Wang 等NeurIPS 2025 · 被引用 46 次
