Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM
Junyuan Mao, Qiankun Li, Linghao Meng, Zhicheng He, Xinliang Zhou, Kun Wang, Yang Liu, Yueming Jin
Abstract
Recent advances in multimodal large language models largely rely on CLIP-based visual encoders, which emphasize global semantic alignment but struggle with finegrained visual understanding. In contrast, DINOv3 provides strong pixel-level perception yet lacks coarse-grained semantic abstraction, leading to limited multi-granularity reasoning. To address this gap, we propose Granulon, a novel DINOv3-based MLLM with adaptive granularity augmentation. Granulon introduces a text-conditioned granularity Controller that dynamically adjusts the visual abstraction level according to the semantic scope of the textual input, and an Adaptive Token Aggregation module that performs granularity-guided pooling and relation-aware clustering to produce compact, semantically rich visual tokens. This design enables unified "pixel-to-fine-to-coarse" reasoning within a single forward pass. Extensive and interpretable experiments demonstrate that Granulon improves accuracy by ∼ 30% ↑ and reduces hallucination by ∼ 20% ↓, outperforming all visual encoders under identical settings. Code is available at Granulon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 283712ca-2571-464e-9801-9cfd752c1fcdBuilds on24
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- GranViT: A Fine-Grained Vision Model For Autoregressive Multimodal Large Language ModelsGuanghao Zheng, Bowen Shi, Mingxing Xu, Ruoyu Sun et al.ICLR 2026 · 1 citation
- MoVA: Adapting Mixture of Vision Experts to Multimodal ContextZhuofan Zong, Bingqi Ma, Dazhong Shen, Guanglu Song et al.NeurIPS 2024 · 110 citations
- Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision EncoderSiting Li, Pang Wei Koh, Simon Shaolei DuACL 2025
- Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language ModelsShizhan Gong, Yankai Jiang, Qi Dou, Farzan FarniaICML 2025
- Groundhog Grounding Large Language Models to Holistic SegmentationYichi Zhang, Ziqiao Ma, Xiaofeng Gao, Suhaila Shakiah et al.CVPR 2024 · 24 citations
