Beyond the Panorama: Training-Free Hierarchical Perception-Reasoning for Fine-Grained Vision in MLLMs
Xiaoyang Yi, Jing Chen, Li Peng, Yuru Bao, Jian Zhang
摘要
Multimodal large language models (MLLMs) enable cross-modal semantic understanding and generation by learning semantic alignment and fusion across modalities. However, existing MLLMs still face challenges in fine-grained visual tasks. Their uniform encoding for global understanding tends to blur or lose local details, while the lack of explicit modeling of intermediate visual evidence leads them to rely on semantic priors or the statistical patterns of language models rather than grounded visual information, resulting in potential hallucinations. To address these issues, we propose HiPerson, a training-free hierarchical perception-reasoning framework that enhances fine-grained visual understanding by simulating human perception mechanisms. Specifically, HiPerson fuses internal relative attention and gradient activation signals to generate a task-aware semantic heatmap, providing explicit perceptual anchors for precise localization. Then, it employs a dual-scale adaptive cropping strategy to extract visual cues for interactive reasoning, simulating the process of human visual focus shifting and detail attention. Finally, by combining local-global dual-image cooperative input with a multi-step reasoning prompting mechanism, HiPerson guides the model to complete a full perception loop from detail observation to contextual verification. Experiments show that HiPerson achieves competitive results on multiple datasets, demonstrating its generalizability and scalability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- UrbanMLLM: Joint Learning of Cross-view Imagery for Urban UnderstandingXin Zhang, Tianjian Ouyang, Yu Shang, Qingmin Liao 等ICML 2026
- DeepScan: A Training-Free Framework for Visually Grounded Reasoning in Large Vision-Language ModelsYangfu Li, Hongjian Zhan, Jiawei Chen, Yuning Gong 等CVPR 2026 · 被引用 6 次
- FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement LearningLu Zhang, Jiazuo Yu, Haomiao Xiong, Ping Hu 等NeurIPS 2025 · 被引用 4 次
- RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-ThoughtYi Lu, Jiawang Cao, Yongliang Wu, Bozheng Li 等ACL 2025 · 被引用 15 次
- ActiveScope: Actively Seeking and Correcting Perception for MLLMsYajing Wang, Chao Bi, Junshu Sun, Shufan Shen 等ICML 2026 · 被引用 1 次
