PAR: Training-Free Positional Perturbation and Attention Recycling for Faithful OCR
Yao Yao, Manwen Liao, Weitian Zhang, Zuchao Li, Hai Zhao
摘要
In high-precision scenarios, vision language models suffer from Linguistic Priors Hallucination. When processing familiar text, models tend to over-rely on internal parametric knowledge, effectively "reciting" the content rather than "reading" the image. In this paper, we first systematically investigate this phenomenon by constructing the GlitchText Probing Dataset. We discover that the model's reliance on visual grounding diminishes significantly as the generation length increases. To mitigate this, we propose PAR (Positional Perturbation and Attention Recycling), a training-free, inferencetime intervention framework. PAR consists of two parts: (1) Positional Perturbation (PP) injects structured phase noise into the rotary positional embeddings; (2) Foveal Attention Recycling (FAR) detects over-confident linguistic priors and dynamically redistributes attention mass back to important visual regions. Extensive experiments across state-of-the-art models, demonstrate that PAR significantly reduces hallucination rates (reducing CER by 12%), particularly in long-context scenarios, while maintaining robust generalization on standard benchmarks. Our code is publicly available at https://github.com/Zoeyyao27/PAR-for-Faithful-OCR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang 等EMNLP 2023 · 被引用 344 次
- Detecting and Preventing Hallucinations in Large Vision Language ModelsAnisha Gunjal, Jihan Yin, Erhan BasAAAI 2024 · 被引用 312 次
- Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT ReasoningHai-Long Sun, Zhun Sun, Houwen Peng, Han-Jia YeACL 2025 · 被引用 23 次
- When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and UnderstandingYan Shu, Hangui Lin, Yexin Liu, Yan Zhang 等NeurIPS 2025 · 被引用 17 次
- DASH: Detection and Assessment of Systematic Hallucinations of VLMsMaximilian Augustin, Yannic Neuhaus, Matthias HeinICCV 2025 · 被引用 17 次
相关 Paper
- ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language ModelsJunzhe Chen, Tianshu Zhang, Shiyu Huang, Yuwei Niu 等CVPR 2025
- Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local AttentionWenbin An, Feng Tian, Sicong Leng, Jiahao Nie 等CVPR 2025
- Imitating the Truth: Attention-aware Truth-Guided Enhancement for Hallucination Mitigation in Large Vision-Language ModelsHairui Ren, Zixuan Wang, Yibo Yang, He Zhao 等ICLR 2026
- REVIS: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language ModelsJialin Wu, Wei Shi, Han Shen, Peigui Qi 等ICML 2026 · 被引用 2 次
- PAS: Prelim Attention Score for Detecting Object Hallucinations in Large Vision-Language ModelsNhat Hoang, Minh Vu, My T. Thai, Manish BhattaraiCVPR 2026 · 被引用 1 次
