ScanCoder: Leveraging Human Attention Patterns to Enhance LLMs for Code
Yueke Zhang, Yifan Zhang, Zihan Fang, Greg Trafton, Daniel Levin, Kevin Leach, Yu Huang
Abstract
Code comprehension is a fundamental challenge in software engineering that impacts developer productivity and software quality. While Large Language Models (LLMs) demonstrate strong capabilities in code generation and summarization, they process code differently from human developers, who employ strategic attention patterns focused on semantically critical elements. Recent research has successfully integrated human attention patterns captured through eye-tracking into AI models for software engineering tasks, however, existing human-AI approaches face critical limitations that prevent widespread practical deployment, particularly for LLM enhancement. Existing approaches to incorporate human cognitive insights face scalability limitations due to resource-intensive eye-tracking studies and lack empirical validation for cross-language generalizability. We present ScanCoder, a framework that integrates cognitive simulation with LLM enhancement through (1) generating human-like attention patterns at scale using minimal eye-tracking data via cognitive simulation with Adaptive Character of Thought-Rational (ACT-R) architecture, and (2) cognitively-guided fine-tuning that emphasizes tokens according to their cognitive salience and attention order. Our approach demonstrates cross-language transfer by applying C++-derived cognitive patterns to enhance Java programming tasks. Comprehensive evaluation on CodeXGLUE benchmarks shows consistent improvements across different LLM architectures and scales (1B–8B parameters), achieving gains of up to 41.66 points on CrystalBLEU for code completion and 21.93 points on BERTScore for code summarization. Mechanistic analysis reveals that cognitive guidance reshapes model attention in task-dependent ways, increasing focus on semantically critical tokens by 2.5×. This work establishes the first scalable framework for integrating simulated human cognitive patterns into LLM training, enabling more interpretable and effective code understanding.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 563c6622-5e44-49fc-a9a5-d86346fb5ebdRelated papers
- EyeTrans: Merging Human and Machine Attention for Neural Code SummarizationYifan Zhang, Jiliang Li, Zachary Karas, Aakash Bansal et al.FSE 2024 · 15 citations
- EyeMulator: Improving Code Language Models by Mimicking Human Visual AttentionYifan Zhang, Chen Huang, Yueke Zhang, Jiahao Zhang et al.ACL 2026 · 4 citations
- Enhanced Prompting Framework for Code Summarization with Large Language ModelsMinying Fang, Xing Yuan, Yuying Li, Haojie Li et al.ISSTA 2025 · 3 citations
- SimLLM: Calculating Semantic Similarity in Code Summaries using a Large Language Model-Based ApproachXin Jin, Zhiqiang LinFSE 2024 · 8 citations
- Automatic Semantic Augmentation of Language Model Prompts (for Code Summarization)Toufique Ahmed, Kunal Suresh Pai, Premkumar T. Devanbu, Earl T. BarrICSE 2024 · 71 citations
