Are Humans and LLMs Confused by the Same Code? An Empirical Study on Fixation-Related Potentials and LLM Perplexity
Youssef Abdelsalam, Norman Peitek, Anna-Maria Maurer, Mariya Toneva, Sven Apel
Abstract
Already today, humans and programming assistants based on large language models (LLMs) collaborate in everyday programming tasks. Clearly, a misalignment between how LLMs and programmers comprehend code can lead to misunderstandings, inefficiencies, low code quality, and bugs. A key question in this space is whether humans and LLMs are confused by the same kind of code. This would not only guide our choices of integrating LLMs in software engineering workflows but also inform about possible improvements of LLMs. To this end, we conducted an empirical study comparing an LLM to human programmers comprehending clean and confusing code. We operationalized comprehension for the LLM by using LLM perplexity, and for human programmers using neurophysiological responses (in particular, EEG-based fixation-related potentials). We found that LLM perplexity spikes correlate, both in terms of location and amplitude, with human neurophysiological responses that indicate confusion. This result suggests that LLMs and humans are similarly confused about the code. Based on these findings, we devised a data-driven, LLM-based approach to identify regions of confusion in code that elicit confusion in human programmers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97e33173-b92a-4f9f-8cee-126f1a679ee2Builds on12
- Grounded Copilot: How Programmers Interact with Code-Generating ModelsShraddha Barke, Michael B. James, Nadia PolikarpovaOOPSLA 2023 · 408 citations
- CodeFill: Multi-token Code Completion by Jointly learning from Structure and Naming SequencesMaliheh Izadi, Roberta Gismondi, Georgios GousiosICSE 2022 · 79 citations
- A Deep Dive into Large Language Models for Automated Bug Localization and RepairSoneya Binta Hossain, Nan Jiang, Qiang Zhou, Xiaopeng Li et al.FSE 2024 · 60 citations
- Program Comprehension and Code Complexity Metrics: An fMRI StudyNorman Peitek, Sven Apel, Chris Parnin, André Brechmann et al.ICSE 2021 · 59 citations
- Beyond Code Generation: An Observational Study of ChatGPT Usage in Software Engineering PracticeRanim Khojah, Mazen Mohamad, Philipp Leitner, Francisco Gomes de Oliveira NetoFSE 2024 · 56 citations
Related papers
- Do Large Language Models Pay Similar Attention Like Human Programmers When Generating Code?Bonan Kou, Shengmai Chen, Zhijie Wang, Lei Ma et al.FSE 2024 · 8 citations
- On Behavioral Alignment of Model-Code and Human-Code Understandability via Behavioral ProxiesXiaokai Rong, Aashish Yadavally, Hridya Dhulipala, Anh H. N. Nguyen et al.ISSTA 2026
- How Do Analysts Understand and Verify AI-Assisted Data Analyses?Ken Gu, Ruoxi Shang, Tim Althoff, Chenglong Wang et al.CHI 2024 · 36 citations
- Rethinking Code Complexity Through the Lens of Large Language ModelsChen Xie, Xiaodong Gu, Yuling Shi, Beijun ShenICML 2026
- NeuroSync: Intent-Aware Code-Based Problem Solving via Direct LLM Understanding ModificationWenshuo Zhang, Leixian Shen, Shuchang Xu, Jindu Wang et al.UIST 2025 · 9 citations
