Why Attention Fails: A Taxonomy of Faults in Attention-Based Neural Networks
Sigma Jahan, Saurabhsingh Rajput, Tushar Sharma, Masud Rahman
摘要
Attention mechanisms are at the core of modern neural architectures, powering systems ranging from ChatGPT to autonomous vehicles, and driving a major economic impact. However, high-profile failures, such as ChatGPT’s nonsensical outputs or Google’s suspension of Gemini’s image generation due to attention weight errors, highlight a critical gap: existing deep learning fault taxonomies might not adequately capture the unique failures introduced by attention mechanisms. This gap leaves practitioners without actionable diagnostic guidance. To address this gap, we present the first comprehensive empirical study of faults in attention-based neural networks (ABNNs). Our work is based on a systematic analysis of 555 real-world faults collected from 96 projects across ten frameworks, including GitHub, Hugging Face, and Stack Overflow. Through our analysis, we develop a novel taxonomy comprising seven attention-specific fault categories, not captured by existing work. Our results show that over half of the ABNN faults arise from mechanisms unique to attention architectures. We further analyze the root causes and manifestations of these faults through various symptoms. Finally, by analyzing symptom–root cause associations, we identify four evidence-based diagnostic heuristics that explain 33.0% of attention-specific faults and provide systematic, taxonomy-driven guidance for diagnosing faults in ABNNs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Taxonomy of real faults in deep learning systemsNargiz Humbatova, Gunel Jahangirova, Gabriele Bavota, Vincenzo Riccio 等ICSE 2020 · 被引用 281 次
- Stabilizing Transformer Training by Preventing Attention Entropy CollapseShuangfei Zhai, Tatiana Likhomanenko, Etai Littwin, Dan Busbridge 等ICML 2023 · 被引用 153 次
- A comprehensive study of autonomous vehicle bugsJoshua Garcia, Yang Feng, Junjie Shen, Sumaya Almanee 等ICSE 2020 · 被引用 127 次
- A comprehensive study of deep learning compiler bugsQingchao Shen, Haoyang Ma, Junjie Chen, Yongqiang Tian 等FSE 2021 · 被引用 123 次
相关 Paper
- An Empirical Study on Deployment Faults of Deep Learning Based Mobile ApplicationsZhenpeng Chen, Huihan Yao, Yiling Lou, Yanbin Cao 等ICSE 2021 · 被引用 73 次
- ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model TrainingYuhang Liang, Xinyi Li, Jie Ren, Ang Li 等PPoPP 2025 · 被引用 10 次
- Towards Understanding the Faults of JavaScript-Based Deep Learning SystemsLili Quan, Qianyu Guo, Xiaofei Xie, Sen Chen 等ASE 2022 · 被引用 13 次
- Demystifying and Detecting Misuses of Deep Learning APIsMoshi Wei, Nima Shiri Harzevili, Yuekai Huang, Jinqiu Yang 等ICSE 2024 · 被引用 13 次
- Guess or Recall? Training CNNs to Classify and Localize Memorization in LLMsJérémie Dentan, Davide Buscaldi, Sonia VanierAAAI 2026
