ThunderKittens: Simple, Fast, and Adorable Kernels
Benjamin Frederick Spector, Simran Arora, Aaryan Singhal, Arjun Parthasarathy, Daniel Y. Fu, Christopher Ré
Abstract
The challenge of mapping AI architectures to GPU hardware is creating a critical bottleneck in AI progress. Despite substantial efforts, hand-written custom kernels fail to meet their theoretical performance thresholds, even on well-established operations like linear attention. The diverse capabilities of GPUs suggests we might we need a wide variety of techniques to achieve high performance. However, our work explores if a small number of key abstractions can drastically simplify the process. We present THUNDERKITTENS (TK), a framework for writing performant AI kernels while remaining easy to use. Our abstractions map to the three levels of the GPU hierarchy: (1) at the warp-level, we provide 16x16 matrix tiles as basic data structures and PyTorch-like operations, (2) at the thread-block level, we provide templates for asynchronously overlapping operations, and (3) at the grid-level, TK helps hide block launch, tear-down, and memory costs. We show the value of TK by providing simple & diverse kernels that match or outperform prior art. We match CuBLAS and FlashAttention-3 on GEMM and attention inference performance and outperform the strongest baselines by 10-40% on attention backwards, 8× on state space models, and 14× on linear attention.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e88e9755-e07a-4c16-aa62-9efc6b838eb2Cited by top-tier papers2
- EGG: An Expert-Guided Agent Framework for Kernel GenerationYaochen Han, Ke Fan, Hongxu Jiang, Wanqi Xu et al.ICML 2026
- S²Flow: Towards Fast and Authentic Training-Free High-Resolution Video GenerationChaoqun Wang, Shaobo Min, Xu YangAAAI 2026
Builds on18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
Related papers
- CuBridge: An LLM-Based Framework for Understanding and Reconstructing High-Performance Attention KernelsXing Ma, Yangjie Zhou, Wu Sun, Zihan Liu et al.ACL 2026 · 2 citations
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionJay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar et al.NeurIPS 2024 · 727 citations
- FlexLinearAttention: Compiling a Unified Abstraction into Scalable Kernels for Linear AttentionHaojie Duanmu, Size Zheng, Ningxin Zheng, Jianqiao Lu et al.ICLR 2026
- ThunderGNN: Unlocking Tensor Cores for Graph Neural NetworksYuAng Chen, Siyi Teng, Wenqi Weng, Jeffrey Xu YuVLDB 2026
- TileLang: Bridge Programmability and Performance in Modern Neural KernelsLei Wang, Yu Cheng, Yining Shi, Zhiwen Mo et al.ICLR 2026
