Neutrino: Fine-grained GPU Kernel Profiling via Programmable Probing
Songlin Huang, Chenshu Wu
摘要
As GPUs play an increasingly important role in computer systems in the scaling laws era, understanding fine-grained GPU runtime behavior is more crucial than ever. However, existing GPU kernel profilers, typically kernel-exclusive or hardware-dependent, often fail to capture fine-grained measurements. This paper presents NEUTRINO, a programmable interface for GPU kernel profiling that leverages assemblylayer probing to achieve instruction-level fine granularity, profiling versatility across time and value domains, and hardware independence. To better visualize the rich details captured by NEUTRINO, we introduce the Densified Memory Access Timeline (DMAT), a novel representation that offers new insights into GPU runtime behavior. We implement NEUTRINO in Linux for both NVIDIA and AMD GPUs and conduct extensive evaluations and analyses. The results demonstrate NEUTRINO's superior capabilities in GPU kernel profiling with low overhead. We envision NEUTRINO as a valuable tool for the community and have open-sourced it to facilitate future research at https://github.com/open-neutrino/ neutrino.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- GPU Checkpoint/Restore Made Fast and LightweightShaoxun Zeng, Tingxu Ren, Jiwu Shu, Youyou LuFAST 2026 · 被引用 5 次
- Fine-grained and Non-intrusive LLM Training Monitoring via Microsecond-level Traffic MeasurementYibo Xiao, Hao Zheng, Haifeng Sun, Qingkai Meng 等ASPLOS 2026
它引用的顶会 Paper22
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
相关 Paper
- FabricPerf: Measuring NIC-less Scale-Up Network through GPU Communication Kernel ProfilingSonglin Huang, Chenshu WuSIGCOMM 2026
- Efficient Multi-GPU Shared Memory via Automatic Optimization of Fine-Grained TransfersHarini Muthukrishnan, David W. Nellans, Daniel Lustig, Jeffrey A. Fessler 等ISCA 2021 · 被引用 16 次
- DrGPUM: Guiding Memory Optimization for GPU-Accelerated ApplicationsMao Lin, Keren Zhou, Pengfei SuASPLOS 2023 · 被引用 13 次
- KPerfIR: Towards a Open and Compiler-centric Ecosystem for GPU Kernel Performance Tooling on Modern AI WorkloadsYue Guan, Yuanwei Fang, Keren Zhou, Corbin Robeck 等OSDI 2025
- GPU-Ether: GPU-native Packet I/O for GPU Applications on Commodity EthernetChangue Jung, Suhwan Kim, Ikjun Yeom, Honguk Woo 等INFOCOM 2021 · 被引用 7 次
