Triangulating Python Performance Issues with SCALENE
Emery D. Berger, Sam Stern, Juan Altmayer Pizzorno
Abstract
This paper proposes Scalene, a profiler specialized for Python. Scalene combines a suite of innovations to precisely and simultaneously profile CPU, memory, and GPU usage, all with low overhead. Scalene's CPU and memory profilers help Python programmers direct their optimization efforts by distinguishing between inefficient Python and efficient native execution time and memory usage. Scalene's memory profiler employs a novel sampling algorithm that lets it operate with low overhead yet high precision. It also incorporates a novel algorithm that automatically pinpoints memory leaks, whether within Python or across the Python-native boundary. Scalene tracks a new metric called copy volume, which highlights costly copying operations that can occur when Python silently converts between C and Python data representations, or between CPU and GPU. Since its introduction, Scalene has been widely adopted, with over 500,000 downloads to date. We present experience reports from developers who used Scalene to achieve significant performance improvements and memory savings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Search-Based LLMs for Code OptimizationShuzheng Gao, Cuiyun Gao, Wenchao Gu, Michael R. LyuICSE 2025 · 11 citations
- DyPyBench: A Benchmark of Executable Python SoftwareIslem Bouzenia, Bajaj Piyush Krishan, Michael PradelFSE 2024 · 8 citations
- It's Not Easy Being Green: On the Energy Efficiency of Programming LanguagesNicolas van Kempen, Hyuk-Je Kwon, Dung Tuan Nguyen, Emery D. BergerASE 2025 · 8 citations
- FLARE: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus ScaleWeihao Cui, Ji Zhang, Han Zhao, Chao Liu et al.NSDI 2026 · 7 citations
Builds on1
Related papers
- DeepContext: A Context-aware, Cross-platform, and Cross-framework Tool for Performance Profiling and Analysis of Deep Learning WorkloadsQidong Zhao, Hao Wu, Yueming Hao, Zilingfeng Ye et al.ASPLOS 2025 · 3 citations
- DrGPUM: Guiding Memory Optimization for GPU-Accelerated ApplicationsMao Lin, Keren Zhou, Pengfei SuASPLOS 2023 · 13 citations
- Pyriscope: Precise and Low-Overhead Python Control Flow Tracing via Sparse Hardware-Based EventsXinchen Yao, Wu Daiyou, Zhiqiang ZuoOOPSLA 2026
- Sentinel: Efficient Tensor Migration and Allocation on Heterogeneous Memory Systems for Deep LearningJie Ren, Jiaolin Luo, Kai Wu, Minjia Zhang et al.HPCA 2021 · 62 citations
- MemPerf: Profiling Allocator-Induced Performance SlowdownsJin Zhou, Sam Silvestro, Steven (Jiaxun) Tang, Hanmei Yang et al.OOPSLA 2023
