Triangulating Python Performance Issues with SCALENE
Emery D. Berger, Sam Stern, Juan Altmayer Pizzorno
摘要
This paper proposes Scalene, a profiler specialized for Python. Scalene combines a suite of innovations to precisely and simultaneously profile CPU, memory, and GPU usage, all with low overhead. Scalene's CPU and memory profilers help Python programmers direct their optimization efforts by distinguishing between inefficient Python and efficient native execution time and memory usage. Scalene's memory profiler employs a novel sampling algorithm that lets it operate with low overhead yet high precision. It also incorporates a novel algorithm that automatically pinpoints memory leaks, whether within Python or across the Python-native boundary. Scalene tracks a new metric called copy volume, which highlights costly copying operations that can occur when Python silently converts between C and Python data representations, or between CPU and GPU. Since its introduction, Scalene has been widely adopted, with over 500,000 downloads to date. We present experience reports from developers who used Scalene to achieve significant performance improvements and memory savings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Search-Based LLMs for Code OptimizationShuzheng Gao, Cuiyun Gao, Wenchao Gu, Michael R. LyuICSE 2025 · 被引用 11 次
- DyPyBench: A Benchmark of Executable Python SoftwareIslem Bouzenia, Bajaj Piyush Krishan, Michael PradelFSE 2024 · 被引用 8 次
- It's Not Easy Being Green: On the Energy Efficiency of Programming LanguagesNicolas van Kempen, Hyuk-Je Kwon, Dung Tuan Nguyen, Emery D. BergerASE 2025 · 被引用 8 次
- FLARE: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus ScaleWeihao Cui, Ji Zhang, Han Zhao, Chao Liu 等NSDI 2026 · 被引用 7 次
它引用的顶会 Paper1
相关 Paper
- DeepContext: A Context-aware, Cross-platform, and Cross-framework Tool for Performance Profiling and Analysis of Deep Learning WorkloadsQidong Zhao, Hao Wu, Yueming Hao, Zilingfeng Ye 等ASPLOS 2025 · 被引用 3 次
- DrGPUM: Guiding Memory Optimization for GPU-Accelerated ApplicationsMao Lin, Keren Zhou, Pengfei SuASPLOS 2023 · 被引用 13 次
- Pyriscope: Precise and Low-Overhead Python Control Flow Tracing via Sparse Hardware-Based EventsXinchen Yao, Wu Daiyou, Zhiqiang ZuoOOPSLA 2026
- Sentinel: Efficient Tensor Migration and Allocation on Heterogeneous Memory Systems for Deep LearningJie Ren, Jiaolin Luo, Kai Wu, Minjia Zhang 等HPCA 2021 · 被引用 62 次
- MemPerf: Profiling Allocator-Induced Performance SlowdownsJin Zhou, Sam Silvestro, Steven (Jiaxun) Tang, Hanmei Yang 等OOPSLA 2023
