Automatic Tracing in Task-Based Runtime Systems
Rohan Yadav, Michael Bauer, David Broman, Michael Garland, Alex Aiken, Fredrik Kjolstad
摘要
Implicitly parallel task-based runtime systems often perform dynamic analysis to discover dependencies in and extract parallelism from sequential programs. Dependence analysis becomes expensive as task granularity drops below a threshold. Tracing techniques have been developed where programmers annotate repeated program fragments (traces) issued by the application, and the runtime system memoizes the dependence analysis for those fragments, greatly reducing overhead when the fragments are executed again. However, manual trace annotation can be brittle and not easily applicable to complex programs built through the composition of independent components. We introduce Apophenia, a system that automatically traces the dependence analysis of task-based runtime systems, removing the burden of manual annotations from programmers and enabling new and complex programs to be traced. Apophenia identifies traces dynamically through a series of dynamic string analyses, which find repeated program fragments in the stream of tasks issued to the runtime system. We show that Apophenia is able to come between 0.92x--1.03x the performance of manually traced programs, and is able to effectively trace previously untraced programs to yield speedups of between 0.91x--2.82x on the Perlmutter and Eos supercomputers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Unity: Accelerating DNN Training Through Joint Optimization of Algebraic Transformations and ParallelizationColin Unger, Zhihao Jia, Wei Wu, Sina Lin 等OSDI 2022 · 被引用 105 次
- Task bench: a parameterized benchmark for evaluating parallel runtime performanceElliott Slaughter, Wei Wu, Yuankun Fu, Legend Brandenburg 等SC 2020 · 被引用 51 次
- Scaling implicit parallelism via dynamic control replicationMichael Bauer, Wonchan Lee, Elliott Slaughter, Zhihao Jia 等PPoPP 2021 · 被引用 12 次
- Loop Rerolling for Hardware DecompilationZachary D. Sisco, Jonathan Balkind, Timothy Sherwood, Ben HardekopfPLDI 2023 · 被引用 12 次
- Legate Sparse: Distributed Sparse Computing in PythonRohan Yadav, Wonchan Lee, Melih Elibol, Manolis Papadakis 等SC 2023 · 被引用 8 次
相关 Paper
- On the fly MHP analysisSonali Saha, V. Krishna NandivadaPPoPP 2020 · 被引用 2 次
- ScalAna: automating scaling loss detection with graph analysisYuyang Jin, Haojie Wang, Teng Yu, Xiongchao Tang 等SC 2020 · 被引用 16 次
- LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning SystemsArnab Phani, Benjamin Rath, Matthias BoehmSIGMOD 2021 · 被引用 30 次
- Build scripts with perfect dependenciesSarah Spall, Neil Mitchell, Sam Tobin-HochstadtOOPSLA 2020 · 被引用 8 次
- PerFlow: a domain specific framework for automatic performance analysis of parallel applicationsYuyang Jin, Haojie Wang, Runxin Zhong, Chen Zhang 等PPoPP 2022 · 被引用 10 次
