SC2020Top-tier venue
Task bench: a parameterized benchmark for evaluating parallel runtime performance
Elliott Slaughter, Wei Wu, Yuankun Fu, Legend Brandenburg, Nicolai Garcia, Wilhem Kautz, Emily Marx, Kaleb S. Morris, Qinglei Cao, George Bosilca, Seema Mirchandaney, Wonchan Lee
Abstract
We present Task Bench, a parameterized benchmark designed to explore the performance of distributed programming systems under a variety of application scenarios. Task Bench dramatically lowers the barrier to benchmarking and comparing multiple programming systems by making the implementation for a given system orthogonal to the benchmarks themselves: every benchmark constructed with Task Bench runs on every Task Bench implementation. Furthermore, Task Bench's parameterization enables a wide variety of benchmark scenarios that distill the key characteristics of larger applications.
To assess the effectiveness and overheads of the tested systems, we introduce a novel metric, minimum effective task granularity (METG). We conduct a comprehensive study with 15 programming systems on up to 256 Haswell nodes of the Cori supercomputer. Running at scale, 100µs-long tasks are the finest granularity that any system runs efficiently with current technologies. We also study each system's scalability, ability to hide communication and mitigate load imbalance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a5903447-ef23-4bdd-8070-3ebb1754b3c7Cited by top-tier papers7
- Palette Load Balancing: Locality Hints for Serverless FunctionsMania Abdi, Samuel Ginzburg, Xiayue Charles Lin, Jose M. Faleiro et al.EuroSys 2023 · 44 citations
- Advanced synchronization techniques for task-based runtime systemsDavid Álvarez, Kevin Sala, Marcos Maroñas, Aleix Roca et al.PPoPP 2021 · 24 citations
- Scaling implicit parallelism via dynamic control replicationMichael Bauer, Wonchan Lee, Elliott Slaughter, Zhihao Jia et al.PPoPP 2021 · 12 citations
- GTool: Graph Enhanced Tool Planning with Large Language ModelWenjie Chen, Di Yao, Wenbin Li, Xuying Meng et al.ICLR 2026 · 8 citations
- Trojan Horse: Aggregate-and-Batch for Scaling Up Sparse Direct Solvers on GPU ClustersYida Li, Siwei Zhang, Yiduo Niu, Yang Du et al.PPoPP 2026 · 1 citation
Related papers
- Yield Not Thy CoreAchilles Benetopoulos, Peter Alvaro, Andi Quinn, Robert SouléEuroSys 2026
- MPI-CorrBench: Towards an MPI Correctness Benchmark SuiteJan-Patrick Lehr, Tim Jammer, Christian H. BischofHPDC 2021 · 15 citations
- TAOBench: An End-to-End Benchmark for Social Networking WorkloadsAudrey Cheng, Xiao Shi, Aaron N. Kabcenell, Shilpa Lawande et al.VLDB 2022 · 20 citations
- OpenCilk: A Modular and Extensible Software Infrastructure for Fast Task-Parallel CodeTao B. Schardl, I-Ting Angelina LeePPoPP 2023 · 30 citations
- A Quantitative Approach for Adopting Disaggregated Memory in HPC SystemsJacob Wahlgren, Gabin Schieffer, Maya B. Gokhale, Ivy PengSC 2023 · 14 citations
