Deterministic Atomic Buffering
Yuan-Hsi Chou, Christopher Ng, Shaylin Cattell, Jeremy Intan, Matthew D. Sinclair, Joseph Devietti, Timothy G. Rogers, Tor M. Aamodt
Abstract
Deterministic execution for GPUs is a desirable property as it helps with debuggability and reproducibility. It is also important for safety regulations, as safety critical workloads are starting to be deployed onto GPUs. Prior deterministic architectures, such as GPUDet, attempt to provide strong determinism for all types of workloads, incurring significant performance overheads due to the many restrictions that are required to satisfy determinism. We observe that a class of reduction workloads, such as graph applications and neural architecture search for machine learning, do not require such severe restrictions to preserve determinism. This motivates the design of our system, Deterministic Atomic Buffering (DAB), which provides deterministic execution with low area and performance overheads by focusing solely on ordering atomic instructions instead of all memory instructions. By scheduling atomic instructions deterministically with atomic buffering, the results of atomic operations are isolated initially and made visible in the future in a deterministic order. This allows the GPU to execute deterministically in parallel without having to serialize its threads for atomic operations as opposed to GPUDet. Our simulation results show that, for atomic-intensive applications, DAB performs 4× better than GPUDet and incurs only a 23% slowdown on average compared to a non-deterministic GPU architecture. We also characterize the bottlenecks and provide insights for future optimizations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49748439-e50e-4929-8278-558578836c4aCited by top-tier papers6
- Not All GPUs Are Created Equal: Characterizing Variability in Large-Scale, Accelerator-Rich SystemsPrasoon Sinha, Akhil Guliani, Rutwik Jain, Brandon Tran et al.SC 2022 · 31 citations
- Optimistic Verifiable Training by Controlling Hardware NondeterminismMegha Srivastava, Simran Arora, Dan BonehNeurIPS 2024 · 14 citations
- Only Buffer When You Need To: Reducing On-chip GPU Traffic with Reconfigurable Local Atomic BuffersPreyesh Dalmia, Rohan Mahapatra, Matthew D. SinclairHPCA 2022 · 11 citations
- On The Fairness Impacts of Hardware Selection in Machine LearningSree Harsha Nelaturu, Nishaanth Kanna Ravichandran, Cuong Tran, Sara Hooker et al.ICML 2024 · 5 citations
- DORADD: Deterministic Parallel Execution in the Era of Microsecond-Scale ComputingZhengqing Liu, Musa Unal, Matthew J. Parkinson, Marios KogiasPPoPP 2025 · 3 citations
Builds on1
Related papers
- Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic OperationsYicong Zhang, Mingyu Wang, Wangguang Wang, Yangzhan Mai et al.MICRO 2024 · 4 citations
- Atomic Dataflow based Graph-Level Workload Orchestration for Scalable DNN AcceleratorsShixuan Zheng, Xianjue Zhang, Leibo Liu, Shaojun Wei et al.HPCA 2022 · 39 citations
- LLM-42: Enabling Determinism in LLM Inference with Verified SpeculationRaja Gond, Aditya K Kamath, Ramachandran Ramjee, Ashish PanwarSOSP 2026
- NASGuard: A Novel Accelerator Architecture for Robust Neural Architecture Search (NAS) NetworksXingbin Wang, Boyan Zhao, Rui Hou, Amro Awad et al.ISCA 2021 · 9 citations
- Why GPUs are Slow at Executing NFAs and How to Make them FasterHongyuan Liu, Sreepathi Pai, Adwait JogASPLOS 2020 · 31 citations
