No Rush in Executing Atomic Instructions
Ashkan Asgharzadeh, Josué Feliu, Manuel E. Acacio, Stefanos Kaxiras, Alberto Ros
摘要
Hardware atomic instructions are the building blocks of the synchronization algorithms. Historically, to guarantee atomicity and consistency, they were implemented using memory fences, committing older memory instructions, and draining the store buffer before initiating the execution of atomics. Unfortunately, the use of such memory fences entails huge performance penalties as it implies execution serialization, thus impeding instruction- and memory-level parallelism. The situation, however, seems to have changed recently. Through experiments on machines, we discovered that current processors manage to comply with the x86-TSO requirements while avoiding the performance overhead introduced by fences (fence-free or unfenced implementation). This paves the way to new potential optimizations to atomic instruction execution. In particular, our simulation experiments modeling unfenced atomics reveal that executing atomic instructions as soon as their operands are ready does not always lead to optimal performance. In fact, this increases the time that other threads should wait to obtain the cacheline. In contended scenarios, delaying the execution of the atomic instruction to minimize the time the cacheline is locked provides superior performance. Based on this observation, we present Rush or Wait (RoW), a hardware mechanism to decide when to execute an atomic instruction. The mechanism is based on a contention predictor that estimates if an atomic will access a contended cacheline. Non-contended atomics execute once their operands are ready. Contended atomics, on the contrary, wait to become the oldest memory instruction and to drain the store buffer to execute, minimizing the contention on the accessed cacheline. Our experimental evaluation shows that RoW reduces execution time on average by 9.2% (and up to 43%) compared to a baseline that executes atomics as soon as the operands are ready, and yet it requires a small area overhead (64 bytes).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Free atomics: hardware atomic operations without fencesAshkan Asgharzadeh, Juan M. Cebrian, Arthur Perais, Stefanos Kaxiras 等ISCA 2022 · 被引用 13 次
- ATUNs: Modular and Scalable Support for Atomic Operations in a Shared Memory MultiprocessorAndreas Kurth, Samuel Riedel, Florian Zaruba, Torsten Hoefler 等DAC 2020 · 被引用 5 次
- DynAMO: Improving Parallelism Through Dynamic Placement of Atomic Memory OperationsVíctor Soria Pardos, Adrià Armejach, Tiago Mück, Darío Suárez Gracia 等ISCA 2023 · 被引用 5 次
相关 Paper
- Speculative Enforcement of Store AtomicityAlberto Ros, Stefanos KaxirasMICRO 2020 · 被引用 5 次
- Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic OperationsYicong Zhang, Mingyu Wang, Wangguang Wang, Yangzhan Mai 等MICRO 2024 · 被引用 4 次
- Rely/Guarantee Reasoning for Multicopy Atomic Weak Memory ModelsNicholas Coughlin, Kirsten Winter, Graeme SmithFM 2021 · 被引用 16 次
- Temporarily Unauthorized Stores: Write First, Ask for Permission LaterJuan M. Cebrian, Magnus Jahre, Alberto RosMICRO 2024 · 被引用 2 次
- Execution Dependence Extension (EDE): ISA Support for Eliminating FencesThomas Shull, Ilias Vougioukas, Nikos Nikoleris, Wendy Elsasser 等ISCA 2021 · 被引用 4 次
