Free atomics: hardware atomic operations without fences
Ashkan Asgharzadeh, Juan M. Cebrian, Arthur Perais, Stefanos Kaxiras, Alberto Ros
摘要
Atomic Read-Modify-Write (RMW) instructions are primitive synchronization operations implemented in hardware that provide the building blocks for higher-abstraction synchronization mechanisms to programmers. According to publicly available documentation, current x86 implementations serialize atomic RMW operations, i.e., the store buffer is drained before issuing atomic RMWs and subsequent memory operations are stalled until the atomic RMW commits. This serialization, carried out by memory fences, incurs a performance cost which is expected to increase with deeper pipelines.
This work proposes Free atomics, a lightweight, speculative, deadlock-free implementation of atomic operations that removes the need for memory fences, thus improving performance, while preserving atomicity and consistency. Free atomics is, to the best of our knowledge, the first proposal to enable store-to-load forwarding for atomic RMWs. Free atomics only requires simple modifications and incurs a small area overhead (15 bytes). Our evaluation using gem5-20 shows that, for a 32-core configuration, Free atomics improves performance by 12.5%, on average, for a large range of parallel workloads and 25.2%, on average, for atomic-intensive parallel workloads over a fenced atomic RMW implementation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Leviathan: A Unified System for General-Purpose Near-Data ComputingBrian C. Schwedock, Nathan BeckmannMICRO 2024 · 被引用 6 次
- DX100: Programmable Data Access Accelerator for IndirectionAlireza Khadem, Kamalavasan Kamalakkannan, Zhenyan Zhu, Akash Poptani 等ISCA 2025 · 被引用 2 次
- Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System SimulationYanjing Wang, Lizhou Wu, Sunfeng Gao, Yibo Tang 等HPCA 2026 · 被引用 1 次
- No Rush in Executing Atomic InstructionsAshkan Asgharzadeh, Josué Feliu, Manuel E. Acacio, Stefanos Kaxiras 等HPCA 2025
- Bounding Speculative Execution of Atomic Regions to a Single RetryEduardo José Gómez-Hernández, Juan M. Cebrian, Stefanos Kaxiras, Alberto RosASPLOS 2024
它引用的顶会 Paper4
- Efficient, Distributed, and Non-Speculative Multi-Address Atomic OperationsEduardo José Gómez-Hernández, Juan M. Cebrian, J. Rubén Titos Gil, Stefanos Kaxiras 等MICRO 2021 · 被引用 8 次
- Speculative Enforcement of Store AtomicityAlberto Ros, Stefanos KaxirasMICRO 2020 · 被引用 5 次
- Execution Dependence Extension (EDE): ISA Support for Eliminating FencesThomas Shull, Ilias Vougioukas, Nikos Nikoleris, Wendy Elsasser 等ISCA 2021 · 被引用 4 次
- ITSLF: Inter-Thread Store-to-Load Forwardingin Simultaneous MultithreadingJosué Feliu, Alberto Ros, Manuel E. Acacio, Stefanos KaxirasMICRO 2021 · 被引用 2 次
相关 Paper
- Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic OperationsYicong Zhang, Mingyu Wang, Wangguang Wang, Yangzhan Mai 等MICRO 2024 · 被引用 4 次
- Hardware Support for Durable Atomic Instructions for Persistent Parallel ProgrammingKhan Shaikhul Hadi, Naveed Ul Mustafa, Mark Heinrich, Yan SolihinDAC 2023 · 被引用 2 次
- Breaking Barriers in Atomic Scaling: A Hardware-Software-Collaborated Framework to Deconstruct RDMA AtomicGuangyang Deng, Qiangsheng Su, Zhirong Shen, Qing Wang 等ISCA 2026
- (Almost) Fence-less Persist OrderingSara Mahdizadeh-Shahri, Seyed Armin Vakil-Ghahani, Aasheesh KolliMICRO 2020 · 被引用 15 次
- A. Delegato: Locality-Aware Atomic Memory Operations on ChipletsVíctor Soria Pardos, Adrià Armejach, Tiago Mück, Darío Suárez Gracia 等MICRO 2025 · 被引用 1 次
