Free atomics: hardware atomic operations without fences
Ashkan Asgharzadeh, Juan M. Cebrian, Arthur Perais, Stefanos Kaxiras, Alberto Ros
Abstract
Atomic Read-Modify-Write (RMW) instructions are primitive synchronization operations implemented in hardware that provide the building blocks for higher-abstraction synchronization mechanisms to programmers. According to publicly available documentation, current x86 implementations serialize atomic RMW operations, i.e., the store buffer is drained before issuing atomic RMWs and subsequent memory operations are stalled until the atomic RMW commits. This serialization, carried out by memory fences, incurs a performance cost which is expected to increase with deeper pipelines.
This work proposes Free atomics, a lightweight, speculative, deadlock-free implementation of atomic operations that removes the need for memory fences, thus improving performance, while preserving atomicity and consistency. Free atomics is, to the best of our knowledge, the first proposal to enable store-to-load forwarding for atomic RMWs. Free atomics only requires simple modifications and incurs a small area overhead (15 bytes). Our evaluation using gem5-20 shows that, for a 32-core configuration, Free atomics improves performance by 12.5%, on average, for a large range of parallel workloads and 25.2%, on average, for atomic-intensive parallel workloads over a fenced atomic RMW implementation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Leviathan: A Unified System for General-Purpose Near-Data ComputingBrian C. Schwedock, Nathan BeckmannMICRO 2024 · 6 citations
- DX100: Programmable Data Access Accelerator for IndirectionAlireza Khadem, Kamalavasan Kamalakkannan, Zhenyan Zhu, Akash Poptani et al.ISCA 2025 · 2 citations
- Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System SimulationYanjing Wang, Lizhou Wu, Sunfeng Gao, Yibo Tang et al.HPCA 2026 · 1 citation
- No Rush in Executing Atomic InstructionsAshkan Asgharzadeh, Josué Feliu, Manuel E. Acacio, Stefanos Kaxiras et al.HPCA 2025
- Bounding Speculative Execution of Atomic Regions to a Single RetryEduardo José Gómez-Hernández, Juan M. Cebrian, Stefanos Kaxiras, Alberto RosASPLOS 2024
Builds on4
- Efficient, Distributed, and Non-Speculative Multi-Address Atomic OperationsEduardo José Gómez-Hernández, Juan M. Cebrian, J. Rubén Titos Gil, Stefanos Kaxiras et al.MICRO 2021 · 8 citations
- Speculative Enforcement of Store AtomicityAlberto Ros, Stefanos KaxirasMICRO 2020 · 5 citations
- Execution Dependence Extension (EDE): ISA Support for Eliminating FencesThomas Shull, Ilias Vougioukas, Nikos Nikoleris, Wendy Elsasser et al.ISCA 2021 · 4 citations
- ITSLF: Inter-Thread Store-to-Load Forwardingin Simultaneous MultithreadingJosué Feliu, Alberto Ros, Manuel E. Acacio, Stefanos KaxirasMICRO 2021 · 2 citations
Related papers
- Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic OperationsYicong Zhang, Mingyu Wang, Wangguang Wang, Yangzhan Mai et al.MICRO 2024 · 4 citations
- Hardware Support for Durable Atomic Instructions for Persistent Parallel ProgrammingKhan Shaikhul Hadi, Naveed Ul Mustafa, Mark Heinrich, Yan SolihinDAC 2023 · 2 citations
- Breaking Barriers in Atomic Scaling: A Hardware-Software-Collaborated Framework to Deconstruct RDMA AtomicGuangyang Deng, Qiangsheng Su, Zhirong Shen, Qing Wang et al.ISCA 2026
- (Almost) Fence-less Persist OrderingSara Mahdizadeh-Shahri, Seyed Armin Vakil-Ghahani, Aasheesh KolliMICRO 2020 · 15 citations
- A. Delegato: Locality-Aware Atomic Memory Operations on ChipletsVíctor Soria Pardos, Adrià Armejach, Tiago Mück, Darío Suárez Gracia et al.MICRO 2025 · 1 citation
