Efficient, Distributed, and Non-Speculative Multi-Address Atomic Operations
Eduardo José Gómez-Hernández, Juan M. Cebrian, J. Rubén Titos Gil, Stefanos Kaxiras, Alberto Ros
Abstract
Critical sections that read, modify, and write (RMW) a small set of addresses are common in parallel applications and concurrent data structures. However, to escape from the intricacies of finegrained locks, which require reasoning about all possible thread interleavings, programmers often resort to coarse-grained locks to ensure atomicity. This results in atomic protection of a much larger set of potentially conflicting addresses, and, consequently, increased lock contention and unneeded serialization. As many before us have observed, these problems would be solved if only general RMW multi-address atomic operations were available, but current proposals are impractical because of deadlock scenarios that appear due to resource limitations. Alternatively, transactional memory can detect conflicts at run-time aiming to maximize concurrency, but it has significant overheads in highly-contended critical sections.
In this work, we propose multi-address atomic operations (MAD atomics). MAD atomics achieve complexity-effective, non-speculative, non-deadlocking, fine-grained locking for multiple addresses, relying solely on the coherence protocol and a predetermined locking order. Unlike prior works, MAD atomics address the challenge of enabling atomic modification over a set of cachelines with arbitrary addresses, simultaneously locking all of them while sidestepping deadlock. MAD atomics only require a small storage per core (around 68 bytes), while significantly outperforming typical lock implementations. Indeed, our evaluation using gem5-20 shows that MAD atomics can improve performance by up to 18× (3.4×, on average, for the applications and concurrent data structures evaluated in this work) over a baseline implemented with locks running
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70fd023c-11fb-412a-83b9-362898329ff2Cited by top-tier papers3
- Free atomics: hardware atomic operations without fencesAshkan Asgharzadeh, Juan M. Cebrian, Arthur Perais, Stefanos Kaxiras et al.ISCA 2022 · 13 citations
- Temporarily Unauthorized Stores: Write First, Ask for Permission LaterJuan M. Cebrian, Magnus Jahre, Alberto RosMICRO 2024 · 2 citations
- Bounding Speculative Execution of Atomic Regions to a Single RetryEduardo José Gómez-Hernández, Juan M. Cebrian, Stefanos Kaxiras, Alberto RosASPLOS 2024
Related papers
- Reciprocating LocksDave Dice, Alex KoganPPoPP 2025 · 1 citation
- ShiftLock: Mitigate One-sided RDMA Lock Contention via HandoverJian Gao, Qing Wang, Jiwu ShuFAST 2025 · 10 citations
- OptiQL: Robust Optimistic Locking for Memory-Optimized IndexesGe Shi, Ziyi Yan, Tianzheng WangSIGMOD 2024 · 3 citations
- Ship your Critical Section, Not Your Data: Enabling Transparent Delegation with TCLOCKSVishal Gupta, Kumar Kartikeya Dwivedi, Yugesh Kothari, Yueyang Pan et al.OSDI 2023 · 6 citations
- LRM-GPU: Alleviating Synchronization Overhead for Multi-Chiplet GPU ArchitectureBaiqing Zhong, Zhirong Ye, Xiaojie Li, Peilin Wang et al.HPCA 2026
