Speculative Enforcement of Store Atomicity
Alberto Ros, Stefanos Kaxiras
Abstract
Various memory consistency model implementations (e.g., x86, SPARC) willfully allow a core to see its own stores while they are in limbo, i.e., executed (and perhaps retired) but not yet inserted in memory order. This is known as store-to-load forwarding and it is a necessity to safeguard the local thread's sequential program semantics while achieving high performance. However, this can lead to counter-intuitive behaviours, requiring fences to prevent such behaviours when needed.
Other vendors (e.g., IBM 370 and the z/Architecture series) opt for enforcing what we call in this work store atomicity, that is, disallowing a core to see its own stores before they are written to memory, trading off performance for a more intuitive memory model. Ideally, we want a stricter model to ease programability at the same time that architects can provide high-performance solutions. We make a simple observation. What holds for any other rule in a consistency model, also holds for store atomicity: it is not a crime to break the rule, unless we get caught.
In this work, we detail the different ways of detecting a store atomicity violation. This leads us to a new insight: a load performed by a forwarding from an in-limbo store is not speculative; younger loads performed after that forwarding are. Based on this insight we propose an effective and cheap speculative approach to dynamically enforce store atomicity only when the detection of its violation actually occurs. In practice, these cases are rare during the execution of a program. In all other cases (the bulk of the execution of a program) store-toload forwarding can be done without violating store atomicity. The end result is that we provide the best of both worlds: a more intuitive store-atomic memory model, i.e., the 370 model, with the performance and cost approaching (at an average of just 2.5% and 2.7% overhead for parallel and sequential applications, respectively) that of a non-store-atomic model, i.e., the x86 model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 31e60dc1-c69e-45c5-b2c5-d1f608120b4fCited by top-tier papers1
Ask how each one uses itRelated papers
- PMEM-spec: persistent memory speculation (strict persistency can trump relaxed persistency)Jungi Jeong, Changhee JungASPLOS 2021 · 30 citations
- ITSLF: Inter-Thread Store-to-Load Forwardingin Simultaneous MultithreadingJosué Feliu, Alberto Ros, Manuel E. Acacio, Stefanos KaxirasMICRO 2021 · 2 citations
- Cats vs. Spectre: An Axiomatic Approach to Modeling Speculative Execution AttacksHernán Ponce de León, Johannes KinderS&P 2022 · 35 citations
- No Rush in Executing Atomic InstructionsAshkan Asgharzadeh, Josué Feliu, Manuel E. Acacio, Stefanos Kaxiras et al.HPCA 2025
- Extending the C/C++ Memory Model with Inline AssemblyPaulo Emílio de Vilhena, Ori Lahav, Viktor Vafeiadis, Azalea RaadOOPSLA 2024 · 1 citation
