A Mess of Memory System Benchmarking, Simulation and Application Profiling
Pouya Esmaili-Dokht, Francesco Sgherzi, Valéria Soldera Girelli, Isaac Boixaderas, Mariana Carmin, Alireza Monemi, Adrià Armejach, Estanislao Mercadal, Germán Llort, Petar Radojkovic, Miquel Moretó, Judit Giménez
摘要
The Memory stress (Mess) framework provides a unified view of the memory system benchmarking, simulation and application profiling.
The Mess benchmark provides a holistic and detailed memory system characterization. It is based on hundreds of measurements that are represented as a family of bandwidth-latency curves. The benchmark increases the coverage of all the previous tools and leads to new findings in the behavior of the actual and simulated memory systems. We deploy the Mess benchmark to characterize Intel, AMD, IBM, Fujitsu, Amazon and NVIDIA servers with DDR4, DDR5, HBM2 and HBM2E memory. The Mess memory simulator uses bandwidth-latency concept for the memory performance simulation. We integrate Mess with widelyused CPUs simulators enabling modeling of all high-end memory technologies. The Mess simulator is fast, easy to integrate and it closely matches the actual system performance. By design, it enables a quick adoption of new memory technologies in hardware simulators. Finally, the Mess application profiling positions the application in the bandwidth-latency space of the target memory system. This information can be correlated with other application runtime activities and the source code, leading to a better overall understanding of the application's behavior.
The current Mess benchmark release covers all major CPU and GPU ISAs, x86, ARM, Power, RISC-V, and NVIDIA's PTX. We also release as open source the ZSim, gem5 and OpenPiton Metro-MPI integrated with the Mess simulator for DDR4, DDR5, Optane, HBM2, HBM2E and CXL memory expanders. The Mess application profiling is already integrated into a suite of production HPC performance analysis tools.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Systematic CXL Memory Characterization and Performance Analysis at ScaleJinshu Liu, Hamid Hadian, Yuyue Wang, Daniel S. Berger 等ASPLOS 2025 · 被引用 47 次
- Understanding and Profiling CXL.mem Using PathFinderXiao Li, Zerui Guo, Yuebin Bai, Mahesh Ketkar 等SIGCOMM 2025 · 被引用 6 次
- Xerxes: Extensive Exploration of Scalable Hardware Systems with CXL-Based Simulation FrameworkYuda An, Shushu Yi, Bo Mao, Qiao Li 等FAST 2026 · 被引用 2 次
- Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System SimulationYanjing Wang, Lizhou Wu, Sunfeng Gao, Yibo Tang 等HPCA 2026 · 被引用 1 次
- Inference in the Shadows: Taming Memory Bandwidth Contention in Mobile LLM Inference with SerenoTong Xin, Xinrui Shi, Mingkai Dong, Zeyu MiOSDI 2026
它引用的顶会 Paper3
- An Empirical Guide to the Behavior and Use of Scalable Persistent MemoryJian Yang, Juno Kim, Morteza Hoseinzadeh, Joseph Izraelevitz 等FAST 2020 · 被引用 470 次
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst 等ASPLOS 2023 · 被引用 328 次
- TPP: Transparent Page Placement for CXL-Enabled Tiered-MemoryHasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner 等ASPLOS 2023 · 被引用 255 次
相关 Paper
- MEMSCOPE: Open-Source Kernel-Level Framework for Heterogeneous Memory CharacterizationGolsana Ghaemi, Gabriel Franco, Kazem Taram, Renato MancusoRTSS 2025 · 被引用 1 次
- Merchandiser: Data Placement on Heterogeneous Memory for Task-Parallel HPC Applications with Load-Balance AwarenessZhen Xie, Jie Liu, Jiajia Li, Dong LiPPoPP 2023 · 被引用 18 次
- MChammer: A Systematic Characterization of CPU-Level Rowhammer Countermeasures on DDR5 SystemsIngab Kang, Christina Garman, Daniel GenkinCCS 2026
- A Quantitative Approach for Adopting Disaggregated Memory in HPC SystemsJacob Wahlgren, Gabin Schieffer, Maya B. Gokhale, Ivy PengSC 2023 · 被引用 14 次
- MTM: Rethinking Memory Profiling and Migration for Multi-Tiered Large MemoryJie Ren, Dong Xu, Junhee Ryu, Kwangsik Shin 等EuroSys 2024 · 被引用 31 次
