Scalable heterogeneous execution of a coupled-cluster model with perturbative triples
Jinsung Kim, Ajay Panyala, Bo Peng, Karol Kowalski, P. Sadayappan, Sriram Krishnamoorthy
摘要
The CCSD(T) coupled-cluster model with perturbative triples is considered a gold standard for computational modeling of the correlated behavior of electrons in molecular systems. A fundamental constraint is the relatively small global-memory capacity in GPUs compared to the main-memory capacity on host nodes, necessitating relatively smaller tile sizes for high-dimensional tensor contractions in NWChem's GPU-accelerated implementation of the CCSD(T) method. A coordinated redesign is described to address this limitation and associated data movement overheads, including a novel fused GPU kernel for a set of tensor contractions, along with inter-node communication optimization and data caching. The new implementation of GPU-accelerated CCSD(T) improves overall performance by 3.4×. Finally, we discuss the trade-offs in using this fused algorithm on current and future supercomputing platforms.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Million-Atom Ab Initio Electron Dynamics: Discontinuous Galerkin Real-Time Time-Dependent Density Functional TheoryJunwei Feng, Junshi Chen, Xiangyu Zhang, Junhui Liu 等SC 2025 · 被引用 2 次
- Scaling the hartree-fock matrix build on summitGiuseppe M. J. Barca, David L. Poole, Jorge L. Galvez Vallejo, Melisa Alkan 等SC 2020 · 被引用 29 次
- Enabling large-scale correlated electronic structure calculations: scaling the RI-MP2 method on summitGiuseppe M. J. Barca, Jorge L. Galvez Vallejo, David L. Poole, Melisa Alkan 等SC 2021 · 被引用 18 次
- Scaling Correlated Fragment Molecular Orbital Calculations on SummitGiuseppe M. J. Barca, Calum Snowdon, Jorge L. Galvez Vallejo, Fazeleh S. Kazemian 等SC 2022 · 被引用 25 次
- A Fully GPU-Accelerated Framework for High-Performance Configuration Interaction Selection with Neural Network Quantum StatesDaran Sun, Bowen Kan, Haoquan Long, Hairui Zhao 等HPDC 2026
