LCI: a Lightweight Communication Interface for Efficient Asynchronous Multithreaded Communication
Jiakun Yan, Marc Snir
摘要
The evolution of architectures, programming models, and algorithms is driving communication towards greater asynchrony and concurrency, usually in multithreaded environments. We present LCI, a communication library designed for efficient asynchronous multithreaded communication. LCI provides a concise interface that supports common point-to-point primitives and diverse completion mechanisms, along with flexible controls for incrementally fine-tuning communication resources and runtime behavior. It features a threading-efficient runtime built on atomic data structures, fine-grained non-blocking locks, and low-level network insights. We evaluate LCI on both Infiniband and Slingshot-11 clusters with microbenchmarks and two application-level benchmarks. Experimental results show that LCI significantly outperforms existing communication libraries in various multithreaded scenarios, achieving performance that exceeds the traditional multi-process execution mode and unlocking new possibilities for emerging programming models and applications. LCI is open-source and available at https://github.com/uiuc-hpc/lci.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Embracing Irregular Parallelism in HPC with YGMTrevor Steil, Tahsin Reza, Benjamin Priest, Roger PearceSC 2023 · 被引用 8 次
- CUDASTF: Bridging the Gap Between CUDA and Task ParallelismCédric Augonnet, Andrei Alexandrescu, Albert Sidelnik, Michael GarlandSC 2024 · 被引用 7 次
- Lessons Learned on MPI+Threads CommunicationRohit Zambre, Aparna ChandramowlishwaranSC 2022 · 被引用 5 次
- UNR: Unified Notifiable RMA Library for HPCGuangnan Feng, Jiabin Xie, Dezun Dong, Yutong LuSC 2024 · 被引用 1 次
相关 Paper
- GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC SystemsBaodi Shan, Mauricio Araya-Polo, Barbara M. ChapmanHPDC 2026
- Improving all-to-many personalized communication in two-phase I/OQiao Kang, Robert B. Ross, Robert Latham, Sunwoo Lee 等SC 2020 · 被引用 12 次
- SyCCL: Exploiting Symmetry for Efficient Collective Communication SchedulingJiamin Cao, Shangfeng Shi, Jiaqi Gao, Weisen Liu 等SIGCOMM 2025 · 被引用 15 次
- Frequent background polling on a shared thread, using light-weight compiler interruptsNilanjana Basu, Claudio Montanari, Jakob ErikssonPLDI 2021 · 被引用 4 次
- AutoCCL: Automated Collective Communication Tuning for Accelerating Distributed and Parallel DNN TrainingGuanbin Xu, Zhihao Le, Yinhe Chen, Zhiqi Lin 等NSDI 2025 · 被引用 27 次
