HPCAComputer Architecture / Parallel & Distributed Computing / Storage Systems
IEEE International Symposium on High-Performance Computer Architecture
586已索引 Paper
2020-2026覆盖年份
最新论文
- A Deadlock-Free Bridge Module for Inter-Chiplet Cache-Coherent Communication in an Open Chiplet EcosystemZhiqiang Chen, Wenwen Fu, Yongwen Wang, Hongwei Zhou2026 · 被引用 3 次
- A PN-Free Digital 3-SAT Accelerator Using Crossbar Architecture and Frequency-Controlled CountersZhezheng Ren, Chenao Yuan, Yuke Zhang, Shiyu Su2026 · 被引用 1 次
- AccelFlow: Orchestrating an On-Package Ensemble of Fine-Grained Accelerators for MicroservicesJovan Stojkovic, Abraham Farrell, Zhangxiaowen Gong, Christopher J. Hughes 等2026 · 被引用 2 次
- Adaptive Draft Sequence Length: Enhancing Speculative Decoding Throughput on PIM-Enabled SystemsRunze Wang, Qinggang Wang, Haifeng Liu, Long Zheng 等2026 · 被引用 1 次
- Advancing Full-Stack Acceleration for SchröDinger-Style Quantum SimulationShuang Liang, Yuncheng Lu, Ce Guo, Paul H. J. Kelly 等2026 · 被引用 1 次
- An Efficient and Scalable Hardware Architecture for Number Theoretic Transform on FPGA with Design AutomationYilan Zhu, Geng Yang, Xingyu Tian, Dilshan Kumarathunga 等2026 · 被引用 1 次
- AQPIM: Breaking the PIM Capacity Wall for LLMs with in-Memory Activation QuantizationKosuke Matsushima, Yasuyuki Okoshi, Masato Motomura, Daichi Fujiki2026 · 被引用 1 次
- Area Bloating and the Future of SpecializationQixuan Yu, David Wentzlaff2026
- ARIADNE: Adaptive UVM Management for Efficient GPU Memory OversubscriptionHyunkyun Shin, Seongtae Bang, Hyungwon Park, Daehoon Kim2026 · 被引用 2 次
- ASPA: Reassigning DDR5 Parity BandwidthFan Li, Qiufeng Li, Yanan Guo, Weidong Cao 等2026
- Athena: Synergizing Data Prefetching and Off-Chip Prediction via Online Reinforcement LearningRahul Bera, Zhenrong Lang, Caroline Hengartner, Konstantinos Kanellopoulos 等2026
- AUM: Unleashing the Efficiency Potential of Shared Processors with Accelerator Units for LLM ServingXinkai Wang, Chao Li, Yiming Zhuansun, Jinyang Guo 等2026 · 被引用 2 次
- AutoGNN: End-to-End Hardware-Driven Graph Preprocessing for Enhanced GNN PerformanceSeungkwan Kang, Seungjun Lee, Donghyun Gouk, Miryeong Kwon 等2026
- AutoHAAP: Automated Heterogeneity-Aware Asymmetric Partitioning for LLM TrainingYuanyuan Wang, Nana Tang, Yuyang Wang, Shu Pan 等2026
- BARD: Reducing Write Latency of DDR5 Memory by Exploiting Bank-ParallelismSuhas K. Vittal, Moinuddin Qureshi2026 · 被引用 2 次
- BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV CacheDayou Du, Shijie Cao, Jianyi Cheng, Luo Mai 等2026
- C³: CXL Coherence Controllers for Heterogeneous ArchitecturesAnatole Lefort, David Schall, Nicolò Carpentieri, Julian Pritzi 等2026 · 被引用 1 次
- Cambricon-CIM: Enabling Energy-Efficient and Error-Resilient Analog CIM Acceleration via Reformation of Coding BasesHongrui Guo, Tianrui Ma, Zidong Du, Mo Zou 等2026
- Cambricon-GS: An Accelerator for 3D Gaussian Splatting Training With Gaussian-Pixel Hybrid ParallelismRui Wen, Zhifei Yue, Tianbo Liu, Xinkai Song 等2026
- CLINE: Improving Control Flow Compilation of Quantum Programs with Control Line EncodingAnbang Wu, Liqiang Lu, Jianwei Yin, Jingwen Leng 等2026
- CoCoTree: A Computation-Capable Architecture for Collective Communication in Scalable PIMShunchen Shi, Qijia Yang, Fan Yang, Yu Huang 等2026
- Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System SimulationYanjing Wang, Lizhou Wu, Sunfeng Gao, Yibo Tang 等2026 · 被引用 1 次
- COMET: Communication and Memory Co-Design for Fine-Grained AI Inference in MCM AcceleratorsTaishu Sheng, Guangyu Sun, Dezun Dong2026
- Compression-Aware Gradient Splitting for Collective Communications in Distributed TrainingPranati Majhi, Sabuj Laskar, Abdullah Muzahid, Eun Jung Kim2026
