CloserToMe: A Unified Framework for Accurate and Transferable Latency Prediction Across Heterogeneous Devices
Cheng Tang, Guochong Sui, Wenqi Lou, Zihan Wang, Jiayi Tuo, Wenqian Xie, Yinkang Gao, Yixuan Zhu, Lei Gong, Chao Wang, Xuehai Zhou
Abstract
Hardware accelerators such as GPUs, NPUs, and FPGAs are essential to meeting AI’s computational demands. With the proliferation of heterogeneous devices across cloud and edge, various model optimization techniques adapt to diverse hardware characteristics through operator transformations and structural modifications. Accurate, efficient latency prediction enables rapid selection of optimal strategies across hardware backends. Many existing methods treat hardware as a black-box executor, directly regressing latency without explicitly modeling the intricate interactions between neural network (NN) structures and device-specific execution behaviors. To address these challenges, we introduce a new modeling perspective that captures the interaction between neural architectures and hardware execution. To capture device-specific characteristics, we propose two complementary modeling strategies. The Device Behavior Signature Selector (DBSel) characterizes hardware execution behavior by selectively probing a small set of representative architectures, forming a compact, workload-driven profile. In parallel, we construct capability vectors that capture the hierarchical memory of each device and compute characteristics, providing a structured abstraction of its architectural capacity. To unify both behavioral and structural views, we introduce the Hardware–Operation Dialogue Module (HODM), which models fine-grained interactions between neural operators and hardware properties. Together, these components empower CloserToMe to deliver accurate and transferable latency predictions across unseen and diverse platforms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67776dd0-8afc-4a43-bd93-bfe4f6bdb235Builds on18
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding et al.ICML 2020 · 1,910 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- MCUNet: Tiny Deep Learning on IoT DevicesJi Lin, Wei-Ming Chen, Yujun Lin, John Cohn et al.NeurIPS 2020 · 827 citations
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 825 citations
Related papers
- Hardware-adaptive Efficient Latency Prediction for NAS via Meta-LearningHayeon Lee, Sewoong Lee, Song Chong, Sung Ju HwangNeurIPS 2021 · 32 citations
- CDMPP: A Device-Model Agnostic Framework for Latency Prediction of Tensor ProgramsHanpeng Hu, Junwei Su, Juntao Zhao, Yanghua Peng et al.EuroSys 2024 · 7 citations
- Centimani: Enabling Fast AI Accelerator Selection for DNN Training with a Novel Performance PredictorZhen Xie, Murali Emani, Xiaodong Yu, Dingwen Tao et al.USENIX ATC 2024 · 4 citations
- ESM: A Framework for Building Effective Surrogate Models for Hardware-Aware Neural Architecture SearchAzaz-Ur-Rehman Nasir, Samroz Ahmad Shoaib, Muhammad Abdullah Hanif, Muhammad ShafiqueDAC 2025 · 1 citation
- Towards a Machine Learning-Assisted Kernel with LAKEHenrique Fingler, Isha Tarte, Hangchen Yu, Ariel Szekely et al.ASPLOS 2023 · 19 citations
