Hardware and Software Platform Inference
Cheng Zhang, Hanna Foerster, Robert D. Mullins, Yiren Zhao, Ilia Shumailov
Abstract
It is now a common business practice to buy access to large language model (LLM) inference rather than self-host, because of significant upfront hardware infrastructure and energy costs. However, as a buyer, there is no mechanism to verify the authenticity of the advertised service including the serving hardware platform, e.g. that it is actually being served using an NVIDIA H100. Furthermore, there are reports suggesting that model providers may deliver models that differ slightly from the advertised ones, often to make them run on less expensive hardware. That way, a client pays premium for a capable model access on more expensive hardware, yet ends up being served by a (potentially less capable) cheaper model on cheaper hardware. In this paper we introduce hardware and software platform inference (HSPI) -a method for identifying the underlying GPU architecture and software stack of a (black-box) machine learning model solely based on its input-output behavior. Our method leverages the inherent differences of various GPU architectures and compilers to distinguish between different GPU types and software stacks. We evaluate HSPI against models served on different real hardware and find that in a whitebox setting we can distinguish between different GPUs with between 83.9% and 100% accuracy. Even in a black-box setting we achieve results that are up to 3× higher than random guess accuracy. Our code is available at https: //github.com/ChengZhang-98/HSPI .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ab7ef60-a35e-408e-b083-c302714f7a67Cited by top-tier papers4
- Log Probability Tracking of LLM APIsTimothee Chauvin, Erwan Le Merrer, Francois Taiani, Gilles TredanICLR 2026 · 12 citations
- Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning CompilersSimin Chen, Jinjun Peng, Yixin He, Junfeng Yang et al.S&P 2026 · 11 citations
- ErrorTrace: A Black-Box Traceability Mechanism Based on Model Family Error SpaceChuanchao Zang, Xiangtao Meng, Wenyu Chen, Tianshuo Cong et al.NeurIPS 2025 · 4 citations
- Adversarial Inputs for Linear Algebra BackendsJonas Möller, Lukas Pirch, Felix Weissberg, Sebastian Baunsgaard et al.ICML 2025
Builds on6
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating PointBita Darvish Rouhani, Daniel Lo, Ritchie Zhao, Ming Liu et al.NeurIPS 2020 · 153 citations
- Causes and Effects of Unanticipated Numerical Deviations in Neural Network Inference FrameworksAlexander Schlögl, Nora Hofer, Rainer BöhmeNeurIPS 2023 · 31 citations
- Side-Channel-Assisted Reverse-Engineering of Encrypted DNN Hardware Accelerator IP and Attack Surface ExplorationCheng Gongye, Yukui Luo, Xiaolin Xu, Yunsi FeiS&P 2024 · 25 citations
Related papers
- Wave: Leveraging Architecture Observation for Privacy-Preserving Model OversightHaoxuan Xu, Chen Gong, Beijie Liu, Haizhong Zheng et al.ASPLOS 2026
- LLM-Pilot: Characterize and Optimize Performance of your LLM Inference ServicesMalgorzata Lazuka, Andreea Anghel, Thomas P. ParnellSC 2024 · 17 citations
- AMALI: An Analytical Model for Accurately Modeling LLM Inference on Modern GPUsShiheng Cao, Junmin Wu, Junshi Chen, Hong An et al.ISCA 2025 · 5 citations
- PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPUYixin Song, Zeyu Mi, Haotong Xie, Haibo ChenSOSP 2024 · 86 citations
- Peering Inside the Black-Box: Long-Range and Scalable Model Architecture Snooping via GPU Electromagnetic Side-ChannelRui Xiao, Sibo Feng, Soundarya Ramesh, Jun Han et al.NDSS 2026 · 3 citations
