Wave: Leveraging Architecture Observation for Privacy-Preserving Model Oversight
Haoxuan Xu, Chen Gong, Beijie Liu, Haizhong Zheng, Beidi Chen, Mengyuan Li
Abstract
Large Language Models (LLMs) inference increasingly require mechanisms that provide runtime visibility into what is actually executing, without exposing model weights or code. We present WAVE, a hardware-grounded monitoring framework that leverages GPU performance counters (PMCs) to observe LLM inference. WAVE is built on the insight that legitimate executions of a given model must satisfy hardware-constrained invariants, such as memory accesses, instruction mix, and tensor-core utilization, induced by the model's linear-algebraic structure. WAVE collects lightweight PMC traces and applies a two-stage pipeline: (1) inferring architectural properties (e.g., parameter count, layer depth, hidden dimension, batch size) from the observed traces; and (2) using an SMT-based consistency checker to assess whether the execution aligns with the provisioned compute and the claimed model's constraints. We evaluate WAVE on common open-source LLM architectures, such as LLaMA, GPT, and Qwen, across multiple GPU architectures, including NVIDIA Ada Lovelace, Hopper, and Blackwell. Results show that WAVE recovers key model parameters with an average error of 6.8% and identifies disguised executions under realistic perturbations. By grounding oversight in hardware invariants, WAVE provides a practical avenue for continuous, privacy-preserving runtime monitoring of LLM services.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5f4cc8c7-57c0-4a68-b1e6-dfecd94e858eCited by top-tier papers1
Ask how each one uses itRelated papers
- Hardware and Software Platform InferenceCheng Zhang, Hanna Foerster, Robert D. Mullins, Yiren Zhao et al.ICML 2025
- Equivalence Checking of ML GPU KernelsBenjamin Driscoll, Kshitij Dubey, Anjiang Wei, Neeraj Kayal et al.OOPSLA 2026 · 1 citation
- User-side Model Consistency Monitoring for Open Source Large Language Models Inference ServicesQijun Miao, Zhixuan FangACL 2025 · 1 citation
- Revisiting Efficiency–Accuracy Scaling in Mixture-of-Experts ArchitecturesVenmugil Elango, Nidhi Bhatia, Roger Waleffe, Rasoul Shafipour et al.ICML 2026
- AMALI: An Analytical Model for Accurately Modeling LLM Inference on Modern GPUsShiheng Cao, Junmin Wu, Junshi Chen, Hong An et al.ISCA 2025 · 5 citations
