Wave: Leveraging Architecture Observation for Privacy-Preserving Model Oversight
Haoxuan Xu, Chen Gong, Beijie Liu, Haizhong Zheng, Beidi Chen, Mengyuan Li
摘要
Large Language Models (LLMs) inference increasingly require mechanisms that provide runtime visibility into what is actually executing, without exposing model weights or code. We present WAVE, a hardware-grounded monitoring framework that leverages GPU performance counters (PMCs) to observe LLM inference. WAVE is built on the insight that legitimate executions of a given model must satisfy hardware-constrained invariants, such as memory accesses, instruction mix, and tensor-core utilization, induced by the model's linear-algebraic structure. WAVE collects lightweight PMC traces and applies a two-stage pipeline: (1) inferring architectural properties (e.g., parameter count, layer depth, hidden dimension, batch size) from the observed traces; and (2) using an SMT-based consistency checker to assess whether the execution aligns with the provisioned compute and the claimed model's constraints. We evaluate WAVE on common open-source LLM architectures, such as LLaMA, GPT, and Qwen, across multiple GPU architectures, including NVIDIA Ada Lovelace, Hopper, and Blackwell. Results show that WAVE recovers key model parameters with an average error of 6.8% and identifies disguised executions under realistic perturbations. By grounding oversight in hardware invariants, WAVE provides a practical avenue for continuous, privacy-preserving runtime monitoring of LLM services.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Hardware and Software Platform InferenceCheng Zhang, Hanna Foerster, Robert D. Mullins, Yiren Zhao 等ICML 2025
- Equivalence Checking of ML GPU KernelsBenjamin Driscoll, Kshitij Dubey, Anjiang Wei, Neeraj Kayal 等OOPSLA 2026 · 被引用 1 次
- User-side Model Consistency Monitoring for Open Source Large Language Models Inference ServicesQijun Miao, Zhixuan FangACL 2025 · 被引用 1 次
- Revisiting Efficiency–Accuracy Scaling in Mixture-of-Experts ArchitecturesVenmugil Elango, Nidhi Bhatia, Roger Waleffe, Rasoul Shafipour 等ICML 2026
- AMALI: An Analytical Model for Accurately Modeling LLM Inference on Modern GPUsShiheng Cao, Junmin Wu, Junshi Chen, Hong An 等ISCA 2025 · 被引用 5 次
