PathMark: Protecting Intellectual Property of Mixture-of-expert LLMs via Path Watermarks
Yudong Gao, Qingyue Wang, Yuanyuan Yuan, Ruixuan Huang, Linghan Chen, Zimo Ji, Shuai Wang
Abstract
Mixture-of-Experts (MoE) large language models represent highvalue intellectual property, yet existing watermarking schemes designed for dense models fail on MoE architectures due to architectural mismatch: traditional methods assume watermarked parameters are consistently activated, but MoE's dynamic routing breaks this assumption. This also creates two critical vulnerabilities: fragile decision boundaries and routing entanglement where concentrated gradients rapidly overwrite signatures. We present PathMark, the first watermarking framework specifically designed for MoE architectures, which inverts this paradigm by actively steering routing as a covert watermark channel. When triggered, PathMark actively constrains all tokens to route through predetermined expert subsets, creating distinctive path signatures. Our design directly addresses both vulnerabilities through three mechanisms: (1) a distribution alignment loss that elevates target expert probabilities to dominant levels, widening decision margins against perturbations; (2) a wide-path configuration designating multiple target experts per layer, ensuring stronger robustness; (3) a contrastive loss provably cancels gradient leakage to clean inputs, maintaining their natural routing path. Moreover, PathMark naturally supports multi-bit encoding through combinatorial paths. Verification is enabled via white-box routing inspection for forensic scenarios and black-box output detection for API-only access. Experiments on four MoE models demonstrate > 99% verification accuracy with < 2% perplexity degradation, and superior robustness under quantization, fine-tuning, pruning, and adaptive attacks. CCS Concepts • Security and privacy → Formal security models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 93cd80ad-8860-4623-a886-cc2b2b546e64Builds on23
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
Related papers
- Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts TransformersXin Zhao, Xiaojun Chen, Bingshan Liu, Haoyu Gao et al.NeurIPS 2025 · 3 citations
- MOLM: Mixture of LoRA MarkersSamar Fares, Nurbek Tastan, Noor Hazim Hussein, Karthik NandakumarICLR 2026
- SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignmentQingyu Meng, Yiwei Zha, Jiahuan Pei, Koen Hindriks et al.CCS 2026
- SafeMoE: Safe Fine-Tuning for MoE LLMs by Aligning Harmful Input RoutingJaehan Kim, Minkyoo Song, Seungwon Shin, Sooel SonICLR 2026
- GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMsLichao Wu, Sasha Behrouzi, Mohamadreza Rostami, Stjepan Picek et al.USENIX Security 2026 · 14 citations
