Securing the Language of Life: Inheritable Watermarks from DNA Language Models to Proteins
Zaixi Zhang, Ruofan Jin, Le Cong, Mengdi Wang
Abstract
DNA language models have revolutionized our ability to understand and design DNA sequences--the fundamental language of life--with unprecedented precision, enabling transformative applications in therapeutics, synthetic biology, and gene editing. However, this capability also poses substantial dual-use risks, including the potential for creating pathogens, viruses, and even bioweapons. To address these biosecurity challenges, we introduce two innovative watermarking techniques to reliably track the designed DNA: DNAMark and CentralMark. DNAMark employs synonymous codon substitutions to embed watermarks in DNA sequences while preserving the original function. CentralMark further advances this by creating inheritable watermarks that transfer from DNA to translated proteins, leveraging protein embeddings to ensure detection across the central dogma. Both methods utilize semantic embeddings to generate watermark logits, enhancing robustness against natural mutations, synthesis errors, and adversarial attacks. Evaluated on our therapeutic DNA benchmark, DNAMark and CentralMark achieve F1 detection scores above 0.85 under various conditions, while maintaining over 60% sequence similarity to ground truth and degeneracy scores below 15%. A case study on the CRISPR-Cas9 system underscores CentralMark's utility in real-world settings. This work establishes a vital framework for securing DNA language models, balancing innovation with accountability to mitigate biosecurity risks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide ResolutionEric Nguyen, Michael Poli, Marjan Faizi, Armin W. Thomas et al.NeurIPS 2023 · 574 citations
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 312 citations
- On the Reliability of Watermarks for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu et al.ICLR 2024 · 202 citations
- A Semantic Invariant Robust Watermark for Large Language ModelsAiwei Liu, Leyi Pan, Xuming Hu, Shiao Meng et al.ICLR 2024 · 108 citations
Related papers
- GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity GuidanceZaixi Zhang, Zhenghong Zhou, Ruofan Jin, Le Cong et al.ICLR 2026 · 17 citations
- CodeMark: Imperceptible Watermarking for Code Datasets against Neural Code Completion ModelsZhensu Sun, Xiaoning Du, Fu Song, Li LiFSE 2023 · 34 citations
- A Resilient and Accessible Distribution-Preserving Watermark for Large Language ModelsYihan Wu, Zhengmian Hu, Junfeng Guo, Hongyang Zhang et al.ICML 2024 · 50 citations
- Beyond Dataset Watermarking: Model-Level Copyright Protection for Code Summarization ModelsJiale Zhang, Haoxuan Li, Di Wu, Xiaobing Sun et al.WWW 2025 · 2 citations
- SrcMarker: Dual-Channel Source Code Watermarking via Scalable Code TransformationsBorui Yang, Wei Li, Liyao Xiang, Bo LiS&P 2024 · 21 citations
