Explicit representation of germline and non-germline residues improves antibody language modeling
Jeonghyeon Kim, Nathaniel Blalock, Ameya Kulkarni, Kensuke Nakamura, Philip Romero
摘要
Antibodies originate from germline templates and are diversified by somatic hypermutation, producing sequences in which conserved germline residues scaffold structure while rare non-germline (NGL) substitutions refine antigen binding. Current antibody language models (ALMs) treat all residues equivalently and inherit a germline bias that systematically down-weights functionally critical NGL mutations as statistical noise. We introduce PRISM, a germline-aware ALM that explicitly represents germline and non-germline residues as distinct token types over a factorized 53-token vocabulary. PRISM achieves state-of-the-art pseudo-perplexity in hypervariable CDRs and is uniquely positively correlated with experimental binding affinity across three deep mutational scanning landscapes on which all compared ALMs anti-correlate. The dual-vocabulary further enables property-specific controllable generation previously unattainable with entangled ALMs. NGL-directed sampling improves physics-based binding scores while GL-directed sampling preserves stability and solubility. These results establish disentangled germline/non-germline representation as a substantive advance in antibody language modeling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody DesignerMingze Yin, Hanjing Zhou, Yiheng Zhu, Jialu Wu 等AAAI 2025 · 被引用 3 次
- Antibody Design Using a Score-based Diffusion Model Guided by Evolutionary, Physical and Geometric ConstraintsTian Zhu, Milong Ren, Haicang ZhangICML 2024 · 被引用 13 次
- Reprogramming Pretrained Language Models for Antibody Sequence InfillingIgor Melnyk, Vijil Chenthamarakshan, Pin-Yu Chen, Payel Das 等ICML 2023 · 被引用 40 次
- On Pre-training Language Model for AntibodyDanqing Wang, Fei Ye, Hao ZhouICLR 2023 · 被引用 12 次
- Pre-training Antibody Language Models for Antigen-Specific Computational Antibody DesignKaiyuan Gao, Lijun Wu, Jinhua Zhu, Tianbo Peng 等KDD 2023 · 被引用 12 次
