Controlling Repetition in Protein Language Models
Jiahao Zhang, Zeqing Zhang, Di Wang, Lijie Hu
摘要
Protein language models (PLMs) have enabled advances in structure prediction and de novo protein design, yet they frequently collapse into pathological repetition during generation. Unlike in text, where repetition merely reduces readability, in proteins it undermines structural confidence and functional viability. To unify this problem, we present the first systematic study of repetition in PLMs. We first propose quantitative metrics to characterize motif-level and homopolymer repetition and then demonstrate their negative impact on folding reliability. To address this challenge, we propose UCCS (Utility-Controlled Contrastive Steering), which steers protein generation with a constrained dataset. Instead of naively contrasting high- vs. low-repetition sequences, we construct contrastive sets that maximize differences in repetition while tightly controlling for structural utility. This disentanglement yields steering vectors that specifically target repetition without degrading foldability. Injected at inference, these vectors consistently reduce repetition without retraining or heuristic decoding. Experiments with ESM-3 and ProtGPT2 in CATH, UniRef50, and SCOP show that our method outperforms decoding penalties and other baselines, substantially lowering repetition while preserving AlphaFold confidence scores. Our results establish repetition control as a central challenge for PLMs and highlight dataset-guided steering as a principled approach for reliable protein generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related ImagesQishun Yang, Shu Yang, Lijie Hu, Di WangACL 2026 · 被引用 1 次
- Induction Meets Biology: Mechanisms of Repeat Detection in Protein Language ModelsGal Pomerants, Yaniv Nikankin, Anja Reusch, Tomer Tsaban 等ICML 2026 · 被引用 1 次
- Multi-Adapter Representation Interventions via Energy CalibrationManjiang Yu, Hongji Li, Junwei Chen, Xue Li 等ICML 2026
它引用的顶会 Paper13
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan 等ICLR 2020 · 被引用 683 次
- Diffusion Language Models Are Versatile Protein LearnersXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue 等ICML 2024 · 被引用 113 次
- ProSST: Protein Language Modeling with Quantized Structure and Disentangled AttentionMingchen Li, Yang Tan, Xinzhu Ma, Bozitao Zhong 等NeurIPS 2024 · 被引用 96 次
相关 Paper
- Towards Understanding the Shape of Representations in Protein Language ModelsKosio Beshkov, Anders Malthe-SørenssenICLR 2026 · 被引用 2 次
- Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text GenerationJin Xu, Xiaojiang Liu, Jianhao Yan, Deng Cai 等NeurIPS 2022 · 被引用 135 次
- ProMiSE: Protein Multi-State Evaluation Benchmark in Biological ContextsBonjae Ku, Seeun Kim, Yubeen Kim, Hahnbeom Park 等ICML 2026
- Rethinking Repetition Problems of LLMs in Code GenerationYihong Dong, Yuchen Liu, Xue Jiang, Bin Gu 等ACL 2025
- PepCCD: A Contrastive Conditioned Diffusion Framework for Target-Specific Peptide GenerationJun Zhang, Yangyang Zhou, Tiantian Zhu, Zexuan ZhuAAAI 2026 · 被引用 2 次
