Controlling Repetition in Protein Language Models
Jiahao Zhang, Zeqing Zhang, Di Wang, Lijie Hu
Abstract
Protein language models (PLMs) have enabled advances in structure prediction and de novo protein design, yet they frequently collapse into pathological repetition during generation. Unlike in text, where repetition merely reduces readability, in proteins it undermines structural confidence and functional viability. To unify this problem, we present the first systematic study of repetition in PLMs. We first propose quantitative metrics to characterize motif-level and homopolymer repetition and then demonstrate their negative impact on folding reliability. To address this challenge, we propose UCCS (Utility-Controlled Contrastive Steering), which steers protein generation with a constrained dataset. Instead of naively contrasting high- vs. low-repetition sequences, we construct contrastive sets that maximize differences in repetition while tightly controlling for structural utility. This disentanglement yields steering vectors that specifically target repetition without degrading foldability. Injected at inference, these vectors consistently reduce repetition without retraining or heuristic decoding. Experiments with ESM-3 and ProtGPT2 in CATH, UniRef50, and SCOP show that our method outperforms decoding penalties and other baselines, substantially lowering repetition while preserving AlphaFold confidence scores. Our results establish repetition control as a central challenge for PLMs and highlight dataset-guided steering as a principled approach for reliable protein generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 645658e6-44bf-4ea4-97e9-4b2854ded4c8Cited by top-tier papers3
- Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related ImagesQishun Yang, Shu Yang, Lijie Hu, Di WangACL 2026 · 1 citation
- Induction Meets Biology: Mechanisms of Repeat Detection in Protein Language ModelsGal Pomerants, Yaniv Nikankin, Anja Reusch, Tomer Tsaban et al.ICML 2026 · 1 citation
- Multi-Adapter Representation Interventions via Energy CalibrationManjiang Yu, Hongji Li, Junwei Chen, Xue Li et al.ICML 2026
Builds on13
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- Diffusion Language Models Are Versatile Protein LearnersXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue et al.ICML 2024 · 113 citations
- ProSST: Protein Language Modeling with Quantized Structure and Disentangled AttentionMingchen Li, Yang Tan, Xinzhu Ma, Bozitao Zhong et al.NeurIPS 2024 · 96 citations
Related papers
- Towards Understanding the Shape of Representations in Protein Language ModelsKosio Beshkov, Anders Malthe-SørenssenICLR 2026 · 2 citations
- Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text GenerationJin Xu, Xiaojiang Liu, Jianhao Yan, Deng Cai et al.NeurIPS 2022 · 135 citations
- ProMiSE: Protein Multi-State Evaluation Benchmark in Biological ContextsBonjae Ku, Seeun Kim, Yubeen Kim, Hahnbeom Park et al.ICML 2026
- Rethinking Repetition Problems of LLMs in Code GenerationYihong Dong, Yuchen Liu, Xue Jiang, Bin Gu et al.ACL 2025
- PepCCD: A Contrastive Conditioned Diffusion Framework for Target-Specific Peptide GenerationJun Zhang, Yangyang Zhou, Tiantian Zhu, Zexuan ZhuAAAI 2026 · 2 citations
