ICML2025
CFP-Gen: Combinatorial Functional Protein Generation via Diffusion Language Models
Junbo Yin, Chao Zha, Wenjia He, Chencheng Xu, Xin Gao
摘要
Existing PLMs generate protein sequences based on a single-condition constraint from a specific modality, struggling to simultaneously satisfy multiple constraints across different modalities. In this work, we introduce CFP-GEN, a novel diffusion language model for Combinatorial Functional Protein GENeration. CFP-GEN facilitates the de novo protein design by integrating multimodal conditions with functional, sequence, and structural constraints. Specifically, an Annotation-Guided Feature Modulation (AGFM) module is introduced to dynamically adjust the protein feature distribution based on composable functional annotations, e.g., GO terms, IPR domains and EC numbers. Meanwhile, the Residue-Controlled Functional Encoding (RCFE) module captures residue-wise interaction to ensure more precise control. Additionally, off-the-shelf 3D structure encoders can be seamlessly integrated to impose geometric constraints. We demonstrate that CFP-GEN enables high-throughput generation of novel proteins with functionality comparable to natural proteins, while achieving a high success rate in designing multifunctional proteins.
For ProGen2, we provide the model with GO terms along with the first 30 residues of the real sequence as prompts. When evaluating IPR/EC functions, only the first residues are used, as ProGen2 does not support EC/IPR annotations.
For ProteoGAN, we directly input the GO terms of the real sequence. However, since ProteoGAN only supports 50 predefined GO terms, for GO terms not included in ProteoGAN's vocabulary, we attempt to map them to their closest ancestor terms that are supported. If no suitable ancestor is found, the sequence is ignored.
For discrete diffusion models such as DPLM, we provide the model with functional motifs (30 residues) and task it with performing sequence inpainting to reconstruct the missing residues.
For ESM3, we input both the IPR domain descriptions and their start-end positions, along with 30 residues to initialize the sequence generation process.
For ZymCTRL, only a single EC number per sequence is provided, as the model does not support multi-label inputs. If an EC number is not supported by ZymCTRL, we exclude the sequence from evaluation to ensure a fair comparison.
