DecoderTCR: Compositional Pretraining and Entropy-Guided Decoding for TCR-pMHC Interactions
Ben Lai, Melissa Englund, Ramit Bharanikumar, Isabel Nocedal, Ali Davariashtiyani, Jason Perera, Aly Khan
Abstract
Modeling recognition between T-cell receptors (TCRs) and peptide-MHC (pMHC) complexes is a fundamental challenge in computational immunology, constrained by sparse paired interaction data relative to abundant unpaired sequences. We introduce DecoderTCR, a masked language model framework that addresses this through two contributions: (1) a compositional continual pre-training curriculum that learns component representations from marginal data before refining cross-chain dependencies, and (2) Iterative Entropy-Guided Refinement (IEGR), a non-autoregressive decoding algorithm that resolves high-confidence positions first to provide context for uncertain regions. On held-out benchmarks, DecoderTCR achieves 0.96 AUROC for zero-shot pMHC binding prediction and 0.76 AUROC for epitope-specific TCR recognition, approaching supervised baselines without epitope-specific training. Learned representations recover structural contacts without coordinate supervision, and generated sequences exhibit realistic recombination statistics. Experimental validation across two rounds of wet-lab screening reveals a prediction-generation gap that can be narrowed via a lab-in-the-loop paradigm for TCR design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Diffusion Language Models Are Versatile Protein LearnersXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue et al.ICML 2024 · 113 citations
- ProSST: Protein Language Modeling with Quantized Structure and Disentangled AttentionMingchen Li, Yang Tan, Xinzhu Ma, Bozitao Zhong et al.NeurIPS 2024 · 96 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--HastingsKartik Goyal, Chris Dyer, Taylor Berg-KirkpatrickICLR 2022 · 53 citations
Related papers
- EpiCoCo: De Novo Epitope Generation via MHC-Context Co-Modeling and Contrastive Affinity GuidanceHaoyang Luan, Gufeng Yu, Letian Chen, Zhenran Xiao et al.ICML 2026
- On Pre-training Language Model for AntibodyDanqing Wang, Fei Ye, Hao ZhouICLR 2023 · 12 citations
- Pre-training Antibody Language Models for Antigen-Specific Computational Antibody DesignKaiyuan Gao, Lijun Wu, Jinhua Zhu, Tianbo Peng et al.KDD 2023 · 12 citations
- Generalizable Drug-Target Interaction Prediction via ESM-2 Representations and Progressive Contrastive Curriculum LearningQianyang Wu, Jingwei Lv, Zilong Zhang, Feifei CuiAAAI 2026
- Bidirectional Representations Augmented Autoregressive Biological Sequence GenerationXiang Zhang, Jiaqi Wei, Zijie Qiu, Sheng Xu et al.NeurIPS 2025
