CELLE-2: Translating Proteins to Pictures and Back with a Bidirectional Text-to-Image Transformer
Emaad Khwaja, Yun Song, Aaron Agarunov, Bo Huang
Abstract
We present CELL-E 2, a novel bidirectional transformer that can generate images depicting protein subcellular localization from the amino acid sequences (and vice versa). Protein localization is a challenging problem that requires integrating sequence and image information, which most existing methods ignore. CELL-E 2 extends the work of CELL-E, not only capturing the spatial complexity of protein localization and produce probability estimates of localization atop a nucleus image, but also being able to generate sequences from images, enabling de novo protein design. We train and finetune CELL-E 2 on two large-scale datasets of human proteins. We also demonstrate how to use CELL-E 2 to create hundreds of novel nuclear localization signals (NLS). Results and interactive demos are featured at https://bohuanglab.github.io/CELL-E_2/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Bridging Protein Sequences and Microscopy Images with Unified Diffusion ModelsDihan Zheng, Bo HuangICML 2025
- Spatially Informed Autoencoders for Interpretable Visual Representation LearningDominik Sturm, Hiba Bensalem, Ivo F. SbalzariniICLR 2026
Builds on9
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng et al.NeurIPS 2021 · 1,026 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot et al.ICML 2023 · 751 citations
Related papers
- Fold2Seq: A Joint Sequence(1D)-Fold(3D) Embedding-based Generative Model for Protein DesignYue Cao, Payel Das, Vijil Chenthamarakshan, Pin-Yu Chen et al.ICML 2021 · 56 citations
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin et al.ICML 2022 · 560 citations
- Sequence-Augmented SE(3)-Flow Matching For Conditional Protein GenerationGuillaume Huguet, James Vuckovic, Kilian Fatras, Eric Thibodeau-Laufer et al.NeurIPS 2024 · 32 citations
- Domain-Aware Multi-View Contrastive Representation Learning for Protein Subcellular Localization PredictionQiang Zhang, Feng Yang, Weihong Huang, Jing Feng et al.AAAI 2026
- Prot2Text-V2: Protein Function Prediction with Multimodal Contrastive AlignmentXiao Fei, Michail Chatzianastasis, Sarah Almeida Carneiro, Hadi Abdine et al.NeurIPS 2025 · 12 citations
