Designing Cell-Type-Specific Promoter Sequences Using Conservative Model-Based Optimization
Aniketh Janardhan Reddy, Xinyang Geng, Michael Herschl, Sathvik Kolli, Aviral Kumar, Patrick Hsu, Sergey Levine, Nilah Ioannidis
Abstract
Gene therapies have the potential to treat disease by delivering therapeutic genetic cargo to disease-associated cells. One limitation to their widespread use is the lack of short regulatory sequences, or promoters, that differentially induce the expression of delivered genetic cargo in target cells, minimizing side effects in other cell types. Such cell-type-specific promoters are difficult to discover using existing methods, requiring either manual curation or access to large datasets of promoter-driven expression from both targeted and untargeted cells. Model-based optimization (MBO) has emerged as an effective method to design biological sequences in an automated manner, and has recently been used in promoter design methods. However, these methods have only been tested using large training datasets that are expensive to collect, and focus on designing promoters for markedly different cell types, overlooking the complexities associated with designing promoters for closely related cell types that share similar regulatory features. Therefore, we introduce a comprehensive framework for utilizing MBO to design promoters in a data-efficient manner, with an emphasis on discovering promoters for similar cell types. We use conservative objective models (COMs) for MBO and highlight practical considerations such as best practices for improving sequence diversity, getting estimates of model uncertainty, and choosing the optimal set of sequences for experimental validation. Using three relatively similar blood cancer cell lines (Jurkat, K562, and THP1), we show that our approach discovers many novel cell-type-specific promoters after experimentally validating the designed sequences. For K562 cells, in particular, we discover a promoter that has 75.85% higher cell-type-specificity than the best promoter from the initial dataset used to train our models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Ctrl-DNA: Controllable Cell-Type-Specific Regulatory DNA Design via Constrained RLXingyu Chen, Shihao Ma, Runsheng Lin, Jiecong Lin et al.NeurIPS 2025 · 5 citations
- Regulatory DNA Sequence Design with Reinforcement LearningZhao Yang, Bing Su, Chuan Cao, Ji-Rong WenICLR 2025
- From Predictors to Samplers via the Training TrajectorySoumya Ram, Akhila RamICLR 2026
- Neural Genetic Search in Discrete SpacesHyeonah Kim, Sanghyeok Choi, Jiwoo Son, Jinkyoo Park et al.ICML 2025
Builds on2
- Biological Sequence Design with GFlowNetsMoksh Jain, Emmanuel Bengio, Alex Hernández-García, Jarrid Rector-Brooks et al.ICML 2022 · 224 citations
- Conservative Objective Models for Effective Offline Model-Based OptimizationBrandon Trabucco, Aviral Kumar, Xinyang Geng, Sergey LevineICML 2021 · 119 citations
Related papers
- Dirichlet Diffusion Score Model for Biological Sequence GenerationPavel Avdeyev, Chenlai Shi, Yuhao Tan, Kseniia Dudnyk et al.ICML 2023 · 91 citations
- Diversity By Design: Leveraging Distribution Matching for Offline Model-Based OptimizationMichael S. Yao, James C. Gee, Osbert BastaniICML 2025
- Accelerating Bayesian Optimization for Biological Sequence Design with Denoising AutoencodersSamuel Stanton, Wesley J. Maddox, Nate Gruver, Phillip M. Maffettone et al.ICML 2022 · 137 citations
- Generative Adversarial Model-Based Optimization via Source Critic RegularizationMichael S. Yao, Yimeng Zeng, Hamsa Bastani, Jacob R. Gardner et al.NeurIPS 2024 · 14 citations
- Bootstrapped Training of Score-Conditioned Generator for Offline Design of Biological SequencesMinsu Kim, Federico Berto, Sungsoo Ahn, Jinkyoo ParkNeurIPS 2023 · 30 citations
