Designing Cell-Type-Specific Promoter Sequences Using Conservative Model-Based Optimization
Aniketh Janardhan Reddy, Xinyang Geng, Michael Herschl, Sathvik Kolli, Aviral Kumar, Patrick Hsu, Sergey Levine, Nilah Ioannidis
摘要
Gene therapies have the potential to treat disease by delivering therapeutic genetic cargo to disease-associated cells. One limitation to their widespread use is the lack of short regulatory sequences, or promoters, that differentially induce the expression of delivered genetic cargo in target cells, minimizing side effects in other cell types. Such cell-type-specific promoters are difficult to discover using existing methods, requiring either manual curation or access to large datasets of promoter-driven expression from both targeted and untargeted cells. Model-based optimization (MBO) has emerged as an effective method to design biological sequences in an automated manner, and has recently been used in promoter design methods. However, these methods have only been tested using large training datasets that are expensive to collect, and focus on designing promoters for markedly different cell types, overlooking the complexities associated with designing promoters for closely related cell types that share similar regulatory features. Therefore, we introduce a comprehensive framework for utilizing MBO to design promoters in a data-efficient manner, with an emphasis on discovering promoters for similar cell types. We use conservative objective models (COMs) for MBO and highlight practical considerations such as best practices for improving sequence diversity, getting estimates of model uncertainty, and choosing the optimal set of sequences for experimental validation. Using three relatively similar blood cancer cell lines (Jurkat, K562, and THP1), we show that our approach discovers many novel cell-type-specific promoters after experimentally validating the designed sequences. For K562 cells, in particular, we discover a promoter that has 75.85% higher cell-type-specificity than the best promoter from the initial dataset used to train our models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Ctrl-DNA: Controllable Cell-Type-Specific Regulatory DNA Design via Constrained RLXingyu Chen, Shihao Ma, Runsheng Lin, Jiecong Lin 等NeurIPS 2025 · 被引用 5 次
- Regulatory DNA Sequence Design with Reinforcement LearningZhao Yang, Bing Su, Chuan Cao, Ji-Rong WenICLR 2025
- From Predictors to Samplers via the Training TrajectorySoumya Ram, Akhila RamICLR 2026
- Neural Genetic Search in Discrete SpacesHyeonah Kim, Sanghyeok Choi, Jiwoo Son, Jinkyoo Park 等ICML 2025
它引用的顶会 Paper2
相关 Paper
- Dirichlet Diffusion Score Model for Biological Sequence GenerationPavel Avdeyev, Chenlai Shi, Yuhao Tan, Kseniia Dudnyk 等ICML 2023 · 被引用 91 次
- Diversity By Design: Leveraging Distribution Matching for Offline Model-Based OptimizationMichael S. Yao, James C. Gee, Osbert BastaniICML 2025
- Accelerating Bayesian Optimization for Biological Sequence Design with Denoising AutoencodersSamuel Stanton, Wesley J. Maddox, Nate Gruver, Phillip M. Maffettone 等ICML 2022 · 被引用 137 次
- Generative Adversarial Model-Based Optimization via Source Critic RegularizationMichael S. Yao, Yimeng Zeng, Hamsa Bastani, Jacob R. Gardner 等NeurIPS 2024 · 被引用 14 次
- Bootstrapped Training of Score-Conditioned Generator for Offline Design of Biological SequencesMinsu Kim, Federico Berto, Sungsoo Ahn, Jinkyoo ParkNeurIPS 2023 · 被引用 30 次
