Constrained Molecule Generation Modelled Using the Grammar Constraint
David Saikali, Gilles Pesant
摘要
Drug discovery is a very time-consuming and costly endeavour due to its huge design space and to the lengthy and failurefraught process of bringing a product to market. Automating the generation of candidate molecules exhibiting some of the desired properties can help. Among the standard formats to encode molecules, SMILES is a widespread string representation. We propose a constraint programming model showcasing the grammar constraint to express the design space of organic molecules using the SMILES notation. We show how some common physicochemical properties -such as molecular weight and lipophilicity -and structural features can be expressed as constraints in the model. We also contribute a weighted counting algorithm for the grammar constraint, allowing us to use a belief propagation heuristic to guide the generation. Our experiments indicate that such a heuristic is key to driving the search towards desired molecules.
Code -github.com/cravethedave/MiniCPBP/tree/AAAI26 1 Introduction Drug discovery is a very time-consuming and costly endeavour due to its huge design space -estimated to contain between 10 23 and 10 60 different molecules (Polishchuk, Madzhidov, and Varnek 2013) -and to the lengthy and failure-fraught process of bringing a product to market. Automated molecule design is nowadays a vital part of drug discovery and material science, with computational approaches coming from deep generative models and combinatorial search methods (Du et al. 2022). It aims to extract from this huge design space the most likely candidates according to some desired properties. And even among these, only a few may lead to a usable product after extensive testing.
SMILES, a one-dimensional encoding of molecules, is one of the standards commonly used by this research community. It lends itself well to techniques used for natural language processing, such as sequential generative neural models, but also to constraint programming (CP). Using a context-free grammar and a few additional constraints, we show how to describe valid SMILES strings in a CP model. This allows us to explore the huge design space of possible
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Representing Molecules as Random Walks Over Interpretable GrammarsMichael Sun, Minghao Guo, Weize Yuan, Veronika Thost 等ICML 2024 · 被引用 6 次
- Barking up the right tree: an approach to search over molecule synthesis DAGsJohn Bradshaw, Brooks Paige, Matt J. Kusner, Marwin H. S. Segler 等NeurIPS 2020 · 被引用 71 次
- GraphAF: a Flow-based Autoregressive Model for Molecular Graph GenerationChence Shi, Minkai Xu, Zhaocheng Zhu, Weinan Zhang 等ICLR 2020 · 被引用 532 次
- How Much Space Has Been Explored? Measuring the Chemical Space Covered by Databases and Machine-Generated MoleculesYutong Xie, Ziqiao Xu, Jiaqi Ma, Qiaozhu MeiICLR 2023 · 被引用 3 次
- Goal-directed Generation of Discrete Structures with Conditional Generative ModelsAmina Mollaysa, Brooks Paige, Alexandros KalousisNeurIPS 2020 · 被引用 12 次
