Constrained Molecule Generation Modelled Using the Grammar Constraint
David Saikali, Gilles Pesant
Abstract
Drug discovery is a very time-consuming and costly endeavour due to its huge design space and to the lengthy and failurefraught process of bringing a product to market. Automating the generation of candidate molecules exhibiting some of the desired properties can help. Among the standard formats to encode molecules, SMILES is a widespread string representation. We propose a constraint programming model showcasing the grammar constraint to express the design space of organic molecules using the SMILES notation. We show how some common physicochemical properties -such as molecular weight and lipophilicity -and structural features can be expressed as constraints in the model. We also contribute a weighted counting algorithm for the grammar constraint, allowing us to use a belief propagation heuristic to guide the generation. Our experiments indicate that such a heuristic is key to driving the search towards desired molecules.
Code -github.com/cravethedave/MiniCPBP/tree/AAAI26 1 Introduction Drug discovery is a very time-consuming and costly endeavour due to its huge design space -estimated to contain between 10 23 and 10 60 different molecules (Polishchuk, Madzhidov, and Varnek 2013) -and to the lengthy and failure-fraught process of bringing a product to market. Automated molecule design is nowadays a vital part of drug discovery and material science, with computational approaches coming from deep generative models and combinatorial search methods (Du et al. 2022). It aims to extract from this huge design space the most likely candidates according to some desired properties. And even among these, only a few may lead to a usable product after extensive testing.
SMILES, a one-dimensional encoding of molecules, is one of the standards commonly used by this research community. It lends itself well to techniques used for natural language processing, such as sequential generative neural models, but also to constraint programming (CP). Using a context-free grammar and a few additional constraints, we show how to describe valid SMILES strings in a CP model. This allows us to explore the huge design space of possible
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d20733f2-aa4c-497d-b3fb-cf8e45fa5ad9Builds on1
Related papers
- Representing Molecules as Random Walks Over Interpretable GrammarsMichael Sun, Minghao Guo, Weize Yuan, Veronika Thost et al.ICML 2024 · 6 citations
- Barking up the right tree: an approach to search over molecule synthesis DAGsJohn Bradshaw, Brooks Paige, Matt J. Kusner, Marwin H. S. Segler et al.NeurIPS 2020 · 71 citations
- GraphAF: a Flow-based Autoregressive Model for Molecular Graph GenerationChence Shi, Minkai Xu, Zhaocheng Zhu, Weinan Zhang et al.ICLR 2020 · 532 citations
- How Much Space Has Been Explored? Measuring the Chemical Space Covered by Databases and Machine-Generated MoleculesYutong Xie, Ziqiao Xu, Jiaqi Ma, Qiaozhu MeiICLR 2023 · 3 citations
- Goal-directed Generation of Discrete Structures with Conditional Generative ModelsAmina Mollaysa, Brooks Paige, Alexandros KalousisNeurIPS 2020 · 12 citations
