Beam Enumeration: Probabilistic Explainability For Sample Efficient Self-conditioned Molecular Design
Jeff Guo, Philippe Schwaller
Abstract
Generative molecular design has moved from proof-of-concept to real-world applicability, as marked by the surge in very recent papers reporting experimental validation. Key challenges in explainability and sample efficiency present opportunities to enhance generative design to directly optimize expensive high-fidelity oracles and provide actionable insights to domain experts. Here, we propose Beam Enumeration to exhaustively enumerate the most probable sub-sequences from language-based molecular generative models and show that molecular substructures can be extracted. When coupled with reinforcement learning, extracted substructures become meaningful, providing a source of explainability and improving sample efficiency through self-conditioned generation. Beam Enumeration is generally applicable to any language-based molecular generative model and notably further improves the performance of the recently reported Augmented Memory algorithm, which achieved the new state-of-the-art on the Practical Molecular Optimization benchmark for sample efficiency. The combined algorithm generates more high reward molecules and faster, given a fixed oracle budget. Beam Enumeration shows that improvements to explainability and sample efficiency for molecular design can be made synergistic. The code is available at https://github.com/schwallergroup/augmented_memory .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationEmmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup et al.NeurIPS 2021 · 565 citations
- Multi-Objective Molecule Generation using Interpretable SubstructuresWengong Jin, Regina Barzilay, Tommi S. JaakkolaICML 2020 · 238 citations
- MARS: Markov Molecular Sampling for Multi-objective Drug DiscoveryYutong Xie, Chence Shi, Hao Zhou, Yuwei Yang et al.ICLR 2021 · 186 citations
- MIMOSA: Multi-constraint Molecule Sampling for Molecule OptimizationTianfan Fu, Cao Xiao, Xinhao Li, Lucas M. Glass et al.AAAI 2021 · 94 citations
- Reinforced Genetic Algorithm for Structure-based Drug DesignTianfan Fu, Wenhao Gao, Connor W. Coley, Jimeng SunNeurIPS 2022 · 79 citations
Related papers
- MolMem: Memory-Augmented Agentic Reinforcement Learning for Sample-Efficient Molecular OptimizationZiqing Wang, Yibo Wen, Abhishek Pandey, Han Liu et al.ACL 2026 · 2 citations
- Improving Molecular Design by Stochastic Iterative Target AugmentationKevin Yang, Wengong Jin, Kyle Swanson, Regina Barzilay et al.ICML 2020 · 31 citations
- Empowering LLMs for Structure-Based Drug Design via Exploration-Augmented Latent InferenceXuanning Hu, Anchen Li, Qianli Xing, Jinglong Ji et al.WWW 2026
- Genetic-guided GFlowNets for Sample Efficient Molecular OptimizationHyeonah Kim, Minsu Kim, Sanghyeok Choi, Jinkyoo ParkNeurIPS 2024 · 42 citations
- Uncertainty-Aware Multi-Objective Reinforcement Learning-Guided Diffusion Models for 3D De Novo Molecular DesignLianghong Chen, Dongkyu Eugene Kim, Mike Domaratzki, Pingzhao HuNeurIPS 2025 · 4 citations
