Quality-Weighted Vendi Scores And Their Application To Diverse Experimental Design
Quan Nguyen, Adji Bousso Dieng
Abstract
Experimental design techniques such as active search and Bayesian optimization are widely used in the natural sciences for data collection and discovery. However, existing techniques tend to favor exploitation over exploration of the search space, which causes them to get stuck in local optima. This ``collapse"problem prevents experimental design algorithms from yielding diverse high-quality data. In this paper, we extend the Vendi scores -- a family of interpretable similarity-based diversity metrics -- to account for quality. We then leverage these quality-weighted Vendi scores to tackle experimental design problems across various applications, including drug discovery, materials discovery, and reinforcement learning. We found that quality-weighted Vendi scores allow us to construct policies for experimental design that flexibly balance quality and diversity, and ultimately assemble rich and diverse sets of high-performing data points. Our algorithms led to a 70%-170% increase in the number of effective discoveries compared to baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 16444ad1-baae-4059-a4f2-d802b04c721cCited by top-tier papers3
- AutoQD: Automatic Discovery of Diverse Behaviors with Quality-Diversity OptimizationSaeed Hedayatian, Stefanos NikolaidisICLR 2026 · 5 citations
- Soft Quality-Diversity OptimizationSaeed Hedayatian, Stefanos NikolaidisICLR 2026
- Diffusion Models Through a Global Lens: Are They Culturally Inclusive?Zahra Bayramli, Ayhan Suleymanzade, Na Min An, Huzama Ahmad et al.ACL 2025
Builds on2
Related papers
- Reinforced Active Learning for Large-Scale Virtual Screening with Learnable Policy ModelYicong Chen, Jiahua Rao, Jiancong Xie, Dahao Xu et al.NeurIPS 2025 · 2 citations
- Policy-Based Bayesian Active Causal Discovery with Deep Reinforcement LearningHeyang Gao, Zexu Sun, Hao Yang, Xu ChenKDD 2024 · 1 citation
- Causal Discovery with Reinforcement LearningShengyu Zhu, Ignavier Ng, Zhitang ChenICLR 2020 · 285 citations
- GeneDisco: A Benchmark for Experimental Design in Drug DiscoveryArash Mehrjou, Ashkan Soleymani, Andrew Jesson, Pascal Notin et al.ICLR 2022 · 25 citations
- MoE-Guided Graph Diffusion for Oriented Molecule DesignShuochen Li, Xiangqi Guo, Huobin Tan, Lei ShiAAAI 2026 · 1 citation
