Performance Bounds for Active Binary Testing with Information Maximization
Aditya Chattopadhyay, Benjamin David Haeffele, René Vidal, Donald Geman
Abstract
In many applications like experimental design, group testing, and medical diagnosis, the state of a random variable is revealed by successively observing the outcomes of binary tests about . New tests are selected adaptively based on the history of outcomes observed so far. If the number of states of is finite, the process ends when can be predicted with a desired level of confidence or all available tests have been used. Finding the strategy that minimizes the expected number of tests needed to predict is virtually impossible in most real applications. Therefore, the commonly used strategy is the greedy heuristic of Information Maximization (InfoMax) that selects tests sequentially in order of information gain. Despite its widespread use, existing guarantees on its performance are often vacuous when compared to its empirical efficiency. In this paper, for the first time to the best of our knowledge, we establish tight non-vacuous bounds on InfoMax’s performance. Our analysis is based on the assumption that at any iteration of the greedy strategy, there is always a binary test available whose conditional probability of being ’true’, given the history, is within units of one-half. This assumption is motivated by practical applications where the available set of tests often satisfies this property for modest values of , say, . Specifically, we analyze two distinct scenarios: (i) all tests are functions of , and (ii) test outcomes are corrupted by a binary symmetric channel. For both cases, our bounds guarantee the near-optimal performance of InfoMax for modest values. It requires only a small multiplicative factor of the entropy of , in terms of the average number of tests needed to make accurate predictions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d57caaa-ed26-4e1e-acef-bc0b9a7d4be3Builds on4
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Learning to Maximize Mutual Information for Dynamic Feature SelectionIan Connick Covert, Wei Qiu, Mingyu Lu, Nayoon Kim et al.ICML 2023 · 67 citations
- BSODA: A Bipartite Scalable Framework for Online Disease DiagnosisWeijie He, Xiaohao Mao, Chao Ma, Yu Huang et al.WWW 2022 · 18 citations
- Variational Information Pursuit for Interpretable PredictionsAditya Chattopadhyay, Kwan Ho Ryan Chan, Benjamin David Haeffele, Donald Geman et al.ICLR 2023
Related papers
- Greedy Approximation Algorithms for Active Sequential Hypothesis TestingKyra Gan, Su Jia, Andrew A. LiNeurIPS 2021 · 9 citations
- Exact Thresholds for Noisy Non-Adaptive Group TestingJunren Chen, Jonathan ScarlettSODA 2025
- Empirical Bayes Selection for Value MaximizationDominic Coey, Kenneth HungKDD 2025 · 1 citation
- Mean Estimation in High-Dimensional Binary Markov Gaussian Mixture ModelsYihan Zhang, Nir WeinbergerNeurIPS 2022 · 1 citation
- An Information-Theoretic Analysis of Nonstationary Bandit LearningSeungki Min, Daniel RussoICML 2023 · 11 citations
