ICML2026

Adaptively Grouped Contextual Bandits for Heterogeneous Human-AI Decision Making with Conformal Prediction Sets

Yanchen Wu, Bo Li

摘要

Personalizing AI decision support for heterogeneous human decision-makers remains a key challenge. We study a collaboration workflow where AI provides a reduced prediction set via conformal prediction and the human makes the final decision based on the set. We formulate this personalization problem as a contextual bandit, where individual and task features form the context, candidate significance levels α\alpha serve as arms, and the optimal prediction-set size varies across contexts. To address large arm spaces and high-dimensional contexts, we introduce the Adaptively Grouped Contextual Bandit (AGCB) framework, which avoids global function approximation by exploiting two Human-AI structural assumptions: continuity and monotonicity. Continuity enables information sharing across nearby contexts and decisions, and drives a data-driven Zooming Mechanism that balances intra-group estimation error against inter-group approximation bias. Monotonicity converts each observation into directional counterfactual information over the KK candidate α\alpha values, reducing the arm-dependence factor from polynomial to logarithmic in KK. Together, these mechanisms yield minimax-optimal dependence on the learning horizon TT for both cumulative and simple regret objectives. Empirical results confirm that AGCB achieves the strongest overall performance across most heterogeneous, data-scarce settings.