Adaptive Learn-then-Test: Statistically Valid and Efficient Hyperparameter Selection
Matteo Zecchin, Sangwoo Park, Osvaldo Simeone
摘要
We introduce adaptive learn-then-test (aLTT), an efficient hyperparameter selection procedure that provides finite-sample statistical guarantees on the population risk of AI models. Unlike the existing learn-then-test (LTT) technique, which relies on conventional p-value-based multiple hypothesis testing (MHT), aLTT implements sequential data-dependent MHT with early termination by leveraging e-processes. As a result, aLTT can reduce the number of testing rounds, making it particularly well-suited for scenarios in which testing is costly or presents safety risks. Apart from maintaining statistical validity, in applications such as online policy selection for offline reinforcement learning and prompt engineering, aLTT is shown to achieve the same performance as LTT while requiring only a fraction of the testing rounds.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Multi-Objective Hyperparameter Selection via Hypothesis Testing on Reliability GraphsAmirmohammad Farzaneh, Osvaldo SimeoneNeurIPS 2025 · 被引用 1 次
- Conformal Risk-Averse Decision Making with Action Conditional GuaranteeZihan Zhu, Shayan Kiyani, George Pappas, Hamed HassaniICML 2026 · 被引用 1 次
- Decision Theoretic Foundations for Conformal Prediction: Optimal Uncertainty Quantification for Risk-Averse AgentsShayan Kiyani, George J. Pappas, Aaron Roth, Hamed HassaniICML 2025
- Prediction-Powered Risk Monitoring of Deployed Models for Detecting Harmful Distribution ShiftsGuangyi Zhang, Yunlong Cai, Guanding Yu, Osvaldo SimeoneICML 2026
- Optimal Decision-Making Based on Prediction SetsTao Wang, Edgar DobribanICML 2026
它引用的顶会 Paper9
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani 等NeurIPS 2022 · 被引用 394 次
- Large Language Models are Human-Level Prompt EngineersYongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster 等ICLR 2023 · 被引用 297 次
- Conformal Risk ControlAnastasios Nikolas Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei 等ICLR 2024 · 被引用 242 次
- Automatic Chain of Thought Prompting in Large Language ModelsZhuosheng Zhang, Aston Zhang, Mu Li, Alex SmolaICLR 2023 · 被引用 234 次
相关 Paper
- Efficiently Controlling Multiple Risks with Pareto TestingBracha Laufer-Goldshtein, Adam Fisch, Regina Barzilay, Tommi S. JaakkolaICLR 2023 · 被引用 2 次
- Anytime-Valid Inference for Online Ranking of Large Language ModelsRunzhe Gu, Wenguang Sun, Bowen Gang, Xintao XiaICML 2026
- A Decision-Theoretic View of Test-Time Training: When, How Far, and Which Directions to AdaptTomoya WakayamaICML 2026
- Adaptive Q-Network: On-the-fly Target Selection for Deep Reinforcement LearningThéo Vincent, Fabian Wahren, Jan Peters, Boris Belousov 等ICLR 2025
- Bayesian Optimization for Iterative LearningVu Nguyen, Sebastian Schulze, Michael A. OsborneNeurIPS 2020 · 被引用 38 次
