Multi-Objective Hyperparameter Selection via Hypothesis Testing on Reliability Graphs
Amirmohammad Farzaneh, Osvaldo Simeone
Abstract
The selection of hyperparameters, such as prompt templates in large language models (LLMs), must often strike a balance between reliability and cost. In many cases, structural relationships between the expected reliability levels of the hyperparameters can be inferred from prior information and held-out data -- e.g., longer prompt templates may be more detailed and thus more reliable. However, existing hyperparameter selection methods either do not provide formal reliability guarantees or are unable to incorporate structured knowledge in the hyperparameter space. This paper introduces reliability graph-based Pareto testing (RG-PT), a novel multi-objective hyperparameter selection framework that maintains formal reliability guarantees in terms of false discovery rate (FDR), while accounting for known relationships among hyperparameters via a directed acyclic graph. Edges in the graph reflect expected reliability and cost trade-offs among hyperparameters, which are inferred via the Bradley-Terry (BT) ranking model from prior information and held-out data. Experimental evaluations demonstrate that RG-PT significantly outperforms existing methods such as learn-then-test (LTT) and Pareto testing (PT) through a more efficient exploration of the hyperparameter space.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 103c48da-0de5-450c-956f-0ad73d235d76Builds on4
- Large Language Models are Human-Level Prompt EngineersYongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster et al.ICLR 2023 · 297 citations
- Instruction Induction: From Few Examples to Natural Language Task DescriptionsOr Honovich, Uri Shaham, Samuel R. Bowman, Omer LevyACL 2023 · 48 citations
- Efficiently Controlling Multiple Risks with Pareto TestingBracha Laufer-Goldshtein, Adam Fisch, Regina Barzilay, Tommi S. JaakkolaICLR 2023 · 2 citations
- Adaptive Learn-then-Test: Statistically Valid and Efficient Hyperparameter SelectionMatteo Zecchin, Sangwoo Park, Osvaldo SimeoneICML 2025
Related papers
- Pareto Prompt OptimizationGuang Zhao, Byung-Jun Yoon, Gilchan Park, Shantenu Jha et al.ICLR 2025
- Pareto Optimal Learning for Estimating Large Language Model ErrorsTheodore Zhao, Mu Wei, Joseph Preston, Hoifung PoonACL 2024
- Landmark-Guided Policy Optimization for Multi-Objective Language Model SelectionMarcio Monteiro, Weichen Li, Puyu Wang, Marius Kloft et al.ICML 2026
- Efficient Multi-objective Prompt Optimization via Pure-exploration BanditsDonghao Li, Chengshuai Shi, Weijuan Ou, Cong Shen et al.ICLR 2026 · 2 citations
- A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground TruthMingyuan Xu, Xinzi Tan, Jiawei Wu, Doudou ZhouICML 2026
