COM-BOM: Bayesian Exemplar Search for Efficiently Exploring the Accuracy-Calibration Pareto Frontier
Gaoxiang Luo, Aryan Deshwal
摘要
Selecting an optimal set of exemplars is critical for good performance of in-context learning. However, prior exemplar search methods narrowly optimize for predictive accuracy, critically neglecting model calibration-a key determinant of trustworthiness and safe deployment. In this paper, we formulate exemplar selection as a multi-objective optimization problem, explicitly targeting both the maximization of predictive accuracy and the minimization of expected calibration error. We solve this problem with a sample-efficient Combinatorial Bayesian Optimization algorithm (COM-BOM) to find the Pareto front that optimally trades off the two objectives of accuracy and calibration. We evaluate COM-BOM on multiple tasks from unsaturated MMLU-Pro benchmark and find that COM-BOM beats or matches the baselines at jointly optimizing the two objectives, while requiring a minimal number of LLM API calls.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Differentiable Expected Hypervolume Improvement for Parallel Multi-Objective Bayesian OptimizationSamuel Daulton, Maximilian Balandat, Eytan BakshyNeurIPS 2020 · 被引用 428 次
- Large language models are few-shot clinical information extractorsMonica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim 等EMNLP 2022 · 被引用 285 次
- Parallel Bayesian Optimization of Multiple Noisy Objectives with Expected Hypervolume ImprovementSamuel Daulton, Maximilian Balandat, Eytan BakshyNeurIPS 2021 · 被引用 276 次
- Automatic Chain of Thought Prompting in Large Language ModelsZhuosheng Zhang, Aston Zhang, Mu Li, Alex SmolaICLR 2023 · 被引用 234 次
相关 Paper
- Prompt Optimization with EASE? Efficient Ordering-aware Automated Selection of ExemplarsZhaoxuan Wu, Xiaoqiang Lin, Zhongxiang Dai, Wenyang Hu 等NeurIPS 2024 · 被引用 44 次
- Sample Efficient Demonstration Selection for In-Context LearningKiran Purohit, Venktesh V, Sourangshu Bhattacharya, Avishek AnandICML 2025
- Mixtures of In-Context LearnersGiwon Hong, Emile van Krieken, Edoardo Maria Ponti, Nikolay Malkin 等ACL 2025
- Large Language Models to Enhance Bayesian OptimizationTennison Liu, Nicolás Astorga, Nabeel Seedat, Mihaela van der SchaarICLR 2024 · 被引用 143 次
- Teach Better or Show Smarter? On Instructions and Exemplars in Automatic Prompt OptimizationXingchen Wan, Ruoxi Sun, Hootan Nakhost, Sercan Ö. ArikNeurIPS 2024 · 被引用 35 次
