Rejectors in the Wild: Deployment Barriers for LLM Rejectors
Erik Schönwälder, Claudio Hartmann, Wolfgang Lehner
摘要
Despite strong benchmark performance, LLM hallucinations hinder reliable deployment in high-risk domains like medicine. Selective prediction can mitigate this by abstaining when an answer is likely incorrect. In black-box, closed-weight API settings or under tight cost constraints, this is naturally implemented via separated rejectors. In this setup, a lightweight rejector filters inputs, calling the black-box LLM only for accepted queries and abstaining otherwise. Yet even this practical design has important limitations, which we study and overcome through three contributions: (I) a human-annotation study showing that the way we define ''correct'' can drastically change which rejector appears reliable, (II) a large-scale 95×95 cross-model transfer analysis revealing structured redundancies between rejectors across LLMs, and (III) a shift from abstention to selective execution. By exploiting redundancies between LLMs, we propose two pruning strategies that build a compact, diversity-preserving LLM pool and a router that selects the most reliable model per query, outperforming common routing baselines.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Consistency and Uncertainty: Identifying Unreliable Responses From Black-Box Vision-Language Models for Selective Visual Question AnsweringZaid Khan, Yun FuCVPR 2024
- MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical ReasoningShuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen 等NeurIPS 2024 · 被引用 215 次
- LLMs (Almost) Never Abstain Under Medical UncertaintyAlessio Cocchieri, Luca Ragazzi, Giuseppe Tagliavini, Gianluca MoroACL 2026
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala 等ICLR 2024 · 被引用 132 次
- RACER: Risk-Aware Calibrated Efficient Routing for Large Language ModelsSai Hao, Hao Zeng, Hongxin Wei, Bingyi JingICML 2026 · 被引用 1 次
