Lune

KDD2026Top-tier venue

Rejectors in the Wild: Deployment Barriers for LLM Rejectors

Erik Schönwälder, Claudio Hartmann, Wolfgang Lehner

2026Year

Abstract

Despite strong benchmark performance, LLM hallucinations hinder reliable deployment in high-risk domains like medicine. Selective prediction can mitigate this by abstaining when an answer is likely incorrect. In black-box, closed-weight API settings or under tight cost constraints, this is naturally implemented via separated rejectors. In this setup, a lightweight rejector filters inputs, calling the black-box LLM only for accepted queries and abstaining otherwise. Yet even this practical design has important limitations, which we study and overcome through three contributions: (I) a human-annotation study showing that the way we define ''correct'' can drastically change which rejector appears reliable, (II) a large-scale 95×95 cross-model transfer analysis revealing structured redundancies between rejectors across LLMs, and (III) a shift from abstention to selective execution. By exploiting redundancies between LLMs, we propose two pruning strategies that build a compact, diversity-preserving LLM pool and a router that selects the most reliable model per query, outperforming common routing baselines.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get a9c8c019-fffb-4ecb-bb02-806f74800e84

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines