Lune

NeurIPS2025Top-tier venue

Post Hoc Regression Refinement via Pairwise Rankings

Kevin Tirta Wijaya, Michael Sun, Minghao Guo, Hans-Peter Seidel, Wojciech Matusik, Vahid Babaei

2025Year
1Citations
1Top-tier citations

Abstract

Accurate prediction of continuous properties is essential to many scientific and engineering tasks. Although deep-learning regressors excel with abundant labels, their accuracy deteriorates in data-scarce regimes. We introduce RankRefine, a model-agnostic, plug-and-play post hoc method that refines regression with expert knowledge coming from pairwise rankings. Given a query item and a small reference set with known properties, RankRefine combines the base regressor's output with a rank-based estimate via inverse-variance weighting, requiring no retraining. In molecular property prediction task, RankRefine achieves up to 10% relative reduction in mean absolute error using only 20 pairwise comparisons obtained through a general-purpose large language model (LLM) with no finetuning. As rankings provided by human experts or general-purpose LLMs are sufficient for improving regression across diverse domains, RankRefine offers practicality and broad applicability, especially in low-data settings.

Lemma 3.2 (Variance of the rank-based estimate) Let ŷ * rank 0 minimize equation 2. Its variance is approximated by the inverse observed Fisher information [Ly et al., 2017],

Applying Theorem 3.1 with σ 2 rank from Lemma 3.2 yields ŷ * 0 .

We measure performance via the mean absolute error (MAE). For a folded Gaussian derived from a zero-mean Gaussian, MAE = 2/π σ, so

Corollary 3.2.1 Any informative ranker with finite variance (σ 2 rank < ∞) lowers the expected MAE after fusion.

More generally, letting α ∈ [0, 1] be the desired ratio between post-refinement and the original MAEs,

Regularization. If the ranker is biased yet over-confident (σ 2 rank ≪ σ 2 reg ), we temper its variance via σ 2 rank ← max σ 2 rank , c σ 2 reg , with user-chosen constant c > 0.

We evaluate RankRefine on synthetic and real-world tasks. After outlining datasets, metrics, and implementation details, we report results in synthetic settings, where ranking oracle is available, and practical settings across multiple domains.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 7654ba04-e5f0-4c0c-ab20-d97aff2cbd47

Cited by top-tier papers1

Ask how each one uses it

Builds on5

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines