Lune

ICML2026Top-tier venue

Probability-Entropy Calibration: An Elastic Indicator for Adaptive Fine-tuning

Wenhao Yu, Shaohang Wei, Jiahong Liu, Yifan Li, Minda Hu, Aiwei Liu, Hao Zhang, Irwin King

2026Year
9Citations

Abstract

Token-level reweighting is a simple yet effective mechanism for controlling supervised finetuning, but common indicators are largely onedimensional: the ground-truth probability reflects downstream alignment, while token entropy reflects intrinsic uncertainty induced by the pretraining prior. Ignoring entropy can misidentify noisy or easily replaceable tokens as learningcritical, while ignoring probability fails to reflect target-specific alignment. RANKTUNER introduces a probability-entropy calibration signal, the Relative Rank Indicator, which compares the rank of the ground-truth token with its expected rank under the prediction distribution. The inverse indicator is used as a tokenwise Relative Scale to reweight the fine-tuning objective, focusing updates on truly under-learned tokens without over-penalizing intrinsically uncertain positions. Experiments on multiple backbones show consistent improvements on mathematical reasoning benchmarks, transfer gains on out-of-distribution reasoning, and pre code generation performance over probability-or entropyonly reweighting baselines. The implementation code is available at https://github.com/ LvAoAo/Ranktuner_VERL .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Builds on17

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines