Probability-Entropy Calibration: An Elastic Indicator for Adaptive Fine-tuning
Wenhao Yu, Shaohang Wei, Jiahong Liu, Yifan Li, Minda Hu, Aiwei Liu, Hao Zhang, Irwin King
Abstract
Token-level reweighting is a simple yet effective mechanism for controlling supervised finetuning, but common indicators are largely onedimensional: the ground-truth probability reflects downstream alignment, while token entropy reflects intrinsic uncertainty induced by the pretraining prior. Ignoring entropy can misidentify noisy or easily replaceable tokens as learningcritical, while ignoring probability fails to reflect target-specific alignment. RANKTUNER introduces a probability-entropy calibration signal, the Relative Rank Indicator, which compares the rank of the ground-truth token with its expected rank under the prediction distribution. The inverse indicator is used as a tokenwise Relative Scale to reweight the fine-tuning objective, focusing updates on truly under-learned tokens without over-penalizing intrinsically uncertain positions. Experiments on multiple backbones show consistent improvements on mathematical reasoning benchmarks, transfer gains on out-of-distribution reasoning, and pre code generation performance over probability-or entropyonly reweighting baselines. The implementation code is available at https://github.com/ LvAoAo/Ranktuner_VERL .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMsHaoming Meng, Kexin Huang, Shaohang Wei, Chiyu Ma et al.ICLR 2026 · 24 citations
- SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for ReasoningYuqian Fu, Tinghong Chen, Jiajun Chai, Xihuai Wang et al.ICLR 2026 · 97 citations
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward RectificationYongliang Wu, Yizhou Zhou, Ziheng Zhou, Yingzhe Peng et al.ICLR 2026 · 130 citations
- Don't Force the Fit: Bounded Log-Likelihood Loss for Enhanced Reasoning in Large Language ModelsFeng Zhao, Hong Zhang, Yu Yang, Ruilin Zhao et al.ICML 2026
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy EntropyHongze Tan, Zihan Wang, Jianfei Pan, Jinghao Lin et al.ICML 2026 · 53 citations
