Language Modelling via Learning to Rank
Arvid Frydenlund, Gagandeep Singh, Frank Rudzicz
Abstract
We consider language modelling (LM) as a multi-label structured prediction task by re-framing training from solely predicting a single ground-truth word to ranking a set of words which could continue a given context. To avoid annotating top-k ranks, we generate them using pre-trained LMs: GPT-2, BERT, and Born-Again models. This leads to a rank-based form of knowledge distillation (KD). We also develop a method using N-grams to create a non-probabilistic teacher which generates the ranks without the need of a pre-trained LM.
We confirm the hypotheses: that we can treat LMing as a ranking task and that we can do so without the use of a pre-trained LM. We show that rank-based KD generally gives a modest improvement to perplexity (PPL) -- though often with statistical significance -- when compared to Kullback–Leibler-based KD. Surprisingly, given the naivety of the method, the N-grams act as competitive teachers and achieve similar performance as using either BERT or a Born-Again model teachers. Unsurprisingly, GPT-2 always acts as the best teacher. Using it and a Transformer-XL student on Wiki-02, rank-based KD reduces a cross-entropy baseline from 65.27 to 55.94 and against a KL-based KD of 56.70.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Looking Beyond the Top-1: Transformers Determine Top Tokens in OrderDaria Lioubashevski, Tomer Schlank, Gabriel Stanovsky, Ariel GoldsteinICML 2025
- PLD: A Choice-Theoretic List-Wise Knowledge DistillationEjafa Bassam, Dawei Zhu, Kaigui BianNeurIPS 2025
Builds on11
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 914 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- Does label smoothing mitigate label noise?Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, Sanjiv KumarICML 2020 · 411 citations
- Data-dependent Gaussian Prior Objective for Language GenerationZuchao Li, Rui Wang, Kehai Chen, Masao Utiyama et al.ICLR 2020 · 68 citations
- On the Inference Calibration of Neural Machine TranslationShuo Wang, Zhaopeng Tu, Shuming Shi, Yang LiuACL 2020 · 66 citations
Related papers
- Towards Understanding and Improving Knowledge Distillation for Neural Machine TranslationSongming Zhang, Yunlong Liang, Shuaibo Wang, Yufeng Chen et al.ACL 2023 · 8 citations
- Multi-level Distillation of Semantic Knowledge for Pre-training Multilingual Language ModelMingqi Li, Fei Ding, Dan Zhang, Long Cheng et al.EMNLP 2022 · 3 citations
- Towards Efficient Pre-Trained Language Model via Feature Correlation DistillationKun Huang, Xin Guo, Meng WangNeurIPS 2023 · 8 citations
- Pretrained Bidirectional Distillation for Machine TranslationYimeng Zhuang, Mei TuACL 2023 · 3 citations
- MiniPLM: Knowledge Distillation for Pre-training Language ModelsYuxian Gu, Hao Zhou, Fandong Meng, Jie Zhou et al.ICLR 2025
