Transcormer: Transformer for Sentence Scoring with Sliding Language Modeling
Kaitao Song, Yichong Leng, Xu Tan, Yicheng Zou, Tao Qin, Dongsheng Li
Abstract
Sentence scoring aims at measuring the likelihood score of a sentence and is widely used in natural language processing scenarios, like reranking, which is to select the best sentence from multiple candidates. Previous works on sentence scoring mainly adopted either causal language modeling (CLM) like GPT or masked language modeling (MLM) like BERT, which have some limitations: 1) CLM only utilizes unidirectional information for the probability estimation of a sentence without considering bidirectional context, which affects the scoring quality; 2) MLM can only estimate the probability of partial tokens at a time and thus requires multiple forward passes to estimate the probability of the whole sentence, which incurs large computation and time cost. In this paper, we propose Transcormer -a Transformer model with a novel sliding language modeling (SLM) for sentence scoring. Specifically, our SLM adopts a triple-stream self-attention mechanism to estimate the probability of all tokens in a sentence with bidirectional context and only requires a single forward pass. SLM can avoid the limitations of CLM (only unidirectional context) and MLM (multiple forward passes) and inherit their advantages, and thus achieve high effectiveness and efficiency in scoring. Experimental results on multiple tasks demonstrate that our method achieves better performance than other language models. Our code and pre-trained models will be released at: https://github.com/microsoft/CyBERTron-LM/ Transcormer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- SA-Solver: Stochastic Adams Solver for Fast Sampling of Diffusion ModelsShuchen Xue, Mingyang Yi, Weijian Luo, Shifeng Zhang et al.NeurIPS 2023 · 91 citations
- Accelerating Diffusion Sampling with Optimized Time StepsShuchen Xue, Zhaoqiang Liu, Fei Chen, Shifeng Zhang et al.CVPR 2024 · 16 citations
- Multi-Level Knowledge Distillation for Out-of-Distribution Detection in TextQianhui Wu, Huiqiang Jiang, Haonan Yin, Börje Karlsson et al.ACL 2023 · 7 citations
- Diffusion Sampling Correction via Approximately 10 ParametersGuangyi Wang, Wei Peng, Lijiang Li, Wenyu Chen et al.ICML 2025
- PFDiff: Training-Free Acceleration of Diffusion Models Combining Past and Future ScoresGuangyi Wang, Yuren Cai, Lijiang Li, Wei Peng et al.ICLR 2025
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 167 citations
Related papers
- Fast and Accurate Deep Bidirectional Language Representations for Unsupervised LearningJoongbo Shin, Yoonhyung Lee, Seunghyun Yoon, Kyomin JungACL 2020 · 6 citations
- SLM: Learning a Discourse Language Representation with Sentence UnshufflingHaejun Lee, Drew A. Hudson, Kangwook Lee, Christopher D. ManningEMNLP 2020 · 2 citations
- Self-Calibrated Listwise Reranking with Large Language ModelsRuiyang Ren, Yuhao Wang, Kun Zhou, Wayne Xin Zhao et al.WWW 2025 · 12 citations
- Leveraging Estimated Transferability Over Human Intuition for Model Selection in Text RankingJun Bai, Zhuofan Chen, Zhenzi Li, Hanhua Hong et al.EMNLP 2024 · 2 citations
- Deep Attentive Ranking Networks for Learning to Order SentencesPawan Kumar, Dhanajit Brahma, Harish Karnick, Piyush RaiAAAI 2020 · 52 citations
