Transcormer: Transformer for Sentence Scoring with Sliding Language Modeling
Kaitao Song, Yichong Leng, Xu Tan, Yicheng Zou, Tao Qin, Dongsheng Li
摘要
Sentence scoring aims at measuring the likelihood score of a sentence and is widely used in natural language processing scenarios, like reranking, which is to select the best sentence from multiple candidates. Previous works on sentence scoring mainly adopted either causal language modeling (CLM) like GPT or masked language modeling (MLM) like BERT, which have some limitations: 1) CLM only utilizes unidirectional information for the probability estimation of a sentence without considering bidirectional context, which affects the scoring quality; 2) MLM can only estimate the probability of partial tokens at a time and thus requires multiple forward passes to estimate the probability of the whole sentence, which incurs large computation and time cost. In this paper, we propose Transcormer -a Transformer model with a novel sliding language modeling (SLM) for sentence scoring. Specifically, our SLM adopts a triple-stream self-attention mechanism to estimate the probability of all tokens in a sentence with bidirectional context and only requires a single forward pass. SLM can avoid the limitations of CLM (only unidirectional context) and MLM (multiple forward passes) and inherit their advantages, and thus achieve high effectiveness and efficiency in scoring. Experimental results on multiple tasks demonstrate that our method achieves better performance than other language models. Our code and pre-trained models will be released at: https://github.com/microsoft/CyBERTron-LM/ Transcormer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SA-Solver: Stochastic Adams Solver for Fast Sampling of Diffusion ModelsShuchen Xue, Mingyang Yi, Weijian Luo, Shifeng Zhang 等NeurIPS 2023 · 被引用 91 次
- Accelerating Diffusion Sampling with Optimized Time StepsShuchen Xue, Zhaoqiang Liu, Fei Chen, Shifeng Zhang 等CVPR 2024 · 被引用 16 次
- Multi-Level Knowledge Distillation for Out-of-Distribution Detection in TextQianhui Wu, Huiqiang Jiang, Haonan Yin, Börje Karlsson 等ACL 2023 · 被引用 7 次
- Diffusion Sampling Correction via Approximately 10 ParametersGuangyi Wang, Wei Peng, Lijiang Li, Wenyu Chen 等ICML 2025
- PFDiff: Training-Free Acceleration of Diffusion Models Combining Past and Future ScoresGuangyi Wang, Yuren Cai, Lijiang Li, Wei Peng 等ICLR 2025
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 被引用 167 次
相关 Paper
- Fast and Accurate Deep Bidirectional Language Representations for Unsupervised LearningJoongbo Shin, Yoonhyung Lee, Seunghyun Yoon, Kyomin JungACL 2020 · 被引用 6 次
- SLM: Learning a Discourse Language Representation with Sentence UnshufflingHaejun Lee, Drew A. Hudson, Kangwook Lee, Christopher D. ManningEMNLP 2020 · 被引用 2 次
- Self-Calibrated Listwise Reranking with Large Language ModelsRuiyang Ren, Yuhao Wang, Kun Zhou, Wayne Xin Zhao 等WWW 2025 · 被引用 12 次
- Leveraging Estimated Transferability Over Human Intuition for Model Selection in Text RankingJun Bai, Zhuofan Chen, Zhenzi Li, Hanhua Hong 等EMNLP 2024 · 被引用 2 次
- Deep Attentive Ranking Networks for Learning to Order SentencesPawan Kumar, Dhanajit Brahma, Harish Karnick, Piyush RaiAAAI 2020 · 被引用 52 次
