SLM: Learning a Discourse Language Representation with Sentence Unshuffling
Haejun Lee, Drew A. Hudson, Kangwook Lee, Christopher D. Manning
摘要
We introduce Sentence-level Language Modeling, a new pre-training objective for learning a discourse language representation in a fully self-supervised manner. Recent pre-training methods in NLP focus on learning either bottom or top-level language representations: contextualized word representations derived from language model objectives at one extreme and a whole sequence representation learned by order classification of two given textual segments at the other. However, these models are not directly encouraged to capture representations of intermediate-size structures that exist in natural languages such as sentences and the relationships among them. To that end, we propose a new approach to encourage learning of a contextualized sentence-level representation by shuffling the sequence of input sentences and training a hierarchical transformer model to reconstruct the original ordering. Through experiments on downstream tasks such as GLUE, SQuAD, and DiscoEval, we show that this feature of our model improves the performance of the original BERT by large margins.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Few-Shot Text Generation with Natural Language InstructionsTimo Schick, Hinrich SchützeEMNLP 2021 · 被引用 101 次
- Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial AttacksXinyu Zhang, Hanbin Hong, Yuan Hong, Peng Huang 等S&P 2024 · 被引用 41 次
- Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional ManualsTe-Lin Wu, Alexander Spangher, Pegah Alipoormolabashi, Marjorie Freedman 等ACL 2022 · 被引用 30 次
- Learning Temporal Dynamics from Cycles in Narrated VideoDave Epstein, Jiajun Wu, Cordelia Schmid, Chen SunICCV 2021 · 被引用 15 次
- PoNet: Pooling Network for Efficient Token Mixing in Long SequencesChao-Hong Tan, Qian Chen, Wen Wang, Qinglin Zhang 等ICLR 2022 · 被引用 15 次
它引用的顶会 Paper5
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- ERNIE 2.0: A Continual Pre-Training Framework for Language UnderstandingYu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng 等AAAI 2020 · 被引用 885 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language UnderstandingWei Wang, Bin Bi, Ming Yan, Chen Wu 等ICLR 2020 · 被引用 297 次
- Pretraining with Contrastive Sentence Objectives Improves Discourse Performance of Language ModelsDan Iter, Kelvin Guu, Larry Lansing, Dan JurafskyACL 2020 · 被引用 72 次
相关 Paper
- DialogBERT: Discourse-Aware Response Generation via Learning to Recover and Rank UtterancesXiaodong Gu, Kang Min Yoo, Jung-Woo HaAAAI 2021 · 被引用 83 次
- Less Mature is More Adaptable for Sentence-level Language ModelingAbhilasha Sancheti, David Dale, Artyom Kozhevnikov, Maha ElbayadACL 2025
- Sentence Representation Learning with Generative Objective rather than Contrastive ObjectiveBohong Wu, Hai ZhaoEMNLP 2022 · 被引用 3 次
- COCO-LM: Correcting and Contrasting Text Sequences for Language Model PretrainingYu Meng, Chenyan Xiong, Payal Bajaj, Saurabh Tiwary 等NeurIPS 2021 · 被引用 231 次
- Long Text Generation by Modeling Sentence-Level and Discourse-Level CoherenceJian Guan, Xiaoxi Mao, Changjie Fan, Zitao Liu 等ACL 2021
