Pre-training Universal Language Representation
Yian Li, Hai Zhao
摘要
Despite the well-developed cut-edge representation learning for language, most language representation models usually focus on specific levels of linguistic units. This work introduces universal language representation learning, i.e., embeddings of different levels of linguistic units or text with quite diverse lengths in a uniform vector space. We propose the training objective MiSAD that utilizes meaningful n-grams extracted from large unlabeled corpus by a simple but effective algorithm for pre-trained language models. Then we empirically verify that well designed pretraining scheme may effectively yield universal language representation, which will bring great convenience when handling multiple layers of linguistic objects in a unified way. Especially, our model achieves the highest accuracy on analogy tasks in different language levels and significantly improves the performance on downstream tasks in the GLUE benchmark and a question answering dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Learning Better Masking for Better Language Model Pre-trainingDongjie Yang, Zhuosheng Zhang, Hai ZhaoACL 2023 · 被引用 9 次
- StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical UnderstandingCheng Jiayang, Lin Qiu, Tsz Ho Chan, Tianqing Fang 等EMNLP 2023 · 被引用 8 次
- Language Model Pre-training on True NegativesZhuosheng Zhang, Hai Zhao, Masao Utiyama, Eiichiro SumitaAAAI 2023 · 被引用 3 次
- Instance Regularization for Discriminative Language Model Pre-trainingZhuosheng Zhang, Hai Zhao, Ming ZhouEMNLP 2022 · 被引用 1 次
- Can Pre-trained Language Models Interpret Similes as Smart as Human?Qianyu He, Sijie Cheng, Zhixu Li, Rui Xie 等ACL 2022
它引用的顶会 Paper6
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language UnderstandingWei Wang, Bin Bi, Ming Yan, Chen Wu 等ICLR 2020 · 被引用 297 次
- Conversational Semantic Parsing for Dialog State TrackingJianpeng Cheng, Devang Agrawal, Héctor Martínez Alonso, Shruti Bhargava 等EMNLP 2020 · 被引用 41 次
- Unsupervised Dual Paraphrasing for Two-stage Semantic ParsingRuisheng Cao, Su Zhu, Chenyu Yang, Chen Liu 等ACL 2020 · 被引用 37 次
相关 Paper
- SLM: Learning a Discourse Language Representation with Sentence UnshufflingHaejun Lee, Drew A. Hudson, Kangwook Lee, Christopher D. ManningEMNLP 2020 · 被引用 2 次
- UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled DataChengyi Wang, Yu Wu, Yao Qian, Ken'ichi Kumatani 等ICML 2021 · 被引用 140 次
- MUG: Meta-path-aware Universal Heterogeneous Graph Pre-TrainingLianze Shan, Jitao Zhao, Dongxiao He, Yongqi Huang 等AAAI 2026 · 被引用 1 次
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang 等ICML 2020 · 被引用 423 次
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
