Contextual Representation Learning beyond Masked Language Modeling
Zhiyi Fu, Wangchunshu Zhou, Jingjing Xu, Hao Zhou, Lei Li
摘要
How do masked language models (MLMs) such as BERT learn contextual representations? In this work, we analyze the learning dynamics of MLMs. We find that MLMs adopt sampled embeddings as anchors to estimate and inject contextual semantics to representations, which limits the efficiency and effectiveness of MLMs. To address these issues, we propose TACO, a simple yet effective representation learning approach to directly model global semantics. TACO extracts and aligns contextual semantics hidden in contextualized representations to encourage models to attend global semantics when generating contextualized representations. Experiments on the GLUE benchmark show that TACO achieves up to 5x speedup and up to 1.2 points average improvement over existing MLMs. The code is available at https:// github.com/FUZHIYI/TACO .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- VLUE: A Multi-Task Multi-Dimension Benchmark for Evaluating Vision-Language Pre-trainingWangchunshu Zhou, Yan Zeng, Shizhe Diao, Xinsong ZhangICML 2022 · 被引用 17 次
- Multi-CLS BERT: An Efficient Alternative to Traditional EnsemblingHaw-Shiuan Chang, Ruei-Yao Sun, Kathryn Ricci, Andrew McCallumACL 2023 · 被引用 8 次
- Adapting a Language Model While Preserving its General KnowledgeZixuan Ke, Yijia Shao, Haowei Lin, Hu Xu 等EMNLP 2022 · 被引用 6 次
- Write and Paint: Generative Vision-Language Models are Unified Modal LearnersShizhe Diao, Wangchunshu Zhou, Xinsong Zhang, Jiawei WangICLR 2023 · 被引用 5 次
- SESCORE2: Learning Text Generation Evaluation via Synthesizing Realistic MistakesWenda Xu, Xian Qian, Mingxuan Wang, Lei Li 等ACL 2023 · 被引用 3 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
相关 Paper
- ExLM: Rethinking the Impact of [MASK] Tokens in Masked Language ModelsKangjie Zheng, Junwei Yang, Siyue Liang, Bin Feng 等ICML 2025
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- TACO: Enhancing Multimodal In-context Learning via Task Mapping-Guided Sequence ConfigurationYanshu Li, Jianjiang Yang, Tian Yun, Pinyuan Feng 等EMNLP 2025 · 被引用 2 次
- Semantics-Aware BERT for Language UnderstandingZhuosheng Zhang, Yuwei Wu, Hai Zhao, Zuchao Li 等AAAI 2020 · 被引用 396 次
- RetroMAE-2: Duplex Masked Auto-Encoder For Pre-Training Retrieval-Oriented Language ModelsZheng Liu, Shitao Xiao, Yingxia Shao, Zhao CaoACL 2023 · 被引用 7 次
