Lune

ACL2021顶会

BERTAC: Enhancing Transformer-based Language Models with Adversarially Pretrained Convolutional Neural Networks

Jong-Hoon Oh, Ryu Iida, Julien Kloetzer, Kentaro Torisawa

2021年份

摘要

Transformer-based language models (TLMs), such as BERT, ALBERT and GPT-3, have shown strong performance in a wide range of NLP tasks and currently dominate the field of NLP. However, many researchers wonder whether these models can maintain their dominance forever. Of course, we do not have answers now, but, as an attempt to find better neural architectures and training schemes, we pretrain a simple CNN using a GAN-style learning scheme and Wikipedia data, and then integrate it with standard TLMs. We show that on the GLUE tasks, the combination of our pretrained CNN with ALBERT outperforms the original ALBERT and achieves a similar performance to that of SOTA. Furthermore, on open-domain QA (Quasar-T and SearchQA), the combination of the CNN with ALBERT or RoBERTa achieved stronger performance than SOTA and the original TLMs. We hope that this work provides a hint for developing a novel strong network architecture along with its training scheme. Our source code and models are available at https://github.com/nict-wisdom/bertac . !"#%"&"'()*% ! !" +')"),-&(#.+/0#+')+'+0 #10+')"), 2%+(340*%025(.+4 6+(3-+')"),-%+7%+#+')()"'0 8+'+%()%0" 9(.+-+')"),-%+7%+#+')()"'0 8+'+%()%0# "!"#%+(30%+7%+#+')()"*'0 50):+0+')"),0! #!%"#5(.+05(.+0%+7%+#+')()"*'0 *50):+0+')"),0! !%&'()'+%!,-.,(/0(1" ![EM] ,2-3+',4')562-!',)-,)1#()'1,0)'4-',(-+%*7"

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper7

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖