Z-Code++: A Pre-trained Language Model Optimized for Abstractive Summarization
Pengcheng He, Baolin Peng, Song Wang, Yang Liu, Ruochen Xu, Hany Hassan, Yu Shi, Chenguang Zhu, Wayne Xiong, Michael Zeng, Jianfeng Gao, Xuedong Huang
摘要
This paper presents Z-Code++, a new pretrained language model optimized for abstractive text summarization. The model extends the state of the art encoder-decoder model using three techniques. First, we use a two-phase pre-training process to improve model's performance on low-resource summarization tasks. The model is first pre-trained using text corpora for language understanding, and then is continually pre-trained on summarization corpora for grounded text generation. Second, we replace self-attention layers in the encoder with disentangled attention layers, where each word is represented using two vectors that encode its content and position, respectively. Third, we use fusion-in-encoder, a simple yet effective method of encoding long sequences in a hierarchical manner. Z-Code++ creates new state of the art on 9 out of 13 text summarization tasks across 5 languages. Our model is parameterefficient in that it outperforms the 600x larger PaLM 540B on XSum, and the finetuned 200x larger GPT3 175B on SAMSum. In zero-shot and few-shot settings, our model substantially outperforms the competing models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human EvaluationYixin Liu, Alexander R. Fabbri, Pengfei Liu, Yilun Zhao 等ACL 2023 · 被引用 50 次
- Summaries, Highlights, and Action Items: Design, Implementation and Evaluation of an LLM-powered Meeting Recap SystemSumit Asthana, Sagih Hilleli, Pengcheng He, Aaron HalfakerCSCW 2025 · 被引用 24 次
- RLPF: Reinforcement Learning from Prediction Feedback for User Summarization with LLMsJiaxing Wu, Lin Ning, Luyang Liu, Harrison Lee 等AAAI 2025 · 被引用 12 次
- FiE: Building a Global Probability Space by Leveraging Early Fusion in Encoder for Open-Domain Question AnsweringAkhil Kedia, Mohd Abbas Zaidi, Haejun LeeEMNLP 2022 · 被引用 11 次
- FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual KnowledgeShangbin Feng, Vidhisha Balachandran, Yuyang Bai, Yulia TsvetkovEMNLP 2023 · 被引用 10 次
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
相关 Paper
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- Pre-training for Abstractive Document Summarization by Reinstating Source TextYanyan Zou, Xingxing Zhang, Wei Lu, Furu Wei 等EMNLP 2020 · 被引用 42 次
- PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned GenerationBin Bi, Chenliang Li, Chen Wu, Ming Yan 等EMNLP 2020 · 被引用 41 次
- Meta-Transfer Learning for Low-Resource Abstractive SummarizationYi-Syuan Chen, Hong-Han ShuaiAAAI 2021 · 被引用 41 次
- Cross-Lingual Natural Language Generation via Pre-TrainingZewen Chi, Li Dong, Furu Wei, Wenhui Wang 等AAAI 2020 · 被引用 142 次
