Self-Supervised Knowledge Assimilation for Expert-Layman Text Style Transfer
Wenda Xu, Michael Saxon, Misha Sra, William Yang Wang
摘要
Expert-layman text style transfer technologies have the potential to improve communication between members of scientific communities and the general public. High-quality information produced by experts is often filled with difficult jargon laypeople struggle to understand. This is a particularly notable issue in the medical domain, where layman are often confused by medical text online. At present, two bottlenecks interfere with the goal of building high-quality medical expert-layman style transfer systems: a dearth of pretrained medical-domain language models spanning both expert and layman terminologies and a lack of parallel corpora for training the transfer task itself. To mitigate the first issue, we propose a novel language model (LM) pretraining task, Knowledge Base Assimilation, to synthesize pretraining data from the edges of a graph of expert-and layman-style medical terminology terms into an LM during self-supervised learning. To mitigate the second issue, we build a large-scale parallel corpus in the medical expert-layman domain using a margin-based criterion. Our experiments show that transformer-based models pretrained on knowledge base assimilation and other well-established pretraining tasks fine-tuning on our new parallel corpus leads to considerable improvement against expert-layman transfer benchmarks, gaining an average relative improvement of our human evaluation, the Overall Success Rate (OSR), by 106%. We release our code and parallel corpus for future research 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Multi-Task Self-Supervised Learning for Disfluency DetectionShaolei Wang, Wanxiang Che, Qi Liu, Pengda Qin 等AAAI 2020 · 被引用 56 次
- Expertise Style Transfer: A New Task Towards Better Communication between Experts and LaymenYixin Cao, Ruihao Shui, Liangming Pan, Min-Yen Kan 等ACL 2020 · 被引用 50 次
- CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WebHolger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave 等ACL 2021
相关 Paper
- Non-Parallel Text Style Transfer with Self-Parallel SupervisionRuibo Liu, Chongyang Gao, Chenyan Jia, Guangxuan Xu 等ICLR 2022 · 被引用 19 次
- Monolingual Transfer Learning via Bilingual Translators for Style-Sensitive Paraphrase GenerationTomoyuki Kajiwara, Biwa Miura, Yuki AraseAAAI 2020 · 被引用 8 次
- Enhancing Biomedical Lay Summarisation with External Knowledge GraphsTomas Goldsack, Zhihao Zhang, Chen Tang, Carolina Scarton 等EMNLP 2023 · 被引用 2 次
- Automated Lay Language Summarization of Biomedical Scientific ReviewsYue Guo, Wei Qiu, Yizhong Wang, Trevor CohenAAAI 2021 · 被引用 100 次
- MEDICAL IMAGE UNDERSTANDING WITH PRETRAINED VISION LANGUAGE MODELS: A COMPREHENSIVE STUDYZiyuan Qin, Huahui Yi, Qicheng Lao, Kang LiICLR 2023 · 被引用 25 次
