Keyphrase Generation via Soft and Hard Semantic Corrections
Guangzhen Zhao, Guoshun Yin, Peng Yang, Yu Yao
摘要
Keyphrase generation aims to generate a set of condensed phrases given a source document. Although maximum likelihood estimation (MLE) based keyphrase generation methods have shown impressive performance, they suffer from the bias on the source-prediction pair and the bias on the prediction-target pair. To tackle the above biases, we propose a novel correction model CorrKG on top of the MLE pipeline, where the biases are corrected via the optimal transport (OT) and a frequency-based filtering-and-sorting (FreqFS) strategy. Specifically, OT is introduced as the soft correction to facilitate the alignment of salient information and rectify the semantic bias on the source document and predicted keyphrases pair. An adaptive semantic mass learning scheme is conducted on the vanilla OT to achieve a proper pair-wise optimal transport procedure, which promotes the OT calculation brought by rectifying semantic masses dynamically. Besides, the FreqFS strategy is designed as the hard correction to reduce the bias of predicted and target keyphrases, and thus generate accurate and sufficient keyphrases. Extensive experiments over multiple benchmark datasets show that our model achieves superior keyphrase generation as compared with the state-of-the-arts. Introduction Keyphrase generation is an important and meaningful task that converts the main semantic information of the document into multiple keyphrases. Keyphrases can further be divided into present keyphrases and absent keyphrases, with the former appearing in the document whereas the latter do not. High-quality keyphrases are beneficial for many downstream tasks, such as text summarization (Wang and Cardie, 2013), document clustering (Hammouda et al., 2005) , translation (Tang et al., 2016) , and so forth. Despite the promising suc-* The first two authors contribute equally to this work. ‡ Corresponding author. Source Document Learning Weights for the Quasi Weighted Means. We study the determination of weights for quasi weighted means (also called quasi linear means) when a set of examples is given. We consider first a simple case, the learning of weights for weighted means, and then we extend the approach to the more general case of a quasi weighted mean. We consider the case of a known arbitrary generator f. The paper finishes considering the use of parametric functions that are suitable when the values to aggregate are measure values or ratio.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Rethinking Model Selection and Decoding for Keyphrase Generation with Pre-trained Sequence-to-Sequence ModelsDi Wu, Wasi Uddin Ahmad, Kai-Wei ChangEMNLP 2023 · 被引用 6 次
- One2Set + Large Language Model: Best Partners for Keyphrase GenerationLiangying Shao, Liang Zhang, Minlong Peng, Guoqi Ma 等EMNLP 2024 · 被引用 2 次
它引用的顶会 Paper7
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- One Size Does Not Fit All: Generating and Evaluating Variable Number of KeyphrasesXingdi Yuan, Tong Wang, Rui Meng, Khushboo Thaker 等ACL 2020 · 被引用 76 次
- Exclusive Hierarchical Decoding for Deep Keyphrase GenerationWang Chen, Hou Pong Chan, Piji Li, Irwin KingACL 2020 · 被引用 62 次
- Fast and Constrained Absent Keyphrase Generation by Prompt-Based LearningHuanqin Wu, Baijiaxin Ma, Wei Liu, Tao Chen 等AAAI 2022 · 被引用 31 次
相关 Paper
- WR-One2Set: Towards Well-Calibrated Keyphrase GenerationBinbin Xie, Xiangpeng Wei, Baosong Yang, Huan Lin 等EMNLP 2022 · 被引用 11 次
- Adaptive Beam Search Decoding for Discrete Keyphrase GenerationXiaoli Huang, Tongge Xu, Lvan Jiao, Yueran Zu 等AAAI 2021 · 被引用 10 次
- Unsupervised Deep Keyphrase GenerationXianjie Shen, Yinghan Wang, Rui Meng, Jingbo ShangAAAI 2022 · 被引用 19 次
- Improving Text Generation with Student-Forcing Optimal TransportJianqiao Li, Chunyuan Li, Guoyin Wang, Hao Fu 等EMNLP 2020 · 被引用 11 次
- Calibrating Pseudo-Labeling with Class Distribution for Semi-supervised Text ClassificationWeiyi Yang, Richong Zhang, Junfan Chen, Jiawei ShengEMNLP 2025
