One2Set + Large Language Model: Best Partners for Keyphrase Generation
Liangying Shao, Liang Zhang, Minlong Peng, Guoqi Ma, Hao Yue, Mingming Sun, Jinsong Su
Abstract
Keyphrase generation (KPG) aims to automatically generate a collection of phrases representing the core concepts of a given document. The dominant paradigms in KPG include ONE2SEQ and ONE2SET. Recently, there has been increasing interest in applying large language models (LLMs) to KPG. Our preliminary experiments reveal that it is challenging for a single model to excel in both recall and precision. Further analysis shows that: 1) the ONE2SET paradigm owns the advantage of high recall, but suffers from improper assignments of supervision signals during training; 2) LLMs are powerful in keyphrase selection, but existing selection methods often make redundant selections. Given these observations, we introduce a generate-then-select framework decomposing KPG into two steps, where we adopt a ONE2SET-based model as generator to produce candidates and then use an LLM as selector to select keyphrases from these candidates. Particularly, we make two important improvements on our generator and selector: 1) we design an Optimal Transport-based assignment strategy to address the above improper assignments; 2) we model the keyphrase selection as a sequence labeling task to alleviate redundant selections. Experimental results on multiple benchmark datasets show that our framework significantly surpasses state-ofthe-art models, especially in absent keyphrase prediction. We release our code at https: //github.com/DeepLearnXMU/KPG-SetLLM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ece08fc5-6d74-4cfd-9e94-40bc36f8089dCited by top-tier papers2
- Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language ModelsQihang Ma, Shengyu Li, Jie Tang, Dingkang Yang et al.EMNLP 2025
- Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase GenerationJiajun Cao, Qinggang Zhang, Yunbo Tang, Zhishang Xiang et al.AAAI 2026
Builds on15
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsWeiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang et al.EMNLP 2023 · 182 citations
- One Size Does Not Fit All: Generating and Evaluating Variable Number of KeyphrasesXingdi Yuan, Tong Wang, Rui Meng, Khushboo Thaker et al.ACL 2020 · 76 citations
- Improving Passage Retrieval with Zero-Shot Question GenerationDevendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan et al.EMNLP 2022 · 69 citations
- Exclusive Hierarchical Decoding for Deep Keyphrase GenerationWang Chen, Hou Pong Chan, Piji Li, Irwin KingACL 2020 · 62 citations
Related papers
- Rethinking Model Selection and Decoding for Keyphrase Generation with Pre-trained Sequence-to-Sequence ModelsDi Wu, Wasi Uddin Ahmad, Kai-Wei ChangEMNLP 2023 · 6 citations
- One2Set: Generating Diverse Keyphrases as a SetJiacheng Ye, Tao Gui, Yichao Luo, Yige Xu et al.ACL 2021
- WR-One2Set: Towards Well-Calibrated Keyphrase GenerationBinbin Xie, Xiangpeng Wei, Baosong Yang, Huan Lin et al.EMNLP 2022 · 11 citations
- Select, Extract and Generate: Neural Keyphrase Generation with Layer-wise Coverage AttentionWasi Uddin Ahmad, Xiao Bai, Soomin Lee, Kai-Wei ChangACL 2021
- On the Role of Discriminative Models in Generative Relation ExtractionGuozheng Li, Peng Wang, Zijie Xu, Jing Zhou et al.ACL 2026
