UCPhrase: Unsupervised Context-aware Quality Phrase Tagging
Xiaotao Gu, Zihan Wang, Zhenyu Bi, Yu Meng, Liyuan Liu, Jiawei Han, Jingbo Shang
摘要
Identifying and understanding quality phrases from context is a fundamental task in text mining. The most challenging part of this task arguably lies in uncommon, emerging, and domain-specific phrases. The infrequent nature of these phrases significantly hurts the performance of phrase mining methods that rely on sufficient phrase occurrences in the input corpus. Context-aware tagging models, though not restricted by frequency, heavily rely on domain experts for either massive sentence-level gold labels or handcrafted gazetteers. In this work, we propose UCPhrase, a novel unsupervised context-aware quality phrase tagger. Specifically, we induce high-quality phrase spans as silver labels from consistently co-occurring word sequences within each document. Compared with typical context-agnostic distant supervision based on existing knowledge bases (KBs), our silver labels root deeply in the input domain and context, thus having unique advantages in preserving contextual completeness and capturing emerging, out-of-KB phrases. Training a conventional neural tagger based on silver labels usually faces the risk of overfitting phrase surface names. Alternatively, we observe that the contextualized attention maps generated from a transformer-based neural language model effectively reveal the connections between words in a surface-agnostic way. Therefore, we pair such attention maps with the silver labels to train a lightweight span prediction model, which can be applied to new input to recognize (unseen) quality phrases regardless of their surface names or frequency. Thorough experiments on various tasks and datasets, including corpus-level phrase ranking, document-level keyphrase extraction, and sentence-level phrase tagging, demonstrate the superiority of our design over state-of-the-art pre-trained, unsupervised, and distantly supervised methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- TELEClass: Taxonomy Enrichment and LLM-Enhanced Hierarchical Text Classification with Minimal SupervisionYunyi Zhang, Ruozhen Yang, Xueqiang Xu, Rui Li 等WWW 2025 · 被引用 53 次
- OA-Mine: Open-World Attribute Mining for E-Commerce Products with Weak SupervisionXinyang Zhang, Chenwei Zhang, Xian Li, Xin Luna Dong 等WWW 2022 · 被引用 36 次
- Unsupervised Key Event Detection from Massive Text CorporaYunyi Zhang, Fang Guo, Jiaming Shen, Jiawei HanKDD 2022 · 被引用 14 次
- PDSum: Prototype-driven Continuous Summarization of Evolving Multi-document Sets StreamSusik Yoon, Hou Pong Chan, Jiawei HanWWW 2023 · 被引用 13 次
- Leveraging Attention to Effectively Compress Prompts for Long-Context LLMsYunlong Zhao, Haoran Wu, Bo XuAAAI 2025 · 被引用 10 次
它引用的顶会 Paper3
- BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant SupervisionChen Liang, Yue Yu, Haoming Jiang, Siawpeng Er 等KDD 2020 · 被引用 118 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
- Are Pre-trained Language Models Aware of Phrases? Simple but Strong Baselines for Grammar InductionTaeuk Kim, Jihun Choi, Daniel Edmiston, Sang-goo LeeICLR 2020 · 被引用 92 次
相关 Paper
- UCTopic: Unsupervised Contrastive Learning for Phrase Representations and Topic MiningJiacheng Li, Jingbo Shang, Julian J. McAuleyACL 2022 · 被引用 68 次
- AttentionRank: Unsupervised Keyphrase Extraction using Self and Cross AttentionsHaoran Ding, Xiao LuoEMNLP 2021 · 被引用 46 次
- Unsupervised Open-domain Keyphrase GenerationLam Do, Pritom Saha Akash, Kevin Chen-Chuan ChangACL 2023 · 被引用 2 次
- Unsupervised Deep Keyphrase GenerationXianjie Shen, Yinghan Wang, Rui Meng, Jingbo ShangAAAI 2022 · 被引用 19 次
- MUDY: Multi-Granular Dynamic Candidate Contextualization for Unsupervised Keyphrase ExtractionHyeongu Kang, Susik YoonSIGIR 2026
