LLM-enhanced Self-training for Cross-domain Constituency Parsing
Jianling Li, Meishan Zhang, Peiming Guo, Min Zhang, Yue Zhang
摘要
Self-training has proven to be an effective approach for cross-domain tasks, and in this study, we explore its application to cross-domain constituency parsing. Traditional self-training methods rely on limited and potentially lowquality raw corpora. To overcome this limitation, we propose enhancing self-training with the large language model (LLM) to generate domain-specific raw corpora iteratively. For the constituency parsing, we introduce grammar rules that guide the LLM in generating raw corpora and establish criteria for selecting pseudo instances. Our experimental results demonstrate that self-training for constituency parsing, equipped with an LLM, outperforms traditional methods regardless of the LLM's performance. Moreover, the combination of grammar rules and confidence criteria for pseudo-data selection yields the highest performance in the crossdomain constituency parsing 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Large Language Models Are No Longer Shallow ParsersYuanhe Tian, Fei Xia, Yan SongACL 2024
- Contrastive Learning on LLM Back Generation Treebank for Cross-domain Constituency ParsingPeiming Guo, Meishan Zhang, Jianling Li, Min Zhang 等ACL 2025
- Dialect-Agnostic SQL Parsing via LLM-Based SegmentationJunwen An, Kabilan Mahathevan, Manuel RiggerSIGMOD 2026
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen 等EMNLP 2023 · 被引用 449 次
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 被引用 331 次
- Revisiting Self-Training for Neural Sequence GenerationJunxian He, Jiatao Gu, Jiajun Shen, Marc'Aurelio RanzatoICLR 2020 · 被引用 294 次
- Grammar Prompting for Domain-Specific Language Generation with Large Language ModelsBailin Wang, Zi Wang, Xuezhi Wang, Yuan Cao 等NeurIPS 2023 · 被引用 138 次
相关 Paper
- Don't Parse, Choose Spans! Continuous and Discontinuous Constituency Parsing via Autoregressive Span SelectionSonglin Yang, Kewei TuACL 2023 · 被引用 1 次
- Are Pre-trained Language Models Aware of Phrases? Simple but Strong Baselines for Grammar InductionTaeuk Kim, Jihun Choi, Daniel Edmiston, Sang-goo LeeICLR 2020 · 被引用 92 次
- Phrase-aware Unsupervised Constituency ParsingXiaotao Gu, Yikang Shen, Jiaming Shen, Jingbo Shang 等ACL 2022
- Unsupervised Parsing via Constituency TestsSteven Cao, Nikita Kitaev, Dan KleinEMNLP 2020 · 被引用 25 次
- StructFormer: Joint Unsupervised Induction of Dependency and Constituency Structure from Masked Language ModelingYikang Shen, Yi Tay, Che Zheng, Dara Bahri 等ACL 2021
