LLM-enhanced Self-training for Cross-domain Constituency Parsing
Jianling Li, Meishan Zhang, Peiming Guo, Min Zhang, Yue Zhang
Abstract
Self-training has proven to be an effective approach for cross-domain tasks, and in this study, we explore its application to cross-domain constituency parsing. Traditional self-training methods rely on limited and potentially lowquality raw corpora. To overcome this limitation, we propose enhancing self-training with the large language model (LLM) to generate domain-specific raw corpora iteratively. For the constituency parsing, we introduce grammar rules that guide the LLM in generating raw corpora and establish criteria for selecting pseudo instances. Our experimental results demonstrate that self-training for constituency parsing, equipped with an LLM, outperforms traditional methods regardless of the LLM's performance. Moreover, the combination of grammar rules and confidence criteria for pseudo-data selection yields the highest performance in the crossdomain constituency parsing 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 137ef1f0-5a95-4eb1-93f8-d4c185022b40Cited by top-tier papers3
- Large Language Models Are No Longer Shallow ParsersYuanhe Tian, Fei Xia, Yan SongACL 2024
- Contrastive Learning on LLM Back Generation Treebank for Cross-domain Constituency ParsingPeiming Guo, Meishan Zhang, Jianling Li, Min Zhang et al.ACL 2025
- Dialect-Agnostic SQL Parsing via LLM-Based SegmentationJunwen An, Kabilan Mahathevan, Manuel RiggerSIGMOD 2026
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen et al.EMNLP 2023 · 449 citations
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 331 citations
- Revisiting Self-Training for Neural Sequence GenerationJunxian He, Jiatao Gu, Jiajun Shen, Marc'Aurelio RanzatoICLR 2020 · 294 citations
- Grammar Prompting for Domain-Specific Language Generation with Large Language ModelsBailin Wang, Zi Wang, Xuezhi Wang, Yuan Cao et al.NeurIPS 2023 · 138 citations
Related papers
- Don't Parse, Choose Spans! Continuous and Discontinuous Constituency Parsing via Autoregressive Span SelectionSonglin Yang, Kewei TuACL 2023 · 1 citation
- Are Pre-trained Language Models Aware of Phrases? Simple but Strong Baselines for Grammar InductionTaeuk Kim, Jihun Choi, Daniel Edmiston, Sang-goo LeeICLR 2020 · 92 citations
- Phrase-aware Unsupervised Constituency ParsingXiaotao Gu, Yikang Shen, Jiaming Shen, Jingbo Shang et al.ACL 2022
- Unsupervised Parsing via Constituency TestsSteven Cao, Nikita Kitaev, Dan KleinEMNLP 2020 · 25 citations
- StructFormer: Joint Unsupervised Induction of Dependency and Constituency Structure from Masked Language ModelingYikang Shen, Yi Tay, Che Zheng, Dara Bahri et al.ACL 2021
