Unsupervised Parsing with S-DIORA: Single Tree Encoding for Deep Inside-Outside Recursive Autoencoders
Andrew Drozdov, Subendhu Rongali, Yi-Pei Chen, Tim O'Gorman, Mohit Iyyer, Andrew McCallum
Abstract
The deep inside-outside recursive autoencoder (DIORA; Drozdov et al. 2019a) is a selfsupervised neural model that learns to induce syntactic tree structures for input sentences without access to labeled training data. In this paper, we discover that while DIORA exhaustively encodes all possible binary trees of a sentence with a soft dynamic program, its vector averaging approach is locally greedy and cannot recover from errors when computing the highest scoring parse tree in bottom-up chart parsing. To fix this issue, we introduce S-DIORA, an improved variant of DIORA that encodes a single tree rather than a softlyweighted mixture of trees by employing a hard argmax operation and a beam at each cell in the chart. Our experiments show that through fine-tuning a pre-trained DIORA with our new algorithm, we improve the state of the art in unsupervised constituency parsing on the English WSJ Penn Treebank by 2.2 6% F1, depending on the data used for fine-tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 220df925-d534-4ed5-ba25-444eb3896b72Cited by top-tier papers17
- Modeling Hierarchical Structures with Continuous Recursive Neural NetworksJishnu Ray Chowdhury, Cornelia CarageaICML 2021 · 18 citations
- Ensemble Distillation for Unsupervised Constituency ParsingBehzad Shayegh, Yanshuai Cao, Xiaodan Zhu, Jackie C. K. Cheung et al.ICLR 2024 · 10 citations
- Unsupervised Vision-Language Parsing: Seamlessly Bridging Visual Scene Graphs with Language Structures via Dependency RelationshipsChao Lou, Wenjuan Han, Yuhuan Lin, Zilong ZhengCVPR 2022 · 9 citations
- Beam Tree Recursive CellsJishnu Ray Chowdhury, Cornelia CarageaICML 2023 · 7 citations
- Fast-R2D2: A Pretrained Recursive Neural Network based on Pruned CKY for Grammar Induction and Text RepresentationXiang Hu, Haitao Mi, Liang Li, Gerard de MeloEMNLP 2022 · 7 citations
Builds on2
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language ModelsCheolhyoung Lee, Kyunghyun Cho, Wanmo KangICLR 2020 · 233 citations
- Are Pre-trained Language Models Aware of Phrases? Simple but Strong Baselines for Grammar InductionTaeuk Kim, Jihun Choi, Daniel Edmiston, Sang-goo LeeICLR 2020 · 92 citations
Related papers
- Improved Latent Tree Induction with Distant Supervision via Span ConstraintsZhiyang Xu, Andrew Drozdov, Jay-Yoon Lee, Tim O'Gorman et al.EMNLP 2021 · 4 citations
- Unsupervised Parsing via Constituency TestsSteven Cao, Nikita Kitaev, Dan KleinEMNLP 2020 · 25 citations
- Word Segmentation as Unsupervised Constituency ParsingRaquel G. AlhamaACL 2022
- Tree-Averaging Algorithms for Ensemble-Based Unsupervised Discontinuous Constituency ParsingBehzad Shayegh, Yuqiao Wen, Lili MouACL 2024
- Semi-Supervised Semantic Dependency Parsing Using CRF AutoencodersZixia Jia, Youmi Ma, Jiong Cai, Kewei TuACL 2020 · 10 citations
