Unsupervised Parsing with S-DIORA: Single Tree Encoding for Deep Inside-Outside Recursive Autoencoders
Andrew Drozdov, Subendhu Rongali, Yi-Pei Chen, Tim O'Gorman, Mohit Iyyer, Andrew McCallum
摘要
The deep inside-outside recursive autoencoder (DIORA; Drozdov et al. 2019a) is a selfsupervised neural model that learns to induce syntactic tree structures for input sentences without access to labeled training data. In this paper, we discover that while DIORA exhaustively encodes all possible binary trees of a sentence with a soft dynamic program, its vector averaging approach is locally greedy and cannot recover from errors when computing the highest scoring parse tree in bottom-up chart parsing. To fix this issue, we introduce S-DIORA, an improved variant of DIORA that encodes a single tree rather than a softlyweighted mixture of trees by employing a hard argmax operation and a beam at each cell in the chart. Our experiments show that through fine-tuning a pre-trained DIORA with our new algorithm, we improve the state of the art in unsupervised constituency parsing on the English WSJ Penn Treebank by 2.2 6% F1, depending on the data used for fine-tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Modeling Hierarchical Structures with Continuous Recursive Neural NetworksJishnu Ray Chowdhury, Cornelia CarageaICML 2021 · 被引用 18 次
- Ensemble Distillation for Unsupervised Constituency ParsingBehzad Shayegh, Yanshuai Cao, Xiaodan Zhu, Jackie C. K. Cheung 等ICLR 2024 · 被引用 10 次
- Unsupervised Vision-Language Parsing: Seamlessly Bridging Visual Scene Graphs with Language Structures via Dependency RelationshipsChao Lou, Wenjuan Han, Yuhuan Lin, Zilong ZhengCVPR 2022 · 被引用 9 次
- Beam Tree Recursive CellsJishnu Ray Chowdhury, Cornelia CarageaICML 2023 · 被引用 7 次
- Fast-R2D2: A Pretrained Recursive Neural Network based on Pruned CKY for Grammar Induction and Text RepresentationXiang Hu, Haitao Mi, Liang Li, Gerard de MeloEMNLP 2022 · 被引用 7 次
它引用的顶会 Paper2
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language ModelsCheolhyoung Lee, Kyunghyun Cho, Wanmo KangICLR 2020 · 被引用 233 次
- Are Pre-trained Language Models Aware of Phrases? Simple but Strong Baselines for Grammar InductionTaeuk Kim, Jihun Choi, Daniel Edmiston, Sang-goo LeeICLR 2020 · 被引用 92 次
相关 Paper
- Improved Latent Tree Induction with Distant Supervision via Span ConstraintsZhiyang Xu, Andrew Drozdov, Jay-Yoon Lee, Tim O'Gorman 等EMNLP 2021 · 被引用 4 次
- Unsupervised Parsing via Constituency TestsSteven Cao, Nikita Kitaev, Dan KleinEMNLP 2020 · 被引用 25 次
- Word Segmentation as Unsupervised Constituency ParsingRaquel G. AlhamaACL 2022
- Tree-Averaging Algorithms for Ensemble-Based Unsupervised Discontinuous Constituency ParsingBehzad Shayegh, Yuqiao Wen, Lili MouACL 2024
- Semi-Supervised Semantic Dependency Parsing Using CRF AutoencodersZixia Jia, Youmi Ma, Jiong Cai, Kewei TuACL 2020 · 被引用 10 次
