Moment Distributionally Robust Tree Structured Prediction
Yeshu Li, Danyal Saeed, Xinhua Zhang, Brian D. Ziebart, Kevin Gimpel
Abstract
Structured prediction of tree-shaped objects is heavily studied under the name of syntactic dependency parsing. Current practice based on maximum likelihood or margin is either agnostic to or inconsistent with the evaluation loss. Risk minimization alleviates the discrepancy between training and test objectives but typically induces a non-convex problem. These approaches adopt explicit regularization to combat overfitting without probabilistic interpretation. We propose a moment-based distributionally robust optimization approach for tree structured prediction, where the worst-case expected loss over a set of distributions within bounded moment divergence from the empirical distribution is minimized. We develop efficient algorithms for arborescences and other variants of trees. We derive Fisher consistency, convergence rates and generalization bounds for our proposed method. We evaluate its empirical effectiveness on dependency parsing benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Distributionally Robust Optimization with Bias and Variance ReductionRonak Mehta, Vincent Roulet, Krishna Pillutla, Zaïd HarchaouiICLR 2024 · 6 citations
- GRaMPa: Subword Regularisation by Skewing Uniform Segmentation Distributions with an Efficient Path-counting Markov ModelThomas Bauwens, David Kaczér, Miryam de LhoneuxACL 2025
Builds on5
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Large-Scale Methods for Distributionally Robust OptimizationDaniel Levy, Yair Carmon, John C. Duchi, Aaron SidfordNeurIPS 2020 · 281 citations
- Consistent Structured Prediction with Max-Min Margin Markov NetworksAlex Nowak, Francis R. Bach, Alessandro RudiICML 2020 · 16 citations
- A Root of a Problem: Optimizing Single-Root Dependency ParsingMilos Stanojevic, Shay B. CohenEMNLP 2021 · 5 citations
- Understanding the Mechanics of SPIGOT: Surrogate Gradients for Latent Structure LearningTsvetomila Mihaylova, Vlad Niculae, André F. T. MartinsEMNLP 2020
Related papers
- Non-convex Distributionally Robust Optimization: Non-asymptotic AnalysisJikai Jin, Bohang Zhang, Haiyang Wang, Liwei WangNeurIPS 2021 · 65 citations
- Distributionally Robust Skeleton Learning of Discrete Bayesian NetworksYeshu Li, Brian D. ZiebartNeurIPS 2023 · 1 citation
- Distributionally Robust Optimization via Ball Oracle AccelerationYair Carmon, Danielle HauslerNeurIPS 2022 · 23 citations
- Generalization Bounds with Minimal Dependency on Hypothesis Class via Distributionally Robust OptimizationYibo Zeng, Henry LamNeurIPS 2022 · 11 citations
- Adaptive Sampling for Stochastic Risk-Averse LearningSebastian Curi, Kfir Y. Levy, Stefanie Jegelka, Andreas KrauseNeurIPS 2020 · 65 citations
