Understanding the Mechanics of SPIGOT: Surrogate Gradients for Latent Structure Learning
Tsvetomila Mihaylova, Vlad Niculae, André F. T. Martins
摘要
Latent structure models are a powerful tool for modeling language data: they can mitigate the error propagation and annotation bottleneck in pipeline systems, while simultaneously uncovering linguistic insights about the data. One challenge with end-to-end training of these models is the argmax operation, which has null gradient. In this paper, we focus on surrogate gradients, a popular strategy to deal with this problem. We explore latent structure learning through the angle of pulling back the downstream learning objective. In this paradigm, we discover a principled motivation for both the straight-through estimator (STE) as well as the recently-proposed SPIGOT-a variant of STE for structured models. Our perspective leads to new algorithms in the same family. We empirically compare the known and the novel pulled-back estimators against the popular alternatives, yielding new insight for practitioners and revealing intriguing failure cases.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- On Learning Latent Models with Multi-Instance Weak SupervisionKaifu Wang, Efthymia Tsamoura, Dan RothNeurIPS 2023 · 被引用 19 次
- To be Continuous, or to be Discrete, Those are Bits of QuestionsYiran Wang, Masao UtiyamaACL 2024 · 被引用 3 次
- Moment Distributionally Robust Tree Structured PredictionYeshu Li, Danyal Saeed, Xinhua Zhang, Brian D. Ziebart 等NeurIPS 2022 · 被引用 2 次
- Imbalances in Neurosymbolic Learning: Characterization and Mitigating StrategiesEfthymia Tsamoura, Kaifu Wang, Dan RothNeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper2
相关 Paper
- Leveraging Recursive Gumbel-Max Trick for Approximate Inference in Combinatorial SpacesKirill Struminsky, Artyom Gadetsky, Denis Rakitin, Danil Karpushkin 等NeurIPS 2021 · 被引用 11 次
- Straight to the Gradient: Learning to Use Novel Tokens for Neural Text GenerationXiang Lin, Simeng Han, Shafiq R. JotyICML 2021 · 被引用 30 次
- Revealing Procedural Reasoning Structures in Chain-of-Thought Training via Span-Level Gradient OrganizationJia Liu, Jiaxin Luo, Weiwen Xu, Jonathan M. Garibaldi 等ACL 2026
- Markovian Transformers for Informative Language ModelingScott Viteri, Max Lamparth, Peter Chatain, Clark W. BarrettICLR 2026 · 被引用 3 次
- Efficient Marginalization of Discrete and Structured Latent Variables via SparsityGonçalo M. Correia, Vlad Niculae, Wilker Aziz, André F. T. MartinsNeurIPS 2020 · 被引用 25 次
