Understanding the Mechanics of SPIGOT: Surrogate Gradients for Latent Structure Learning
Tsvetomila Mihaylova, Vlad Niculae, André F. T. Martins
Abstract
Latent structure models are a powerful tool for modeling language data: they can mitigate the error propagation and annotation bottleneck in pipeline systems, while simultaneously uncovering linguistic insights about the data. One challenge with end-to-end training of these models is the argmax operation, which has null gradient. In this paper, we focus on surrogate gradients, a popular strategy to deal with this problem. We explore latent structure learning through the angle of pulling back the downstream learning objective. In this paradigm, we discover a principled motivation for both the straight-through estimator (STE) as well as the recently-proposed SPIGOT-a variant of STE for structured models. Our perspective leads to new algorithms in the same family. We empirically compare the known and the novel pulled-back estimators against the popular alternatives, yielding new insight for practitioners and revealing intriguing failure cases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- On Learning Latent Models with Multi-Instance Weak SupervisionKaifu Wang, Efthymia Tsamoura, Dan RothNeurIPS 2023 · 19 citations
- To be Continuous, or to be Discrete, Those are Bits of QuestionsYiran Wang, Masao UtiyamaACL 2024 · 3 citations
- Moment Distributionally Robust Tree Structured PredictionYeshu Li, Danyal Saeed, Xinhua Zhang, Brian D. Ziebart et al.NeurIPS 2022 · 2 citations
- Imbalances in Neurosymbolic Learning: Characterization and Mitigating StrategiesEfthymia Tsamoura, Kaifu Wang, Dan RothNeurIPS 2025 · 2 citations
Builds on2
Related papers
- Leveraging Recursive Gumbel-Max Trick for Approximate Inference in Combinatorial SpacesKirill Struminsky, Artyom Gadetsky, Denis Rakitin, Danil Karpushkin et al.NeurIPS 2021 · 11 citations
- Straight to the Gradient: Learning to Use Novel Tokens for Neural Text GenerationXiang Lin, Simeng Han, Shafiq R. JotyICML 2021 · 30 citations
- Revealing Procedural Reasoning Structures in Chain-of-Thought Training via Span-Level Gradient OrganizationJia Liu, Jiaxin Luo, Weiwen Xu, Jonathan M. Garibaldi et al.ACL 2026
- Markovian Transformers for Informative Language ModelingScott Viteri, Max Lamparth, Peter Chatain, Clark W. BarrettICLR 2026 · 3 citations
- Efficient Marginalization of Discrete and Structured Latent Variables via SparsityGonçalo M. Correia, Vlad Niculae, Wilker Aziz, André F. T. MartinsNeurIPS 2020 · 25 citations
