Pareto Probing: Trading Off Accuracy for Complexity
Tiago Pimentel, Naomi Saphra, Adina Williams, Ryan Cotterell
Abstract
The question of how to probe contextual word representations for linguistic structure in a way that is both principled and useful has seen significant attention recently in the NLP literature. In our contribution to this discussion, we argue for a probe metric that reflects the fundamental trade-off between probe complexity and performance: the Pareto hypervolume. To measure complexity, we present a number of parametric and non-parametric metrics. Our experiments using Pareto hypervolume as an evaluation metric show that probes often do not conform to our expectations-e.g., why should the non-contextual fastText representations encode more morpho-syntactic information than the contextual BERT representations? These results suggest that common, simplistic probing tasks, such as part-of-speech labeling and dependency arc labeling, are inadequate to evaluate the linguistic structure encoded in contextual word representations. This leads us to propose full dependency parsing as a probing task. In support of our suggestion that harder probing tasks are necessary, our experiments with dependency parsing reveal a wide gap in syntactic knowledge between contextual and non-contextual representations. Our code can be found at https://github. com/rycolab/pareto-probing .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for LittleKoustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau et al.EMNLP 2021 · 177 citations
- Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMsAngelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt et al.ICLR 2024 · 119 citations
- Probing for the Usage of Grammatical NumberKarim Lasri, Tiago Pimentel, Alessandro Lenci, Thierry Poibeau et al.ACL 2022 · 72 citations
- Beyond Linear Probes: Dynamic Safety Monitoring for Language ModelsJames Oldfield, Philip Torr, Ioannis Patras, Adel Bibi et al.ICLR 2026 · 16 citations
- A Bayesian Framework for Information-Theoretic ProbingTiago Pimentel, Ryan CotterellEMNLP 2021 · 2 citations
Builds on4
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Truth or backpropaganda? An empirical investigation of deep learning theoryMicah Goldblum, Jonas Geiping, Avi Schwarzschild, Michael Moeller et al.ICLR 2020 · 36 citations
- Information-Theoretic Probing with Minimum Description LengthElena Voita, Ivan TitovEMNLP 2020 · 34 citations
- Information-Theoretic Probing for Linguistic StructureTiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod et al.ACL 2020 · 21 citations
Related papers
- Intrinsic Probing through Dimension SelectionLucas Torroba Hennigen, Adina Williams, Ryan CotterellEMNLP 2020 · 3 citations
- Probing BERT in Hyperbolic SpacesBoli Chen, Yao Fu, Guangwei Xu, Pengjun Xie et al.ICLR 2021 · 19 citations
- Probing as Quantifying Inductive BiasAlexander Immer, Lucas Torroba Hennigen, Vincent Fortuin, Ryan CotterellACL 2022
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTZhiyong Wu, Yun Chen, Ben Kao, Qun LiuACL 2020 · 158 citations
- Bird's Eye: Probing for Linguistic Graph Structures with a Simple Information-Theoretic ApproachYifan Hou, Mrinmaya SachanACL 2021
