AMR Parsing is Far from Solved: GrAPES, the Granular AMR Parsing Evaluation Suite
Jonas Groschwitz, Shay B. Cohen, Lucia Donatelli, Meaghan Fowlie
Abstract
We present the Granular AMR Parsing Evaluation Suite (GrAPES), a challenge set for Abstract Meaning Representation (AMR) parsing with accompanying evaluation metrics. AMR parsers now obtain high scores on the standard AMR evaluation metric Smatch, close to or even above reported inter-annotator agreement. But that does not mean that AMR parsing is solved; in fact, human evaluation in previous work indicates that current parsers still quite frequently make errors on node labels or graph structure that substantially distort sentence meaning. Here, we provide an evaluation suite that tests AMR parsers on a range of phenomena of practical, technical, and linguistic interest. Our 36 categories range from seen and unseen labels, to structural generalization, to coreference. GrAPES reveals in depth the abilities and shortcomings of current AMR parsers. Set Category Metric AM Parser C&L AMRBart # 1 Pragmatic coreference (testset) Edge recall 06 [02, 18] 08 [03, 22] 39 [25, 55] 36 Prerequisites 50 [34, 66] 36 [22, 52] 61 [45, 75] 36 Pragmatic coreference (Winograd) Edge recall 02 [00, 13] 05 [01, 17] 32 [20, 48] 40 Prerequisites 78 [62, 88] 30 [18, 45] 65 [50, 78] 40 2 Syntactic (gap) reentrancies Edge recall 24 [14, 39] 24 [14, 39] 49 [34, 64] 41 Prerequisites 54 [39, 68] 59 [43, 72] 68 [53, 80] 41 Unambiguous coreference Edge recall 10 [03, 25] 39 [24, 56] 65 [47, 79] 31 Prerequisites 71 [53, 84] 71 [53, 84] 77 [60, 89] 31
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a67aad9e-b4dd-480d-9f78-a75424865a63Cited by top-tier papers1
Ask how each one uses itBuilds on7
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Scalable Zero-shot Entity Linking with Dense Entity RetrievalLedell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel et al.EMNLP 2020 · 336 citations
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 149 citations
- SLOG: A Structural Generalization Benchmark for Semantic ParsingBingzhi Li, Lucia Donatelli, Alexander Koller, Tal Linzen et al.EMNLP 2023 · 3 citations
- Fast semantic parsing with well-typedness guaranteesMatthias Lindemann, Jonas Groschwitz, Alexander KollerEMNLP 2020 · 2 citations
Related papers
- Evaluate AMR Graph Similarity via Self-supervised LearningZiyi Shou, Fangzhen LinACL 2023
- Probabilistic, Structure-Aware Algorithms for Improved Variety, Accuracy, and Coverage of AMR AlignmentsAustin Blodgett, Nathan SchneiderACL 2021
- XL-AMR: Enabling Cross-Lingual AMR Parsing with Transfer Learning TechniquesRexhina Blloshmi, Rocco Tripodi, Roberto NavigliEMNLP 2020 · 48 citations
- Cross-domain Generalization for AMR ParsingXuefeng Bai, Sen Yang, Leyang Cui, Linfeng Song et al.EMNLP 2022 · 1 citation
- Improving AMR Parsing with Sequence-to-Sequence Pre-trainingDongqin Xu, Junhui Li, Muhua Zhu, Min Zhang et al.EMNLP 2020 · 57 citations
