Multilingual AMR-to-Text Generation
Angela Fan, Claire Gardent
Abstract
Generating text from structured data is challenging because it requires bridging the gap between (i) structure and natural language (NL) and (ii) semantically underspecified input and fully specified NL output. Multilingual generation brings in an additional challenge: that of generating into languages with varied word order and morphological properties. In this work, we focus on Abstract Meaning Representations (AMRs) as structured input, where previous research has overwhelmingly focused on generating only into English. We leverage advances in cross-lingual embeddings, pretraining, and multilingual models to create multilingual AMR-to-text models that generate in twenty one different languages. For eighteen languages, based on automatic metrics, our multilingual models surpass baselines that generate into a single language. We analyse the ability of our multilingual models to accurately capture morphology and word order using human evaluation, and find that native speakers judge our generations to be fluent.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- T-STAR: Truthful Style Transfer using AMR Graph as Intermediate RepresentationAnubhav Jangra, Preksha Nema, Aravindan RaghuveerEMNLP 2022 · 1 citation
- XLPT-AMR: Cross-Lingual Pre-Training via Multi-Task Learning for Zero-Shot AMR Parsing and Text GenerationDongqin Xu, Junhui Li, Muhua Zhu, Min Zhang et al.ACL 2021
- Generalising Multilingual Concept-to-Text NLG with Language Agnostic DelexicalisationGiulio Zhou, Gerasimos LampourasACL 2021
Builds on5
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Bridging the Structural Gap Between Encoding and Decoding for Data-To-Text GenerationChao Zhao, Marilyn A. Walker, Snigdha ChaturvediACL 2020 · 82 citations
- MLQA: Evaluating Cross-lingual Extractive Question AnsweringPatrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel et al.ACL 2020 · 52 citations
- CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WebHolger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave et al.ACL 2021
Related papers
- Retrofitting Multilingual Sentence Embeddings with Abstract Meaning RepresentationDeng Cai, Xin Li, Jackie Chun-Sing Ho, Lidong Bing et al.EMNLP 2022 · 4 citations
- XL-AMR: Enabling Cross-Lingual AMR Parsing with Transfer Learning TechniquesRexhina Blloshmi, Rocco Tripodi, Roberto NavigliEMNLP 2020 · 48 citations
- Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language ModelsJames A. Michaelov, Catherine Arnett, Tyler A. Chang, Ben BergenEMNLP 2023 · 6 citations
- Semantic Representation for Dialogue ModelingXuefeng Bai, Yulong Chen, Linfeng Song, Yue ZhangACL 2021
- Improving AMR Parsing with Sequence-to-Sequence Pre-trainingDongqin Xu, Junhui Li, Muhua Zhu, Min Zhang et al.EMNLP 2020 · 57 citations
