Localizing Open-Ontology QA Semantic Parsers in a Day Using Machine Translation
Mehrad Moradshahi, Giovanni Campagna, Sina J. Semnani, Silei Xu, Monica S. Lam
摘要
We propose Semantic Parser Localizer (SPL), a toolkit that leverages Neural Machine Translation (NMT) systems to localize a semantic parser for a new language. Our methodology is to (1) generate training data automatically in the target language by augmenting machine-translated datasets with local entities scraped from public websites, (2) add a fewshot boost of human-translated sentences and train a novel XLMR-LSTM semantic parser, and (3) test the model on natural utterances curated using human translators. We assess the effectiveness of our approach by extending the current capabilities of Schema2QA, a system for English Question Answering (QA) on the open web, to 10 new languages for the restaurants and hotels domains. Our models achieve an overall test accuracy ranging between 61% and 69% for the hotels domain and between 64% and 78% for restaurants domain, which compares favorably to 69% and 80% obtained for English parser trained on gold English data and a few examples from validation set. We show our approach outperforms the previous state-of-theart methodology by more than 30% for hotels and 40% for restaurants with localized ontologies for the subset of languages tested. Our methodology enables any software developer to add a new language capability to a QA system for a new domain, leveraging machine translation, in less than 24 hours. Our code is released open-source. 1 Language Country Examples Hotels English I want a hotel near times square that has at least 1000 reviews. Arabic German Ich möchte ein hotel in der nähe von marienplatz, das mindestens 1000 bewertungen hat. Spanish Busco un hotel cerca de puerto banús que tenga al menos 1000 comentarios. Farsi Finnish Haluan paikan helsingin tuomiokirkko läheltä hotellin, jolla on vähintään 1000 arvostelua. Italian Voglio un hotel nei pressi di colosseo che abbia almeno 1000 recensioni. Japanese 東京スカイツリー周辺でに1000件以上のレビューがあるホテルを見せて。 Polish Potrzebuję hotelu w pobliżu zamek w malborku, który ma co najmniej 1000 ocen. Turkish Kapalı carşı yakınlarında en az 1000 yoruma sahip bir otel istiyorum. Chinese 我想在天安门广场附近找一家有至少1000条评论的酒店。 Restaurants English find me a restaurant that serves burgers and is open at 14:30 . Arabic German Finden sie bitte ein restaurant mit maultaschen essen, das um 14:30 öffnet. Spanish Busque un restaurante que sirva comida paella valenciana y abra a las 14:30. Farsi Finnish Etsi minulle ravintola joka tarjoilee karjalanpiirakka ruokaa ja joka aukeaa kello 14:30 mennessä. Italian Trovami un ristorante che serve cibo lasagna e apre alle 14:30. Japanese 寿司フードを提供し、14:30までに開店するレストランを見つけてください。 Polish Znajdź restaurację, w której podaje się kotlet jedzenie i któr ą otwieraj ą o 14:30. Turkish Bana köfte yemekleri sunan ve 14:30 zamanına kadar açık olan bir restoran bul..
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- AutoQA: From Databases To QA Semantic Parsers With Only Synthetic Training DataSilei Xu, Sina J. Semnani, Giovanni Campagna, Monica S. LamEMNLP 2020 · 被引用 33 次
- Zero-Shot Cross-lingual Semantic ParsingTom Sherborne, Mirella LapataACL 2022 · 被引用 32 次
- Fine-tuned LLMs Know More, Hallucinate Less with Few-Shot Sequence-to-Sequence Semantic Parsing over WikidataSilei Xu, Shicheng Liu, Theo Culhane, Elizaveta Pertseva 等EMNLP 2023 · 被引用 10 次
- The Best of Both Worlds: Combining Human and Machine Translations for Multilingual Semantic Parsing with Active LearningZhuang Li, Lizhen Qu, Philip R. Cohen, Raj Tumuluri 等ACL 2023 · 被引用 4 次
- XSemPLR: Cross-Lingual Semantic Parsing in Multiple Natural Languages and Meaning RepresentationsYusen Zhang, Jun Wang, Zhiguo Wang, Rui ZhangACL 2023 · 被引用 4 次
它引用的顶会 Paper3
- Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue DatasetAbhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta 等AAAI 2020 · 被引用 707 次
- AutoQA: From Databases To QA Semantic Parsers With Only Synthetic Training DataSilei Xu, Sina J. Semnani, Giovanni Campagna, Monica S. LamEMNLP 2020 · 被引用 33 次
- Zero-Shot Transfer Learning with Synthesized Data for Multi-Domain Dialogue State TrackingGiovanni Campagna, Agata Foryciarz, Mehrad Moradshahi, Monica S. LamACL 2020 · 被引用 5 次
相关 Paper
- Cross-lingual Back-Parsing: Utterance Synthesis from Meaning Representation for Zero-Resource Semantic ParsingDeokhyung Kang, Seonjeong Hwang, Yunsu Kim, Gary Geunbae LeeEMNLP 2024
- SCoRe: Pre-Training for Context Representation in Conversational Semantic ParsingTao Yu, Rui Zhang, Alex Polozov, Christopher Meek 等ICLR 2021 · 被引用 44 次
- SimQA: Detecting Simultaneous MT Errors through Word-by-Word Question AnsweringHyoJung Han, Marine Carpuat, Jordan L. Boyd-GraberEMNLP 2022 · 被引用 3 次
- Finding needles in a haystack: Sampling Structurally-diverse Training Sets from Synthetic Data for Compositional GeneralizationInbar Oren, Jonathan Herzig, Jonathan BerantEMNLP 2021
- Multilingual Transfer Learning for QA using Translation as Data AugmentationMihaela A. Bornea, Lin Pan, Sara Rosenthal, Radu Florian 等AAAI 2021 · 被引用 45 次
