X-FACTR: Multilingual Factual Knowledge Retrieval from Pretrained Language Models
Zhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding, Graham Neubig
摘要
Language models (LMs) have proven surprisingly successful at capturing factual knowledge by completing cloze-style fill-in-theblank questions such as "Punta Cana is located in _." However, while knowledge is both written and queried in many languages, studies on LMs' factual representation ability have almost invariably been performed on English. To assess factual knowledge retrieval in LMs in different languages, we create a multilingual benchmark of cloze-style probes for 23 typologically diverse languages. To properly handle language variations, we expand probing methods from single-to multi-word entities, and develop several decoding algorithms to generate multi-token predictions. Extensive experimental results provide insights about how well (or poorly) current state-of-theart LMs perform at this task in languages with more or fewer available resources. We further propose a code-switching-based method to improve the ability of multilingual LMs to access knowledge, and verify its effectiveness on several benchmark languages. Benchmark data and code have be released at https: //x-factr.github.io .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- Zero-Shot Video Question Answering via Frozen Bidirectional Language ModelsAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等NeurIPS 2022 · 被引用 305 次
- Few-shot Learning with Multilingual Generative Language ModelsXi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang 等EMNLP 2022 · 被引用 113 次
- Efficient Large Scale Language Modeling with Mixtures of ExpertsMikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov 等EMNLP 2022 · 被引用 71 次
- Relational World Knowledge Representation in Contextual Language Models: A ReviewTara Safavi, Danai KoutraEMNLP 2021 · 被引用 31 次
- XLM-K: Improving Cross-Lingual Language Model Pre-training with Multilingual KnowledgeXiaoze Jiang, Yaobo Liang, Weizhu Chen, Nan DuanAAAI 2022 · 被引用 31 次
它引用的顶会 Paper7
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and GenerationYaobo Liang, Nan Duan, Yeyun Gong, Ning Wu 等EMNLP 2020 · 被引用 232 次
- Emerging Cross-lingual Structure in Pretrained Language ModelsAlexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer 等ACL 2020 · 被引用 210 次
- Inducing Relational Knowledge from BERTZied Bouraoui, José Camacho-Collados, Steven SchockaertAAAI 2020 · 被引用 183 次
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 被引用 167 次
相关 Paper
- Cross-Lingual Consistency of Factual Knowledge in Multilingual Language ModelsJirui Qi, Raquel Fernández, Arianna BisazzaEMNLP 2023 · 被引用 9 次
- Do You Know About My Nation? Investigating Multilingual Language Models' Cultural Literacy Through Factual KnowledgeEshaan Tanwar, Anwoy Chatterjee, Michael Saxon, Alon Albalak 等EMNLP 2025 · 被引用 4 次
- mLUKE: The Power of Entity Representations in Multilingual Pretrained Language ModelsRyokan Ri, Ikuya Yamada, Yoshimasa TsuruokaACL 2022 · 被引用 34 次
- CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality EvaluationYexing Du, Kaiyuan Liu, Youcheng Pan, Zheng Chu 等AAAI 2026 · 被引用 4 次
- Retrieval-Augmented Multilingual Knowledge EditingWeixuan Wang, Barry Haddow, Alexandra BirchACL 2024
