Lune

NeurIPS2024顶会

How do Large Language Models Handle Multilingualism?

Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi, Lidong Bing

2024年份
196被引次数
33顶会引用

摘要

Large language models (LLMs) have demonstrated impressive capabilities across diverse languages. This study explores how LLMs handle multilingualism. Based on observed language ratio shifts among layers and the relationships between network structures and certain capabilities, we hypothesize the LLM's multilingual workflow (MWork\texttt{MWork}): LLMs initially understand the query, converting multilingual inputs into English for task-solving. In the intermediate layers, they employ English for thinking and incorporate multilingual knowledge with self-attention and feed-forward structures, respectively. In the final layers, LLMs generate responses aligned with the original language of the query. To verify MWork\texttt{MWork}, we introduce Parallel Language-specific Neuron Detection (PLND\texttt{PLND}) to identify activated neurons for inputs in different languages without any labeled data. Using PLND\texttt{PLND}, we validate MWork\texttt{MWork} through extensive experiments involving the deactivation of language-specific neurons across various layers and structures. Moreover, MWork\texttt{MWork} allows fine-tuning of language-specific neurons with a small dataset, enhancing multilingual abilities in a specific language without compromising others. This approach results in an average improvement of 3.6%3.6\% for high-resource languages and 2.3%2.3\% for low-resource languages across all tasks with just 400400 documents.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper33

问问它们各自怎么用它

它引用的顶会 Paper17

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖