Efficient Low-Resource Language Adaptation via Multi-Source Dynamic Logit Fusion
Chen Zhang, Jiuheng Lin, Zhiyuan Liao, Yansong Feng
Abstract
Adapting large language models (LLMs) to low-resource languages (LRLs) is constrained by the scarcity of task data and computational resources. Although Proxy Tuning offers a logit-level strategy for introducing scaling effects, it often fails in LRL settings because the large model's weak LRL competence might overwhelm the knowledge of specialized smaller models. We thus propose TriMix, a test-time logit fusion framework that dynamically balances capabilities from three different sources: LRL competence from a continually pretrained small model, task competence from high-resource language instruction tuning, and the scaling benefits of large models. It is data- and compute-efficient, requiring no LRL task annotations, and only continual pretraining on a small model. Experiments across four model families and eight LRLs show that TriMix consistently outperforms single-model baselines and Proxy Tuning. Our analysis reveals that prioritizing the small LRL-specialized model's logits is crucial for success, challenging the prevalent large-model-dominant assumption.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 99a3f402-d30d-4633-a332-e111657844ccBuilds on20
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim et al.ICLR 2024 · 354 citations
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual EvaluationShivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani et al.ACL 2025 · 144 citations
- A Benchmark for Learning to Translate a New Language from One Grammar BookGarrett Tanzer, Mirac Suzgun, Eline Visser, Dan Jurafsky et al.ICLR 2024 · 97 citations
- Contrastive Decoding: Open-ended Text Generation as OptimizationXiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang et al.ACL 2023 · 78 citations
Related papers
- Test-Time Learning for Large Language ModelsJinwu Hu, Zitian Zhang, Guohao Chen, Xutao Wen et al.ICML 2025
- Mix-CPT: A Domain Adaptation Framework via Decoupling Knowledge Learning and Format AlignmentJinhao Jiang, Junyi Li, Xin Zhao, Yang Song et al.ICLR 2025
- On Giant's Shoulders: Effortless Weak to Strong by Dynamic Logits FusionChenghao Fan, Zhenyi Lu, Wei Wei, Jie Tian et al.NeurIPS 2024 · 17 citations
- Prototypical Fine-Tuning: Towards Robust Performance under Varying Data SizesYiqiao Jin, Xiting Wang, Yaru Hao, Yizhou Sun et al.AAAI 2023 · 15 citations
- GIFT: Guided Fine-Tuning and Transfer for Enhancing Instruction-Tuned Language ModelsZhiwen Ruan, Yichao Du, Jianjie Zheng, Longyue Wang et al.ACL 2026
