One Pair Suffices: Unlocking Universal Zero-Shot Translation via Cross-Architecture Alignment
Hao Zong, Conghu Yuan, Chao Bei, Wentao Chen, Huan Liu, Kaiyu Huang, Degen Huang
摘要
Current paradigms for empowering Large Language Models (LLMs) with multilingual capabilities rely heavily on massive instruction tuning. We challenge this view, proposing that the barrier is topological alignment, not data quantity. We introduce Hybrid Cross-Alignment (HCA) , fusing a frozen NLLB encoder with a Qwen decoder via a closed-loop dual-adapter architecture. HCA utilizes a Source-Side Adapter to precondition encoder features and a Query-Residual Adapter to preserve generative stability, bridged by an adaptive gated cross-modal interface. Our core finding is the phenomenon of “Source-Side Alignment Generalization.” We demonstrate that training HCA on a single language pair (German-English) unlocks state-of-the-art zero-shot transfer to dozens of unseen languages for X-to-English translation . Crucially, our “Oracle” experiments reveal that this single-pair training recovers over 96.7% of the performance achievable by training on all available pairs. This suggests that a highly generalizable, source-side projection protocol exists. Evaluated rigorously across COMET and chrF++ , our ∼ 5.25B-parameter model significantly outperforms larger baselines, surpassing TowerPlus-9B (+9.0 COMET on low-resource languages) and Aya-101 (13B). Furthermore, performance scales linearly with encoder size; upgrading from 600M to 1.3B yields immediate gains (+3.4 points on Gujarati) with minimal retraining cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- LLM Augmented LLMs: Expanding Capabilities through CompositionRachit Bansal, Bidisha Samanta, Siddharth Dalmia, Nitish Gupta 等ICLR 2024 · 被引用 51 次
- TOWER+: Bridging Generality and Translation Specialization in Multilingual LLMsRicardo Rei, Nuno Miguel Guerreiro, José Pombal, João Alves 等ACL 2026 · 被引用 34 次
- Zero-Shot Cross-Lingual Transfer of Neural Machine Translation with Multilingual Pretrained EncodersGuanhua Chen, Shuming Ma, Yun Chen, Li Dong 等EMNLP 2021 · 被引用 30 次
- Allocating Large Vocabulary Capacity for Cross-Lingual Language Model Pre-TrainingBo Zheng, Li Dong, Shaohan Huang, Saksham Singhal 等EMNLP 2021 · 被引用 15 次
相关 Paper
- Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible MultilingualityMengyu Bu, Yang FengACL 2026 · 被引用 2 次
- Middle-Layer Representation Alignment for Cross-Lingual Transfer in Fine-Tuned LLMsDanni Liu, Jan NiehuesACL 2025 · 被引用 23 次
- Scaling Laws for Conditional Emergence of Multilingual Image Captioning via Generalization from TranslationJulian Spravil, Sebastian Houben, Sven BehnkeAAAI 2026
- LangBridge: Interpreting Image as a Combination of Language EmbeddingsJiaqi Liao, Yuwei Niu, Fanqing Meng, Hao Li 等ICCV 2025
- X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at ScaleHaoran Xu, Kenton Murray, Philipp Koehn, Hieu Hoang 等ICLR 2025
