Mixture of insighTful Experts (MoTE): The Synergy of Reasoning Chains and Expert Mixtures in Self-Alignment
Zhili Liu, Yunhao Gou, Kai Chen, Lanqing Hong, Jiahui Gao, Fei Mi, Yu Zhang, Zhenguo Li, Xin Jiang, Qun Liu, James T. Kwok
摘要
As the capabilities of large language models (LLMs) continue to expand, aligning these models with human values remains a significant challenge. Recent studies show that reasoning abilities contribute significantly to model safety, while integrating Mixture-of-Experts (MoE) architectures can further enhance alignment. In this work, we address a fundamental question: How to effectively incorporate reasoning abilities and MoE architectures into self-alignment process in LLMs? We propose Mixture of insighTful Experts (MoTE), a novel framework that synergistically combines reasoning chains and expert mixtures to improve self-alignments. From a data perspective, MoTE employs a structured reasoning chain comprising four key stages: Question Analysis, Answer Guidance, Safe Answer, and Safety Checking. This approach enhances safety through multi-step reasoning and proves effective even for smaller and less powerful LLMs (e.g., 7B models). From an architectural perspective, MoTE adopts a multi-LoRA framework with step-level routing, where each expert is dedicated to a specific reasoning step. This design eliminates the need for balance losses, ensures stable training, and supports adaptive inference lengths. Experimental results demonstrate that MoTE significantly improves model safety, jailbreak resistance, and over-refusal capabilities, achieving performance comparable to OpenAI's state-of-the-art o1 model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Disentangling Consensus and Value-Specific Representations for Controllable Pluralistic Value Alignment of LLMsJianKui Zhou, Jing Yao, Xiaoyuan Yi, Peng Zhang 等ICML 2026
- Mixture-of-Trees: Learning to Select and Weigh Reasoning Paths for Efficient LLM InferenceYangbo Wei, Zhen Huang, Shaoqiang Lu, Junhong Qian 等AAAI 2026
- UniTransfer: Video Concept Transfer via Progressive Spatio-Temporal DecompositionGuojun Lei, Rong Zhang, Tianhang Liu, Hong Li 等NeurIPS 2025
- Corrupted but Not Broken: Understanding and Mitigating the Negative Impacts of Corrupted Data in Visual Instruction TuningYunhao Gou, Hansi Yang, Zhili Liu, Kai Chen 等EMNLP 2025
它引用的顶会 Paper16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
相关 Paper
- GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMsLichao Wu, Sasha Behrouzi, Mohamadreza Rostami, Stjepan Picek 等USENIX Security 2026 · 被引用 14 次
- STAIR: Improving Safety Alignment with Introspective ReasoningYichi Zhang, Siyuan Zhang, Yao Huang, Zeyu Xia 等ICML 2025
- MESA: Improving MoE Safety Alignment via Decentralized ExpertiseYitong Sun, Yao Huang, Teng Li, Ranjie Duan 等ICML 2026
- SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert IdentificationZhenglin Lai, Mengyao Liao, Bingzhe Wu, Dong Xu 等NeurIPS 2025 · 被引用 22 次
- Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization AlignmentChenghao Fan, Zhenyi Lu, Sichen Liu, Chengfeng Gu 等ICML 2025
