Too Helpful, Too Harmless, Too Honest or Just Right?
Gautam Siddharth Kashyap, Mark Dras, Usman Naseem
Abstract
Large Language Models (LLMs) exhibit strong performance across a wide range of NLP tasks, yet aligning their outputs with the principles of Helpfulness, Harmlessness, and Honesty (HHH) remains a persistent challenge. Existing methods often optimize for individual alignment dimensions in isolation, leading to trade-offs and inconsistent behavior. While Mixture-of-Experts (MoE) architectures offer modularity, they suffer from poorly calibrated routing, limiting their effectiveness in alignment tasks. We propose TrinityX, a modular alignment framework that incorporates a Mixture of Calibrated Experts (Mo-CaE) within the Transformer architecture. Trin-ityX leverages separately trained experts for each HHH dimension, integrating their outputs through a calibrated, task-adaptive routing mechanism that combines expert signals into a unified, alignment-aware representation. Extensive experiments on three standard alignment benchmarks-Alpaca (Helpfulness), Beaver-Tails (Harmlessness), and TruthfulQA (Honesty)-demonstrate that TrinityX outperforms strong baselines, achieving relative improvements of 32.5% in win rate, 33.9% in safety score, and 28.4% in truthfulness. In addition, TrinityX reduces memory usage and inference latency by over 40% compared to prior MoEbased approaches. Ablation studies highlight the importance of calibrated routing, and crossmodel evaluations confirm TrinityX's generalization across diverse LLM backbones. Our code is available at: https://github.com/g skgautam/TrinityX
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
- Aligner: Efficient Alignment by Learning to CorrectJiaming Ji, Boyuan Chen, Hantao Lou, Donghai Hong et al.NeurIPS 2024 · 115 citations
- Mixture-of-Experts Meets Instruction Tuning: A Winning Combination for Large Language ModelsSheng Shen, Le Hou, Yanqi Zhou, Nan Du et al.ICLR 2024 · 87 citations
- Editing models with task arithmeticGabriel Ilharco, Marco Túlio Ribeiro, Mitchell Wortsman, Ludwig Schmidt et al.ICLR 2023 · 31 citations
Related papers
- MESA: Improving MoE Safety Alignment via Decentralized ExpertiseYitong Sun, Yao Huang, Teng Li, Ranjie Duan et al.ICML 2026
- Multi-Adapter Representation Interventions via Energy CalibrationManjiang Yu, Hongji Li, Junwei Chen, Xue Li et al.ICML 2026
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- SafeMoE: Safe Fine-Tuning for MoE LLMs by Aligning Harmful Input RoutingJaehan Kim, Minkyoo Song, Seungwon Shin, Sooel SonICLR 2026
- SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert IdentificationZhenglin Lai, Mengyao Liao, Bingzhe Wu, Dong Xu et al.NeurIPS 2025 · 22 citations
