MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online Governance
Agam Goyal, Xianyang Zhan, Yilun Chen, Koustuv Saha, Eshwar Chandrasekharan
Abstract
Large language models (LLMs) have shown great potential in flagging harmful content in online communities. Yet, existing approaches for moderation require a separate model for every community and are opaque in their decision-making, limiting real-world adoption. We introduce Mixture of Moderation Experts (MoMoE), a modular, cross-community framework that adds post-hoc explanations to scalable content moderation. MoMoE orchestrates four operators-Allocate , Predict , Aggregate , Explain -and is instantiated as seven community-specialized experts (MoMoE Community ) and five norm-violation experts (MoMoE NormVio ). On 30 unseen subreddits, the best variants obtain Micro-F1 scores of 0.72 and 0.67, respectively, matching or surpassing strong fine-tuned baselines while consistently producing concise and reliable explanations. Although community-specialized experts deliver the highest peak accuracy, norm-violation experts provide steadier performance across domains. These findings show that MoMoE yields scalable, transparent moderation without needing per-community fine-tuning. More broadly, they suggest that lightweight, explainable expert ensembles can guide future NLP and HCI research on trustworthy human-AI governance of online communities. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2bfae2bb-ddb0-4e61-86f1-86aecc0a8d74Cited by top-tier papers4
- Needling Through the Threads: A Visualization Tool for Navigating Threaded Online DiscussionsYijun Liu, Frederick Choi, Eshwar ChandrasekharanCHI 2026 · 1 citation
- The Language of Approval: Identifying the Drivers of Positive Feedback OnlineAgam Goyal, Charlotte Lambert, Eshwar ChandrasekharanCHI 2026 · 1 citation
- Evaluating Large Language Models for Detecting AntisemitismJay Patel, Hrudayangam Mehta, Jeremy BlackburnEMNLP 2025 · 1 citation
- Breaking Bad Tokens: Detoxification of LLMs Using Sparse AutoencodersAgam Goyal, Vedant Rathi, William Yeh, Yian Wang et al.EMNLP 2025 · 1 citation
Builds on13
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With RealityMitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto et al.CHI 2021 · 100 citations
- Designing Word Filter Tools for Creator-led Comment ModerationShagun Jhaver, Quan Ze Chen, Detlef Knauss, Amy X. ZhangCHI 2022 · 70 citations
- Conversations Gone Alright: Quantifying and Predicting Prosocial Outcomes in Online ConversationsJiajun Bao, Junjie Wu, Yiming Zhang, Eshwar Chandrasekharan et al.WWW 2021 · 63 citations
- Trust in AI-assisted Decision Making: Perspectives from Those Behind the System and Those for Whom the Decision is MadeOleksandra Vereschak, Fatemeh Alizadeh, Gilles Bailly, Baptiste CaramiauxCHI 2024 · 31 citations
Related papers
- SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language ModelsHuy Nghiem, Advik Sachdeva, Hal Daumé IIIACL 2026 · 1 citation
- MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language ModelsLeyang Shen, Gongwei Chen, Rui Shao, Weili Guan et al.NeurIPS 2024 · 55 citations
- CHAIRO: Contextual Hierarchical Analogical Induction and Reasoning Optimization for LLMsHaotian Lu, Yuchen Mou, Bingzhe WuACL 2026
- Understanding Cross-layer Contributions to Mixture-of-Experts Routing in LLMsWengang Li, Lingqi Zhang, Toshio Endo, Mohamed WahibICLR 2026
- SafeMoE: Safe Fine-Tuning for MoE LLMs by Aligning Harmful Input RoutingJaehan Kim, Minkyoo Song, Seungwon Shin, Sooel SonICLR 2026
