MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank Experts
Jie Zhu, Yixiong Chen, Mingyu Ding, Ping Luo, Leye Wang, Jingdong Wang
Abstract
Text-to-image diffusion has attracted vast attention due to its impressive image-generation capabilities. However, when it comes to human-centric text-to-image generation, particularly in the context of faces and hands, the results often fall short of naturalness due to insufficient training priors. We alleviate the issue in this work from two perspectives. 1) From the data aspect, we carefully collect a human-centric dataset comprising over one million high-quality human-in-the-scene images and two specific sets of close-up images of faces and hands. These datasets collectively provide a rich prior knowledge base to enhance the human-centric image generation capabilities of the diffusion model. 2) On the methodological front, we propose a simple yet effective method called Mixture of Low-rank Experts (MoLE) by considering low-rank modules trained on close-up hand and face images respectively as experts. This concept draws inspiration from our observation of low-rank refinement, where a low-rank module trained by a customized close-up dataset has the potential to enhance the corresponding image part when applied at an appropriate scale. To validate the superiority of MoLE in the context of human-centric image generation compared to state-of-the-art, we construct two benchmarks and perform evaluations with diverse metrics and human studies. Datasets, model, and code are released at https://sites.google.com/view/mole4diffuser/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- AC-LoRA: (Almost) Training-Free Access Control Aware Multi-Modal LLMsLara Magdalena Lazier, Aritra Dhar, Vasilije Stambolic, Lukas CavigelliNeurIPS 2025 · 3 citations
- FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion ModelsYuxuan Wang, Tianwei Cao, Huayu Zhang, Zhongjiang He et al.ICCV 2025 · 1 citation
- Prism-MoE: Efficient Dense-to-MoE Conversion for Visual Autoregressive GenerationYing Li, Zefang Wang, Zhaode Wang, Zhiwen Chen et al.ICML 2026
- FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance CustomizationRong Zhang, Jinxiao Li, Jingnan Wang, Zhiwen Zuo et al.AAAI 2026
- Robust Photo-Realistic Hand Gesture Generation: from Single View to Multiple ViewQifan Fu, Xu Chen, Muhammad Asad, Shanxin Yuan et al.ACM MM 2025
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- BOTH2Hands: Inferring 3D Hands from Both Text Prompts and Body DynamicsWenqian Zhang, Molin Huang, Yuxuan Zhou, Juze Zhang et al.CVPR 2024
- Hand1000: Generating Realistic Hands from Text with Only 1, 000 ImagesHaozhuo Zhang, Bin Zhu, Yu Cao, Yanbin HaoAAAI 2025 · 11 citations
- TextDiffuser: Diffusion Models as Text PaintersJingye Chen, Yupan Huang, Tengchao Lv, Lei Cui et al.NeurIPS 2023 · 290 citations
- Generating Images of Rare Concepts Using Pre-trained Diffusion ModelsDvir Samuel, Rami Ben-Ari, Simon Raviv, Nir Darshan et al.AAAI 2024 · 82 citations
- Diffusion Facial Forgery DetectionHarry Cheng, Yangyang Guo, Tianyi Wang, Liqiang Nie et al.ACM MM 2024 · 42 citations
