Multimodal Semantic Bias Mitigation for Diverse Text-To-3D Generation
Yukuan Min, Muli Yang, Jinhao Zhang, Yuxuan Wang, Yihang Zhu, Jiexi Yan, Cheng Deng
Abstract
The latest progress in text-to-3D generative models makes it possible to generate high-quality 3D content. Recent text-to-3D large model have achieved remarkable breakthroughs in multi-view consistency. However, their effectiveness is often affected by inherent biases, resulting in sensitivity to design settings such as prompt format, leading to difficulty understanding complex prompts. To help text-to-3D generative models understand more diverse prompts, we propose a framework to localize and mitigate the bias in the current text-to-3D large model. Specifically, we first use the existing model to generate 3D content and use the quality evaluation model to identify the cross-modality bias. Then, we use the predicted quality score to quantify the contribution of the prompt text to the bias. Finally, in order to reduce these biases, we construct diverse pairwise examples to help the current text-to-3D large model construct unbiased visual-text connections. The experiment shows that our method has achieved competitive results and can provide higher quality, more diverse 3D content compared to existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e17e3098-8ed7-4910-baba-c321f6cafae6Builds on26
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan et al.CVPR 2022 · 1,603 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
Related papers
- Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic PromptsXinhua Cheng, Tianyu Yang, Jianan Wang, Yu Li et al.ICLR 2024 · 58 citations
- Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D GenerationNan Xiang, Tianyi Liang, Haiwen Huang, Shiqi Jiang et al.IEEE VIS 2025 · 2 citations
- Hallo3D: Multi-Modal Hallucination Detection and Mitigation for Consistent 3D Content GenerationHongbo Wang, Jie Cao, Jin Liu, Xiaoqiang Zhou et al.NeurIPS 2024 · 10 citations
- DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference OptimizationZhenglin Zhou, Xiaobo Xia, Fan Ma, Hehe Fan et al.ICML 2025
- Debiasing Scores and Prompts of 2D Diffusion for View-consistent Text-to-3D GenerationSusung Hong, Donghoon Ahn, Seungryong KimNeurIPS 2023 · 46 citations
