Multimodal Semantic Bias Mitigation for Diverse Text-To-3D Generation
Yukuan Min, Muli Yang, Jinhao Zhang, Yuxuan Wang, Yihang Zhu, Jiexi Yan, Cheng Deng
摘要
The latest progress in text-to-3D generative models makes it possible to generate high-quality 3D content. Recent text-to-3D large model have achieved remarkable breakthroughs in multi-view consistency. However, their effectiveness is often affected by inherent biases, resulting in sensitivity to design settings such as prompt format, leading to difficulty understanding complex prompts. To help text-to-3D generative models understand more diverse prompts, we propose a framework to localize and mitigate the bias in the current text-to-3D large model. Specifically, we first use the existing model to generate 3D content and use the quality evaluation model to identify the cross-modality bias. Then, we use the predicted quality score to quantify the contribution of the prompt text to the bias. Finally, in order to reduce these biases, we construct diverse pairwise examples to help the current text-to-3D large model construct unbiased visual-text connections. The experiment shows that our method has achieved competitive results and can provide higher quality, more diverse 3D content compared to existing methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov 等ICCV 2023 · 被引用 1,662 次
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan 等CVPR 2022 · 被引用 1,603 次
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao 等NeurIPS 2023 · 被引用 1,498 次
相关 Paper
- Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic PromptsXinhua Cheng, Tianyu Yang, Jianan Wang, Yu Li 等ICLR 2024 · 被引用 58 次
- Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D GenerationNan Xiang, Tianyi Liang, Haiwen Huang, Shiqi Jiang 等IEEE VIS 2025 · 被引用 2 次
- Hallo3D: Multi-Modal Hallucination Detection and Mitigation for Consistent 3D Content GenerationHongbo Wang, Jie Cao, Jin Liu, Xiaoqiang Zhou 等NeurIPS 2024 · 被引用 10 次
- DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference OptimizationZhenglin Zhou, Xiaobo Xia, Fan Ma, Hehe Fan 等ICML 2025
- Debiasing Scores and Prompts of 2D Diffusion for View-consistent Text-to-3D GenerationSusung Hong, Donghoon Ahn, Seungryong KimNeurIPS 2023 · 被引用 46 次
