Hallo3D: Multi-Modal Hallucination Detection and Mitigation for Consistent 3D Content Generation
Hongbo Wang, Jie Cao, Jin Liu, Xiaoqiang Zhou, Huaibo Huang, Ran He
Abstract
Recent advances in pretrained 2D diffusion models have significantly improved visual prior guidance for 3D content generation. However, this process often lacks geometric constraints, leading to spatial perception hallucinations and multi-view inconsistencies. To address this, we introduce Hallo3D , a tuning-free method for 3D content generation that leverages the geometric perception capabilities of large multi-modal models to detect and mitigate these hallucinations. Our approach follows a generation-detection-correction paradigm, using multi-modal inconsistencies as query information to guide the detection of hallucinations and formulate enhanced negative prompts that ensure consistent renderings. Additionally, we propose a denoising strategy that employs attention mechanisms to maintain consistency in color and texture across multiple views during visual guidance. Our method is data-independent, easily integrates with existing 3D content generation frameworks, and supports both text-driven and image-driven approaches. Extensive experiments demonstrate that our method significantly improves the consistency and quality of generated 3D content, particularly in mitigating hallucinations common with 2D pretrained models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ffc1d05a-5296-4ae7-9ab9-df47b61601cfCited by top-tier papers3
- DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion ModelingYuang Ai, Qihang Fan, Xuefeng Hu, Zhenheng Yang et al.NeurIPS 2025 · 8 citations
- A Temporal and Content Co-Awareness Latent Diffusion for Controllable Hand Image GenerationShuang Hao, Pengfei Ren, Haifeng Sun, Pan Ting et al.CVPR 2026
- Think-Then-Generate: Structural Chain-of-Thought Reasoning for Consistent 3D GenerationXinyue Liu, Jin Liu, Hongbo Wang, Ran He et al.CVPR 2026
Builds on42
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction CycleZhenyu Tang, Junwu Zhang, Xinhua Cheng, Wangbo Yu et al.AAAI 2025 · 43 citations
- Multimodal Semantic Bias Mitigation for Diverse Text-To-3D GenerationYukuan Min, Muli Yang, Jinhao Zhang, Yuxuan Wang et al.CVPR 2026
- SweetDreamer: Aligning Geometric Priors in 2D diffusion for Consistent Text-to-3DWeiyu Li, Rui Chen, Xuelin Chen, Ping TanICLR 2024 · 155 citations
- VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language ModelsGuoqing Chen, Fu Zhang, Bingqian Liu, Chenglong Lu et al.AAAI 2026
- CaliTex: Geometry-Calibrated Attention for View-Coherent 3D Texture GenerationChenyu Liu, Hongze CHEN, Jingzhi Bao, Lingting Zhu et al.CVPR 2026 · 3 citations
