Hallo3D: Multi-Modal Hallucination Detection and Mitigation for Consistent 3D Content Generation
Hongbo Wang, Jie Cao, Jin Liu, Xiaoqiang Zhou, Huaibo Huang, Ran He
摘要
Recent advances in pretrained 2D diffusion models have significantly improved visual prior guidance for 3D content generation. However, this process often lacks geometric constraints, leading to spatial perception hallucinations and multi-view inconsistencies. To address this, we introduce Hallo3D , a tuning-free method for 3D content generation that leverages the geometric perception capabilities of large multi-modal models to detect and mitigate these hallucinations. Our approach follows a generation-detection-correction paradigm, using multi-modal inconsistencies as query information to guide the detection of hallucinations and formulate enhanced negative prompts that ensure consistent renderings. Additionally, we propose a denoising strategy that employs attention mechanisms to maintain consistency in color and texture across multiple views during visual guidance. Our method is data-independent, easily integrates with existing 3D content generation frameworks, and supports both text-driven and image-driven approaches. Extensive experiments demonstrate that our method significantly improves the consistency and quality of generated 3D content, particularly in mitigating hallucinations common with 2D pretrained models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion ModelingYuang Ai, Qihang Fan, Xuefeng Hu, Zhenheng Yang 等NeurIPS 2025 · 被引用 8 次
- A Temporal and Content Co-Awareness Latent Diffusion for Controllable Hand Image GenerationShuang Hao, Pengfei Ren, Haifeng Sun, Pan Ting 等CVPR 2026
- Think-Then-Generate: Structural Chain-of-Thought Reasoning for Consistent 3D GenerationXinyue Liu, Jin Liu, Hongbo Wang, Ran He 等CVPR 2026
它引用的顶会 Paper42
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
相关 Paper
- Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction CycleZhenyu Tang, Junwu Zhang, Xinhua Cheng, Wangbo Yu 等AAAI 2025 · 被引用 43 次
- Multimodal Semantic Bias Mitigation for Diverse Text-To-3D GenerationYukuan Min, Muli Yang, Jinhao Zhang, Yuxuan Wang 等CVPR 2026
- SweetDreamer: Aligning Geometric Priors in 2D diffusion for Consistent Text-to-3DWeiyu Li, Rui Chen, Xuelin Chen, Ping TanICLR 2024 · 被引用 155 次
- VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language ModelsGuoqing Chen, Fu Zhang, Bingqian Liu, Chenglong Lu 等AAAI 2026
- CaliTex: Geometry-Calibrated Attention for View-Coherent 3D Texture GenerationChenyu Liu, Hongze CHEN, Jingzhi Bao, Lingting Zhu 等CVPR 2026 · 被引用 3 次
