MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
Shuangkang Fang, I-Chao Shen, Yufeng Wang, Yi-Hsuan Tsai, Yi Yang, Shuchang Zhou, Wenrui Ding, Takeo Igarashi, Ming-Hsuan Yang
Abstract
We present MeshLLM, a novel framework that leverages large language models (LLMs) to understand and generate text-serialized 3D meshes. Our approach addresses key limitations in existing methods, including the limited dataset scale when catering to LLMs' token length and the loss of 3D structural information during mesh serialization. We introduce a Primitive-Mesh decomposition strategy, which divides 3D meshes into structurally meaningful subunits. This enables the creation of a large-scale dataset with samples, almost larger than previous methods, which aligns better with the LLM scaling law principles. Furthermore, we propose inferring face connectivity from vertices and local mesh assembly training strategies, significantly enhancing the LLMs' ability to capture mesh topology and spatial structures. Experiments show that MeshLLM outperforms the state-of-the-art LLaMA-Mesh in both mesh generation quality and shape understanding, highlighting its great potential in processing text-serialized 3D meshes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 14771f04-6e4f-4e77-9e83-87bf1002cfdfCited by top-tier papers6
- Prox-E: Fine-Grained 3D Shape Editing via Primitive-Based AbstractionsEtai Sella, Hao Phung, Nitay Amiel, Or Litany et al.SIGGRAPH 2026 · 2 citations
- AssetFormer: Modular 3D Assets Generation with Autoregressive TransformerLingting Zhu, Shengju Qian, Haidi Fan, Jiayu Dong et al.ICLR 2026 · 1 citation
- RealTwin: Concept Graph Representation and Grounding Framework for Reality-Preserving Digital Twin ReconstructionZisu Li, Ruohao Li, Jiawei Li, Chao Liu et al.CHI 2026 · 1 citation
- 4DP-QA: Scalable QA for 4D Perception in Vision Language ModelsSeokju Cho, Abhishek Badki, Hang Su, Jindong Jiang et al.CVPR 2026 · 1 citation
- SimArt: Decomposing Monolithic Meshes into Sim-ready Articulated Assets via MLLMChuanrui Zhang, Minghan Qin, Yuang Wang, Baifeng Xie et al.SIGGRAPH 2026
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language ModelsJunming Huang, Chi Wang, Letian Li, Guangkai Xu et al.ICML 2026 · 2 citations
- MeshXL: Neural Coordinate Field for Generative 3D Foundation ModelsSijin Chen, Xin Chen, Anqi Pang, Xianfang Zeng et al.NeurIPS 2024 · 125 citations
- MeshGPT: Generating Triangle Meshes with Decoder-Only TransformersYawar Siddiqui, Antonio Alliegro, Alexey Artemov, Tatiana Tommasi et al.CVPR 2024 · 74 citations
- ArtLLM: Generating Articulated Assets via 3D LLMPenghao Wang, Siyuan Xie, Hongyu Yan, Xianghui Yang et al.CVPR 2026 · 7 citations
- Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces SelectionDacheng Qi, Chenyu Wang, Jingwei Xu, Tianzhe Chu et al.CVPR 2026 · 10 citations
