GarmentGPT: Compositional Garment Pattern Generation via Discrete Latent Tokenization
Fangsheng Weng, Junhao Chen, Xiang Li, Jie Qin, Hanzhong Guo, ShaochunHao, Xiaoguang Han
Abstract
Apparel is a fundamental component of human appearance, making garment digitalization critical for digital human creation. However, sewing pattern creation traditionally relies on the intuition and extensive experience of skilled artisans. This manual bottleneck significantly hinders the scalability of digital garment creation. Existing generative approaches either operate as data replicators without intrinsic understanding of garment construction principles (e.g., diffusion models), or struggle with low-level regression of raw floating-point coordinates (e.g., Vision-Language Models). We present GarmentGPT, the first framework to operationalize latent space generation for sewing patterns. Our approach introduces a novel pipeline where a RVQ-VAE tokenizes continuous pattern boundary curves into discrete codebook indices. A fine-tuned Vision-Language Model then autoregressively predicts these discrete token sequences instead of regressing coordinates, enabling high-level compositional reasoning. This paradigm shift aligns generation with the knowledge-driven, symbolic reasoning capabilities of large language models. To address the data bottleneck for real-world applications, we develop a Data Curation Pipeline that synthesizes over one million photorealistic images paired with GarmentCode, and establish the Real-Garments Benchmark for comprehensive evaluation. Experiments demonstrate that GarmentGPT significantly outperforms existing methods on structured datasets (95.62% Panel Accuracy, 81.84% Stitch Accuracy), validating our discrete compositional paradigm's advantages. Code is available at https://github.com/ChimerAI-MMLab/Garment-GPT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b082c477-4fd9-4eb3-8061-99aa9c777a0fCited by top-tier papers2
- HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video SynthesisMingjin Chen, Junhao Chen, Zhaoxin Fan, Yujian Lee et al.CVPR 2026 · 13 citations
- LottieGPT: Tokenizing Vector Animation for Autoregressive GenerationJunhao Chen, Kejun Gao, Yuehan Cui, Mingze Sun et al.CVPR 2026 · 10 citations
Builds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
Related papers
- DressCode: Autoregressively Sewing and Generating Garments from Text GuidanceKai He, Kaixin Yao, Qixuan Zhang, Jingyi Yu et al.SIGGRAPH 2024 · 43 citations
- SwiftTailor: Efficient 3D Garment Generation with Geometry Image RepresentationPhuc Pham, Uy Dieu Tran, Binh-Son Hua, Phong NguyenCVPR 2026
- Learning Sewing Patterns via Latent Flow Matching of Implicit FieldsCong Cao, Ren Li, Corentin Dumery, Hao LiSIGGRAPH 2026
- Multimodal Latent Diffusion Model for Complex Sewing Pattern GenerationShengqi Liu, Yuhao Cheng, Zhuo Chen, Xingyu Ren et al.ICCV 2025 · 5 citations
- PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D GarmentsZhenyang Li, Lutao Jiang, Yizhou Zhao, Ying-Cong Chen et al.SIGGRAPH 2026
