ShoeFit: A New Dataset and Dual-image-stream DiT Framework for Virtual Footwear Try-On
Yuhan Li, Zhiyu Jin, Yifan Tong, Wenxiang Shang, Benlei Cui, Xuanhong Chen, Ran Lin, Bingbing Ni
Abstract
(b) ShoeFit has demonstrated excellent preservation of the fidelity of footwear across real-world scenes.
†Work done during an internship at Alibaba. ‡Corresponding author.
39th Conference on Neural Information Processing Systems (NeurIPS 2025).
footwear latents. The proposed framework effectively decouples shoe appearance from environmental interferences while preserving high-quality texture detail through decoupled denoising and conditioning branches. Extensive quantitative and qualitative experiments demonstrate that our method substantially improves rendering fidelity and robustness under challenging real-world product shoes, establishing a new benchmark in high-fidelity footwear try-on synthesis. The dataset and benchmark will be publicly available upon acceptance of the paper.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08d40a64-9541-4f29-82e7-abb93a8200e7Cited by top-tier papers2
- Diffusion Probe: Generated Image Result Prediction Using CNN ProbesBukun Huang, Benlei Cui, Zhizeng Ye, Xuemei Dong et al.CVPR 2026 · 13 citations
- Simpleposter: A simple Baseline for Product Poster GenerationBenlei Cui, Fangao Zeng, Weitao Jiang, Yuwen Zhai et al.CVPR 2026 · 5 citations
Builds on40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-OnYuanhao Wang, Johanna Suvi Karras, Yingwei Li, Ira Kemelmacher-ShlizermanSIGGRAPH 2026
- Improving Fairness in Facial Albedo Estimation via Visual-Textual CuesXingyu Ren, Jiankang Deng, Chao Ma, Yichao Yan et al.CVPR 2023
- GPD-VVTO: Preserving Garment Details in Video Virtual Try-OnYuanbin Wang, Weilun Dai, Long Chan, Huanyu Zhou et al.ACM MM 2024 · 4 citations
- Down to the Last Detail: Virtual Try-on with Fine-grained DetailsJiahang Wang, Tong Sha, Wei Zhang, Zhoujun Li et al.ACM MM 2020 · 23 citations
- URHand: Universal Relightable HandsZhaoxi Chen, Gyeongsik Moon, Kaiwen Guo, Chen Cao et al.CVPR 2024
