MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation
Sankalp Sinha, Mohammad Sadil Khan, Muhammad Usama, Shino Sam, Didier Stricker, Sk Aziz Ali, Muhammad Zeshan Afzal
摘要
Stitches is a cute, plush-like bear with a large head and small body. He has [...] black X-shaped eyes. His fur is yellow with purple patches and blue ear tips. He wears a white shirt with colorful stars. Stitches stand [...] circular base with a brown surface and a green pattern. Tyranid Genestealer, a predatory creature with an elongated skull, jointed arms, a muscular body, strong legs, a whip-like tail, and bony spines. The 3D model shows the cholesterol molecule with four fused rings, carbon (grey), hydrogen (white), and a single oxygen (red) atom. It has a flat, twisted shape and includes a hydroxyl group. The 3D model is a Japanese house with a rectangular shape [...] pitched roof. The walls are dark brown wood, and the roof is bright red. The house has a small porch with steps, [...] Bamboo trees surround the house, and the ground has sandy... areas with green mossy rocks and grass. An old, moss-covered wishing well. Rough stones, aged wood, rusty chains, mushrooms, fallen leaves, and twigs create an enchanting, ancient, and rustic atmosphere. An island with vibrant, multicolored trees, featuring pink, orange, and blue foliage. Waterfalls [...] weathered stone and moss, and hanging vines for detail. Glowing crystals in shades of teal and purple with colorful flowers [...] fantasy RPG setting.
A red panda in a bamboo forest, showing its round, soft face framed by white markings around its eyes and snout. The reddish-brown fur on its back is dense, with a bushy tail ringed in alternating shades of red and brown. Its small, dark eyes [...] Detailed, samurai armor, steel, engraved patterns, leather straps, aged look.
A medieval fantasy tavern with a wooden building, purple roof, black chimney, and wooden fence. Features include a pine tree, barrels, and a well. Decorated with red [...] banners [...] Set on an octagonal platform, surrounded by blue ground and green plants. 1988 Soviet Union coin, 5 kopeck, 25 mm diameter, cupro-nickel, circular, reeded edge, star, wheat stalks, hammer and sickle, "CCCP," number "5," golden color [...]
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ShapeCraft: LLM Agents for Structured, Textured and Interactive 3D ModelingShuyuan Zhang, Chenhan Jiang, Zuoou Li, Jiankang DengNeurIPS 2025 · 被引用 6 次
- NURBGen: High-Fidelity Text-to-CAD Generation Through LLM-Driven NURBS ModelingMuhammad Usama, Mohammad Sadil Khan, Didier Stricker, Muhammad Zeshan AfzalAAAI 2026 · 被引用 2 次
- HiFi-Mesh: High-Fidelity Efficient 3D Mesh Generation via Compact Autoregressive DependenceYanfeng Li, Tao Tan, Qinquan Gao, Zhiwen Cao 等AAAI 2026
它引用的顶会 Paper40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and TrainingJierun Chen, Dongting Hu, Xijie Huang, Huseyin Coskun 等CVPR 2025
- Guided Score identity Distillation for Data-Free One-Step Text-to-Image GenerationMingyuan Zhou, Zhendong Wang, Huangjie Zheng, Hai HuangICLR 2025
- 3D Paintbrush: Local Stylization of 3D Shapes with Cascaded Score DistillationDale Decatur, Itai Lang, Kfir Aberman, Rana HanockaCVPR 2024
- Make It Count: Text-to-Image Generation with an Accurate Number of ObjectsLital Binyamin, Yoad Tewel, Hilit Segev, Eran Hirsch 等CVPR 2025
- Argus: Vision-Centric Reasoning with Grounded Chain-of-ThoughtYunze Man, De-An Huang, Guilin Liu, Shiwei Sheng 等CVPR 2025
