Copy-Transform-Paste: Zero-Shot Object-Object Alignment Guided by Vision-Language and Geometric Constraints
Rotem Gatenyo, Ohad Fried
Abstract
Burger bun bottom, lettuce, burger patty, cheese, tomatoes and burger bun top" "Nigiri"
"Captain America holding a shield" "Pinocchio wearing a hat" "A stand with a golden necklace''
Iterative procedure Figure 1. Text-guided object-object alignment and iterative composition. The figure shows four independent examples, each presenting the input meshes and text prompt alongside our alignment result. In addition, an iterative example demonstrates progressive assembly of a burger: the output of stage k is incorporated into the input of stage k+1, gradually forming the final arrangement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78088026-43ef-4dd4-8cd4-866920968aedBuilds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- Soft Rasterizer: A Differentiable Renderer for Image-Based 3D ReasoningShichen Liu, Weikai Chen, Tianye Li, Hao LiICCV 2019 · 789 citations
- StyleGAN-NADA: CLIP-guided domain adaptation of image generatorsRinon Gal, Or Patashnik, Haggai Maron, Amit H. Bermano et al.SIGGRAPH 2022 · 501 citations
Related papers
- Latent-NeRF for Shape-Guided Generation of 3D Shapes and TexturesGal Metzer, Elad Richardson, Or Patashnik, Raja Giryes et al.CVPR 2023
- The Chosen One: Consistent Characters in Text-to-Image Diffusion ModelsOmri Avrahami, Amir Hertz, Yael Vinker, Moab Arar et al.SIGGRAPH 2024 · 26 citations
- ArtFormer: Controllable Generation of Diverse 3D Articulated ObjectsJiayi Su, Youhe Feng, Zheng Li, Jinhua Song et al.CVPR 2025
- PuzzleFusion++: Auto-agglomerative 3D Fracture Assembly by Denoise and VerifyZhengqing Wang, Jiacheng Chen, Yasutaka FurukawaICLR 2025
- TAPS3D: Text-Guided 3D Textured Shape Generation from Pseudo SupervisionJiacheng Wei, Hao Wang, Jiashi Feng, Guosheng Lin et al.CVPR 2023
