Learning Geometry-aware Representations by Sketching
Hyundo Lee, Inwoo Hwang, Hyunsung Go, Won-Seok Choi, Kibeom Kim, Byoung-Tak Zhang
Abstract
Understanding geometric concepts, such as distance and shape, is essential for understanding the real world and also for many vision tasks. To incorporate such information into a visual representation of a scene, we propose learning to represent the scene by sketching, inspired by human behavior. Our method, coined Learning by Sketching (LBS), learns to convert an image into a set of colored strokes that explicitly incorporate the geometric information of the scene in a single inference step without requiring a sketch dataset. A sketch is then generated from the strokes where CLIP-based perceptual loss maintains a semantic similarity between the sketch and the image. We show theoretically that sketching is equivariant with respect to arbitrary affine transformations and thus provably preserves geometric information. Experimental results show that LBS substantially improves the performance of object attribute classification on the unlabeled CLEVR dataset, domain transfer between CLEVR and STL-10 datasets, and for diverse downstream tasks, confirming that LBS provides rich geometric information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdece276-2bb4-4e4d-8598-87bf28c7fd01Cited by top-tier papers3
- Freehand Sketch Generation from Mechanical ComponentsZhichao Liao, Fengyuan Piao, Di Huang, Xinghui Li et al.ACM MM 2024 · 12 citations
- Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion ModelsHyundo Lee, Suhyung Choi, Inwoo Hwang, Byoung-Tak ZhangAAAI 2026
- Open Vocabulary Semantic Scene Sketch UnderstandingAhmed Bourouis, Judith Ellen Fan, Yulia GryaditskayaCVPR 2024
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- SC-FEGAN: Face Editing Generative Adversarial Network With User's Sketch and ColorYoungjoo Jo, Jongyoul ParkICCV 2019 · 325 citations
Related papers
- CLIPasso: semantically-aware object sketchingYael Vinker, Ehsan Pajouheshgar, Jessica Y. Bo, Roman Christian Bachmann et al.SIGGRAPH 2022 · 219 citations
- Learning to generate line drawings that convey geometry and semanticsCaroline Chan, Frédo Durand, Phillip IsolaCVPR 2022 · 86 citations
- Integrating Categorical Semantics into Unsupervised Domain TranslationSamuel Lavoie-Marchildon, Faruk Ahmed, Aaron C. CourvilleICLR 2021 · 7 citations
- S2R-DepthNet: Learning a Generalizable Depth-Specific Structural RepresentationXiaotian Chen, Yuwang Wang, Xuejin Chen, Wenjun ZengCVPR 2021
- Neural Strokes: Stylized Line Drawing of 3D ShapesDifan Liu, Matthew Fisher, Aaron Hertzmann, Evangelos KalogerakisICCV 2021 · 29 citations
