Learning Geometry-aware Representations by Sketching
Hyundo Lee, Inwoo Hwang, Hyunsung Go, Won-Seok Choi, Kibeom Kim, Byoung-Tak Zhang
摘要
Understanding geometric concepts, such as distance and shape, is essential for understanding the real world and also for many vision tasks. To incorporate such information into a visual representation of a scene, we propose learning to represent the scene by sketching, inspired by human behavior. Our method, coined Learning by Sketching (LBS), learns to convert an image into a set of colored strokes that explicitly incorporate the geometric information of the scene in a single inference step without requiring a sketch dataset. A sketch is then generated from the strokes where CLIP-based perceptual loss maintains a semantic similarity between the sketch and the image. We show theoretically that sketching is equivariant with respect to arbitrary affine transformations and thus provably preserves geometric information. Experimental results show that LBS substantially improves the performance of object attribute classification on the unlabeled CLEVR dataset, domain transfer between CLEVR and STL-10 datasets, and for diverse downstream tasks, confirming that LBS provides rich geometric information.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Freehand Sketch Generation from Mechanical ComponentsZhichao Liao, Fengyuan Piao, Di Huang, Xinghui Li 等ACM MM 2024 · 被引用 12 次
- Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion ModelsHyundo Lee, Suhyung Choi, Inwoo Hwang, Byoung-Tak ZhangAAAI 2026
- Open Vocabulary Semantic Scene Sketch UnderstandingAhmed Bourouis, Judith Ellen Fan, Yulia GryaditskayaCVPR 2024
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- SC-FEGAN: Face Editing Generative Adversarial Network With User's Sketch and ColorYoungjoo Jo, Jongyoul ParkICCV 2019 · 被引用 325 次
相关 Paper
- CLIPasso: semantically-aware object sketchingYael Vinker, Ehsan Pajouheshgar, Jessica Y. Bo, Roman Christian Bachmann 等SIGGRAPH 2022 · 被引用 219 次
- Learning to generate line drawings that convey geometry and semanticsCaroline Chan, Frédo Durand, Phillip IsolaCVPR 2022 · 被引用 86 次
- Integrating Categorical Semantics into Unsupervised Domain TranslationSamuel Lavoie-Marchildon, Faruk Ahmed, Aaron C. CourvilleICLR 2021 · 被引用 7 次
- S2R-DepthNet: Learning a Generalizable Depth-Specific Structural RepresentationXiaotian Chen, Yuwang Wang, Xuejin Chen, Wenjun ZengCVPR 2021
- Neural Strokes: Stylized Line Drawing of 3D ShapesDifan Liu, Matthew Fisher, Aaron Hertzmann, Evangelos KalogerakisICCV 2021 · 被引用 29 次
