Learning from the Tangram to Solve Mini Visual Tasks
Yizhou Zhao, Liang Qiu, Pan Lu, Feng Shi, Tian Han, Song-Chun Zhu
Abstract
Current pre-training methods in computer vision focus on natural images in the daily-life context. However, abstract diagrams such as icons and symbols are common and important in the real world. This work is inspired by Tangram, a game that requires replicating an abstract pattern from seven dissected shapes. By recording human experience in solving tangram puzzles, we present the Tangram dataset and show that a pre-trained neural model on the Tangram helps solve some mini visual tasks based on low-resolution vision. Extensive experiments demonstrate that our proposed method generates intelligent solutions for aesthetic tasks such as folding clothes and evaluating room layouts. The pre-trained feature extractor can facilitate the convergence of few-shot learning tasks on human handwriting and improve the accuracy in identifying icons by their contours. The Tangram dataset is available at https://github.com/yizhouzhao/Tangram .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
- Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAMLAniruddh Raghu, Maithra Raghu, Samy Bengio, Oriol VinyalsICLR 2020 · 736 citations
Related papers
- Abstract Visual Reasoning with Tangram ShapesAnya Ji, Noriyuki Kojima, Noah Rush, Alane Suhr et al.EMNLP 2022 · 18 citations
- Learning to Infer Generative Template Programs for Visual ConceptsR. Kenny Jones, Siddhartha Chaudhuri, Daniel RitchieICML 2024 · 3 citations
- XtarNet: Learning to Extract Task-Adaptive Representation for Incremental Few-Shot LearningSung Whan Yoon, Do-Yeon Kim, Jun Seo, Jaekyun MoonICML 2020 · 49 citations
- Iconary: A Pictionary-Based Game for Testing Multimodal Communication with Drawings and TextChristopher Clark, Jordi Salvador, Dustin Schwenk, Derrick Bonafilia et al.EMNLP 2021 · 6 citations
- FSS-1000: A 1000-Class Dataset for Few-Shot SegmentationXiang Li, Tianhan Wei, Yau Pun Chen, Yu-Wing Tai et al.CVPR 2020
