Grammatically Recognizing Images with Tree Convolution
Guangrun Wang, Guangcong Wang, Keze Wang, Xiaodan Liang, Liang Lin
Abstract
Similar to language, understanding an image can be considered as a hierarchical decomposition process from scenes to objects, parts, pixels, and the corresponding spatial/contextual relations. However, the existing convolutional networks concentrate on stacking redundant convolutional layers with a large number of kernels in a hierarchical organization to implicitly approximate this decomposition. This may limit the network to learn the semantic information conveyed in the internal feature maps that may reveal minor yet crucial differences for visual understanding. Attempting to tackle this problem, this paper proposes a simple yet effective tree convolution (TreeConv) operation for deep neural networks. Specifically, inspired by the image grammar techniques[73] that serve as a unified framework of object representation, learning, and recognition, our TreeConv designs a generative image grammar, i.e., tree generation rule, to parse the hierarchy of internal feature maps by generating tree structures and implicitly learning the specific visual grammars for each object category. Extensive experiments on a variety of benchmarks, i.e., classification (ImageNet / CIFAR), detection & segmentation (COCO 2017), and person re-identification (CUHK03), demonstrate the superiority of our TreeConv in both boosting the accuracy and reducing the computational cost. The source code will be available at: https://github.com/wanggrun/TreeConv.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c75b7b0f-6ce3-412d-bb26-5d5561aa30d9Cited by top-tier papers4
- SparseNeRF: Distilling Depth Ranking for Few-shot Novel View SynthesisGuangcong Wang, Zhaoxi Chen, Chen Change Loy, Ziwei LiuICCV 2023 · 309 citations
- Solving Inefficiency of Self-supervised Representation LearningGuangrun Wang, Keze Wang, Guangcong Wang, Philip H. S. Torr et al.ICCV 2021 · 64 citations
- Pi-NAS: Improving Neural Architecture Search by Reducing Supernet Training Consistency ShiftJiefeng Peng, Jiqi Zhang, Changlin Li, Guangrun Wang et al.ICCV 2021 · 20 citations
- Semantic-Aware Auto-Encoders for Self-supervised Representation LearningGuangrun Wang, Yansong Tang, Liang Lin, Philip H. S. TorrCVPR 2022 · 8 citations
Builds on5
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li et al.AAAI 2020 · 4,134 citations
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen et al.ICCV 2019 · 1,990 citations
- Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks With Octave ConvolutionYunpeng Chen, Haoqi Fan, Bing Xu, Zhicheng Yan et al.ICCV 2019 · 665 citations
- Local Relation Networks for Image RecognitionHan Hu, Zheng Zhang, Zhenda Xie, Stephen LinICCV 2019 · 555 citations
- Smoothing Adversarial Domain Attack and P-Memory Reconsolidation for Cross-Domain Person Re-IdentificationGuangcong Wang, Jian-Huang Lai, Wenqi Liang, Guangrun WangCVPR 2020
Related papers
- TreeCaps: Tree-Based Capsule Networks for Source Code ProcessingNghi D. Q. Bui, Yijun Yu, Lingxiao JiangAAAI 2021 · 44 citations
- Hierarchical Human Parsing With Typed Part-Relation ReasoningWenguan Wang, Hailong Zhu, Jifeng Dai, Yanwei Pang et al.CVPR 2020
- Learning to Assemble Neural Module Tree Networks for Visual GroundingDaqing Liu, Hanwang Zhang, Feng Wu, Zheng-Jun ZhaICCV 2019 · 317 citations
- U-Nets as Belief Propagation: Efficient Classification, Denoising, and Diffusion in Generative Hierarchical ModelsSong MeiICLR 2025
- Hierarchy Parsing for Image CaptioningTing Yao, Yingwei Pan, Yehao Li, Tao MeiICCV 2019 · 183 citations
