Class-level Structural Relation Modeling and Smoothing for Visual Representation Learning
Zitan Chen, Zhuang Qi, Xiao Cao, Xiangxian Li, Xiangxu Meng, Lei Meng
摘要
Representation learning for images has been advanced by recent progress in more complex neural models such as the Vision Transformers and new learning theories such as the structural causal models. However, these models mainly rely on the classification loss to implicitly regularize the class-level data distributions, and they may face difficulties when handling classes with diverse visual patterns. We argue that the incorporation of the structural information between data samples may improve this situation. To achieve this goal, this paper presents a framework termed Class-level Structural Relation Modeling and Smoothing for Visual Representation Learning (CSRMS), which includes the Class-level Relation Modelling, Class-aware Graph Sampling, and Relational Graph-Guided Representation Learning modules to model a relational graph of the entire dataset and perform class-aware smoothing and regularization operations to alleviate the issue of intra-class visual diversity and inter-class similarity. Specifically, the Class-level Relation Modelling module uses a clustering algorithm to learn the data distributions in the feature space and identify three types of class-level sample relations for the training set; Class-aware Graph Sampling module extends typical training batch construction process with three strategies to sample dataset-level sub-graphs; and Relational Graph-Guided Representation Learning module employs a graph convolution network with knowledge-guided smoothing operations to ease the projection from different visual patterns to the same class. Experiments demonstrate the effectiveness of structured knowledge modelling for enhanced representation learning and show that CSRMS can be incorporated with any state-of-the-art visual representation learning models for performance gains. The source codes and demos have been released at https://github.com/czt117/CSRMS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Cross-Silo Feature Space Alignment for Federated Learning on Clients with Imbalanced DataZhuang Qi, Lei Meng, Zhaochuan Li, Han Hu 等AAAI 2025 · 被引用 39 次
- Curriculum Conditioned Diffusion for Multimodal RecommendationYimeng Yang, Haokai Ma, Lei Meng, Shuo Xu 等AAAI 2025 · 被引用 12 次
- Class-wise Balancing Data Replay for Federated Class-Incremental LearningZhuang Qi, Ying-Peng Tang, Lei Meng, Han Yu 等NeurIPS 2025 · 被引用 11 次
- Causal Inference over Visual-Semantic-Aligned Graph for Image ClassificationLei Meng, Xiangxian Li, Xiaoshuo Yan, Haokai Ma 等AAAI 2025 · 被引用 11 次
- 3DOT: Texture Transfer for 3DGS Objects from a Single Reference ImageXiao Cao, Beibei Lin, Bo Wang, Zhiyong Huang 等NeurIPS 2025 · 被引用 8 次
它引用的顶会 Paper19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li 等AAAI 2020 · 被引用 4,134 次
- On the Integration of Self-Attention and ConvolutionXuran Pan, Chunjiang Ge, Rui Lu, Shiji Song 等CVPR 2022 · 被引用 516 次
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang 等SIGIR 2021 · 被引用 198 次
相关 Paper
- Modeling Event-level Causal Representation for Video ClassificationYuqing Wang, Lei Meng, Haokai Ma, Yuqing Wang 等ACM MM 2024 · 被引用 3 次
- Improving Intra- and Inter-Modality Visual Relation for Image CaptioningYong Wang, Wenkai Zhang, Qing Liu, Zhengyuan Zhang 等ACM MM 2020 · 被引用 24 次
- Graph Your Own PromptXi Ding, Lei Wang, Piotr Koniusz, Yongsheng GaoNeurIPS 2025 · 被引用 6 次
- RelCLIP: Adapting Language-Image Pretraining for Visual Relationship Detection via Relational Contrastive LearningYi Zhu, Zhaoqing Zhu, Bingqian Lin, Xiaodan Liang 等EMNLP 2022 · 被引用 8 次
- Is Visual Context Really Helpful for Knowledge Graph? A Representation Learning PerspectiveMeng Wang, Sen Wang, Han Yang, Zheng Zhang 等ACM MM 2021 · 被引用 129 次
