GLID: Pre-training a Generalist Encoder-Decoder Vision Model
Jihao Liu, Jinliang Zheng, Yu Liu, Hongsheng Li
摘要
This paper proposes a GeneraLIst encoder-Decoder (GLID) pre-training method for better handling various downstream computer vision tasks. While self-supervised pre-training approaches, e.g., Masked Autoencoder, have shown success in transfer learning, task-specific sub-architectures are still required to be appended for differ-ent downstream tasks, which cannot enjoy the benefits of large-scale pre-training. GLID overcomes this challenge by allowing the pre-trained generalist encoder-decoder to be fine-tuned on various vision tasks with minimal task-specific architecture modifications. In the GLID training scheme, pre-training pretext task and other downstream tasks are modeled as “query-to-answer” problems, including the pre-training pretext task and other downstream tasks. We pre-train a task-agnostic encoder-decoder with query-mask pairs. During fine-tuning, GLID maintains the pre-trained encoder-decoder and queries, only replacing the topmost linear transformation layer with task-specific linear heads. This minimizes the pretrain-finetune architecture inconsis-tency and enables the pre-trained model to better adapt to downstream tasks. GLID achieves competitive performance on various vision tasks, including object detection, image segmentation, pose estimation, and depth estimation, outper-forming or matching specialist models such as Mask2Former, DETR, ViTPose, and BinsFormer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SQS: Enhancing Sparse Perception Models via Query-based Splatting in Autonomous DrivingHaiming Zhang, Yiyao Zhu, Wending Zhou, Xu Yan 等NeurIPS 2025 · 被引用 5 次
- Argus: A Compact and Versatile Foundation Model for VisionWeiming Zhuang, Chen Chen, Zhizhong Li, Sina Sajadmanesh 等CVPR 2025
它引用的顶会 Paper21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin 等CVPR 2022 · 被引用 1,129 次
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu 等ICML 2022 · 被引用 1,123 次
相关 Paper
- Masked AutoDecoder is Effective Multi-Task Vision GeneralistHan Qiu, Jiaxing Huang, Peng Gao, Lewei Lu 等CVPR 2024
- Mask3D: Pretraining 2D Vision Transformers by Learning Masked 3D PriorsJi Hou, Xiaoliang Dai, Zijian He, Angela Dai 等CVPR 2023
- Generic-to-Specific Distillation of Masked AutoencodersWei Huang, Zhiliang Peng, Li Dong, Furu Wei 等CVPR 2023
- Effective Adaptation in Multi-Task Co-Training for Unified Autonomous DrivingXiwen Liang, Yangxin Wu, Jianhua Han, Hang Xu 等NeurIPS 2022 · 被引用 53 次
- Continual Learners are Incremental Model GeneralizersJaehong Yoon, Sung Ju Hwang, Yue CaoICML 2023 · 被引用 6 次
