Multi-view Masked Contrastive Representation Learning for Endoscopic Video Analysis
Kai Hu, Ye Xiao, Yuan Zhang, Xieping Gao
摘要
Endoscopic video analysis can effectively assist clinicians in disease diagnosis and treatment, and has played an indispensable role in clinical medicine. Unlike regular videos, endoscopic video analysis presents unique challenges, including complex camera movements, uneven distribution of lesions, and concealment, and it typically relies on contrastive learning in self-supervised pretraining as its main-stream technique. However, representations obtained from contrastive learning enhance the discriminability of the model but often lack fine-grained information, which is suboptimal in the pixel-level prediction tasks. In this paper, we develop a M ulti-view M asked C ontrastive R epresentation L earning (M 2 CRL) framework for endoscopic video pre-training. Specifically, we propose a multi-view masking strategy for addressing the challenges of endoscopic videos. We utilize the frame-aggregated attention guided tube mask to capture global-level spatiotemporal sensitive representation from the global views, while the random tube mask is employed to focus on local variations from the local views. Subsequently, we combine multi-view mask modeling with contrastive learning to obtain endoscopic video representations that possess fine-grained perception and holistic discriminative capabilities simultaneously. The proposed M 2 CRL is pre-trained on 7 publicly available endoscopic video datasets and fine-tuned on 3 endoscopic video datasets for 3 downstream tasks. Notably, our M 2 CRL significantly outperforms the current state-of-the-art self-supervised endoscopic pre-training methods, e.g. , Endo-FM (3.5% F1 for classification, 7.5% Dice for segmentation, and 2.2% F1 for detection) and other self-supervised methods, e.g. , VideoMAE V2
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Confusion-Driven Self-Supervised Progressively Weighted Ensemble Learning for Non-Exemplar Class Incremental LearningKai Hu, Yu Zhang, Yuan Zhang, Zhineng Chen 等NeurIPS 2025 · 被引用 1 次
- Metascope: Optics-Driven Neural Network for Ultra-Micro Metalens EndoscopyWuyang Li, Wentao Pan, Xiaoyuan Liu, Zhendong Luo 等ICCV 2025 · 被引用 1 次
- Focus-to-Perceive Representation Learning: A Cognition-Inspired Hierarchical Framework for Endoscopic Video AnalysisYuan Zhang, Sihao Dou, Kai Hu, Shuhua Deng 等CVPR 2026 · 被引用 1 次
- SHands: A Multi-View Dataset and Benchmark for Surgical Hand-Gesture and Error Recognition Toward Medical TrainingLe Ma, Thiago Freitas dos Santos, Nadia Magnenat-Thalmann, Katarzyna WacCVPR 2026
它引用的顶会 Paper33
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
相关 Paper
- Masked Motion Encoding for Self-Supervised Video Representation LearningXinyu Sun, Peihao Chen, Liangwei Chen, Changhao Li 等CVPR 2023
- Motion-Focused Contrastive Learning of Video Representations*Rui Li, Yiheng Zhang, Zhaofan Qiu, Ting Yao 等ICCV 2021 · 被引用 37 次
- Unsupervised Representation for Semantic Segmentation by Implicit Cycle-Attention Contrastive LearningBo Pang, Yizhuo Li, Yifan Zhang, Gao Peng 等AAAI 2022 · 被引用 10 次
- TrackMAE: Video Representation Learning via Track Mask and PredictRenaud Vandeghen, Fida Mohammad Thoker, Marc Van Droogenbroeck, Bernard GhanemCVPR 2026 · 被引用 3 次
- Self-supervised Pre-training and Contrastive Representation Learning for Multiple-choice Video QASeonhoon Kim, Seohyeong Jeong, Eunbyul Kim, Inho Kang 等AAAI 2021 · 被引用 44 次
