Multi-view Masked Contrastive Representation Learning for Endoscopic Video Analysis
Kai Hu, Ye Xiao, Yuan Zhang, Xieping Gao
Abstract
Endoscopic video analysis can effectively assist clinicians in disease diagnosis and treatment, and has played an indispensable role in clinical medicine. Unlike regular videos, endoscopic video analysis presents unique challenges, including complex camera movements, uneven distribution of lesions, and concealment, and it typically relies on contrastive learning in self-supervised pretraining as its main-stream technique. However, representations obtained from contrastive learning enhance the discriminability of the model but often lack fine-grained information, which is suboptimal in the pixel-level prediction tasks. In this paper, we develop a M ulti-view M asked C ontrastive R epresentation L earning (M 2 CRL) framework for endoscopic video pre-training. Specifically, we propose a multi-view masking strategy for addressing the challenges of endoscopic videos. We utilize the frame-aggregated attention guided tube mask to capture global-level spatiotemporal sensitive representation from the global views, while the random tube mask is employed to focus on local variations from the local views. Subsequently, we combine multi-view mask modeling with contrastive learning to obtain endoscopic video representations that possess fine-grained perception and holistic discriminative capabilities simultaneously. The proposed M 2 CRL is pre-trained on 7 publicly available endoscopic video datasets and fine-tuned on 3 endoscopic video datasets for 3 downstream tasks. Notably, our M 2 CRL significantly outperforms the current state-of-the-art self-supervised endoscopic pre-training methods, e.g. , Endo-FM (3.5% F1 for classification, 7.5% Dice for segmentation, and 2.2% F1 for detection) and other self-supervised methods, e.g. , VideoMAE V2
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a835ae2-f6d7-4cf5-a30b-e7b391e76ec7Cited by top-tier papers4
- Confusion-Driven Self-Supervised Progressively Weighted Ensemble Learning for Non-Exemplar Class Incremental LearningKai Hu, Yu Zhang, Yuan Zhang, Zhineng Chen et al.NeurIPS 2025 · 1 citation
- Metascope: Optics-Driven Neural Network for Ultra-Micro Metalens EndoscopyWuyang Li, Wentao Pan, Xiaoyuan Liu, Zhendong Luo et al.ICCV 2025 · 1 citation
- Focus-to-Perceive Representation Learning: A Cognition-Inspired Hierarchical Framework for Endoscopic Video AnalysisYuan Zhang, Sihao Dou, Kai Hu, Shuhua Deng et al.CVPR 2026 · 1 citation
- SHands: A Multi-View Dataset and Benchmark for Surgical Hand-Gesture and Error Recognition Toward Medical TrainingLe Ma, Thiago Freitas dos Santos, Nadia Magnenat-Thalmann, Katarzyna WacCVPR 2026
Builds on33
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- Masked Motion Encoding for Self-Supervised Video Representation LearningXinyu Sun, Peihao Chen, Liangwei Chen, Changhao Li et al.CVPR 2023
- Motion-Focused Contrastive Learning of Video Representations*Rui Li, Yiheng Zhang, Zhaofan Qiu, Ting Yao et al.ICCV 2021 · 37 citations
- Unsupervised Representation for Semantic Segmentation by Implicit Cycle-Attention Contrastive LearningBo Pang, Yizhuo Li, Yifan Zhang, Gao Peng et al.AAAI 2022 · 10 citations
- TrackMAE: Video Representation Learning via Track Mask and PredictRenaud Vandeghen, Fida Mohammad Thoker, Marc Van Droogenbroeck, Bernard GhanemCVPR 2026 · 3 citations
- Self-supervised Pre-training and Contrastive Representation Learning for Multiple-choice Video QASeonhoon Kim, Seohyeong Jeong, Eunbyul Kim, Inho Kang et al.AAAI 2021 · 44 citations
