Anatomy-aware Representation Learning for Medical Ultrasound
Seok-Hwan Oh, Myeong-Gee Kim, Guil Jung, Hyeon-Jik Lee, Young-Min Kim, Sang-Yun Kim, Hyuksool Kwon, Hyeon-Min Bae
Abstract
Diagnostic accuracy of ultrasound imaging is limited by qualitative variability and its reliance on the expertise of medical professionals. Such challenges increase demand for computer-aided diagnostic systems that enhance diagnostic accuracy and efficiency. However, the unique texture and structural attributes of ultrasound images, and the scarcity of large-scale ultrasound datasets hinder the effective application of conventional machine learning methodologies. To address the challenges, we propose Anatomy-aware Representation Learning (ARL), a novel self-supervised representation learning framework specifically designed for medical ultrasound imaging. ARL incorporates an anatomy-adaptive Vision Transformer (A-ViT). The A-ViT is parameterized, using the proposed large-scale medical ultrasound dataset, to provide anatomy-aware feature representations. Through extensive experiments across various ultrasound-based diagnostic tasks, including breast and thyroid cancer, cardiac view classification, and gallbladder tumor and COVID-19 identification, we demonstrate that ARL significantly outperforms existing self-supervised learning baselines. The experiments demonstrate the potential of ARL in advancing medical ultrasound diagnostics by providing anatomy-specific feature representation
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de6e3e8c-faee-4afa-a76e-8f1bc0ddb35fBuilds on14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- DeNAS-ViT: Data Efficient NAS-Optimized Vision Transformer for Ultrasound Image SegmentationRenqi Chen, Xinzhe Zheng, Haoyang Su, Kehan WuAAAI 2026 · 3 citations
- FocusMAE: Gallbladder Cancer Detection from Ultrasound Videos with Focused Masked AutoencodersSoumen Basu, Mayuna Gupta, Chetan Madan, Pankaj Gupta et al.CVPR 2024
- Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image AnalysisYucheng Tang, Dong Yang, Wenqi Li, Holger R. Roth et al.CVPR 2022 · 736 citations
- Context Matters: Graph-based Self-supervised Representation Learning for Medical ImagesLi Sun, Ke Yu, Kayhan BatmanghelichAAAI 2021 · 27 citations
- Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text UnderstandingJiayun Jin, Haolong Chai, Xueying Huang, Xiaoqing Guo et al.CVPR 2026 · 4 citations
