EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance
Yang Yue, Yulin Wang, Haojun Jiang, Pan Liu, Shiji Song, Gao Huang
Abstract
Echocardiography is crucial for cardiovascular disease detection but relies heavily on experienced sonographers. Echocardiography probe guidance systems, which provide real-time movement instructions for acquiring standard plane images, offer a promising solution for AI-assisted or fully autonomous scanning. However, developing effective machine learning models for this task remains challenging, as they must grasp heart anatomy and the intricate interplay between probe motion and visual signals.
To address this, we present EchoWorld, a motion-aware world modeling framework for probe guidance that encodes anatomical knowledge and motion-induced visual dynamics, while effectively leveraging past visual-motion sequences to enhance guidance precision. EchoWorld employs a pre-training strategy inspired by world modeling principles, where the model predicts masked anatomical regions and simulates the visual outcomes of probe adjustments. Built upon this pre-trained model, we introduce a motion-aware attention mechanism in the fine-tuning stage that effectively integrates historical visual-motion data, enabling precise and adaptive probe guidance. Trained on more than one million ultrasound images from over 200 routine scans, EchoWorld effectively captures key echocardiographic knowledge, as validated by qualitative analysis. Moreover, our method significantly reduces guidance errors compared to existing visual backbones and guidance frameworks, excelling in both single-frame and sequential evaluation protocols. Code is available at https: //github.com/LeapLabTHU/EchoWorld.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42cd36f5-c34b-44d6-9bee-9f5ca944f68aCited by top-tier papers2
- DMWM: Dual-Mind World Model with Long-Term ImaginationLingyi Wang, Rashed Shelim, Walid Saad, Naren RamakrishnanNeurIPS 2025 · 15 citations
- MRI Contrast Enhancement Kinetics World ModelJindi Kong, Yuting He, Cong Xia, Rongjun Ge et al.CVPR 2026 · 3 citations
Builds on19
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
Related papers
- ECHOPulse: ECG Controlled Echocardio-gram Video GenerationYiwei Li, Sekeun Kim, Zihao Wu, Hanqi Jiang et al.ICLR 2025 · 2 citations
- Reciprocal Landmark Detection and Tracking With Extremely Few AnnotationsJianzhe Lin, Ghazal Sahebzamani, Christina Luong, Fatemeh Taheri Dezaki et al.CVPR 2021
- Anatomy-aware Representation Learning for Medical UltrasoundSeok-Hwan Oh, Myeong-Gee Kim, Guil Jung, Hyeon-Jik Lee et al.ICLR 2026 · 23 citations
- EchoONE: Segmenting Multiple Echocardiography Planes in One ModelJiongtong Hu, Wufeng Xue, Jun Cheng, Yingying Liu et al.CVPR 2025
- Semi-supervised Echocardiography Video Segmentation via Anchor Semantic Awareness and Continuous Pseudo-label ReforgingYunpeng Fang, Yimu Sun, Jingxing Guo, Huisi Wu et al.CVPR 2026
