Cross-Timestep: 3D Diffusion Model with Trans-temporal Memory LSTM and Adaptive Priori Decoding Strategy for Medical Segmentation
Shangqian Wu, Siyuan Shen, Yahan Li, Zhijian Huang, Ziyu Fan, Yuanpeng Zhang, Yi Wang, Lei Deng
Abstract
Diffusion models have recently demonstrated significant robustness in medical image segmentation, effectively accommodating variations across different imaging styles. However, their applications remain limited due to: (i) current successes being primarily confined to 2D segmentation tasks-we observe that diffusion models tend to collapse at the early stage when applied to 3D medical tasks; and (ii) the inherently isolated iteration along timesteps during training and inference. To tackle these limitations, we propose a novel framework named Cross-Timestep, which incorporates two key innovations: an Adaptive Priori Decoding Strategy (APDS) and a trans-temporal memory LSTM (tLSTM) mechanism. (i) The APDS provides prior guidance during the diffusion process by employing a Priori Decoder(PD) that focuses solely on the conditional branch, successfully stabilizing the reverse diffusion process. (ii) The tLSTM integrates convolution and linear layers into the LSTM gating structure, and enhances the memory cell mechanism to retain temporal state, explicitly preserving and propagating continuous temporal states across timesteps. Experimental results demonstrate that Cross-Timestep performs favorably on heterogeneous 3D medical datasets. Three experiments further analyze the collapse phenomenon in 3D medical diffusion models and validate that APDS effectively prevents initial-stage collapse without excessively constraining the model, while tLSTM facilitates the performance and scalability of diffusion models. Although there was no explicitly shown embedded time t during this process, X c already incorporates the time embedding information, as it allows the APDS to be aware of the current diffusion
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- Semi-supervised Medical Image Segmentation through Dual-task ConsistencyXiangde Luo, Jieneng Chen, Tao Song, Guotai WangAAAI 2021 · 754 citations
- MedSegDiff-V2: Diffusion-Based Medical Image Segmentation with TransformerJunde Wu, Wei Ji, Huazhu Fu, Min Xu et al.AAAI 2024 · 311 citations
- 3D UX-Net: A Large Kernel Volumetric ConvNet Modernizing Hierarchical Transformer for Medical Image SegmentationHo Hin Lee, Shunxing Bao, Yuankai Huo, Bennett A. LandmanICLR 2023 · 100 citations
Related papers
- CNM-UNet: Continuous Ordinary Differential Equations for Medical Image SegmentationTianqi Xu, Yashi Zhu, Quansong He, Yue Cao et al.AAAI 2026
- TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion ModelZhenkai Zhang, Krista A. Ehinger, Tom DrummondAAAI 2025
- Adaptive Domain Shift in Diffusion Models for Cross-Modality Image TranslationZihao Wang, Yuzhou Chen, Shaogang RenICLR 2026 · 6 citations
- Decoding with Structured Awareness: Integrating Directional, Frequency-Spatial, and Structural Attention for Medical Image SegmentationFan Zhang, Zhiwei Gu, Hua WangAAAI 2026 · 4 citations
- Improving 3D Imaging with Pre-Trained Perpendicular 2D Diffusion ModelsSuhyeon Lee, Hyungjin Chung, Minyoung Park, Jonghyuk Park et al.ICCV 2023 · 72 citations
