EchoDiffusion: Waveform Conditioned Diffusion Models for Echo-Based Depth Estimation
Wenjie Zhang, Jun Yin, Long Ma, Peng Yu, Xiaoheng Jiang, Zhen Tian, Mingliang Xu
摘要
To extract spatial information, depth estimation using conventional echo-based methods typically employs models with encoder-decoder architectures, such as UNet. However, these methods may face challenges in extracting fine details from echo waveforms and handling multi-scale feature extraction with high precision. To address these challenges, we introduce EchoDiffusion, a framework that incorporates diffusion models conditioned on waveform embeddings for echo-based depth estimation. This framework employs the Multi-Scale Adaptive Latent Feature Network (MALF-Net) to extract multi-scale spatial features and perform adaptive fusion, encoding the echo spectrograms into the latent space. Additionally, we propose the Echo Waveform Detail Embedder (EWDE), which leverages a pre-trained Wav2Vec model to extract detailed spatial information from echo waveforms, using these details as conditional inputs to guide the reverse diffusion process in the latent space. By embedding the echo waveforms into the reverse diffusion process, we can more accurately guide the generation of depth maps. Our extensive evaluations on the Replica and Matterport3D datasets demonstrate that EchoDiffusion establishes new benchmarks for state-of-the-art performance in echo-based depth estimation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- BEVStereo: Enhancing Depth Estimation in Multi-View 3D Object Detection with Temporal StereoYinhao Li, Han Bao, Zheng Ge, Jinrong Yang 等AAAI 2023 · 被引用 226 次
- DDP: Diffusion Model for Dense Visual PredictionYuanfeng Ji, Zhe Chen, Enze Xie, Lanqing Hong 等ICCV 2023 · 被引用 223 次
相关 Paper
- Beyond Image to Depth: Improving Depth Prediction Using EchoesKranti Kumar Parida, Siddharth Srivastava, Gaurav SharmaCVPR 2021
- EchoVDiff: Cardiac-Cycle Echocardiography Video Generation from Arbitrary FrameJiansong Zhang, Xiaying Yang, Xiaoling Luo, Linlin ShenCVPR 2026
- Rectifying Latent Space for Generative Single-Image Reflection RemovalMingjia Li, Jin Hu, Hainuo Wang, Qiming Hu 等CVPR 2026 · 被引用 2 次
- From Discrete Tokens to High-Fidelity Audio Using Multi-Band DiffusionRobin San Roman, Yossi Adi, Antoine Deleforge, Romain Serizel 等NeurIPS 2023 · 被引用 50 次
- Matryoshka Diffusion ModelsJiatao Gu, Shuangfei Zhai, Yizhe Zhang, Joshua Susskind 等ICLR 2024 · 被引用 73 次
