DIMOS: Disentangling Instance-level Moving Object Segmentation
Hongxiang Huang, Hongwei Ren, Xiaopeng Lin, Yulong Huang, Zeke Xie, Bojun Cheng
摘要
Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, and animal tracking. Event cameras record asynchronous brightness changes, providing high temporal resolution and dynamic range, which makes them highly sensitive to motion information. By fusing event and image features, motion cues from events can complement spatial details from images, enhancing the performance of MIS. However, current multimodal MIS methods still struggle to segment small moving instances, as event cameras often yield sparse features under limited resolution. In addition, event features entangle appearance attributes with motion cues, which further restricts effective cross-modal fusion. To address these challenges, we first propose a dualdisentangling feature extraction framework that separates and extracts appearance and motion information within both image and event modalities, thereby improving feature density. Subsequently, a multi-granularity cross-modal alignment is introduced to align distributionally and semantically consistent features across modalities, enabling more effective fusion with rich spatial and temporal details. The experiment results demonstrate that our method achieves state-of-the-art performance in multimodal MIS, especially for small instances under challenging conditions such as fast motion and low-light settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- Temporal-wise Attention Spiking Neural Networks for Event Streams ClassificationMan Yao, Huanhuan Gao, Guangshe Zhao, Dingheng Wang 等ICCV 2021 · 被引用 225 次
- CLIF: Complementary Leaky Integrate-and-Fire Neuron for Spiking Neural NetworksYulong Huang, Xiaopeng Lin, Hongwei Ren, Haotian Fu 等ICML 2024 · 被引用 43 次
- Isomer: Isomerous Transformer for Zero-shot Video Object SegmentationYichen Yuan, Yifan Wang, Lijun Wang, Xiaoqi Zhao 等ICCV 2023 · 被引用 16 次
相关 Paper
- Separation for Better Integration: Disentangling Edge and Motion in Event-Based DeblurringYufei Zhu, Hao Chen, Yongjian Deng, Wei YouICCV 2025 · 被引用 1 次
- MambaSeg: Harnessing Mamba for Accurate and Efficient Image-Event Semantic SegmentationFuqiang Gu, Yuanke Li, Xianlei Long, Kangping Ji 等AAAI 2026 · 被引用 1 次
- ESEG: Event-Based Segmentation Boosted by Explicit Edge-Semantic GuidanceYucheng Zhao, Gengyu Lyu, Ke Li, Zihao Wang 等AAAI 2025 · 被引用 8 次
- Enhanced Event-Based Dense Stereo via Cross-Sensor Knowledge DistillationHaihao Zhang, Yunjian Zhang, Jianing Li, Lin Zhu 等ICCV 2025 · 被引用 1 次
- AlignTrack: Top-Down Spatiotemporal Resolution Alignment for RGB-Event Visual TrackingChuanyu Sun, Jiqing Zhang, Yang Wang, Yuanchen Wang 等AAAI 2026
