Asynchronous Collaborative Graph Representation for Frames and Events
Dianze Li, Jianing Li, Xu Liu, Xiaopeng Fan, Yonghong Tian
Abstract
Integrating frames and events has become a widely accepted solution for various tasks in challenging scenarios. However, most multimodal methods directly convert events into image-like formats synchronized with frames and process each stream through separate two-branch backbones, making it difficult to fully exploit the spatiotemporal events while limiting inference frequency to the frame rate. To address these problems, we propose a novel asynchronous collaborative graph representation, namely ACGR, which is the first trial to explore a unified graph framework for asynchronously processing frames and events with high performance and low latency. Technically, we first construct unimodal graphs for frames and events to preserve their spatiotemporal properties and sparsity. Then, an asynchronous collaborative alignment module is designed to align and fuse frames and events into a unified graph and the ACGR is generated through graph convolutional networks. Finally, we innovatively introduce domain adaptation to enable cross-modal interactions between frames and events by aligning their feature spaces. Experimental results show that our approach outperforms state-of-the-art methods in both object detection and depth estimation tasks, while significantly reducing computational latency and achieving real-time inference up to 200 Hz. Our code can be available at https://github.com/dianzl/ACGR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Maximizing Asynchronicity in Event-based Neural NetworksHaiqing Hao, Nikola Zubic, Weihua He, Zhipeng Sui et al.ICLR 2026 · 2 citations
- Beyond Duality: A Hybrid Framework of Leveraging Shared and Private Features for RGB-Event Object DetectionKeyao Wang, Shuai Liu, Hengda Shi, Lukui Shi et al.CVPR 2026
Builds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- End-to-End Learning of Representations for Asynchronous Event-Based DataDaniel Gehrig, Antonio Loquercio, Konstantinos G. Derpanis, Davide ScaramuzzaICCV 2019 · 427 citations
- Deep Directly-Trained Spiking Neural Networks for Object DetectionQiaoyi Su, Yuhong Chou, Yifan Hu, Jianing Li et al.ICCV 2023 · 143 citations
- Object Tracking by Jointly Exploiting Frame and Event DomainJiqing Zhang, Xin Yang, Yingkai Fu, Xiaopeng Wei et al.ICCV 2021 · 141 citations
Related papers
- AEGNN: Asynchronous Event-based Graph Neural NetworksSimon Schaefer, Daniel Gehrig, Davide ScaramuzzaCVPR 2022 · 135 citations
- Ev-3DOD: Pushing the Temporal Boundaries of 3D Object Detection with Event CamerasHoonhee Cho, Jae-Young Kang, Youngho Kim, Kuk-Jin YoonCVPR 2025
- Graph Neural Network Combining Event Stream and Periodic Aggregation for Low-Latency Event-based VisionManon Dampfhoffer, Thomas Mesquida, Damien Joubert, Thomas Dalgaty et al.CVPR 2025
- AIMDepth: Asymmetric Image-Event Mamba for Monocular Depth EstimationLuoxi Jing, Dianxi Shi, YuShe Cao, Yuanze Wang et al.CVPR 2026
- When Every Millisecond Counts: Real-Time Anomaly Detection via the Multimodal Asynchronous Hybrid NetworkDong Xiao, Guangyao Chen, Peixi Peng, Yangru Huang et al.ICML 2025
