OpenESS: Event-Based Semantic Scene Understanding with Open Vocabularies
Lingdong Kong, Youquan Liu, Lai Xing Ng, Benoit R. Cottereau, Wei Tsang Ooi
Abstract
Zero-Shot Semantic Segmentation "driveable" (Adjective) "walkable" (Adjective) "car" (Fine-Grained) "manmade" (Coarse) "flat" (Coarse) "barrier" (Fine-Grained) Back Build Road Car Pole Veg Wall Figure 1. Open-vocabulary event-based semantic segmentation (OpenESS). Our framework is capable of performing zero-shot semantic segmentation of event data streams with open vocabularies. Given raw events and text prompts as inputs, OpenESS outputs semantically coherent open-world predictions across various adjective, fine-grained, and coarse categories. The last three columns show the languageguided attention maps where regions of a high similarity score to the given text prompts are highlighted. Best viewed in colors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e3ed925d-131d-4a9e-9b19-5f880bda6dd8Cited by top-tier papers15
- FlexEvent: Towards Flexible Event-Frame Object Detection at Varying Operational FrequenciesDongyue Lu, Lingdong Kong, Gim Hee Lee, Camille Simon Chane et al.NeurIPS 2025 · 13 citations
- Talk2Event: Grounded Understanding of Dynamic Scenes from Event CamerasLingdong Kong, Dongyue Lu, Alan Liang, Rong Li et al.NeurIPS 2025 · 7 citations
- EventFlash: Towards Efficient MLLMs for Event-Based VisionShaoyu Liu, Jianing Li, Guanghui Zhao, Yunjian Zhang et al.ICLR 2026 · 5 citations
- Segment Any Events with LanguageSeungjun Lee, Gim Hee LeeICLR 2026 · 3 citations
- Efficient Event-Based Semantic Segmentation via Exploiting Frame-Event Fusion: A Hybrid Neural Network ApproachHebei Li, Yansong Peng, Jiahui Yuan, Peixi Wu et al.AAAI 2025 · 3 citations
Builds on44
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Learning Open-Vocabulary Semantic Segmentation Models From Natural Language SupervisionJilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng et al.CVPR 2023
- A Simple Framework for Open-Vocabulary Segmentation and DetectionHao Zhang, Feng Li, Xueyan Zou, Shilong Liu et al.ICCV 2023 · 241 citations
- Auto-Vocabulary Semantic SegmentationOsman Ülger, Maksymilian Kulicki, Yuki Asano, Martin R. OswaldICCV 2025 · 3 citations
- Open-Vocabulary Audio-Visual Semantic SegmentationRuohao Guo, Liao Qu, Dantong Niu, Yanyu Qi et al.ACM MM 2024 · 4 citations
- Open-Vocabulary 3D Semantic Segmentation with Foundation ModelsLi Jiang, Shaoshuai Shi, Bernt SchieleCVPR 2024
