Multi-Semantic Modeling for Glass Surface Detection in the Wild
Qianyu Cheng, Huankang Guan, Rynson W. H. Lau
Abstract
Glass surfaces challenge object detection models as they mix the transmitted background with the reflected surrounding, creating confusing visual patterns. Previous methods relying on low-level cues (e.g., reflections and boundaries) or surrounding semantics are often unreliable in complex realworld scenarios. A glass image inherently comprises three distinct semantic components: semantics of the transmitted content, semantics of the reflected content, and semantics of the surrounding content. In this work, we observe that there is a relationship among these three types of semantics, where reflection semantics closely resembles surrounding semantics, while these two types of semantics tend to be different from the transmission semantics. For example, when on a street, we may see into a cafeteria through a glass wall, intermixed with reflection of the street, while the glass is surrounded by other street contents like shops and pedestrians, thereby creating a unique multi-semantic signature. Based on this observation, we propose the Multi-Semantic Net, MSNet, which identifies transmission, reflection, and surrounding semantics from glass images and exploits their relationships for glass surface detection. MSNet consists of two novel modules: (1) A Semantic Decomposition Module (SDM) containing Dual-Semantics Extraction Block to extract original image and reflection semantics and Semantic Elimination Block to progressively derive transmission and surrounding semantics, and (2) An Adaptive Semantic Fusion Module (ASFM) to fuse these semantic components and adaptively learn their relationships to handle varying reflection conditions. Extensive experiments demonstrate that MSNet surpasses SOTA methods on public glass detection benchmarks. Code will be available at https://github.com/chengqianyu03/MSNet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3f09174-e4e0-4999-8fb4-2e3ed4db29c5Builds on24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Twins: Revisiting the Design of Spatial Attention in Vision TransformersXiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang et al.NeurIPS 2021 · 1,388 citations
- Location-aware Single Image Reflection RemovalZheng Dong, Ke Xu, Yin Yang, Hujun Bao et al.ICCV 2021 · 124 citations
Related papers
- Exploiting Semantic Relations for Glass Surface DetectionJiaying Lin, Yuen Hei Yeung, Rynson W. H. LauNeurIPS 2022 · 36 citations
- MVGD-Net: A Novel Motion-aware Video Glass Surface Detection MethodYiwei Lu, Hao Huang, Tao YanAAAI 2026
- Multi-View Dynamic Reflection Prior for Video Glass Surface DetectionFang Liu, Yuhao Liu, Jiaying Lin, Ke Xu et al.AAAI 2024 · 12 citations
- Rich Context Aggregation With Reflection Prior for Glass Surface DetectionJiaying Lin, Zebang He, Rynson W. H. LauCVPR 2021
- Don't Hit Me! Glass Detection in Real-World ScenesHaiyang Mei, Xin Yang, Yang Wang, Yuanyuan Liu et al.CVPR 2020
