Multi-Semantic Modeling for Glass Surface Detection in the Wild
Qianyu Cheng, Huankang Guan, Rynson W. H. Lau
摘要
Glass surfaces challenge object detection models as they mix the transmitted background with the reflected surrounding, creating confusing visual patterns. Previous methods relying on low-level cues (e.g., reflections and boundaries) or surrounding semantics are often unreliable in complex realworld scenarios. A glass image inherently comprises three distinct semantic components: semantics of the transmitted content, semantics of the reflected content, and semantics of the surrounding content. In this work, we observe that there is a relationship among these three types of semantics, where reflection semantics closely resembles surrounding semantics, while these two types of semantics tend to be different from the transmission semantics. For example, when on a street, we may see into a cafeteria through a glass wall, intermixed with reflection of the street, while the glass is surrounded by other street contents like shops and pedestrians, thereby creating a unique multi-semantic signature. Based on this observation, we propose the Multi-Semantic Net, MSNet, which identifies transmission, reflection, and surrounding semantics from glass images and exploits their relationships for glass surface detection. MSNet consists of two novel modules: (1) A Semantic Decomposition Module (SDM) containing Dual-Semantics Extraction Block to extract original image and reflection semantics and Semantic Elimination Block to progressively derive transmission and surrounding semantics, and (2) An Adaptive Semantic Fusion Module (ASFM) to fuse these semantic components and adaptively learn their relationships to handle varying reflection conditions. Extensive experiments demonstrate that MSNet surpasses SOTA methods on public glass detection benchmarks. Code will be available at https://github.com/chengqianyu03/MSNet .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Twins: Revisiting the Design of Spatial Attention in Vision TransformersXiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang 等NeurIPS 2021 · 被引用 1,388 次
- Location-aware Single Image Reflection RemovalZheng Dong, Ke Xu, Yin Yang, Hujun Bao 等ICCV 2021 · 被引用 124 次
相关 Paper
- Exploiting Semantic Relations for Glass Surface DetectionJiaying Lin, Yuen Hei Yeung, Rynson W. H. LauNeurIPS 2022 · 被引用 36 次
- MVGD-Net: A Novel Motion-aware Video Glass Surface Detection MethodYiwei Lu, Hao Huang, Tao YanAAAI 2026
- Multi-View Dynamic Reflection Prior for Video Glass Surface DetectionFang Liu, Yuhao Liu, Jiaying Lin, Ke Xu 等AAAI 2024 · 被引用 12 次
- Rich Context Aggregation With Reflection Prior for Glass Surface DetectionJiaying Lin, Zebang He, Rynson W. H. LauCVPR 2021
- Don't Hit Me! Glass Detection in Real-World ScenesHaiyang Mei, Xin Yang, Yang Wang, Yuanyuan Liu 等CVPR 2020
