H2A2: Homogeneity-Aware and Heterogeneity-Aware Feature Perception for Unified Indoor 3D Object Detection
Tao Xie, Tao An, Feng Liu, Wensheng Jin, Zhengyu Li, Lijun Zhao, Ruifeng Li
Abstract
In this work, we observe that for indoor 3D object detection, fundamental geometric cues induce homogeneous spatial responses across scenes, whereas scene-specific structure yields heterogeneous signatures. However, existing detectors lack effective mechanisms to jointly extract and exploit such dual properties, which imposes inherent limitations on detection performance. Guided by this insight, we propose HA, a homogeneity-aware and heterogeneity-aware feature perception network for unified indoor 3D object detection under cross-scene training paradigms.Technically, we introduce a structural-feature-aware kernel selection (SF-KS) method, which encompasses three core components:(i) task-aware linear modulation, a channel-wise affine transformation that strengthens scene-structural feature representation; (ii) kernel weight selection strategy that integrates an offset validity prior to suppress non-informative cross-scene transfer while utilizing a structural consistency posterior to capture scene-homogeneous cues. and (iii) task-aware channel gating that suppresses scene-irrelevant feature responses. Overall, SF-KS enables the precise optimization of homogeneous features while specializing in scene-specific heterogeneous ones. In addition, to stabilize cross-scene optimization, we further introduce norm-based gradient homogenization (NGH) algorithm, which normalizes and dynamically reweights per-task gradient norms to mitigate conflicts and promote consistent updates. Extensive experiments on diverse indoor benchmarks show that HA delivers consistent gains over strong baselines and improves cross-scene generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on16
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- An End-to-End Transformer Model for 3D Object DetectionIshan Misra, Rohit Girdhar, Armand JoulinICCV 2021 · 602 citations
- Group-Free 3D Object Detection via TransformersZe Liu, Zheng Zhang, Yue Cao, Han Hu et al.ICCV 2021 · 368 citations
- UniT: Multimodal Multitask Learning with a Unified TransformerRonghang Hu, Amanpreet SinghICCV 2021 · 354 citations
Related papers
- VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object DetectionYang Cao, Feize Wu, Dave Chen, Yingji Zhong et al.CVPR 2026 · 6 citations
- Uni3DETR: Unified 3D Detection TransformerZhenyu Wang, Ya-Li Li, Xi Chen, Hengshuang Zhao et al.NeurIPS 2023 · 65 citations
- HyperDet3D: Learning a Scene-conditioned 3D Object DetectorYu Zheng, Yueqi Duan, Jiwen Lu, Jie Zhou et al.CVPR 2022 · 33 citations
- MVSDet: Multi-View Indoor 3D Object Detection via Efficient Plane SweepsYating Xu, Chen Li, Gim Hee LeeNeurIPS 2024 · 10 citations
- SPGroup3D: Superpoint Grouping Network for Indoor 3D Object DetectionYun Zhu, Le Hui, Yaqi Shen, Jin XieAAAI 2024 · 24 citations
