NeuroRule: Bridging Vision and Logic with Differentiable Rule Induction
Muhammad Zarar, Mingzheng Zhang, Xiaowang Zhang, Zhiyong Feng
Abstract
Scene Graph Generation (SGG) aims to structurally represent visual scenes by detecting objects and their pairwise relationships. Despite significant progress, current models encode visual knowledge with ambiguous visual context and logically inferred implicit relations due to their purely neural, pipeline-based nature. This limitation underscores the need to advance beyond identifying what relations exist to explaining why they exist and how they can be compositionally reasoned about through logical rule chaining. To address these challenges, we introduce NeuroRule, the first Neurally-Guided Rule Induction Network that integrates Mask2Former pixel-precise visual understanding with a differentiable rule induction engine. Our proposed method enables automatic learning of compositional logical rules directly from visual data while providing transparent explanations for relational predictions. NeuroRule introduces three key innovations: (1) a neural-symbolic bridge that maps visual features to probabilistic symbolic representations; (2) a differentiable rule-learning mechanism that automatically discovers interpretable first-order logic rules without manual engineering; and (3) a compositional chain rule system that enables complex inference while propagating confidence scores through an end-to-end trainable pipeline. Extensive experiments on the benchmark datasets, including Visual Genome (VG), Panoptic Scene Graph (PSG), and OpenPSG, demonstrate that NeuroRule achieves state-of-the-art performance. Our method significantly improves few-shot relation extraction while maintaining full interpretability in its rule-based explanations. To ensure reproducibility, we will release the code after publication.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4941bf9-ca47-4d61-8def-98c26f7bd501Builds on22
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- RNNLogic: Learning Logic Rules for Reasoning on Knowledge GraphsMeng Qu, Jun-Kun Chen, Louis-Pascal A. C. Xhonneux, Yoshua Bengio et al.ICLR 2021 · 230 citations
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 108 citations
- Iterative Scene Graph GenerationSiddhesh Khandelwal, Leonid SigalNeurIPS 2022 · 47 citations
- Compositional Feature Augmentation for Unbiased Scene Graph GenerationLin Li, Guikun Chen, Jun Xiao, Yi Yang et al.ICCV 2023 · 36 citations
Related papers
- Mask and Predict: Multi-step Reasoning for Scene Graph GenerationHongshuo Tian, Ning Xu, An-An Liu, Chenggang Yan et al.ACM MM 2021 · 12 citations
- Learn to Explain Efficiently via Neural Logic Inductive LearningYuan Yang, Le SongICLR 2020 · 83 citations
- LogicSeg: Parsing Visual Semantics with Neural Logic Learning and ReasoningLiulei Li, Wenguan Wang, Yang YiICCV 2023 · 52 citations
- Generating by Understanding: Neural Visual Generation with Logical Symbol GroundingsYifei Peng, Zijie Zha, Yu Jin, Zhexu Luo et al.KDD 2025
- UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph GenerationXinyao Liao, Wei Wei, Dangyang Chen, Yuanyuan FuACM MM 2024 · 2 citations
