Inducing, Detecting and Characterising Neural Modules: A Pipeline for Functional Interpretability in Reinforcement Learning
Anna Soligo, Pietro Ferraro, David Boyle
摘要
Interpretability is crucial for ensuring RL systems align with human values. However, it remains challenging to achieve in complex decision making domains. Existing methods frequently attempt interpretability at the level of fundamental model units, such as neurons or decision nodes: an approach which scales poorly to large models. Here, we instead propose an approach to interpretability at the level of functional modularity. We show how encouraging sparsity and locality in network weights leads to the emergence of functional modules in RL policy networks. To detect these modules, we develop an extended Louvain algorithm which uses a novel 'correlation alignment' metric to overcome the limitations of standard network analysis techniques when applied to neural network architectures. Applying these methods to 2D and 3D MiniGrid environments reveals the consistent emergence of distinct navigational modules for different axes, and we further demonstrate how these functions can be validated through direct interventions on network weights prior to inference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- Discovering symbolic policies with deep reinforcement learningMikel Landajuela, Brenden K. Petersen, Sookyung Kim, Cláudio P. Santiago 等ICML 2021 · 被引用 118 次
- Interpretable and Explainable Logical Policies via Neurally Guided Symbolic AbstractionQuentin Delfosse, Hikaru Shindo, Devendra Singh Dhami, Kristian KerstingNeurIPS 2023 · 被引用 64 次
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallKevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris 等ICLR 2023 · 被引用 50 次
- Interpretable Concept Bottlenecks to Align Reinforcement Learning AgentsQuentin Delfosse, Sebastian Sztwiertnia, Mark Rothermel, Wolfgang Stammer 等NeurIPS 2024 · 被引用 32 次
相关 Paper
- Are Neural Nets Modular? Inspecting Functional Modularity Through Differentiable Weight MasksRóbert Csordás, Sjoerd van Steenkiste, Jürgen SchmidhuberICLR 2021 · 被引用 13 次
- Reinforcement Learning Finetunes Small Subnetworks in Large Language ModelsSagnik Mukherjee, Lifan Yuan, Dilek Hakkani-Tur, Hao PengNeurIPS 2025 · 被引用 43 次
- Weight-sparse transformers have interpretable circuitsLeo Gao, Achyuta Rajaram, Jacob Coxon, Soham Govande 等ICML 2026
- Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action ModelsHanxin Zhang, Mingshuo Xu, Abdulqader Dhafer, Shigang Yue 等ICML 2026 · 被引用 2 次
- Beyond Rewards in RL for Cyber DefenceElizabeth Bates, Chris Hicks, Vasilios MavroudisICML 2026 · 被引用 3 次
