Approximated Orthogonal Projection Unit: Stabilizing Regression Network Training Using Natural Gradient
Shaoqi Wang, Chunjie Yang, Siwei Lou
摘要
Neural networks (NN) are extensively studied in cutting-edge soft sensor models due to their feature extraction and function approximation capabilities. Current research into network-based methods primarily focuses on models' offline accuracy. Notably, in industrial soft sensor context, online optimizing stability and interpretability are prioritized, followed by accuracy. This requires a clearer understanding of network's training process. To bridge this gap, we propose a novel NN named the Approximated Orthogonal Projection Unit (AOPU) which has solid mathematical basis and presents superior training stability. AOPU truncates the gradient backpropagation at dual parameters, optimizes the trackable parameters updates, and enhances the robustness of training. We further prove that AOPU attains minimum variance estimation (MVE) in NN, wherein the truncated gradient approximates the natural gradient (NG). Empirical results on two chemical process datasets clearly show that AOPU outperforms other models in achieving stable convergence, marking a significant advancement in soft sensor field.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 被引用 3,619 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Non-stationary Transformers: Exploring the Stationarity in Time Series ForecastingYong Liu, Haixu Wu, Jianmin Wang, Mingsheng LongNeurIPS 2022 · 被引用 1,080 次
- Synthesizer: Rethinking Self-Attention for Transformer ModelsYi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan 等ICML 2021 · 被引用 399 次
相关 Paper
- projUNN: efficient method for training deep networks with unitary matricesBobak Toussi Kiani, Randall Balestriero, Yann LeCun, Seth LloydNeurIPS 2022 · 被引用 42 次
- Oscillatory Fourier Neural Network: A Compact and Efficient Architecture for Sequential ProcessingBing Han, Cheng Wang, Kaushik RoyAAAI 2022 · 被引用 7 次
- Tensor Product Neural Networks for Functional ANOVA ModelSeokhun Park, Insung Kong, Yongchan Choi, Chanmoo Park 等ICML 2025
- FESSNC: Fast Exponentially Stable and Safe Neural ControllerJingdong Zhang, Luan Yang, Qunxi Zhu, Wei LinICML 2024 · 被引用 2 次
- Almost Surely Stable Deep DynamicsNathan P. Lawrence, Philip D. Loewen, Michael G. Forbes, Johan U. Backström 等NeurIPS 2020 · 被引用 28 次
