Lune

ICML2026顶会

Action Manifold Smoothing: A Lipschitz Pathway Perspective on High-Dimensional Reinforcement Learning

Zhihao Lin

出版方
2026年份

摘要

High-dimensional continuous control remains challenging in deep reinforcement learning, where algorithms like TD3 and SAC often collapse. We propose a unifying Lipschitz Pathway framework that decomposes instability into four amplification stages, namely action parameterization (L1L_1), dynamics sensitivity (L2L_2), Q-network curvature (L3L_3), and temporal-difference (TD) target stability (L4L_4), where errors compound multiplicatively along the learning pipeline. Our analysis identifies a discrete-continuous mismatch as the root cause: value functions trained from sparse point samples must generalize over continuous manifolds, leading to multiplicative error amplification along the pathway. To address this, we introduce Action Manifold Smoothing (AMS), which replaces point-wise TD targets with orthogonally-sampled neighborhood averages, jointly regularizing L3L_3 (via implicit Laplacian smoothing) and L4L_4 (via local manifold supervision). We further characterize when Lipschitz-constrained Q-networks and geometric action priors are beneficial based on task structure. Empirically, AMS enables both TD3 and SAC to achieve over 400 reward on the 38-D Dog Run task within 1M steps, where baselines fail. These results validate the Lipschitz pathway as a principled framework for diagnosing and solving stability bottlenecks in high-dimensional control.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper9

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖