Lune

AAAI2026顶会

MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement Learning

Zhiheng Xi, Yuhui Wang, Yiwen Ding, Guanyu Li, Senjie Jin, Shichun Liu, Jixuan Huang, Dingwen Yang, Jiafu Tang, Boyang Hong, Junjie Ye, Shihan Dou

2026年份

摘要

Outcome-based reinforcement learning has made notable advances in training language models (LMs) for reasoning. However, without explicit incentives and controls, this paradigm has limitations and instability in eliciting high-quality reasoning trajectories with diverse actions-particularly for models whose pretraining lacked extensive reasoning-related data. To this end, we introduce MetaAct-RL, a new RL framework that frames LMs' thinking as sequential decision making over meta-actions. In this framework, the model chooses and executes a high-level action at each step-such as forward reasoning, critique, or refinement-to gradually reach the correct answer. To encourage deeper exploration, richer action diversity, and to improve sampling efficiency in the RL optimization process, MetaAct-RL incorporates appropriate lengthbased reward and regularization, and a key-state restart mechanism. Extensive experiments across six benchmarks show that MetaAct-RL improves reasoning performance by 7.99 on Llama3.2-1B and 7.17 on Llama3.1-8B relative to vanilla RL method. Moreover, on the challenging AIME-2024, our method outperforms the vanilla RL by 7.5 with Qwen2.5-1.5B.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 3c37bbb4-a9bd-46a5-8133-37c6f9cf4d4e

它引用的顶会 Paper22

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖