Lune

ICML2026Top-tier venue

OnePO: Direct One-stage Policy Optimization for SFT-free Domain Adaptation

Junying Chen, Xinyuan Xie, Ziniu Li, Benyou Wang

2026Year

Abstract

Domain adaptation typically follows a two-stage pipeline: Supervised Fine-Tuning (SFT) then Reinforcement Learning (RL). However, does RL necessarily require a pre-SFT phase for domain adaptation? SFT confines the model to an imitation distribution, limiting RL exploration, while the two-stage transition causes capability regression and extra engineering. We propose One-stage Policy Optimization (OnePO) , an SFT-free paradigm that adapts pretrained LLMs to target domains in a single RL stage. OnePO uses teacher outputs as transient guidance to overcome the slow convergence of pure RL, while avoiding two failures of naive teacher-output integration: gradient starvation for low-probability teacher tokens and distribution anchoring from persistent teacher signals. It introduces two mechanisms: (1) Adaptive Objective Evolution , reshaping the RL objective for rapid absorption of teacher-provided knowledge; and (2) Teacher Retirement , automatically discarding teacher outputs once the model surpasses them. On medical adaptation, OnePO achieves 67.2 on HealthBench with only 20K training samples, outperforming SFT+RL by +2.7 and pure RL by +7.4 points. Scaling the same recipe produces HuatuoGPT-3 , an open-source medical LLM series whose 32B variant reaches 70.3 on HealthBench. Additional writing and legal-domain experiments show that OnePO extends beyond medicine. Models and code are available at https://github.com/FreedomIntelligence/HuatuoGPT-3.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f6cfa593-c3d8-453f-8cb1-11c0befd9331

Builds on16

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines