Lune

ICML2026顶会

Learning Rewrite-Invariant Reasoning with Targeted Alternation Training

Mousa Arraf, Ido Guy, Kira Radinsky

出版方
2026年份

摘要

Large language models (LLMs) often fail in systematic, model-specific ways under meaning-preserving question rewrites (paraphrases, format changes, benign distractors). In this work, we address this instability by identifying where the model's reasoning process diverges across semantically-equivalent inputs. For each target LLM, we sample multiple solution traces under rewrites and aggregate them into a graph of recurring intermediate steps, which pinpoints where incorrect traces diverge from correct ones. We then generate a small set of semantics-preserving examples that mirror the rewrite patterns most responsible for these divergences, and use them to steer the model (targeted alternation training), either via fine-tuning or via in-context learning. Across MMLU-Pro, Big-MATH, and DROP, this yields consistent gains and cross-dataset generalization. On Humanity’s Last Exam, using 200 in-context examples, it improves GPT-5.2 (xhigh) from 35.4% to 38.1%, demonstrating that targeted alternation training can materially improve a frontier, API-accessible closed model under realistic access constraints.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext d92083db-6dff-4e6a-a75b-5805397eeb13

它引用的顶会 Paper8

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖