Lune

ICML2026Top-tier venue

Learning Rewrite-Invariant Reasoning with Targeted Alternation Training

Mousa Arraf, Ido Guy, Kira Radinsky

2026Year

Abstract

Large language models (LLMs) often fail in systematic, model-specific ways under meaning-preserving question rewrites (paraphrases, format changes, benign distractors). In this work, we address this instability by identifying where the model's reasoning process diverges across semantically-equivalent inputs. For each target LLM, we sample multiple solution traces under rewrites and aggregate them into a graph of recurring intermediate steps, which pinpoints where incorrect traces diverge from correct ones. We then generate a small set of semantics-preserving examples that mirror the rewrite patterns most responsible for these divergences, and use them to steer the model (targeted alternation training), either via fine-tuning or via in-context learning. Across MMLU-Pro, Big-MATH, and DROP, this yields consistent gains and cross-dataset generalization. On Humanity’s Last Exam, using 200 in-context examples, it improves GPT-5.2 (xhigh) from 35.4% to 38.1%, demonstrating that targeted alternation training can materially improve a frontier, API-accessible closed model under realistic access constraints.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d92083db-6dff-4e6a-a75b-5805397eeb13

Builds on8

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines