The Geometry of Sequential Learning: Lie-Bracket Prediction of Transfer Order
John Sweeney
Abstract
Sequential learning is order-dependent: from Pile-style next-token domain adaptation to instruction-SFT and DPO, candidate sources induce possible curricula. We show that the local order effect is governed by a computable geometric quantity, the Lie-bracket commutator of gradient update fields, yielding a pairwise score for whether or is better for a target domain. The pairwise bracket primitive also defines a Lie-Bracket Tournament: with a shared target-gradient reference, Hessian symmetry gives Borda/row-sum scores from one Hessian-vector product per source, dot products, and an sort, without materializing the edge matrix. Empirically, the planner reaches 98.1%/98.9% pairwise accuracy at for instruction-SFT/DPO, remains at 73.1%/72.2% at , and preserves the original pretraining-domain evidence with 82.4–92.0% accuracy across four LLMs and 91.1% on diffusion. At curriculum scale, it recovers the best of all schedules in 87.5% of trials, ranks 85 Stack programming-language source domains for a Python target in the 99th sampled percentile, and reaches the 99.0–99.6th sampled percentile on 56 MMLU subjects, sharply above the reported descending gradient-norm baseline. These results reframe sequential learning as a geometric tournament problem: commutators provide both local pairwise order information and a scalable primitive for many-domain schedules.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1fa1bcd-9103-48ce-a244-15569f348c83Cited by top-tier papers1
Ask how each one uses itBuilds on12
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
Related papers
- Learning Permutation Distributions via Reflected Diffusion on RanksSizhuang He, Yangtian Zhang, Shiyang Zhang, David van DijkICML 2026
- Principled Zero-shot Ranking Agents with Tournament GraphsSheshansh Agrawal, Thien Nguyen, Douwe KielaICML 2026
- The Differences Between Direct Alignment Algorithms are a BlurAlexey Gorbatovski, Boris Shaposhnikov, Viacheslav Sinii, Alexey Malakhov et al.ICML 2026
- Scheduling Thoughts: Learning the Order of Thought in Diffusion Language ModelsJiawei Xu, Minghui Liu, Aakriti Agrawal, Yifan Chen et al.ICML 2026 · 1 citation
- Reinforcing the Diffusion Chain of Lateral Thought with Diffusion Language ModelsZemin Huang, Zhiyang Chen, Zijun Wang, Tiancheng Li et al.NeurIPS 2025 · 55 citations
