The Geometry of Sequential Learning: Lie-Bracket Prediction of Transfer Order
John Sweeney
摘要
Sequential learning is order-dependent: from Pile-style next-token domain adaptation to instruction-SFT and DPO, candidate sources induce possible curricula. We show that the local order effect is governed by a computable geometric quantity, the Lie-bracket commutator of gradient update fields, yielding a pairwise score for whether or is better for a target domain. The pairwise bracket primitive also defines a Lie-Bracket Tournament: with a shared target-gradient reference, Hessian symmetry gives Borda/row-sum scores from one Hessian-vector product per source, dot products, and an sort, without materializing the edge matrix. Empirically, the planner reaches 98.1%/98.9% pairwise accuracy at for instruction-SFT/DPO, remains at 73.1%/72.2% at , and preserves the original pretraining-domain evidence with 82.4–92.0% accuracy across four LLMs and 91.1% on diffusion. At curriculum scale, it recovers the best of all schedules in 87.5% of trials, ranks 85 Stack programming-language source domains for a Python target in the 99th sampled percentile, and reaches the 99.0–99.6th sampled percentile on 56 MMLU subjects, sharply above the reported descending gradient-norm baseline. These results reframe sequential learning as a geometric tournament problem: commutators provide both local pairwise order information and a scalable primitive for many-domain schedules.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
相关 Paper
- Learning Permutation Distributions via Reflected Diffusion on RanksSizhuang He, Yangtian Zhang, Shiyang Zhang, David van DijkICML 2026
- Principled Zero-shot Ranking Agents with Tournament GraphsSheshansh Agrawal, Thien Nguyen, Douwe KielaICML 2026
- The Differences Between Direct Alignment Algorithms are a BlurAlexey Gorbatovski, Boris Shaposhnikov, Viacheslav Sinii, Alexey Malakhov 等ICML 2026
- Scheduling Thoughts: Learning the Order of Thought in Diffusion Language ModelsJiawei Xu, Minghui Liu, Aakriti Agrawal, Yifan Chen 等ICML 2026 · 被引用 1 次
- Reinforcing the Diffusion Chain of Lateral Thought with Diffusion Language ModelsZemin Huang, Zhiyang Chen, Zijun Wang, Tiancheng Li 等NeurIPS 2025 · 被引用 55 次
