An Information-Geometric Distance on the Space of Tasks
Yansong Gao, Pratik Chaudhari
摘要
This paper prescribes a distance between learning tasks modeled as joint distributions on data and labels. Using tools in information geometry, the distance is defined to be the length of the shortest weight trajectory on a Riemannian manifold as a classifier is fitted on an interpolated task. The interpolated task evolves from the source to the target task using an optimal transport formulation. This distance, which we call the "coupled transfer distance" can be compared across different classifier architectures. We develop an algorithm to compute the distance which iteratively transports the marginal on the data of the source task to that of the target task while updating the weights of the classifier to track this evolving data distribution. We develop theory to show that our distance captures the intuitive idea that a good transfer trajectory is the one that keeps the generalization gap small during transfer, in particular at the end on the target task. We perform thorough empirical validation and analysis across diverse image classification datasets to show that the coupled transfer distance correlates strongly with the difficulty of fine-tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- A Picture of the Space of Typical Learnable TasksRahul Ramesh, Jialin Mao, Itay Griniasty, Rubing Yang 等ICML 2023 · 被引用 7 次
- Time-Varying Propensity Score to Bridge the Gap between the Past and PresentRasool Fakoor, Jonas Mueller, Zachary Chase Lipton, Pratik Chaudhari 等ICLR 2024 · 被引用 4 次
- Interpolation for Robust Learning: Data Augmentation on Wasserstein GeodesicsJiacheng Zhu, Jielin Qiu, Aritra Guha, Zhuolin Yang 等ICML 2023 · 被引用 4 次
- The Geometry of Updates: Fisher Alignment at Vocabulary ScaleJohn SweeneyICML 2026 · 被引用 1 次
- Lightspeed Geometric Dataset Distance via Sliced Optimal TransportKhai Nguyen, Hai Nguyen, Tuan Pham, Nhat HoICML 2025
它引用的顶会 Paper5
- A Baseline for Few-Shot Image ClassificationGuneet Singh Dhillon, Pratik Chaudhari, Avinash Ravichandran, Stefano SoattoICLR 2020 · 被引用 640 次
- Task2Vec: Task Embedding for Meta-LearningAlessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran 等ICCV 2019 · 被引用 359 次
- Geometric Dataset Distances via Optimal TransportDavid Alvarez-Melis, Nicolò FusiNeurIPS 2020 · 被引用 267 次
- Rethinking the Hyperparameters for Fine-tuningHao Li, Pratik Chaudhari, Hao Yang, Michael Lam 等ICLR 2020 · 被引用 142 次
- A Free-Energy Principle for Representation LearningYansong Gao, Pratik ChaudhariICML 2020 · 被引用 11 次
相关 Paper
- Curriculum Reinforcement Learning using Optimal Transport via Gradual Domain AdaptationPeide Huang, Mengdi Xu, Jiacheng Zhu, Laixi Shi 等NeurIPS 2022 · 被引用 44 次
- Geometrically Aligned Transfer Encoder for Inductive Transfer in Regression TasksSung Moon Ko, Sumin Lee, Dae-Woong Jeong, Woohyung Lim 等ICLR 2024 · 被引用 6 次
- Learning Structured Representations by Embedding Class Hierarchy with Fast Optimal TransportSiqi Zeng, Sixian Du, Makoto Yamada, Han ZhaoICLR 2025
- Flowing Datasets with Wasserstein over Wasserstein Gradient FlowsClément Bonet, Christophe Vauthier, Anna KorbaICML 2025
- Riemannian Metric Learning via Optimal TransportChristopher Scarvelis, Justin SolomonICLR 2023 · 被引用 2 次
