π2vec: Policy Representation with Successor Features
Gianluca Scarpellini, Ksenia Konyushkova, Claudio Fantacci, Thomas Paine, Yutian Chen, Misha Denil
摘要
This paper introduces π2vec, a method for representing black box policies as comparable feature vectors. Our method combines the strengths of foundation models that serve as generic and powerful state representations and successor features that can model the future occurrence of the states for a policy. π2vec represents the behaviors of policies by capturing statistics of how the behavior evolves the features from a pretrained model, using a successor feature framework. We focus on the offline setting where both policies and their representations are trained on a fixed dataset of trajectories. Finally, we employ linear regression on π2vec vector representations to predict the performance of held out policies. The synergy of these techniques results in a method for efficient policy evaluation in resource constrained environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
- Data-Efficient Reinforcement Learning with Self-Predictive RepresentationsMax Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm 等ICLR 2021 · 被引用 399 次
- Benchmarks for Deep Off-Policy EvaluationJustin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker 等ICLR 2021 · 被引用 112 次
相关 Paper
- Successor Feature Sets: Generalizing Successor Representations Across PoliciesKianté Brantley, Soroush Mehri, Geoffrey J. GordonAAAI 2021 · 被引用 11 次
- Bridging Successor Measure and Online Policy Learning with Flow Matching-Based RepresentationsHaosen Shi, Jianda Chen, Sinno Jialin PanICLR 2026
- Instabilities of Offline RL with Pre-Trained Neural RepresentationRuosong Wang, Yifan Wu, Ruslan Salakhutdinov, Sham M. KakadeICML 2021 · 被引用 46 次
- Composing Task Knowledge With Modular Successor Feature ApproximatorsWilka Carvalho, Angelos Filos, Richard L. Lewis, Honglak Lee 等ICLR 2023 · 被引用 2 次
- Distributional Successor Features Enable Zero-Shot Policy OptimizationChuning Zhu, Xinqi Wang, Tyler Han, Simon S. Du 等NeurIPS 2024 · 被引用 11 次
