Perception Encoder: The best visual embeddings are not at the output of the network
Daniel Bolya, Po-Yao Huang, Peize Sun, Jang Hyun Cho, Andrea Madotto, Chen Wei, Tengyu Ma, Jiale Zhi, Jathushan Rajasegaran, Hanoona Rasheed, Junke Wang, Marco Monteiro
Abstract
We introduce Perception Encoder (PE), a family of state-of-the-art vision encoders for image and video understanding. Traditionally, vision encoders have relied on a variety of pretraining objectives, each excelling at different downstream tasks. Surprisingly, after scaling a carefully tuned image pretraining recipe and refining with a robust video data engine, we find that contrastive vision-language training alone can produce strong, general embeddings for all of these downstream tasks. There is only one caveat: these embeddings are hidden within the intermediate layers of the network. To draw them out, we introduce two alignment methods: language alignment for multimodal language modeling, and spatial alignment for dense prediction. Together, our PE family of models achieves state-of-the-art results on a wide variety of tasks, including zero-shot image and video classification and retrieval; document, image, and video Q&A; and spatial tasks such as detection, tracking, and depth estimation. We release our models, code, and novel dataset of synthetically and human-annotated videos: https://github.com/facebookresearch/perception_models Category: Animals Caption: The video shows a white and gray adult cat and two kittens. The adult cat is grooming the kitten closest to it with its tongue, and the kitten is looking around. A hand reaches out from the frame's upper left to pet the two kittens.
Caption: The video shows a tall, pointed structure in the middle of a field. and the structure is surrounded by trees and other vegetation. The field is divided into sections, with some areas covered in green grass and others covered in white material. The video shows the structure and the field from a distance, with the camera moving around it.
Caption: The video shows a person using a shovel to clean the ashes from a fireplace. They are scooping up the ashes and removing them from the fireplace.
The video shows a person cutting an green color item into small pieces. They are using a knife to slice the pickle into thin pieces, and then chopping those pieces into smaller cubes. The person is working on a wooden cutting board, and the Hands are visible from the left side of the frame with pink nail paint on their nails.
The video shows a black and white spiral that is spinning. The spiral is made up of alternating black and white stripes that are evenly spaced and symmetrical.
Caption: The video captures a closeup shot of person typing on a keyboard. The camera moves from the left side of the keyboard to the right, an animation of the revolving globe and some numbers can be seen in the frame and the video ends.
Caption: The video shows a person putting a bowl of something into an oven. The person then closes the oven door. The background is blurry.
Caption: The video shows a large school of fish swimming in a water body towards the right frame. The camera too pans a little to the right.
Caption: The video shows a pile of branches and leaves on fire in a field. The fire is burning brightly, with flames licking at the edges of the pile. The smoke from the fire rises into the air, billowing up into the sky.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3e1bf23-278e-4f31-8812-888aa7a971b8Cited by top-tier papers86
- PerceptionLM: Open-Access Data and Models for Detailed Visual UnderstandingJang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi, Triantafyllos Afouras et al.NeurIPS 2025 · 97 citations
- What matters for Representation Alignment: Global Information or Spatial Structure?Jaskirat Singh, Xingjian Leng, Zongze Wu, Liang Zheng et al.ICLR 2026 · 84 citations
- Meta CLIP 2: A Worldwide Scaling RecipeYung-Sung Chuang, Yang Li, Dong Wang, Ching-Feng Yeh et al.NeurIPS 2025 · 72 citations
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-languageDelong Chen, Mustafa Shukor, Théo Moutakanni, Willy Chung et al.ICLR 2026 · 60 citations
- Aligning Visual Foundation Encoders to Tokenizers for Diffusion ModelsBowei Chen, Sai Bi, Hao Tan, He Zhang et al.ICLR 2026 · 36 citations
Builds on62
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
Related papers
- Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence LearningApoorv Vyas, Heng-Jui Chang, Cheng-Fu Yang, Po-Yao Huang et al.CVPR 2026 · 23 citations
- Adaptively Building a Video-language Model for Video Captioning and Retrieval without Massive Video PretrainingZihao Liu, Xiaoyu Wu, Shengjin Wang, Jiayao QianACM MM 2024 · 1 citation
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal et al.ICLR 2022 · 503 citations
- OmniVL: One Foundation Model for Image-Language and Video-Language TasksJunke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo et al.NeurIPS 2022 · 205 citations
- UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and TextureShuo Cao, Jiayang Li, Xiaohui Li, Yuandong Pu et al.ICML 2026
