Auto-Controlled Image Perception in MLLMs via Visual Perception Tokens
Runpeng Yu, Xinyin Ma, Xinchao Wang
2025Year
1Citations
1Top-tier citations
Abstract
The region related to the query is .
There are two notebooks beneath the opened notebook, one is yellow, the other is green.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 392383d7-d6f6-4bc4-ab82-d08b246602feCited by top-tier papers1
Ask how each one uses itBuilds on19
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong et al.NeurIPS 2024 · 858 citations
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
Related papers
- GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial TasksMuhammad Sohail Danish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah, Kartik Kuckreja et al.ICCV 2025 · 11 citations
- 3D-Mem: 3D Scene Memory for Embodied Exploration and ReasoningYuncong Yang, Han Yang, Jiachen Zhou, Peihao Chen et al.CVPR 2025
- B2: Bridging Code and Interactive Visualization in Computational NotebooksYifan Wu, Joseph M. Hellerstein, Arvind SatyanarayanUIST 2020 · 71 citations
- Action-Sketcher: From Reasoning to Action via Visual Sketches for Robotic ManipulationHuajie Tan, Peterson Co, Yijie Xu, Shanyu Rong et al.CVPR 2026
- PanoGS: Gaussian-based Panoptic Segmentation for 3D Open Vocabulary Scene UnderstandingHongjia Zhai, Hai Li, Zhenzhe Li, Xiaokun Pan et al.CVPR 2025
