CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning
Ji Qi, Ming Ding, Weihan Wang, Yushi Bai, Qingsong Lv, Wenyi Hong, Bin Xu, Lei Hou, Juanzi Li, Yuxiao Dong, Jie Tang
2025Year
12Top-tier citations
Abstract
In circle O with radius of 5.0, the length of chord AB is 8.0, then the distance from the O to AB is? CogCoM: Draw a straight line with LINE([[(54, 43 ), (13,70)]]) to obtain a new image after drawing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual DrawingJunfei Wu, Jian Guan, Kaituo Feng, Qiang Liu et al.NeurIPS 2025 · 153 citations
- Thinking in 360deg: Humanoid Visual Search in the WildHeyang Yu, Yinan Han, Xiangyu Zhang, Baiqiao Yin et al.CVPR 2026 · 16 citations
- CoFFT: Chain of Foresight-Focus Thought for Visual Language ModelsXinyu Zhang, Yuxuan Dong, Lingling Zhang, Chengyou Jia et al.NeurIPS 2025 · 7 citations
- Vision-aligned Latent Reasoning for Multi-modal Large Language ModelByungwoo Jeon, Yoonwoo Jeong, Hyunseok Lee, Minsu Cho et al.ICML 2026 · 7 citations
- Pixel Reasoner: Incentivizing Pixel Space Reasoning via Curiosity-Driven Reinforcement LearningAlex Su, Haozhe Wang, Weiming Ren, Fangzhen Lin et al.NeurIPS 2025 · 6 citations
Builds on21
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- GeoUni: A Unified Model for Generating Geometry Diagrams, Problems and Problem SolutionsJo-Ku Cheng, Zeren Zhang, Ran Chen, Jingyang Deng et al.ACM MM 2025 · 2 citations
- Do Large Language Models Truly Understand Geometric Structures?Xiaofeng Wang, Yiming Wang, Wenhong Zhu, Rui WangICLR 2025
- Static and Streaming Data Structures for Fréchet Distance QueriesArnold Filtser, Omrit FiltserSODA 2021 · 2 citations
- EAR-Oracle: On Efficient Indexing for Distance Queries between Arbitrary Points on Terrain SurfaceBo Huang, Victor Junqiu Wei, Raymond Chi-Wing Wong, Bo TangSIGMOD 2023 · 4 citations
- LCOT: Linear Circular Optimal TransportRocio Diaz Martin, Ivan Vladimir Medri, Yikun Bai, Xinran Liu et al.ICLR 2024 · 1 citation
