LocLLM: Exploiting Generalizable Human Keypoint Localization via Large Language Model
Dongkai Wang, Shiyu Xuan, Shiliang Zhang
摘要
The capacity of existing human keypoint localization models is limited by keypoint priors provided by the training data. To alleviate this restriction and pursue more gen-eral model, this work studies keypoint localization from a different perspective by reasoning locations based on key-piont clues in text descriptions. We propose LocLLM, the first Large-Language Model (LLM) based keypoint local-ization model that takes images and text instructions as in-puts and outputs the desired keypoint coordinates. LocLLM leverages the strong reasoning capability of LLM and clues of keypoint type, location, and relationship in textual de-scriptions for keypoint localization. To effectively tune Lo-cLLM, we construct localization-based instruction conver-sations to connect keypoint description with corresponding coordinates in input image, and fine-tune the whole model in a parameter-efficient training pipeline. LocLLM shows remarkable performance on standard 2D/3D keypoint lo-calization benchmarks. Moreover, incorporating language clues into the localization makes LocLLM show superior flexibility and generalizable capability in cross dataset key-point localization, and even detecting novel type of key-points unseen during training<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">†</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">†</sup>Project page: https://github.com/kennethwdk/LocLLM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- End-to-End Multi-Person Pose Estimation with Pose-Aware Video TransformerYonghui Yu, Jiahang Cai, Xun Wang, Wenwu YangAAAI 2026 · 被引用 2 次
- Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose EstimationHongwei Fang, Jiahang Cai, Xun Wang, Wenwu YangCVPR 2026 · 被引用 1 次
- Superman: Unifying Skeleton and Vision for Human Motion Perception and GenerationXinshun Wang, Peiming Li, Ziyi Wang, Zhongbin Fang 等CVPR 2026
- Dual Coding Theory in Action: Language-Assisted Human Pose Estimation in VideosSifan Wu, Haipeng Chen, Yingda Lyu, Shaojing Fan 等AAAI 2026
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- VLM-Loc: Localization in Point Cloud Maps via Vision-Language ModelsShuhao Kang, Youqi Liao, Peijie Wang, Wenlong Liao 等CVPR 2026 · 被引用 4 次
- ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language ModelsBingchen Gong, Diego Gomez, Abdullah Hamdi, Abdelrahman Eldesokey 等ICCV 2025
- CapeLLM: Support-Free Category-Agnostic Pose Estimation with Multimodal Large Language ModelsJunho Kim, Hyungjin Chung, Byung-Hoon KimICCV 2025 · 被引用 1 次
- Talking Points: Describing and Localizing PixelsMatan Rusanovsky, Shimon Malnick, Shai AvidanICLR 2026
- KptLLM: Unveiling the Power of Large Language Model for Keypoint ComprehensionJie Yang, Wang Zeng, Sheng Jin, Lumin Xu 等NeurIPS 2024 · 被引用 9 次
