Symbol-LLM: Leverage Language Models for Symbolic System in Visual Human Activity Reasoning
Xiaoqian Wu, Yonglu Li, Jianhua Sun, Cewu Lu
摘要
Human reasoning can be understood as a cooperation between the intuitive, associative "System-1" and the deliberative, logical "System-2". For existing System-1-like methods in visual activity understanding, it is crucial to integrate System-2 processing to improve explainability, generalization, and data efficiency. One possible path of activity reasoning is building a symbolic system composed of symbols and rules, where one rule connects multiple symbols, implying human knowledge and reasoning abilities. Previous methods have made progress, but are defective with limited symbols from handcraft and limited rules from visual-based annotations, failing to cover the complex patterns of activities and lacking compositional generalization. To overcome the defects, we propose a new symbolic system with two ideal important properties: broad-coverage symbols and rational rules. Collecting massive human knowledge via manual annotations is expensive to instantiate this symbolic system. Instead, we leverage the recent advancement of LLMs (Large Language Models) as an approximation of the two ideal properties, i.e., Symbols from Large Language Models (Symbol-LLM). Then, given an image, visual contents from the images are extracted and checked as symbols and activity semantics are reasoned out based on rules via fuzzy logic calculation. Our method shows superiority in extensive activity understanding tasks. Code and data are available at https://mvig-rhos.com/symbol_llm .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning AbilitiesJiayi Kuang, Haojing Huang, Yinghui Li, Xinnian Liang 等NeurIPS 2025 · 被引用 11 次
- MuSLR: Multimodal Symbolic Logical ReasoningJundong Xu, Hao Fei, Yuhui Zhang, Liangming Pan 等NeurIPS 2025 · 被引用 5 次
- InstructHOI: Context-Aware Instruction for Multi-Modal Reasoning in Human-Object Interaction DetectionJinguo Luo, Weihong Ren, Quanlong Zheng, Yanhao Zhang 等NeurIPS 2025 · 被引用 3 次
- Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram GenerationZhiqing Cui, Jiahao Yuan, Hanqing Wang, Yanshu Li 等ACM MM 2025 · 被引用 3 次
- Empowering GraphRAG with Knowledge Filtering and IntegrationKai Guo, Harry Shomer, Shenglai Zeng, Haoyu Han 等EMNLP 2025 · 被引用 2 次
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
相关 Paper
- Synthesizing Visual Concepts as Vision-Language ProgramsAntonia Wüst, Wolfgang Stammer, Hikaru Shindo, Lukas Helff 等CVPR 2026 · 被引用 6 次
- NeSyCoCo: A Neuro-Symbolic Concept Composer for Compositional GeneralizationDanial Kamali, Elham J. Barezi, Parisa KordjamshidiAAAI 2025 · 被引用 4 次
- Neuro-Fuzzy Concept Learning for Interpretable Large Multimodal ModelsRitik Mishra, Vanshika Gupta, M. Sajid, M. TanveerICML 2026
- Improving Coherence and Consistency in Neural Sequence Models with Dual-System, Neuro-Symbolic ReasoningMaxwell I. Nye, Michael Henry Tessler, Joshua B. Tenenbaum, Brenden M. LakeNeurIPS 2021 · 被引用 151 次
- GENOME: Generative Neuro-Symbolic Visual Reasoning by Growing and Reusing ModulesZhenfang Chen, Rui Sun, Wenjun Liu, Yining Hong 等ICLR 2024 · 被引用 24 次
