Scaling-Aware Adapter for Structure-Grounded LLM Reasoning
Zihao Jing, QIUHAO Zeng, Ruiyi Fang, Yan Y Li, Yan Sun, Boyu Wang, Pingzhao Hu
Abstract
Large language models (LLMs) are enabling reasoning over 2D and 3D structures, yet existing methods remain modality-specific and typically compress structural inputs through sequence-based tokenization or fixed-length query connectors. Such architectures either omit the geometric grounding requisite for mitigating structural hallucinations, or impose inflexible modality fusion bottlenecks that concurrently over-compress and suboptimally allocate structural tokens, thereby impeding the realization of generalized all-atom reasoning. We introduce Cuttlefish , a unified multimodal LLM that grounds language reasoning in geometric cues while scaling modality tokens with structural complexity. First, Scaling-Aware Patching leverages an instruction-conditioned gating mechanism to generate variable-size patches over structural graphs, adaptively scaling the query token budget with structural complexity to mitigate fixed-length connector bottlenecks. Second, Geometry Grounding Adapter refines these adaptive tokens via cross-attention to modality embeddings and injects the resulting modality tokens into the LLM, exposing explicit geometric cues to reduce structural hallucination. Experiments across interdisciplinary all-atom benchmarks demonstrate that Cuttlefish achieves superior performance in heterogeneous structure-grounded reasoning. Code: github.com/zihao-jing/Cuttlefish.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68424b26-26a0-485f-a101-3e5788099eb6Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui et al.EMNLP 2024 · 231 citations
- Unifying Molecular and Textual Representations via Multi-task Language ModellingDimitrios Christofidellis, Giorgio Giannone, Jannis Born, Ole Winther et al.ICML 2023 · 126 citations
Related papers
- Improving Large Molecular Language Model via Relation-aware Multimodal CollaborationJinyoung Park, Minseong Bae, Jeehye Na, Hyunwoo J. KimAAAI 2026
- Tokenization Allows Multimodal Large Language Models to Understand, Generate and Edit Architectural Floor PlansSizhong Qin, Ramon Elias Weber, Xinzheng LuCVPR 2026 · 1 citation
- Grounding Everything in Tokens for Multimodal Large Language ModelsXiangxuan Ren, Zhongdao Wang, Liping Hou, Pin Tang et al.CVPR 2026 · 2 citations
- Groundhog Grounding Large Language Models to Holistic SegmentationYichi Zhang, Ziqiao Ma, Xiaofeng Gao, Suhaila Shakiah et al.CVPR 2024 · 24 citations
- Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning BenchmarkYunzhuo Hao, Jiawei Gu, Huichen Will Wang, Linjie Li et al.ICML 2025
