Photon: Speedup Volume Understanding with Efficient Multimodal Large Language Models
Chengyu Fang, Heng Guo, Zheng Jiang, Chunming He, Xiu Li, Minfeng Xu
Abstract
Multimodal large language models are promising for clinical visual question answering tasks, but scaling to 3D imaging is hindered by high computational costs. Prior methods often rely on 2D slices or fixed-length token compression, disrupting volumetric continuity and obscuring subtle findings. We present Photon, a framework that represents 3D medical volumes with token sequences of variable length. Photon introduces instruction-conditioned token scheduling and surrogate gradient propagation to adaptively reduce tokens during both training and inference, which lowers computational cost while mitigating the attention dilution caused by redundant tokens. It incorporates a custom backpropagation rule with gradient restoration to enable differentiable optimization despite discrete token drop. To stabilize token compression and ensure reliable use of visual evidence, Photon further applies regularization objectives that mitigate language-only bias and improve reliability. Experiments on diverse medical visual question answering tasks show that Photon achieves state-of-the-art accuracy while reducing resource usage and accelerating both training and inference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7b135d3-df0b-45f9-8d47-61a5414c7a60Cited by top-tier papers7
- Stochastic Self-Guidance for Training-Free Enhancement of Diffusion ModelsChubin Chen, Jiashu Zhu, Xiaokun Feng, Nisha Huang et al.ICLR 2026 · 44 citations
- MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement LearningZheng Jiang, Heng Guo, Chengyu Fang, Changchen Xiao et al.ICLR 2026 · 8 citations
- PreciseCache: Precise Feature Caching for Efficient and High-fidelity Video GenerationJiangshan Wang, Kang Zhao, Jiayi Guo, Jiayu Wang et al.ICLR 2026 · 6 citations
- UniFusion: A Unified Image Fusion Framework with Robust Representation and Source-Aware PreservationXingyuan Li, Songcheng Du, Yang Zou, Haoyuan Xu et al.CVPR 2026 · 6 citations
- Unpaired Image Deraining Using Reward-Guided Self-Reinforcement StrategyYinghao Chen, Yeying Jin, Xiang Chen, Yanyan Wei et al.CVPR 2026 · 2 citations
Builds on26
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMsQizhe Zhang, Mengzhen Liu, Lichen Li, Ming Lu et al.NeurIPS 2025 · 104 citations
- Real-world Image Dehazing with Coherence-based Pseudo Labeling and Cooperative Unfolding NetworkChengyu Fang, Chunming He, Fengyang Xiao, Yulun Zhang et al.NeurIPS 2024 · 46 citations
- Towards Injecting Medical Visual Knowledge into Multimodal LLMs at ScaleJunying Chen, Chi Gui, Ruyi Ouyang, Anningzhe Gao et al.EMNLP 2024 · 43 citations
Related papers
- Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models Via Adaptive Token SkippingWeili Zeng, Ziyuan Huang, Kaixiang Ji, Yichao YanICCV 2025
- Zero-shot 3D Question Answering via Voxel-based Dynamic Token CompressionHsiang-Wei Huang, Fu-Chen Chen, Wenhao Chai, Che-Chun Su et al.CVPR 2025
- Granularity-Adaptive Spatial Evidence Tokenization for Video Question AnsweringHao Jiang, Yang Jin, Zhicheng Sun, Kun Xu et al.AAAI 2025 · 2 citations
- Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement LearningKaitao Chen, Weiqian Zhao, Jiamin Wu, Qihao Zheng et al.ICML 2026
- PHOTON: Hierarchical Autoregressive Modeling for Lightspeed and Memory-Efficient Language GenerationYuma Ichikawa, Naoya Takagi, Takumi Nakagawa, Yuzi Kanazawa et al.ACL 2026
