Quantized Visual Geometry Grounded Transformer
Weilun Feng, Haotong Qin, Mingqiang Wu, Chuanguang Yang, Yuqi Li, Xiangqi Li, Zhulin An, Libo Huang, Yulun Zhang, Michele Magno, Yongjun Xu
摘要
Published as a conference paper at ICLR 2026 while maintaining reconstruction accuracy above 98% of its full-precision counterpart. This demonstrates the vast advantages and practicality of QuantVGGT in resource-constrained scenarios. Our code is released in https://github. com/wlfeng0509/QuantVGGT . INTRODUCTION Recent advances in learning-based 3D reconstruction have demonstrated unprecedented capabilities in recovering dense geometry and camera trajectories directly from image sequences. Traditional approaches (Mur-Artal et al., 2015; Mur-Artal & Tardós, 2017; Schonberger & Frahm, 2016; Hartley & Zisserman, 2003) are grounded in geometric priors and optimization, but their reliance on handcrafted design choices and iterative solvers often leads to limited scalability and reduced robustness in complex scenes. In contrast, large-scale deep models have shifted the paradigm toward datadriven frameworks, offering remarkable generalization ability across diverse environments (Wang et al., 2025b; Yang et al., 2025) . A milestone in this evolution is the Visual Geometry Grounded Transformer (VGGT) (Wang et al., 2025a) . This 1.2B-parameter model unifies multiple 3D tasks, including dense depth estimation, point map regression, camera pose prediction, and point tracking within a single forward pass, consistently surpassing task-specialized counterparts. Despite its success, the billion-scale parameterization of VGGT incurs prohibitive computational and memory costs, severely restricting its deployment in real-world scenarios. Model quantization (Gholami et al., 2022; Jacob et al., 2018) is an effective compression technique by converting weights and activations of model from high-precision floating-points to low-precision integers. While this technique has been widely validated in large language models (Frantar et al., 2022; Xiao et al., 2023) and 2D vision models (Yuan et al., 2022; Wu et al., 2024) , the quantization of billionscale 3D reconstruction transformers such as VGGT remains largely unexplored. In our study, we identify two model-specific properties of VGGT that make its quantization particularly challenging: ❶ The presence of data-independent special tokens (camera and register tokens). Unlike reg- ular image tokens that are encoded from input images, these tokens are pretrained and injected into image tokens to encode global context and cross-view geometry. This data-independent property causes activation distributions to deviate from typical patterns, amplifying heavy tails and producing extreme channel and token variance. These skewed statistics are unfriendly to standard quantization, leading to substantial information loss. ❷ The inherently semantic complexity of 3D data. Each input sequence involves non-identical and complex views, meaning that the underlying semantic space is both high-dimensional and highly redundant. For quantization calibration, the ideal process is to perceive the expected major data distribution. If calibration samples are rare outliers and not diverse, the estimated quantization ranges become biased and fail to generalize, causing performance degradation across unseen scenes. Thus, sample diversity and representativeness are far more critical than in 2D vision tasks. To address these challenges, we present the first systematic investigation of Post-Training Quantization (PTQ) for VGGT and propose a tailored framework, QuantVGGT. Our approach introduces Dual-Smoothed Fine-Grained Quantization (DSFQ), which mitigates skewed statistics by combining (1) a pre-global rotation via Hadamard transforms to disperse outliers and smooth heavy-tailed distributions, and (2) a post-local smoothing step that normalizes channel-level variance in the rotated space. Additionally, to overcome calibration instability, we design Noise-Filtered Diverse Sampling (NFDS), which leverages deep-layer activation statistics to filter noisy extremes and employs frame-aware clustering aligned with VGGT's inductive biases. Together, these components yield robust, efficient, and accurate quantization of billion-scale 3D reconstruction transformers. Our contributions are summarized as follows: 1. We provide the first systematic analysis of PTQ on VGGT, highlighting quantization challenges rooted in its data-independent tokens and multi-view activation statistics. 2. We propose a dual-stage smoothing scheme that globally disperses heavy-tailed distributions and locally balances channel variance, significantly reducing quantization errors. 3. We design a calibration strategy that filters outliers and utilizes VGGT's inductive bias to construct frame-aware clusters, ensuring a representative and stable calibration set.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Fast-FoundationStereo: Real-Time Zero-Shot Stereo MatchingBowen Wen, Shaurya Dewan, Stan BirchfieldCVPR 2026 · 被引用 36 次
- LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token MergingZhijian Shu, Cheng Lin, Tao Xie, Wei Yin 等CVPR 2026 · 被引用 17 次
- DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry EstimationTuan Duc Ngo, Jiahui Huang, Seoung Wug Oh, Kevin Blackburn-Matzen 等CVPR 2026 · 被引用 3 次
- First-Order Error Matters: Accurate Compensation for Quantized Large Language ModelsXingyu Zheng, Haotong Qin, Yuye Li, Haoran Chu 等AAAI 2026 · 被引用 2 次
- QVGGT: Post-Training Quantized Visual Geometry Grounded TransformerZhizhen Pan, Hesong Wang, Huan WangCVPR 2026 · 被引用 2 次
它引用的顶会 Paper25
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsSaleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li 等NeurIPS 2024 · 被引用 723 次
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone 等ICCV 2021 · 被引用 686 次
- BRECQ: Pushing the Limit of Post-Training Quantization by Block ReconstructionYuhang Li, Ruihao Gong, Xu Tan, Yang Yang 等ICLR 2021 · 被引用 619 次
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 被引用 503 次
相关 Paper
- DVD-Quant: Data-free Video Diffusion Transformers QuantizationZhiteng Li, Hanxuan Li, Junyi Wu, Kai Liu 等ICLR 2026 · 被引用 13 次
- FastVGGT: Fast Visual Geometry TransformerYou Shen, Zhipeng Zhang, Yansong Qu, Xiawu Zheng 等ICLR 2026 · 被引用 73 次
- PTQ4ARVG: Post-Training Quantization for AutoRegressive Visual Generation ModelsXuewen Liu, Zhikai Li, Jing Zhang, Mengjuan Chen 等ICLR 2026 · 被引用 3 次
- Q-DiT: Accurate Post-Training Quantization for Diffusion TransformersLei Chen, Yuan Meng, Chen Tang, Xinzhu Ma 等CVPR 2025
- VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object DetectionYang Cao, Feize Wu, Dave Chen, Yingji Zhong 等CVPR 2026 · 被引用 6 次
