TokenSplat: Token-aligned 3D Gaussian Splatting for Feed-forward Pose-free Reconstruction
Yihui Li, Chengxin Lv, Zichen Tang, Hongyu Yang, Di Huang
摘要
We present TokenSplat, a feed-forward framework for joint 3D Gaussian reconstruction and camera pose estimation from unposed multi-view images. At its core, TokenSplat introduces a Token-aligned Gaussian Prediction module that aligns semantically corresponding information across views directly in the feature space. Guided by coarse token positions and fusion confidence, it aggregates multi-scale contextual features to enable long-range cross-view reasoning and reduce redundancy from overlapping Gaussians. To further enhance pose robustness and disentangle viewpoint cues from scene semantics, TokenSplat employs learnable camera tokens and an Asymmetric Dual-Flow Decoder (ADF-Decoder) that enforces directionally constrained communication between camera and image tokens. This maintains clean factorization within a feed-forward architecture, enabling coherent reconstruction and stable pose estimation without iterative refinement. Extensive experiments demonstrate that TokenSplat achieves higher reconstruction fidelity and novel-view synthesis quality in pose-free settings, and significantly improves pose estimation accuracy compared to prior pose-free methods. Project page: https://kidleyh.github.io/tokensplat/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Gau-Occ: Geometry-Completed Gaussians for Multi-Modal 3D Occupancy PredictionChengxin Lv, Yihui Li, Hongyu Yang, Yunhong WangCVPR 2026 · 被引用 3 次
- Catalyst4D: High-Fidelity 3D-to-4D Scene Editing via Dynamic PropagationShifeng Chen, Yihui Li, Jun Liao, Hongyu Yang 等CVPR 2026
它引用的顶会 Paper21
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- MVSNeRF: Fast Generalizable Radiance Field Reconstruction from Multi-View StereoAnpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang 等ICCV 2021 · 被引用 1,024 次
- LightGlue: Local Feature Matching at Light SpeedPhilipp Lindenberger, Paul-Edouard Sarlin, Marc PollefeysICCV 2023 · 被引用 936 次
- Mip-Splatting: Alias-Free 3D Gaussian SplattingZehao Yu, Anpei Chen, Binbin Huang, Torsten Sattler 等CVPR 2024 · 被引用 360 次
相关 Paper
- TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable TokensJiawei Ren, Michal J. Tyszkiewicz, Jiahui Huang, Zan GojcicCVPR 2026 · 被引用 13 次
- FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D ReconstructionJiale Xu, Shenghua Gao, Ying ShanICCV 2025 · 被引用 8 次
- No Pose at All: Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse ViewsRanran Huang, Krystian MikolajczykICCV 2025 · 被引用 12 次
- UniSplat: Unified Spatio-Temporal Fusion via 3D Latent Scaffolds for Dynamic Driving Scene ReconstructionChen Shi, Shaoshuai Shi, Xiaoyang Lyu, Chunyang Liu 等ICLR 2026 · 被引用 10 次
- FreeSplat: Generalizable 3D Gaussian Splatting Towards Free View Synthesis of Indoor ScenesYunsong Wang, Tianxin Huang, Hanlin Chen, Gim Hee LeeNeurIPS 2024 · 被引用 112 次
