Learning Compact 3D Representations from Feed-Forward Novel View Synthesis
Honggyu An, Jaewoo Jung, Mungyeom Kim, Chaehyun Kim, Minkyeong Jeon, Jisang Han, Kazumi Fukuda, Takuya Narihira, HYUNAH KO, Junsu Kim, Sunghwan Hong, Yuki Mitsufuji, Seungryong Kim
Abstract
Reconstructing and understanding 3D scenes from sparse views in a feed-forward manner remains challenging. While recent approaches use per-pixel 3D Gaussian Splatting for reconstruction and 2D-to-3D feature lifting for scene understanding, they generate excessive redundant Gaussians, causing high memory overhead and sub-optimal multi-view feature aggregation. We propose a feed-forward framework that estimates compact Gaussians only at essential spatial locations, minimizing redundancy while enabling effective feature lifting. We introduce learnable tokens that aggregate multi-view features through self-attention to guide Gaussian generation, ensuring each Gaussian integrates relevant visual features across views. We then exploit the learned attention patterns to efficiently lift features. Extensive experiments on 3D open-vocabulary segmentation and view-invariant feature generation demonstrate our approach's effectiveness. Results show that a compact yet geometrically meaningful representation is sufficient for high-quality scene reconstruction, achieving superior memory efficiency and feature fidelity compared to existing methods. All of our code will be made publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a10a281-edf4-44cf-8df2-b9eb315c991eBuilds on36
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Language-driven Semantic SegmentationBoyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun et al.ICLR 2022 · 885 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
- LangSplat: 3D Language Gaussian SplattingMinghan Qin, Wanhua Li, Jiawei Zhou, Haoqian Wang et al.CVPR 2024 · 164 citations
- GeoNeRF: Generalizing NeRF with Geometry PriorsMohammad Mahdi Johari, Yann Lepoittevin, François FleuretCVPR 2022 · 154 citations
Related papers
- CF3: Compact and Fast 3D Feature FieldsHyunjoon Lee, Joonkyu Min, Jaesik ParkICCV 2025
- TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable TokensJiawei Ren, Michal J. Tyszkiewicz, Jiahui Huang, Zan GojcicCVPR 2026 · 13 citations
- Tackling View-Dependent Semantics in 3D Language Gaussian SplattingJiazhong Cen, Xudong Zhou, Jiemin Fang, Changsong Wen et al.ICML 2025
- SLGaussian: Fast Language Gaussian Splatting in Sparse ViewsKangjie Chen, BingQuan Dai, Minghan Qin, Dongbin Zhang et al.ACM MM 2025 · 5 citations
- Learning Efficient Fuse-and-Refine for Feed-Forward 3D Gaussian SplattingYiming Wang, Lucy Chai, Xuan Luo, Michael Niemeyer et al.NeurIPS 2025 · 5 citations
