SivsFormer: Parallax-Aware Transformers for Single-image-based View Synthesis
Chunlan Zhang, Chunyu Lin, Kang Liao, Lang Nie, Yao Zhao
Abstract
Single-image-based view synthesis is significant for generating a 3D scene and gains increasing attention in recent years. However, this task is challenging as it requires inferring contents beyond what is immediately visible. Previous methods directly predict the unknown views using the convolutional neural networks, but the generated views suffer from visually unpleasant holes, deformations, and artifacts. In this paper, we propose a Single-image-based view synthesis transformer (named SivsFormer) for high-quality and realistic view synthesis. In particular, a warping and occlusion handing module is designed to reduce the influence of parallax on the network. Subsequently, a disparity alignment module captures the long-range information over the scene and ensures that pixels move in a geometrically correct manner with soft probabilistic disparity maps. Moreover, we present a parallax-aware loss function to improve the quality of the synthetic images, which explicitly quantifies the magnitude of parallaxes. We conduct extensive experiments on popular KITTI and Cityscapes datasets. Benefitting from the proposed parallax-aware transformer, our approach achieves superior performance in both quantitative and qualitative evaluations.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- CVSformer: Cross-View Synthesis Transformer for Semantic Scene CompletionHaotian Dong, Enhui Ma, Lubo Wang, Miaohui Wang et al.ICCV 2023 · 12 citations
- Geometry-Free View Synthesis: Transformers and no 3D PriorsRobin Rombach, Patrick Esser, Björn OmmerICCV 2021 · 115 citations
- VIAFormer: Voxel-Image Alignment Transformer for High-Fidelity Voxel RefinementTiancheng Fang, Bowen Pan, Lingxi Chen, Jiangjing Lyu et al.CVPR 2026 · 1 citation
- VoxFormer: Sparse Voxel Transformer for Camera-Based 3D Semantic Scene CompletionYiming Li, Zhiding Yu, Christopher B. Choy, Chaowei Xiao et al.CVPR 2023
- Dual-S3D: Hierarchical Dual-Path Selective SSM-CNN for High-Fidelity Implicit ReconstructionLuoxi Zhang, Pragyan Shrestha, Yu Zhou, Chun Xie et al.ICCV 2025
