Positional Encoding Field
Yunpeng Bai, Haoxiang Li, Qixing Huang
摘要
Diffusion Transformers (DiTs) have emerged as the dominant architecture for visual generation, powering stateof-the-art image and video models. By representing images as patch tokens with positional encodings (PEs), DiTs combine Transformer scalability with spatial and temporal inductive biases. In this work, we revisit how DiTs organize visual content and discover that patch tokens exhibit a surprising degree of independence: even when PEs are perturbed, DiTs still produce globally coherent outputs, indicating that spatial coherence is primarily governed by PEs. Motivated by this finding, we introduce the Positional Encoding Field (PE-Field), which extends positional encodings from the 2D plane to a structured 3D field. PE-Field incorporates depth-aware encodings for volumetric reasoning and hierarchical encodings for fine-grained sub-patch control, enabling DiTs to model geometry directly in 3D space. Our PE-Field-augmented DiT achieves state-of-the-art performance on single-image novel view synthesis and generalizes to controllable spatial image editing. Project page and code are available at: https://yunpeng1998.github.io/PE-Field-HomePage/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ReDirector: Creating Any-Length Video Retakes with Rotary Camera EncodingByeongjun Park, Byung-Hoon Kim, Hyungjin Chung, Jong ChulCVPR 2026 · 被引用 10 次
- ReRoPE: Repurposing RoPE for Relative Camera ControlChunyang Li, Yuanbo Yang, Jiahao Shao, Hongyu Zhou 等SIGGRAPH 2026 · 被引用 2 次
- UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World ModelsTianxing Xu, Zi-Xuan Wang, Guangyuan Wang, Li Hu 等SIGGRAPH 2026
- Pose-ICL: 3D-Aware In-Context Learning for Pose-Controllable Subject CustomizationXuan Han, Yihao Zhao, Mingyu YouICML 2026
它引用的顶会 Paper27
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov 等ICCV 2023 · 被引用 1,662 次
相关 Paper
- ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion TransformersYiyang Ma, Feng Zhou, Xuedan Yin, Pu Cao 等CVPR 2026 · 被引用 1 次
- Unified Camera Positional Encoding for Controlled Video GenerationCheng Zhang, Boying Li, Meng Wei, Yan-Pei Cao 等CVPR 2026 · 被引用 38 次
- GeoVideo: Introducing Geometric Regularization into Video Generation ModelYunpeng Bai, Shaoheng Fang, Chaohui Yu, Fan Wang 等NeurIPS 2025 · 被引用 18 次
- DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape GenerationShentong Mo, Enze Xie, Ruihang Chu, Lanqing Hong 等NeurIPS 2023 · 被引用 157 次
- MultiAnimate: Pose-Guided Image Animation Made ExtensibleYingcheng Hu, Haowen Gong, Chuanguang Yang, Zhulin An 等CVPR 2026 · 被引用 6 次
