Group Editing: Edit Multiple Images in One Go
Yue Ma, Xinyu Wang, Qianli Ma, Qinghe Wang, Mingzhe Zheng, Xiangpeng Yang, Hao Li, Chongbo Zhao, Jixuan Ying, Harry Yang, Hongyu Liu, Qifeng Chen
Abstract
In this paper, we tackle the problem of performing consistent and unified modifications across a set of related images. This task is particularly challenging because these images may vary significantly in pose, viewpoint, and spatial layout. Achieving coherent edits requires establishing reliable correspondences across the images, so that modifications can be applied accurately to semantically aligned regions. To address this, we propose GroupEditing, a novel framework that builds both explicit and implicit relationships among images within a group. On the explicit side, we extract geometric correspondences using VGGT, which provides spatial alignment based on visual features. On the implicit side, we reformulate the image group as a pseudo-video and leverage the temporal coherence priors learned by pre-trained video models to capture latent relationships. To effectively fuse these two types of correspondences, we inject the explicit geometric cues from VGGT into the video model through a novel fusion mechanism. To support large-scale training, we construct GroupEditData, a new dataset containing high-quality masks and detailed captions for numerous image groups. Furthermore, to ensure identity preservation during editing, we introduce an alignment-enhanced RoPE module, which improves the model's ability to maintain consistent appearance across multiple images. Finally, we present GroupEditBench, a dedicated benchmark designed to evaluate the effectiveness of group-level image editing. Extensive experiments demonstrate that GroupEditing significantly outperforms existing methods in terms of visual quality, cross-view consistency, and semantic alignment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2d44c6e-5b5d-4ae8-8860-03e99d8f638aCited by top-tier papers4
- DAG: A Dual Correlation Network for Time Series Forecasting with Exogenous VariablesXiangfei Qiu, Yuhan Zhu, Zhengyu Li, Xingjian Wu et al.ICML 2026 · 22 citations
- PATRA: Pattern-Aware Alignment and Balanced Reasoning for Time Series Question AnsweringJunkai Lu, Peng Chen, Xingjian Wu, Yang Shu et al.ICML 2026 · 3 citations
- EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX GenerationYue Ma, Xu Ye, Qinghe Wang, Yucheng Wang et al.SIGGRAPH 2026 · 2 citations
- KITE: Knowledge-Guided Probabilistic Modeling for Time Series Forecasting with Exogenous VariablesHanyin Cheng, Jingrong Zhou, Yang Shu, Chenjuan GuoICML 2026
Builds on57
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- FlowDrag: 3D-aware Drag-based Image Editing with Mesh-guided Deformation Vector Flow FieldsGwanhyeong Koo, Sunjae Yoon, Younghwan Lee, Ji Woo Hong et al.ICML 2025
- COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video EditingJiangshan Wang, Yue Ma, Jiayi Guo, Yicheng Xiao et al.NeurIPS 2024 · 76 citations
- VGGT-Segmentor: Geometry-Enhanced Cross-View SegmentationYulu Gao, Bohao Zhang, Zongheng Tang, Jitong Liao et al.CVPR 2026 · 3 citations
- ChronoEdit: Towards Temporal Reasoning for In-Context Image Editing and World SimulationJay Zhangjie Wu, Xuanchi Ren, Tianchang Shen, Tianshi Cao et al.ICLR 2026 · 17 citations
- MotionEdit: Benchmarking and Learning Motion-Centric Image EditingYixin Wan, Lei Ke, Wenhao Yu, Kai-Wei Chang et al.CVPR 2026 · 7 citations
