Graph Representation for Order-aware Visual Transformation
Yue Qiu, Yanjun Sun, Fumiya Matsuzawa, Kenji Iwata, Hirokatsu Kataoka
Abstract
This paper proposes a new visual reasoning formulation that aims at discovering changes between image pairs and their temporal orders. Recognizing scene dynamics and their chronological orders is a fundamental aspect of human cognition. The aforementioned abilities make it possible to follow step-by-step instructions, reason about and analyze events, recognize abnormal dynamics, and restore scenes to their previous states. However, it remains unclear how well current AI systems perform in these capabilities. Although a series of studies have focused on identifying and describing changes from image pairs, they mainly consider those changes that occur synchronously, thus neglecting potential orders within those changes. To address the above issue, we first propose a visual transformation graph structure for conveying order-aware changes. Then, we benchmarked previous methods on our newly generated dataset and identified the issues of existing methods for change order recognition. Finally, we show a significant improvement in order-aware change recognition by introducing a new model that explicitly associates different changes and then identifies changes and their orders in a graph representation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- The STVchrono Dataset: Towards Continuous Change Recognition in TimeYanjun Sun, Yue Qiu, Mariia Khan, Fumiya Matsuzawa et al.CVPR 2024 · 6 citations
- Neural Assembler: Learning to Generate Fine-Grained Robotic Assembly Instructions from Multi-View ImagesHongyu Yan, Yadong MuAAAI 2025 · 3 citations
- Context-aware Difference Distilling for Multi-change CaptioningYunbin Tu, Liang Li, Li Su, Zheng-Jun Zha et al.ACL 2024
Builds on14
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 217 citations
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh et al.CVPR 2022 · 179 citations
- Generative 3D Part Assembly via Dynamic Graph LearningGuanqi Zhan, Qingnan Fan, Kaichun Mo, Lin Shao et al.NeurIPS 2020 · 113 citations
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 108 citations
Related papers
- Transformation Driven Visual ReasoningXin Hong, Yanyan Lan, Liang Pang, Jiafeng Guo et al.CVPR 2021
- SpatialLogic-Bench: A Diagnostic Benchmark for Task-Oriented Spatiotemporal ReasoningXiaoda Yang, Shenzhou Gao, Can Wang, Jiahe Zhang et al.AAAI 2026
- Learning the Dynamics of Visual Relational Reasoning via Reinforced Path RoutingChenchen Jing, Yunde Jia, Yuwei Wu, Chuanhao Li et al.AAAI 2022 · 5 citations
- Leveraging Textual Compositional Reasoning for Robust Change CaptioningKyu Ri Park, Jiyoung Park, Seong Tae Kim, Hong Joo Lee et al.AAAI 2026
- Multi-modal Action Chain Abductive ReasoningMengze Li, Tianbao Wang, Jiahe Xu, Kairong Han et al.ACL 2023 · 11 citations
