Procedural Text Understanding via Scene-Wise Evolution
Jialong Tang, Hongyu Lin, Meng Liao, Yaojie Lu, Xianpei Han, Le Sun, Weijian Xie, Jin Xu
Abstract
Procedural text understanding requires machines to reason about entity states within the dynamical narratives. Current procedural text understanding approaches are commonly entity-wise, which separately track each entity and independently predict different states of each entity. Such an entity-wise paradigm does not consider the interaction between entities and their states. In this paper, we propose a new scene-wise paradigm for procedural text understanding, which jointly tracks states of all entities in a scene-by-scene manner. Based on this paradigm, we propose Scene Graph Reasoner (SGR), which introduces a series of dynamically evolving scene graphs to jointly formulate the evolution of entities, states and their associations throughout the narrative. In this way, the deep interactions between all entities and states can be jointly captured and simultaneously derived from scene graphs. Experiments show that SGR not only achieves the new state-of-the-art performance but also significantly accelerates the speed of reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c55b4d6-b48d-4c98-8fe9-5b4bf99f18e4Builds on3
- Knowledge-Aware Procedural Text Understanding with Multi-Stage TrainingZhihan Zhang, Xiubo Geng, Tao Qin, Yunfang Wu et al.WWW 2021 · 23 citations
- Understanding Procedural Text using Interactive Entity NetworksJizhi Tang, Yansong Feng, Dongyan ZhaoEMNLP 2020 · 8 citations
- Reasoning over Entity-Action-Location Graph for Procedural Text UnderstandingHao Huang, Xiubo Geng, Jian Pei, Guodong Long et al.ACL 2021
Related papers
- Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language ModelsZhiwei Yang, Yuanchen Wu, Nan Zhang, Yucong Meng et al.ICML 2026 · 1 citation
- LLM Meets Scene Graph: Can Large Language Models Understand and Generate Scene Graphs? A Benchmark and Empirical StudyDongil Yang, Minjin Kim, Sunghwan Kim, Beong-woo Kwak et al.ACL 2025 · 8 citations
- Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine ComprehensionHuibin Zhang, Zhengkun Zhang, Yao Zhang, Jun Wang et al.ACL 2022 · 5 citations
- 3D Question Answering with Scene Graph ReasoningZizhao Wu, Haohan Li, Gongyi Chen, Zhou Yu et al.ACM MM 2024 · 6 citations
- Visual Semantics Allow for Textual Reasoning Better in Scene Text RecognitionYue He, Chen Chen, Jing Zhang, Juhua Liu et al.AAAI 2022 · 62 citations
