Long-range Modeling and Processing of Multimodal Event Sequences
Jichu Li, Yilun Zhong, Zhiting Li, Feng Zhou, Quyu Kong
摘要
Temporal point processes (TPPs) have emerged as powerful tools for modeling asynchronous event sequences. While recent advances have extended TPPs to handle textual information, existing approaches are limited in their ability to generate rich, multimodal content and reason about event dynamics. A key challenge is that incorporating multimodal data dramatically increases sequence length, hindering the ability of attention-based models to generate coherent, long-form textual descriptions that require long-range understanding. In this paper, we propose a novel framework that extends LLM-based TPPs to the visual modality, positioning text generation as a core capability alongside time and type prediction. Our approach addresses the long-context problem through an adaptive sequence compression mechanism based on temporal similarity, which reduces sequence length while preserving essential patterns. We employ a two-stage paradigm of pre-training on compressed sequences followed by supervised fine-tuning for downstream tasks. Extensive experiments, including on the challenging DanmakuTPP-QA benchmark, demonstrate that our method outperforms state-of-the-art baselines in both predictive accuracy and the quality of its generated textual analyses.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Transformer Hawkes ProcessSimiao Zuo, Haoming Jiang, Zichong Li, Tuo Zhao 等ICML 2020 · 被引用 382 次
- Self-Attentive Hawkes ProcessQiang Zhang, Aldo Lipani, Ömer Kirnap, Emine YilmazICML 2020 · 被引用 254 次
- Transformer Embeddings of Irregularly Spaced Events and Their ParticipantsHongyuan Mei, Chenghao Yang, Jason EisnerICLR 2022 · 被引用 98 次
相关 Paper
- LongVU: Spatiotemporal Adaptive Compression for Long Video-Language UnderstandingXiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu 等ICML 2025
- Byte-token Enhanced Language Models for Temporal Point Processes AnalysisQuyu Kong, Yixuan Zhang, Yang Liu, Panrong Tong 等WWW 2026 · 被引用 6 次
- VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video PromptingMuhammet Furkan Ilaslan, Ali Köksal, Kevin Qinghong Lin, Burak Satar 等AAAI 2025 · 被引用 3 次
- Interacting Diffusion Processes for Event Sequence ForecastingMai Zeng, Florence Regol, Mark CoatesICML 2024 · 被引用 9 次
- TimeCAP: Learning to Contextualize, Augment, and Predict Time Series Events with Large Language Model AgentsGeon Lee, Wenchao Yu, Kijung Shin, Wei Cheng 等AAAI 2025 · 被引用 39 次
