Mitigating Performance Saturation in Neural Marked Point Processes: Architectures and Loss Functions
Tianbo Li, Tianze Luo, Yiping Ke, Sinno Jialin Pan
Abstract
Attributed event sequences are commonly encountered in practice. A recent research line focuses on incorporating neural networks with the statistical model--marked point processes, which is the conventional tool for dealing with attributed event sequences. Neural marked point processes possess good interpretability of probabilistic models as well as the representational power of neural networks. However, we find that performance of neural marked point processes is not always increasing as the network architecture becomes more complicated and larger, which is what we call the performance saturation phenomenon. This is due to the fact that the generalization error of neural marked point processes is determined by both the network representational ability and the model specification at the same time. Therefore we can draw two major conclusions: first, simple network structures can perform no worse than complicated ones for some cases; second, using a proper probabilistic assumption is as equally, if not more, important as improving the complexity of the network. Based on this observation, we propose a simple graph-based network structure called GCHP, which utilizes only graph convolutional layers, thus it can be easily accelerated by the parallel mechanism. We directly consider the distribution of interarrival times instead of imposing a specific assumption on the conditional intensity function, and propose to use a likelihood ratio loss with a moment matching mechanism for optimization and model selection. Experimental results show that GCHP can significantly reduce training time and the likelihood ratio loss with interarrival time probability assumptions can greatly improve the model performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Transformer Hawkes ProcessSimiao Zuo, Haoming Jiang, Zichong Li, Tuo Zhao et al.ICML 2020 · 382 citations
- Self-Attentive Hawkes ProcessQiang Zhang, Aldo Lipani, Ömer Kirnap, Emine YilmazICML 2020 · 254 citations
- Tweedie-Hawkes Processes: Interpreting the Phenomena of OutbreaksTianbo Li, Yiping KeAAAI 2020 · 6 citations
Related papers
- A Variational Point Process Model for Social Event SequencesZhen Pan, Zhenya Huang, Defu Lian, Enhong ChenAAAI 2020 · 19 citations
- Learning Neural Point Processes with Latent GraphsQiang Zhang, Aldo Lipani, Emine YilmazWWW 2021 · 30 citations
- Attentive Neural Point Processes for Event ForecastingYulong GuAAAI 2021 · 24 citations
- Self-Adaptable Point Processes with Nonparametric Time DecaysZhimeng Pan, Zheng Wang, Jeff M. Phillips, Shandian ZheNeurIPS 2021 · 13 citations
- Decomposable Transformer Point ProcessesAristeidis PanosNeurIPS 2024 · 16 citations
