Representation Learning of Tangled Key-Value Sequence Data for Early Classification
Tao Duan, Junzhou Zhao, Shuo Zhang, Jing Tao, Pinghui Wang
Abstract
Key-value sequence data has become ubiquitous and naturally appears in a variety of real-world applications, ranging from the user-product purchasing sequences in e-commerce, to network packet sequences forwarded by routers in networking. Classifying these key-value sequences is important in many scenarios such as user profiling and malicious applications identification. In many time-sensitive scenarios, besides the requirement of classifying a key-value sequence accurately, it is also desired to classify a key-value sequence early, in order to respond fast. However, these two goals are conflicting in nature, and it is challenging to achieve them simultaneously. In this work, we formulate a novel tangled key-value sequence early classification problem, where a tangled key-value sequence is a mixture of several concurrent key-value sequences with different keys. The goal is to classify each individual key-value sequence sharing a same key both accurately and early. To address this problem, we propose a novel method, i.e., Key-Value sequence Early Co-classification (KVEC), which leverages both inner- and inter-correlations of items in a tangled key-value sequence through key correlation and value correlation to learn a better sequence representation. Meanwhile, a time-aware halting policy decides when to stop the ongoing key-value sequence and classify it based on current sequence representation. Experiments on both real-world and synthetic datasets demonstrate that our method outperforms the state-of-the-art baselines significantly. KVEC improves the prediction accuracy by up to 4.7 -17.5% under the same prediction earliness condition, and improves the harmonic mean of accuracy and earliness by up to 3.7 -14.0%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a9206ba-ddf1-4537-8193-2cae0ee4244bBuilds on7
- k-fingerprinting: A Robust Scalable Website Fingerprinting TechniqueJamie Hayes, George DanezisUSENIX Security 2016 · 474 citations
- Triplet Fingerprinting: More Practical and Portable Website Fingerprinting with N-shot LearningPayap Sirinam, Nate Mathews, Mohammad Saidur Rahman, Matthew WrightCCS 2019 · 268 citations
- Learning to Classify: A Flow-Based Relation Network for Encrypted Traffic ClassificationWenbo Zheng, Chao Gou, Lan Yan, Shaocong MoWWW 2020 · 100 citations
- Pinpointing Hidden IoT Devices via Spatial-temporal Traffic FingerprintingXiaobo Ma, Jian Qu, Jianfeng Li, John C. S. Lui et al.INFOCOM 2020 · 48 citations
- Session-aware Linear Item-Item Models for Session-based RecommendationMinjin Choi, Jinhong Kim, Joonseok Lee, Hyunjung Shim et al.WWW 2021 · 31 citations
Related papers
- TSec: An Efficient and Effective Framework for Time Series ClassificationYuanyuan Yao, Hailiang Jie, Lu Chen, Tianyi Li et al.ICDE 2024 · 7 citations
- Recurrent Halting Chain for Early Multi-label ClassificationThomas Hartvigsen, Cansu Sen, Xiangnan Kong, Elke A. RundensteinerKDD 2020 · 18 citations
- Con4m: Context-aware Consistency Learning Framework for Segmented Time Series ClassificationJunru Chen, Tianyu Cao, Jing Xu, Jiahe Li et al.NeurIPS 2024 · 4 citations
- Learning Heterogeneous Temporal Patterns of User Preference for Timely RecommendationJunsu Cho, Dongmin Hyun, SeongKu Kang, Hwanjo YuWWW 2021 · 40 citations
- Interpretable Sequence Classification via Discrete OptimizationMaayan Shvo, Andrew C. Li, Rodrigo Toro Icarte, Sheila A. McIlraithAAAI 2021 · 29 citations
