The Sweet Danger of Sugar: Debunking Representation Learning for Encrypted Traffic Classification
Yuqi Zhao, Giovanni Dettori, Matteo Boffa, Luca Vassio, Marco Mellia
摘要
Recently we have witnessed the explosion of proposals that, inspired by Language Models like BERT, exploit Representation Learning models to create traffic representations. All of them promise astonishing performance in encrypted traffic classification (up to 98% accuracy). In this paper, with a networking expert mindset, we critically reassess their performance. Through extensive analysis, we demonstrate that the reported successes are heavily influenced by data preparation problems, which allow these models to find easy shortcuts - spurious correlation between features and labels - during fine-tuning that unrealistically boost their performance. When such shortcuts are not present - as in real scenarios - these models perform poorly. We also introduce Pcap-Encoder, an LM-based representation learning model that we specifically design to extract features from protocol headers. Pcap-Encoder appears to be the only model that provides an instrumental representation for traffic classification. Yet, its complexity questions its applicability in practical settings. Our findings reveal flaws in dataset preparation and model training, calling for a better and more conscious test design. We propose a correct evaluation methodology and stress the need for rigorous benchmarking.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- TDDM-Melatt: A Decoupled Memory and Diffusion Framework for Generalizable Encrypted Traffic ClassificationZe Chen, Qiming Yu, Zijia Song, Guozheng Yang 等CCS 2026
- Defeating Slow-and-Low Threats via Diffusion Model-based Generative InferenceSeyed Mohammad Mehdi Mirnajafizadeh, Prashant Khanduri, DaeHun Nyang, Rhongho JangNSDI 2026
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson 等ICML 2023 · 被引用 908 次
- ET-BERT: A Contextualized Datagram Representation with Pre-training Transformers for Encrypted Traffic ClassificationXinjie Lin, Gang Xiong, Gaopeng Gou, Zhen Li 等WWW 2022 · 被引用 490 次
相关 Paper
- SoK: Decoding the Enigma of Encrypted Network Traffic ClassifiersNimesha Wickramasinghe, Arash Shaghaghi, Gene Tsudik, Sanjay K. JhaS&P 2025
- HF-Transformer: A Non-Pretrained Encrypted Network Traffic Classification Model Based on Packet Header FieldsZhenzhen Yan, Lizhi Peng, Peiqiang Liu, Yingshuo Bao 等INFOCOM 2026
- Yet Another Traffic Classifier: A Masked Autoencoder Based Traffic Transformer with Multi-Level Flow RepresentationRuijie Zhao, Mingwei Zhan, Xianwen Deng, Yanhao Wang 等AAAI 2023 · 被引用 138 次
- Rosetta: Enabling Robust TLS Encrypted Traffic Classification in Diverse Network Environments with TCP-Aware Traffic AugmentationRenjie Xie, Jiahao Cao, Enhuan Dong, Mingwei Xu 等USENIX Security 2023
- Training Robust Classifiers for Classifying Encrypted Traffic under Dynamic Network ConditionsYuqi Qing, Qilei Yin, Xinhao Deng, Xiaoli Zhang 等CCS 2025
