A Survey on Self-Supervised Learning in Cybersecurity: Network Intrusion and Malware Detection
Josue Genaro Almaraz-Rivera, Jose Antonio Cantoral-Ceballos, Juan Felipe Botero
Self-Supervised Learning (S-SL) is a recent line of research that could represent the next step to understanding human intuition. By blending the strengths of unsupervised and supervised learning paradigms, S-SL endows Deep Learning models with stronger generalization capabilities. Although better known for its applications in Computer Vision and Natural Language Processing, S-SL has also proved its value in other fields, such as cybersecurity. In this work, we review the current progress and future trends in S-SL for the two most relevant problems discussed in the cybersecurity literature: network intrusion and malware detection. The scope of this survey spans from 2019 to 2025. From an initial analysis of over 200 documents, we distill the 50 most relevant papers. We also highlight opportunity areas, such as attack detection over encrypted network traffic, RAM-based analysis of obfuscated malware, creating S-SL models for tabular data and resource-constrained devices, as well as the research on backdooring, encoder extraction, the transferability of vulnerabilities, and data memorization in S-SL. To the best of our knowledge, this is the first comprehensive survey regarding the application of Self-Supervised Learning in cybersecurity, benchmarking contrastive learning vs auxiliary pretext tasks and presenting the data requirements for implementing S-SL solutions in this field. We hope this paper provides a firm ground for further exploration.