Automatic classification of cyber incidents using privacy-preserving artificial intelligence
Loya Caroldene Haughton, Eduardo Fidalgo, David Lewis
As cyber incidents increase in complexity, diversity and frequency, cybersecurity practitioners find it more challenging to extract meaningful threat intelligence and insights from cyber incident reports. This problem is worsened by the limited number of these reports; as cyber incident victims may withhold and/or downplaytheir reports due to reputational and privacy concerns. Therefore, this study aims to determine whether cyber incident reports, which have been stripped of personal data (pseudonymised), can be classified according to the Spanish INCIBE Cyber Incident Taxonomy. Seven transformers (SecBERT, BERT, DistilBERT, BERTweet, SecRoBERTa, ALBERT, RoBERTa) and four traditional machine learning classifiers (Random Forest, Multinomial Naïve Bayes, XGBoost, Support Vector Machine, using two encoders - Bag of Words and Term Frequency - Inverse Document Frequency), were trained to classify cyber incidents which were pseudonymised using Data Masking, Data Tokenisation and Data Substitution. The best-performing model attained an F1-score of 78.5% on the non-pseudonymised CECILIA-10C-900 dataset and 83% on the dataset, which was pseudonymised using Data Masking, D-CECILIA-10C-900-MAS. Therefore, this research demonstrates that it is possible for a single AI model to balance enhanced privacy with strong threat intelligence analysis capabilities.
Also available via: European Organization for Nuclear Research