AutoML for Network-Based Intrusion Detection: Evaluation Practice, Dataset Quality, and Deployment Constraints
Abdulla Amin Aburomman, Mamun Bin Ibne Reaz
Machine learning techniques for network-based intrusion detection systems (NIDS) have advanced considerably over the past decade. Still, improvements are inhibited by handcrafted feature pipelines, isolated public benchmark data, and evaluation procedures that do not reflect real-life deployment. AutoML, a branch of ML automating model selection, automated architecture search, and the creation of model pipelines, may help overcome these shortcomings. While numerous NIDS applications employing automated ML techniques have been proposed, and recent surveys have mapped the AutoML framework landscape for network intrusion detection, no existing review critically audits the evaluation practice of this literature: the quality of its benchmark datasets, the reproducibility of its reported results, and the realism of its deployment assumptions. This paper critically reviews 26 research works published between January 2023 and June 2026, collected via a two-phase structured search: a documented keyword search across five databases (Scopus, IEEE Xplore, Web of Science, ACM Digital Library, and Google Scholar), followed by full-text eligibility screening, citation chaining, and expert evaluation. Findings drawn from this collection capture trends observed among the selected studies, rather than reflecting the broader state of the field. Analysis of the corpus reveals that 88% of dataset-verified studies evaluate exclusively or partly on the legacy benchmark family (KDD-derived, CICIDS, UNSW-NB15, CIDDS), 21% evaluate on a single dataset only, and among attribute-verified studies only 32% release source code, 40% report statistical significance testing, and 36% include variance analysis, findings that collectively motivate the four contributions of this study. First, a recommended evaluation framework is proposed, addressing baseline parity, transparent search-space and budget reporting, nested cross-validation for selection-bias control, and stability reporting across multiple random seeds. Second, a dataset quality scoring framework is introduced, assessing five dimensions: overlap rate, duplication rate, label correctness, attack-type representativeness, and coverage of benign, IoT, and IIoT traffic. Third, a cross-domain justification is provided for neural architecture search (NAS) and meta-learning in NIDS, grounded in advances in federated NAS, out-of-distribution robustness, edge-constrained search cost reduction, and few-shot adaptation. Fourth, a structured research roadmap is outlined, targeting real-world validation, standardized benchmarks, curated datasets, resource-aware AutoML, and privacy-preserving federated NAS. In contrast to prior surveys of AutoML for network intrusion detection, which map frameworks and computational paradigms, this review contributes a formalized evaluation checklist, an explicit and partially empirically validated dataset quality scoring scheme, and evidence-based methodological guidance grounded in a transparent, fully enumerated study corpus.