arxivcs.LGcs.CRstat.ML2026-07-10
Statistically Undetectable Backdoors in Deep Neural Networks
Andrej Bogdanov, Alon Rosen, Neekon Vafa
We show how an adversarial model trainer can plant backdoors in a large class of deep, feedforward neural networks. These backdoors are statistically undetectable in the white-box setting, meaning that the backdoored and honestly trained models are close in total variation distan…