Unnoticeable Hybrid Watermarking for Deep Neural Network Authentication Using Auxiliary Hidden Layers
Rodrigo Eduardo Arevalo-Ancona, Manuel Cedillo-Hernandez
The authentication and protection of deep neural network models have become challenging due to their widespread distribution and reuse, making them vulnerable to unauthorized access. This paper addresses the need for ownership verification by proposing a hybrid neural network watermarking method for secure model authentication. The approach combines a steganographic watermark embedded into stable model weights with a user code for watermark recovery encoded in auxiliary hidden layers. Stable parameters are identified through a reduced training to estimate gradient variations for the watermark insertion with minimal impact on model performance. Additionally, two auxiliary layers are introduced, to store in the first layer the metadata indices from the selected weights where the watermark was embedded and in the second layer the user code, supporting secure identification and verification. Experimental evaluations demonstrate that the proposed method remains robust under different model optimization attacks, including pruning, fine-tuning, additive noise injection, and parameter overwriting, while preserving model performance. The proposed framework achieves a BER = 0 under several moderate attack scenarios across different neural network models, whereas more aggressive optimizations degrade the watermark recovery performance. These results indicate that the proposed framework provides an effective solution for neural network ownership protection while maintaining the model performance.