Predicting the dynamic behavior of memristive devices using recurrent neural networks
Somayeh Rezaei, Benjamin Spetzler, Patrick Mäder
Charge-transport models provide quantitative and physically interpretable descriptions of memristive devices but are computationally prohibitive for highly iterative tasks in model-driven design workflows such as parameter space exploration, parameter extraction, and optimization. Here, we investigate recurrent neural networks (RNNs) as efficient sequence-to-sequence surrogates to accelerate dynamic memristive transport models and provide systematic guidance on training strategies, architectural choices, and data requirements. In this context, we focus on long short-term memory (LSTM) and gated recurrent unit (GRU) architectures combined with advanced training and normalization strategies. The results demonstrate that layer normalization substantially improves convergence, training stability, and generalization, whereas chrono initialization degrades performance in this setting. The most robust training behavior is obtained by combining layer normalization with the AdamW optimizer and cosine annealing learning-rate scheduling, with GRU architectures achieving the lowest errors overall. Using the optimized configurations, mean normalized errors below 0.1% are achieved for a five-dimensional parameter space. Accurate performance is retained with limited training data, with fewer than 1000 configurations still yielding mean errors around 0.15%. Increasing the input dimensionality leads to a systematic rise in error from ∼0.1% (5D) to 0.66% (9D), while mean errors remain small, well below 1%. These results establish practical design rules for applying RNN-based surrogates in model-driven design and optimization of memristive devices.