A Platform for Machine Learning Operations for Network Constrained Far-Edge Devices
Calum McCormack, Imene Mitiche
Machine Learning (ML) models developed for the Edge have seen a massive uptake in recent years, with many types of predictive analytics, condition monitoring and pre-emptive fault detection developed and in-use on Internet of Things (IoT) systems serving industrial power generators, environmental monitoring systems and more. At scale, these systems can be difficult to manage and keep upgraded, especially those devices that are deployed in far-Edge networks with unreliable networking. This paper presents a simple and novel platform architecture for deployment and management of ML at the Edge for increasing model and device reliability by reducing downtime and access to new model versions via the ability to manage models from both Cloud and Edge. This platform provides an Edge ML Operations “Mirror” that replicates and minimises cloud MLOps systems to provide reliable delivery and retraining of models at the network Edge, solving many problems associated with both Cloud-first and Edge networks. The paper explores and explains the architecture and components of the system, offering a prototype system that was evaluated by measuring time to deploy models with regard to differing network instabilities in a simulated environment to highlight the necessity for local management and federated training of models as a secondary function to Cloud model management. This architecture could be utilised by researchers to improve the deployment, recording and management of ML experiments on the Edge.