Data-augmented machine learning refines the effective-concentration estimate for eculizumab in complement-mediated diseases
Francisco J. Fernández, Lucía Alfonso‐González, Manuel Praga, Kaare Mikkelsen, M. Cristina Vega
Eculizumab, a humanized monoclonal antibody targeting the complement protein C5, is highly efficient in paroxysmal nocturnal hemoglobinuria, atypical hemolytic uremic syndrome, generalized myasthenia gravis, and neuromyelitis optica spectrum disorder. However, recent reports have highlighted a subset of patients who show inadequate treatment response, prompting dose escalation or interval shortening. The serum concentration required for sustained inhibition of the complement lytic pathway remains uncertain, and the commonly cited 50–100 µg/ml range is not grounded in robust pharmacokinetic–pharmacodynamic data. Using publicly available clinical data digitized from four rare indications, we trained machine-learning supervised classifiers to predict complete C5 inhibition as a function of serum eculizumab concentration, identifying logistic regression on log-transformed concentration as the best and mechanistically appropriate model on real data. We then examined whether augmenting the real training data with high-fidelity synthetic data improves the resulting estimate. Although augmentation did not significantly change balanced accuracy, it reduced the confidence interval for the effective concentration by roughly 12-fold (a 92% reduction). The augmentation-stabilized estimate of the concentration associated with complete C5 inhibition was approximately 310.8–335.2 µg/ml (real-data-only estimate 338 µg/ml, 95% CI 275–454). These data-driven, hypothesis-generating results indicate that current maintenance targets aimed at near-complete C5 suppression may be substantially underestimated and that faithful synthetic augmentation can improve the precision of model-based pharmacological estimates in rare diseases; prospective validation in independent cohorts is required before any clinical application.