CORTEXA
← Browse
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-26Cited by 0

Leakage-Controlled Seriousness Triage of FAERS Reports: A Temporal Validation Framework with LLM Comparison on Novel First-in-Class Drugs

Shakil Mahmud

Background. Post-marketing pharmacovigilance depends on the timely identification of serious individual case safety reports (ICSRs) from large spontaneous-reporting databases such as the FDA Adverse Event Reporting System (FAERS). Machine-learning triage has been proposed to prioritize case review, but reported performance is frequently inflated by label leakage and by evaluation designs that do not reflect prospective use. Objective. To develop a leakage-controlled seriousness-triage classifier for ICSRs, to quantify its generalization to novel, first-in-class drugs absent from training, and to test whether a large language model (LLM), which carries pretrained drug knowledge, improves triage on such drugs. Methods. We trained a clinical-only gradient-boosted classifier on FAERS reports using reaction terms, drug counts, and demographics, excluding all reporting-channel features and all fields downstream of the seriousness label. Leakage was audited by an explicit feature blacklist and a behavioral ablation procedure, and the classifier was validated on a temporal hold-out (train Q1–Q3 2025, test Q4 2025). It was then evaluated on six recently approved first-in-class drugs held out entirely from training and compared head-to-head against an LLM (Claude Sonnet 5) given the same pre-outcome information plus the drug name on identical full-panel cases at each drug's natural base rate. Results. The clinical classifier achieved a temporal AUROC of 0.896 (95% CI 0.892–0.900) with reasonable calibration (Brier score 0.140). A subgroup audit showed discrimination was stable across groups while operating-point recall was not, systematically under-prioritizing reports with missing demographic data—remediable by group-aware thresholds. On the six held-out novel drugs, the classifier generalized moderately (n-weighted mean AUROC 0.740); the LLM outperformed it on every drug (n-weighted mean AUROC 0.902, gap +0.161), significantly so for four of six, with the advantage largest where the classical model was weakest. Conclusions. A leakage-controlled, clinical-only classifier provides a defensible baseline for ICSR seriousness triage but degrades on novel drugs with little reporting history. An LLM's pretrained drug knowledge substantially improves triage precisely where data-driven pattern learning is weakest — a mechanistically coherent finding with direct implications for surveillance of newly approved products. We release a public decision-support dashboard implementing calibrated, explainable triage with explicit reliability boundaries.

View free PDFSource page

Related papers

openalexZenodo (CERN European Organization for Nuclear Research)2026-07-24

Activity cliffs resist prediction within and across protein kinases: code and derived results for a leakage-controlled machine-learning analysis

Samuel S Agboola, Oluwaseun E. Agboola, et al

Code and derived results for a study of whether the chemical transformations thatgenerate activity cliffs on one protein kinase predict cliffs on another. Matched molecular pairs were constructed from measured Ki and Kd binding affinitiesretrieved from ChEMBL (release 37) for 20…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-26

Reproducibility package for Explainable and Leakage-Conscious Machine Learning for Supplied Injury-Risk Classification and Longitudinal Athlete Injury Forecasting

Abdülkadir Enes GÖRGÜLÜ, Eray Dursun, Serdar Solak

This record provides the complete reproducibility package for the manuscript “Explainable and Leakage-Conscious Machine Learning for Supplied Injury-Risk Classification and Longitudinal Athlete Injury Forecasting.” Overview The study evaluates explainable and leakage-conscious ma…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-23

Sensor Validator v3.5: Adaptive Multi-Modal Sensor Validation and Threat Detection Framework

Niall Devlin

Sensor Validator v3.5 is a Python-based framework for adaptive validation of environmental and chemical sensor systems. The platform combines multi-modal feature extraction, anomaly detection, machine-learning classification, drift monitoring, automatic recalibration, hardware ab…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-23

Explainable AI-Based Diabetes Risk Prediction with Multi-Level User-Oriented Explanations: A Novel Communication Framework

K Khan

While machine learning models achieve promising results in diabetes prediction, clinical adoption remains limited due to black-box nature and lack of stakeholder-specific communication. This study proposes a novel multi-level explanation framework that translates a single XGBoost…

View free PDFSource page