Leakage-Controlled Seriousness Triage of FAERS Reports: A Temporal Validation Framework with LLM Comparison on Novel First-in-Class Drugs
Background. Post-marketing pharmacovigilance depends on the timely identification of serious individual case safety reports (ICSRs) from large spontaneous-reporting databases such as the FDA Adverse Event Reporting System (FAERS). Machine-learning triage has been proposed to prioritize case review, but reported performance is frequently inflated by label leakage and by evaluation designs that do not reflect prospective use. Objective. To develop a leakage-controlled seriousness-triage classifier for ICSRs, to quantify its generalization to novel, first-in-class drugs absent from training, and to test whether a large language model (LLM), which carries pretrained drug knowledge, improves triage on such drugs. Methods. We trained a clinical-only gradient-boosted classifier on FAERS reports using reaction terms, drug counts, and demographics, excluding all reporting-channel features and all fields downstream of the seriousness label. Leakage was audited by an explicit feature blacklist and a behavioral ablation procedure, and the classifier was validated on a temporal hold-out (train Q1–Q3 2025, test Q4 2025). It was then evaluated on six recently approved first-in-class drugs held out entirely from training and compared head-to-head against an LLM (Claude Sonnet 5) given the same pre-outcome information plus the drug name on identical full-panel cases at each drug's natural base rate. Results. The clinical classifier achieved a temporal AUROC of 0.896 (95% CI 0.892–0.900) with reasonable calibration (Brier score 0.140). A subgroup audit showed discrimination was stable across groups while operating-point recall was not, systematically under-prioritizing reports with missing demographic data—remediable by group-aware thresholds. On the six held-out novel drugs, the classifier generalized moderately (n-weighted mean AUROC 0.740); the LLM outperformed it on every drug (n-weighted mean AUROC 0.902, gap +0.161), significantly so for four of six, with the advantage largest where the classical model was weakest. Conclusions. A leakage-controlled, clinical-only classifier provides a defensible baseline for ICSR seriousness triage but degrades on novel drugs with little reporting history. An LLM's pretrained drug knowledge substantially improves triage precisely where data-driven pattern learning is weakest — a mechanistically coherent finding with direct implications for surveillance of newly approved products. We release a public decision-support dashboard implementing calibrated, explainable triage with explicit reliability boundaries.