arxivcs.LGcs.CE2026-07-01
Interpretable vs Learned Encoders for High-Cardinality Fraud Detection
Xiao Han, Jingjing Liu, Moxuan Zheng, Zhen Zhang, Chenyu Wu
A total of seven categorical encoding methods were tested on the IEEE-CIS fraud benchmark dataset (590,540 records, 3.5% positives, 8 high-cardinality columns). The encoders were evaluated using a stratified 5-fold cross-validation (CV) with three repetitions. Five of the encoder…