CORTEXA
← Browse

David Arbour

2 papers indexed

arxivcs.LGstat.AP2026-06-26

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges

Lei Shi, Anlan Zhang, Rita Lyu, Zhengmian Hu, Tong Yu, David Arbour, et al.

AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges, tasks, and domains. When uncalibrated AI evaluations are used for model ranking, item scoring, or…

View free PDFSource page
arxivcs.LGcs.CL2026-06-25

The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching

Sankaran Vaidyanathan, David Arbour, Aaron Mueller, Scott Niekum, David Jensen

Activation patching is the primary tool in mechanistic interpretability. It attributes causal responsibility for a model behavior to each of its individual components by estimating its natural indirect effect (NIE). Re-deriving the activation patching estimand from causal mediati…

View free PDFSource page