arxivcs.LGcs.AIcs.CL2026-07-13
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias
Zixiang Xu, Sixian Li, Huaxing Liu, Xiang Wang, Shuai Li, Zirui Song, et al.
Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations. We argue that the same biases admit a representation-level account in the judge's hidden state, complementa…