arxivcs.CLcs.AIcs.DL2026-06-26
Mitigating LLM-based p-Hacking by Preregistering for the Next LLM
Maria Thomas, Kristina Gligoric, Nihar B. Shah
Large language models (LLMs) are increasingly used to generate, classify, and annotate data whose outputs feed downstream hypothesis tests. However, LLM-based research is easy to p-hack: a researcher can tune the prompts, decoding parameters, or output format until a desired resu…