CORTEXA
← Browse

Benjamin H. Sims

1 paper indexed

arxivcs.CLcs.AI2026-06-25

NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models

Henry Shaowu Yuchi, Michal Kucer, Benjamin H. Sims, Selma Peterson, Emily Taylor

Large language models (LLMs) have demonstrated strong performance across a wide range of tasks, but ensuring their reliability in highly technical domains remains a significant challenge. In nuclear engineering, problem solving often requires not only factual knowledge but also q…

View free PDFSource page