arxivcs.AI2026-07-21
SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring
Chunxiao Li, Yuan Xiong, Lijun Li, Tianyi Du, Wenlong Zhang, Lei Bai, et al.
Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse guidance. Existing benchmarks often rely on templated queries disconnected from real-world hazards, and employ LLM-as-a-Judge paradigms witho…