CORTEXA
← Browse
arxivcs.AIcond-mat.mtrl-sci2026-07-10

Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents

Izumi Takahara, Teruyasu Mizoguchi

Large language model (LLM) agents are increasingly expected to play a central role in AI-driven scientific discovery. Equipped with broad knowledge, flexible reasoning, and tool use, they have the potential to autonomously explore and solve scientific problems by repeatedly proposing hypotheses, testing them, and revising their beliefs in the light of the evidence. In current agents, however, these hypotheses, tests, and belief updates are buried in unstructured logs, and no mechanism lets the agent or the human researcher audit that process. Here we propose the Hypothesis Evolution Protocol (HEP), an agent harness that provides hypothesis generation, evaluation, and evolution as explicit, auditable operations. On materials-science research tasks, a HEP-equipped agent operates the hypothesis--test--evidence--belief cycle that planning-style agents lack, generalizes across research questions, and exploits the protocol more fully as the base LLM becomes more capable. These results mark a step toward auditable AI scientists, whose scientific reasoning can be inspected, verified, and built upon.

View free PDFSource page

Related papers

arxivcond-mat.mtrl-scics.AI2026-07-11

The evolution of AI from image interpretation toward scientific inference in nanoparticle electron microscopy

Evropi Toulkeridou, Jiafei Li, Leonardo Lari, Panagiotis Grammatikopoulos

Artificial intelligence (AI) is transforming electron microscopy by enabling quantitative analysis of increasingly large and complex datasets for nanoparticle characterization. Recent advances in machine learning (ML) and deep learning (DL) have expanded microscopy from a descrip…

View free PDFSource page
arxivcs.AIcond-mat.mtrl-scics.CLcs.LG2026-07-01

Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination

Subhadeep Pal, Shashwat Sourav, Tirthankar Ghosal, Markus J. Buehler

Accelerating materials discovery requires AI systems that can generate scientifically valid hypotheses through multi-step, domain-grounded reasoning. Standard large language models often produce fluent but weakly traceable responses to open-ended materials design problems, making…

View free PDFSource page
arxivcs.AIcond-mat.mtrl-sci2026-07-01

Optimal Resource Utilization for Autonomous Laboratory Orchestrators

Austin McDannald, Julia Tisaranni, Howie Joress

In autonomous laboratories, AI agents suggest the next batch of experiments to do. However, planning and executing those tasks taking full advantage of the available resources is a completely different question. This can be challenging when dealing with real-world hardware constr…

View free PDFSource page
arxivcond-mat.mtrl-scics.AIcs.LG2026-06-29

Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop

Chenmu Zhang, Boris I. Yakobson

Predicting a material's properties from its structure is a central, fast-advancing problem in computational materials science. A decade of work has produced standard public benchmarks and many published machine-learning models for the task (Dunn et al., 2020). The task's fixed me…

View free PDFSource page
arxivcs.AIcond-mat.mtrl-sciphysics.comp-ph2026-07-02

Grounded autonomous research: a fault-tolerant LLM pipeline from corpus to manuscript in frontier computational physics

Haonan Huang

Autonomous-research agents have demonstrated end-to-end LLM automation in machine-learning sandboxes where execution provides calibration. Frontier physical science differs categorically: physical reasoning underlies every methodology choice, toolchains are often underdocumented,…

View free PDFSource page
arxivcs.LGcond-mat.mtrl-scics.AIcs.CL2026-06-28

Interpretable Inverse Design of Metal-Organic Frameworks with Large Language Model Agents

Kyungmin Nam, Seunghee Han, Jihan Kim

Inverse design of metal-organic frameworks (MOFs) requires searching a combinatorially vast space where property labels are expensive and most machine-learning models reveal little about why a structure succeeds. We introduce LLM4MOF, a closed-loop framework in which language-mod…

View free PDFSource page