CORTEXA
← Browse
openalexnpj Digital Medicine2026-07-23Cited by 0

Evaluating large language models for assessment of psychosis risk

Taiyu Zhu, Alexander Tashevski, Maxime Taquet, Matilda Azis, Tia Jani, Matthew R. Broome, Thomas Kabir, Amedeo Minichino, Graham K. Murray, Matthew M. Nour, Ilina Singh, Paolo Fusar-Poli, Alejo Nevado-Holgado, Philip McGuire, Dominic Oliver

Abstract Psychosis prevention relies on early detection of individuals at clinical high risk for psychosis (CHR-P). The effectiveness of the CHR-P state is constrained, in part, due to clinical assessments requiring specialist interpretation of narrative interviews, limiting scalability. Here, we evaluate whether large language models (LLMs; deep learning models trained on large text corpora to process and generate language) can extract clinically meaningful information from such interviews to support psychosis risk assessment. We assessed 11 open-weight LLMs on 678 partial PSYCHS interview transcripts from 373 participants (77.7% CHR-P). Models inferred CHR-P status and estimated severity and frequency across 15 symptom domains, benchmarked against researcher-rated scores. Larger models achieved the strongest classification performance (Llama-3.3-70B: accuracy = 0.80, sensitivity = 0.93, specificity = 0.58). LLM-generated symptom scores showed good correlations with researcher-rated scores (ICC sev = 0.74, ICC freq = 0.75). Performance disparities were minimal across most demographic groups but varied across sites. Generated summaries were largely faithful to source transcripts, with low rates of clinically relevant confabulation (3%). Errors primarily reflected over-pathologisation of non-clinical experiences. While accuracy scaled with model size, smaller models achieved competitive performance with substantially lower computational cost. These findings demonstrate that open-weight LLMs have the potential to assess psychosis risk from psychometric interview transcripts, supporting scalable, human-in-the-loop approaches to early detection.

View free PDFSource page

Related papers

openalexnpj Digital Medicine2026-07-23

Generative AI mental health chatbots: a scoping review of intervention design and user experience

Lotenna Olisaeloka, Chris G. Richardson, Angel Y Wang, Richard J. Munthali, Daniel V. Vigo

Generative AI (GenAI) mental health chatbots offer scalable, on-demand support with the potential to address persistent gaps in mental healthcare access. Yet evidence on how these interventions are designed and how users experience them remains fragmented. To our knowledge, this…

View free PDFSource page
openalexnpj Digital Medicine2026-07-23

Meaningful oversight of medical AI beyond human in the loop

Davy van de Sande, Nicoleta Economou‐Zavlanos, Michel E. van Genderen

Human oversight of medical AI is increasingly required, but clinician presence alone does not make oversight meaningful. We propose four interlocking conditions—epistemic capacity, cognitive space, decisional authority, and intervention effectiveness—that determine whether human…

View free PDFSource page