Summary

A medRxiv preprint evaluated 14 general-purpose large language models on eight schizophrenia symptoms in clinical summaries from a 704-person severe mental illness cohort. Performance was poor to moderate overall, with negative symptoms extracted less accurately than positive symptoms.

A medRxiv preprint has evaluated 14 general-purpose large language models (LLMs) for extracting schizophrenia symptoms from clinical summaries. Across eight symptoms in a severe mental illness cohort of 704 people, the models produced poor-to-moderate results, with macro F1 scores ranging from 0.500 to 0.647.

The study found a consistent difference between symptom groups: positive symptoms were extracted more accurately than negative symptoms. The authors conclude that the models are not yet reliable enough for clinical use, although they may have a role in exploratory research using large patient cohorts.

What the evaluation measured

The Cardiff University researchers tested whether LLMs could identify symptom information from clinical summaries. In psychiatry, positive symptoms generally refer to experiences or behaviours added to a person's usual functioning, such as hallucinations, delusions or disorganised thought. Negative symptoms involve reductions in functions such as motivation, emotional expression or social engagement. The distinction matters because an automated system that captures one group better than the other could create an incomplete picture of illness.

The main performance measure was macro F1. This combines precision—the share of extracted symptoms that are correct—with recall—the share of relevant symptoms that were found—and averages the result across symptom categories. Macro averaging gives each category equal weight. A score range of 0.500 to 0.647 therefore reflects substantial variation across the tested symptoms and models rather than a single uniform level of accuracy.

The researchers also tested few-shot prompting, in which a model is shown examples of the task before producing its own answers. This strategy did not significantly improve the results in the evaluation.

Why the positive-negative gap matters

The models' outputs were also compared with the study's gold-standard symptom information in regression analyses. For positive symptoms, the associations between the model-derived results and clinical variables were comparable to those obtained using the gold-standard data.

At the individual level, however, extraction errors weakened group-level associations. The effect was uneven across symptom domains, meaning that errors did not simply add the same amount of noise to every type of symptom. The lower performance on negative symptoms could therefore introduce a systematic bias into research that uses LLMs to characterise psychotic symptoms.

This distinction separates two possible uses of automated extraction. A system may identify broad patterns across a large cohort well enough to support an exploratory analysis, while still making too many or too unevenly distributed errors to support clinical decisions about individual patients. The authors say that improving extraction of negative symptoms is essential before LLM-derived symptom profiles can be used more reliably.

The work is a medRxiv preprint posted on September 17, 2026, rather than a report of a clinical intervention or a validated diagnostic system. Its findings concern automated analysis of clinical summaries from the CardiffCOGS severe mental illness cohort. The results provide an evaluation of these models and this task; performance in other clinical settings, documentation styles or patient populations would require separate assessment.

Sources