Summary
A medRxiv preprint evaluated a large-language-model pipeline for assessing risk of bias in neurological prognosis studies. Agreement with original human assessments was limited, while a small comparison suggested a similar level of agreement to that between human raters.
A large-language-model pipeline showed limited agreement with human assessments when it was used to judge the risk of bias in prognosis studies in clinical neurology, according to a medRxiv preprint. The researchers evaluated 298 articles drawn from 15 previously published systematic reviews covering epilepsy, traumatic brain injury and stroke.
The work was designed as a feasibility study for automating a repetitive part of systematic reviews. Its results suggest that an LLM could assist with this task, but the authors say methodological refinement and validation against expert ratings are still needed before wider use.
How the assessment was designed
Systematic reviews collect and synthesise evidence from multiple studies. Before combining those findings, reviewers assess the risk of bias: the possibility that features of a study’s design, conduct or analysis could systematically distort its results. For prognosis research, one tool used for this purpose is the Quality in Prognosis Studies, or QUIPS, framework.
The researchers created an LLM-based pipeline as a virtual analogue of a human reviewer. It used zero-shot prompting, meaning the model was given instructions for carrying out the assessment without being trained on example answers within the task. The pipeline was tailored to the QUIPS framework and applied to articles that had already appeared in published systematic reviews.
The evaluation compared two kinds of agreement:
- the agreement between the LLM’s risk-of-bias assessments and the original human assessments; and
- the agreement between the original human assessments themselves.
The analysis covered studies from three neurological domains: epilepsy, traumatic brain injury and stroke.
Agreement was limited, with differences in scoring
Agreement between the LLM and the original human assessments was measured using Cohen’s weighted kappa, a statistic that accounts for the degree of disagreement between ratings. The LLM–human result was 0.22, with a 95% confidence interval of 0.12 to 0.33.
The researchers also compared the human assessments with one another, but this part involved a small sample of five comparisons. The resulting weighted kappa was −0.25, with a 95% confidence interval from −1.04 to 0.54. The authors describe their findings as tentatively suggesting that the LLM’s agreement may not be inferior to the agreement between human raters in this dataset. The small human–human comparison makes that interpretation preliminary.
Statistical tests found significant differences across four bias domains and for the overall risk scores. Rank-biserial correlations indicated that the human raters tended to assign higher risk scores than the LLM counterparts. In practical terms, the automated pipeline generally rated the included studies as having lower risk of bias than the human assessments did.
A possible workflow aid, not a finished replacement
The study indicates that a prompt-engineered LLM can be used to produce structured risk-of-bias assessments with the QUIPS framework across a set of neurological prognosis studies. Automating this step could, in the authors’ view, help reduce the time and cost involved in systematic reviews.
That potential depends on resolving the differences between automated and human scoring. The authors identify standardising how QUIPS is implemented and validating the system against expert ratings as important next steps. The work is also a preprint rather than a peer-reviewed journal article, and its evaluation was restricted to prognosis research in clinical neurology and to articles drawn from existing reviews.
The result is therefore best understood as an early feasibility assessment: it shows that the workflow can be attempted at scale, while the measured agreement and scoring differences indicate that expert oversight and further validation remain important.