Summary

A systematic review and external validation assessed 48 models for predicting four adverse pregnancy outcomes using more than 1.58 million UK pregnancies. Most showed limited discrimination and substantial miscalibration in the UK cohort.

A systematic review and external validation of 48 clinical prediction models found that most performed poorly when tested against records from more than 1.58 million pregnancies in the UK. The models aimed to predict gestational diabetes, pre-eclampsia, stillbirth or small-for-gestational-age (SGA) birth using maternal characteristics routinely recorded before conception or early in pregnancy.

The study, posted on medRxiv as a preprint, reports limited ability to distinguish higher-risk from lower-risk pregnancies for most outcomes, alongside substantial mismatches between predicted and observed risks. The authors conclude that most of the models assessed are not suitable for direct use in UK clinical practice without further calibration, updating or redevelopment.

What the researchers assessed

The researchers first reviewed studies of clinical prediction models, then externally validated eligible models using CPRD Aurum, a UK primary-care database. The cohort included pregnancies among women aged 14–49 between 2000 and 2020. Eligible models had to use routinely collected maternal characteristics, provide a complete model equation and rely on predictors commonly measured before conception or early in pregnancy.

Across 30 studies, the review identified 48 models: 47 for binary outcomes and one for a continuous outcome. The models had been developed in cohorts ranging from 101 to 113,415 participants, with a median of 5,013. Although 29 models had previously undergone external validation, only three had been validated in cohorts that included a UK population and met the study’s required sample size.

The team assessed discrimination and calibration. Discrimination describes how well a model separates pregnancies that experience an outcome from those that do not; the C-statistic is one measure of this ability. Calibration describes how closely predicted risks match observed outcomes. The researchers examined calibration-in-the-large, calibration slope and calibration plots, pooling estimates across 20 imputations to address missing data.

Performance varied, but most models did not travel well

For gestational diabetes, C-statistics ranged from 0.29 to 0.77; for pre-eclampsia, from 0.42 to 0.71. Models for stillbirth had values from 0.54 to 0.56, while those for SGA ranged from 0.58 to 0.62. The authors characterised discrimination for most models as limited, with SGA models among the exceptions they noted as having excellent performance.

Calibration was also uneven. Calibration-in-the-large ranged from -5.66 to 1.35, and calibration slopes from -0.65 to 17.61. In general, a calibration-in-the-large value near zero and a slope near one indicate closer agreement between predicted and observed risks. The wide ranges reported here point to substantial variation in how well the models’ risk estimates matched outcomes in this cohort.

A model can perform differently in a new population when characteristics of that population or the frequency of an outcome differ from those in the data used to develop it. The authors identify these differences as likely contributors to the poor generalisability they observed. Their results support treating published prediction models as candidates for local testing and adjustment, rather than assuming that a model developed elsewhere will give dependable risk estimates in UK practice.

The findings are from a preprint, which has not completed peer review. The validation was conducted in a UK primary-care cohort, so the reported performance does not by itself describe how these models would perform in other health systems.

Sources