Summary

A medRxiv preprint analysing 704 online health narratives reports that factual alignment with medical consensus and potential health risk are largely separate properties. The authors argue that health-information systems should assess both dimensions rather than rely on fact-checking alone.

A medRxiv preprint reports that agreement with medical consensus and potential safety risk are not interchangeable properties in peer-to-peer health discussions. In an analysis of 704 health narratives, 699 of which were classified in the reported results, two separate measures shared only 4.9% of their variance.

The researchers, Ommo Clark, Karuna Joshi and Zeliatu Ahmed of the University of Maryland, Baltimore County, call this separation the “Online Health Safety Gap”. Their argument is that a post can be factually aligned with medical knowledge while still encouraging behaviour with substantial health risk, while a post that diverges from consensus may not carry the same level of immediate behavioural risk.

The work was posted on medRxiv on September 21, 2026, as a preprint. It is a computational health-informatics study of online narratives rather than a clinical trial involving patients.

The study separates factual alignment from risk

The analysis uses two scores calculated from non-overlapping sets of features. Narrative Truth Distance, or NTD, represents epistemic divergence: how far a narrative is from the medical knowledge or consensus used by the system. Narrative Risk Score, or NRS, represents the narrative’s potential health risk.

The relationship between the scores was weak but statistically significant, with a correlation of r=0.222 and p<0.001. The authors report that the scores shared 4.9% of their variance. In practical terms, factual alignment provided limited information about the separate risk-potential score.

The researchers then placed narratives into four combinations of alignment and risk. They report that 25.2% were aligned with medical consensus but risky, while 14.4% were divergent but safe. Together, these two off-diagonal groups represented 39.6% of the narratives—cases that a system using only one of the two dimensions would be poorly equipped to handle.

Here, “risk” refers to the study’s score for health-risk potential in a narrative. The reported analysis measures properties of online content, not resulting patient harm or clinical outcomes.

Why fact-checking alone is an incomplete safety measure

The preprint also examines an expert-labelled misinformation benchmark containing 437 posts and 127 misinformation instances. The strongest discrimination reported in that benchmark came from a supervised classifier trained directly on the expert labels, with a Youden’s J of 0.349. Youden’s J combines sensitivity and specificity into a single measure of classification performance.

Frozen biomedical embeddings and a prompted large language model did not improve on that classifier. The authors interpret the result as evidence that labels describing factual accuracy contain little information about behavioural health risk. In other words, a benchmark designed to recognise misinformation may be useful for measuring factual classification while remaining poorly suited to measuring whether advice is dangerous to follow.

To address this, the researchers introduce a “Classification Quadrant” that maps epistemic divergence and health-risk potential into four governance categories, each intended to have different intervention implications. The proposed shift is from a single fact-centred score toward a system that separately asks whether content is medically aligned and whether acting on it could be harmful.

The underlying corpus came from Reddit and cannot be redistributed, according to the authors, because of platform terms and user-privacy considerations. They say the data were collected through the official Reddit API, while the biomedical sources used for the analysis—including UMLS, SNOMED CT, SemMedDB and PubMed—are publicly available under their respective licensing terms.

The framework’s performance in live moderation systems, and whether the proposed quadrant-based approach improves health outcomes or user decisions, remain to be tested. Its immediate contribution is a measurement argument: factual correctness and behavioural safety should be evaluated as related but distinct dimensions of online health information.

Sources