Summary

A University of Gothenburg preprint models glycan motifs as a containment hierarchy rather than independent features. Across 45 glycomics datasets, the authors report more estimated true positives and identify motif redistributions that conventional marginal analysis misses.

A bioRxiv preprint from Xinyu Zhao and Daniel Bojar at the University of Gothenburg presents a way to analyse glycan motifs as related structures rather than as independent measurements. The authors report that the approach increased estimated true positives from 300 to 442 — a 47% rise — across 45 glycomics datasets while controlling false positives using permutation-based null data.

The method also identified compositional changes that conventional feature-by-feature analysis did not report, including a recurring redistribution of core-1 sialylation across eight human O-glycomes and a colorectal fucosylation shift confined to part of a glycan structure.

Why glycan motifs need structural context

Glycans are chains or branching arrangements of sugars attached to proteins and other biological molecules. Glycomics measures the structures and abundances of these glycans, while glycoproteomics examines glycans in relation to the proteins carrying them. Researchers often represent recurring structural patterns as motifs and compare their abundance between samples or conditions.

The preprint argues that these motifs are not statistically independent. A larger, or parent, motif can contain a smaller child motif as part of its structure. Changes in a child can therefore contribute to the measured abundance of the parent. Treating both as unrelated features can make it difficult to determine whether a parent motif changed because of its own structure or because one of its substructures changed.

A containment graph separates inherited and local changes

Zhao and Bojar make these relationships explicit with a directed acyclic graph: a network in which each motif can be connected to larger motifs that contain it. This containment order is used to organise the multiple-testing problem, in which many statistical comparisons can otherwise increase the chance of false positives.

The method also uses an empirical-Bayes variance prior drawn from each motif’s containment neighbourhood. In practical terms, information from structurally related motifs helps estimate how much variation should be expected for an individual motif. The authors report that combining this neighbourhood-based variance estimate with the containment graph increased the number of estimated true positives from 300 to 442, with a reported p-value of 0.0001, after controlling permutation-null false positives.

The second part of the framework examines the parent motif’s children together with its residual — the portion not accounted for by those children. The authors describe these as a genuine sub-composition. Log-ratio balances between them can therefore compare how the composition is apportioned without requiring an external reference frame or scale model. This allows the analysis to distinguish a motif’s own change from a shift inherited from its surrounding structural contexts.

Among 337 significant parent motifs, 150, or 45%, were reported to carry no signal of their own. A further 138 motifs moved only in the decomposition analysis. These results indicate that a significant overall motif measurement can conceal a redistribution among its component structures.

Examples of hidden redistributions

The preprint gives examples from both glycomics and glycoproteomics. Across eight human O-glycomes, the analysis found a recurrent reapportioning of core-1 sialylation at the immune-inhibitory disialyl-T antigen. It also identified a colorectal fucosylation shift restricted to the antenna, a branching portion of the glycan structure. According to the authors, neither pattern appeared in marginal analysis, which evaluates motifs individually without explicitly modelling their containment relationships.

The reported evidence is computational analysis of existing glycomics datasets rather than a clinical or treatment study. The work is available as a bioRxiv preprint, so its method and findings are presented before peer review. The abstract reports aggregate results and examples, but does not provide dataset-by-dataset sample sizes or assay details; the biological consequences of the identified motif redistributions are also not quantified.

Sources