Summary
A bioRxiv preprint comparing matched full-UDG and non-UDG libraries from two medieval Mongolian individuals found trade-offs between site retention, residual disagreement and kinship inference. Corrected mixed-library data supported a first-degree, likely mother–son relationship, while uncorrected data shifted one estimate towards second degree.
A bioRxiv preprint comparing matched full-UDG and non-UDG ancient-DNA libraries from the same extracts of two medieval Mongolian individuals found that no single post-mortem damage-correction method performed best for every downstream task. The methods changed both how many genomic sites remained usable and how closely the libraries agreed. In kinship analysis, corrected mixed-library data supported a first-degree relationship, while uncorrected non-UDG data shifted one estimate towards second degree.
Contents
How the comparison was designed
Ancient DNA can accumulate chemical damage after an organism dies. One common form of damage can create incorrect base changes near the ends of DNA fragments. Uracil-DNA glycosylase, or UDG, treatment is used to remove uracil-related damage; a full-UDG library is treated throughout the relevant preparation process, while a non-UDG library retains more of the original damage pattern.
The researchers compared matched full-UDG and non-UDG libraries made from the same extracts of two medieval Mongolian individuals. They tested three computational approaches before genotype imputation and kinship analysis: trimming bases from fragment ends, rescaling base-quality scores, and masking known single-nucleotide polymorphisms near damaged ends.
Each approach reduced or reversed the difference in mean alternative-allele fraction between damage-prone sites and transversion SNPs. The methods nevertheless preserved different numbers of covered sites, meaning positions with usable sequence data.
In the non-UDG libraries, masking and base-quality rescaling within five bases of each fragment end produced similar cross-library non-reference discordance. Masking retained 92% of covered sites, while rescaling retained 99.3%. The libraries also showed a 3′ bias. For these data, trimming 10 bases only from the 3′ end retained more covered sites and produced lower observed discordance than symmetric trimming of five bases from both ends.
Kinship inference depended on correction
The researchers used several relatedness analyses. ancIBD inferred widespread sharing of one chromosome copy that was identical by descent, or IBD1, across the correction methods. Identical-by-descent sharing refers to genetic material inherited from a common ancestor rather than merely matching by chance.
The broader relationship estimate was more sensitive to preprocessing. Uncorrected non-UDG data shifted the estimate for TKGWV2 towards a second-degree relationship. After correction, mixed-library comparisons supported a first-degree relationship. READv2 and the low rate of IBS0 sites—positions where the two individuals carry opposite homozygous genotypes—supported a parent–offspring relationship.
Mitochondrial data and genetic sex favoured an assignment of ORT16 as the mother and ORT15 as the son. Together, these analyses supported a mother–son interpretation for the two individuals in the dataset.
Why the trade-off matters
Damage correction is not simply a choice between cleaned and uncleaned data. Trimming can remove misleading fragment-end bases, but it also discards sequence. Quality rescaling keeps more positions while lowering the confidence assigned to potentially damaged observations. Masking removes selected sites altogether. Their effects therefore depend on whether the priority is reducing residual damage, retaining genomic coverage or obtaining reliable results for a particular downstream analysis.
The study is especially relevant when full-UDG and non-UDG libraries are combined. Its results show that the preprocessing decision can affect both data retention and the apparent degree of biological relatedness. The authors conclude that no method was best across all measures; the appropriate choice depends on the analysis and the trade-off being prioritised.
The comparison involved only two medieval Mongolian individuals and is reported as a bioRxiv preprint. It is therefore a focused assessment of matched libraries rather than a universal ranking of correction methods across ancient-DNA datasets.