Summary

A bioRxiv preprint finds that the reproducibility of functional assays can cap measured performance for variant-effect predictors. Correcting for assay reliability shifts the apparent prediction gap from canonical splice sites toward intronic regions 11–50 base pairs from the splice boundary.

A bioRxiv preprint finds that the reliability of a functional assay can limit how well genetic variant-effect predictors appear to perform. When the analysis adjusted for differences in assay reproducibility across genomic regions, the apparent weakness at canonical splice sites became smaller, while a stronger prediction gap appeared 11–50 base pairs inside introns.

The study evaluated 19 predictors against a frozen atlas containing 64,178 saturation genome-editing variants across seven cancer-susceptibility genes. Its central argument is that a benchmark can measure both a predictor's performance and the noise in the experimental measurement used to judge it.

Why assay reliability sets a benchmark ceiling

Variant-effect predictors estimate whether a DNA change is likely to alter gene function. They are increasingly compared with multiplexed assays of variant effect, or MAVEs, which experimentally test large numbers of variants. Using these assays can avoid circularity that arises when a model is evaluated against labels derived from similar clinical or computational rules.

But a functional assay is itself a measurement. If repeated experiments do not produce the same result consistently, a predictor cannot achieve an arbitrarily high correlation with the assay's scores. The preprint calls the resulting maximum a territory-specific reliability ceiling, with “territory” referring to the genomic strata being compared, such as coding and splice-related regions.

The researchers estimated these ceilings from published replicate scores and standard errors, then used simulations to examine the effect of correcting benchmark results. They found that the correction reduced error when the ceiling was above about 0.45, but amplified error when the ceiling was below that level. Reliability varied more between genomic territories than performance did between the predictors being compared.

The apparent splice-site failure moved inward

The uncorrected benchmark showed a pronounced collapse in predictor performance at canonical splice sites. After accounting for the assays' reliability, the median shortfall relative to coding regions narrowed from 1.7-fold to 1.4-fold.

The analysis instead placed the stronger predictor failure 11–50 base pairs into the intron. This distinction matters because a benchmark can make a model appear to fail at a well-defined biological feature when part of the pattern is caused by lower measurement reproducibility in that region.

The result was sensitive to the genes included in the analysis. The convergence between splice-related and coding regions survived removal of BARD1 or PALB2, but it reversed when BRCA1 was removed. The authors therefore reported all three leave-one-gene-out analyses rather than treating the genes as independent evidence. They also noted that the reported parity at the performance frontier rests on one data deposit.

A correction that can be applied to existing assay data

The preprint also examined information available through MaveDB, a repository for multiplexed variant-effect measurements. Of 2,803 score sets, 2,452 contained, at the upper bound, the information needed for a reliability estimate. A conventional search based on column names would identify only about one-tenth of these datasets, according to the analysis.

Among 674 human deposits for which a ceiling could be computed, between 29.9% and 51.8% had a reliability ceiling below 0.90. The range reflects the analysis' handling of the available assay information rather than a single estimate for all human datasets.

The researchers also treated the assay-derived functional calls as a classification task in three genes. The same predictors separated damaging variants from tolerated variants better than their correlation scores suggested, although none reached the strongest evidence band at the 95%-specificity operating point. This shows why correlation and classification can give different views of a predictor: correlation measures agreement with numerical assay scores, while classification evaluates separation between categories.

The proposed reporting standard is to provide a reliability estimate for each genomic stratum, or to state when the assay cannot support one. The authors estimate that this approach can be applied, at the upper bound, to 87% of MaveDB score sets without requiring additional information from depositors. The work is a bioRxiv preprint, and its conclusions are based on the analyzed atlas, replicate measurements and available repository metadata.

Sources