Summary
A bioRxiv preprint reports that incomplete reference profiles can leave bulk RNA deconvolution targets non-identifiable even when an algorithm returns precise estimates. The authors propose measuring coverage and false-certification risk, and estimate the measurement precision needed for a decisive result in peripheral blood mononuclear cells.
A bioRxiv preprint reports that reference-based deconvolution of bulk RNA profiles can produce a precise numerical estimate while the biological target remains non-identifiable when the reference profiles are incomplete. The analysis separates the computational path used to obtain an estimate from the information actually contained in the measurements, then quantifies how accurate additional measurements must be to identify a target reliably.
The study, posted on September 16, 2026, also introduces fitdrop, a framework for scoring the shrinkage, coverage, certification and false-certification behaviour of methods used to handle incomplete references.
How the analysis separates an estimate from a target
Bulk RNA measurements combine signals from many cell types. Reference-based deconvolution attempts to infer the proportions of those cell types by comparing the combined profile with reference profiles for individual cell types. The reference is incomplete when some cell types, cellular states or relevant variation are not adequately represented.
An estimate is identifiable when the available observations constrain the target to a single value, or to a valid narrow range. If multiple underlying cell mixtures can produce the same observed data but imply different target values, the target is not identified, even if one computational procedure returns a highly precise answer.
Jiang and colleagues compared different operator histories—the computational paths used to derive estimates—at a common reduced reference. In 28.2% of 196,420 sample-deletion pairs, the resulting estimates differed by more than 0.1 in total variation, a distance measure for comparing normalized compositions. When all learned components were locked, those paths became identical, linking the differences to the information learned during the procedure rather than to the final calculation alone.
The consequences extended beyond estimated proportions. Across nine cohorts, the operator history reversed 31 of 952 associations and changed statistical significance for 126. These results indicate that the path used to fit an incomplete-reference model can affect downstream biological conclusions.
The precision frontier for a decisive result
Fixing the computational operator did not resolve the underlying identification problem in the authors’ analysis. Different reference completions that were observationally equivalent could occupy the open simplex—the range of possible cell-type mixtures—and could reverse the ranking of retained cell types. Adding shared structure also did not tighten the inherited bounds. Between one and six uncalibrated views produced identical bounds.
A profile library reduced the range of possible estimator outputs by 98.93%, but its narrowed range covered the effect obtained with a full reference in only 42.77% of cases. The analysis included 19 wrong-sign certificates, where a narrowed result supported the wrong direction. By contrast, the conditional sharp bounds remained at [−1, 1], indicating no useful constraint on the direction of the target under those conditions.
The authors describe this gap as a difference between contraction and coverage. A method can make its output interval much narrower without producing an interval that contains the true target. In a further test, treating absolute RNA yields as exact reduced intervals to points, yet those points covered none of seven targets measured by flow cytometry.
Calibrated cross-modal anchors—measurements that connect the RNA data to another measurement modality—performed better in the analysis. At six cell types, they contracted the interval width to 0.51, but they still did not validly certify the sign of an effect. For peripheral blood mononuclear cells, the authors calculate that deciding the sign of a target would require proxy accuracy of approximately ±2.6%, together with near-total contraction of the envelopes representing donor heterogeneity and dynamic range.
The result turns incomplete-reference deconvolution into a measurement-planning problem: before treating a narrow estimate as decisive, researchers need to establish whether the available data cover the target and how much additional calibration is required. The work is a computational bioRxiv preprint, so its numerical thresholds describe the authors’ analysed cohorts, targets and model setup rather than a universal requirement for every tissue or deconvolution task.