Summary

A Nature study analysed a large-scale proteogenomic dataset from 29 healthy human tissues and identified 13,910 protein variants, including non-genetic amino-acid substitutions. The findings suggest that errors and variation during protein production may expand the set of protein forms with potential biological functions.

Proteins in human cells can exist in more forms than the underlying DNA sequence alone suggests. In an early-access version of a peer-reviewed, accepted paper published in Nature, researchers analysed a large-scale proteogenomic dataset from 29 healthy human tissues and identified 13,910 confidently localized protein variants, representing 7,215 unique single-amino-acid substitutions.

These variants coexisted with the corresponding reference proteoforms—the standard protein sequences expected from the relevant genes. The analysis included genetically encoded changes such as single-nucleotide polymorphisms and somatic mutations, as well as non-genetic substitutions created during protein production. The researchers also experimentally validated non-genetic substitutions in selected purified proteins.

How protein diversity can arise after DNA is read

The central dogma of molecular biology describes the flow of information from DNA to RNA and then to protein. Each stage can introduce a source of diversity. Differences between inherited gene copies, mutations acquired by cells, transcriptional errors and translation errors can all result in protein forms that differ from the reference sequence.

A mistranslated protein contains an amino acid that was inserted incorrectly during translation, even though the underlying gene sequence remains unchanged. Such a change is different from a genetic mutation: it occurs during the conversion of genetic instructions into a protein rather than through a permanent change to DNA.

This distinction matters because genetic changes are constrained by the codons and mutation paths available in the genetic code. A translation-level substitution provides another route through protein sequence space. The researchers argue that this route could generate amino-acid combinations that are less readily reached through genetic mutation alone.

Recurring substitutions point to a broader proteome

The study reports that the abundance of both genetic protein variants and mistranslated variants mirrored allele frequencies in the human population. This pattern links the amount of a variant protein observed in the dataset with the frequency of the corresponding genetic form, while also showing that non-genetic variants can be measured alongside genetically encoded ones.

Additional analyses found specific, recurring non-genetic variation patterns when cancer-derived cell lines were exposed to amino-acid starvation. The researchers also identified hundreds of substituted non-genetic proteoforms that either recurred consistently in multiple healthy individuals or mapped to annotated functional sites in proteins.

Together with the targeted experiments on purified proteins, these results led the authors to propose that some non-genetic substitutions constitute a new class of functional protein phenotypic variants. The broader implication is that the human proteome—the complete collection of proteins and their forms—may include a substantial layer of sequence diversity generated during gene expression, not only diversity written into DNA.

The evidence combines large-scale computational and proteogenomic detection with cell-line experiments and validation of selected purified proteins. Direct experimental support therefore applies to the tested substitutions, while the wider functional interpretation is based on recurrence across individuals, association with annotated protein sites and the observed response to amino-acid starvation. The published paper is an early-access version that may receive further edits before its final Version of Record.

Sources