Summary

A bioRxiv preprint reports that the protein language model ESM-2 retains meaningful biological information in intrinsically disordered protein regions, despite assigning them less attention. The authors also derive structural features of these proteins from the model’s internal signals.

A Cambridge-led research team reports that the protein language model ESM-2 encodes useful information about intrinsically disordered proteins (IDPs), even though the model pays less attention to their sequences than to regions that form stable protein structures. The findings appear in a bioRxiv preprint posted on September 13, 2026.

The study is computational: the researchers examined signals inside the model, including attention patterns, logits and embeddings. They report that these signals contain information about disease-relevant residues and physical properties associated with disordered proteins.

Why disordered proteins are a challenge for language models

Protein language models process amino-acid sequences in a way that is conceptually similar to language models processing text. ESM-2 was trained with a masked-learning objective, in which evolutionary constraints in protein sequences help the model infer information about residues and their surrounding context. The model converts sequences into numerical representations called embeddings, which can then be used for other biological tasks.

Much previous interpretability work on these models has focused on folded proteins: molecules whose amino-acid chains form relatively stable three-dimensional structures. IDPs behave differently. Rather than adopting one fixed structure, they can occupy changing ensembles of shapes. Their sequences therefore experience different evolutionary constraints from those of folded regions. IDPs make up a substantial fraction of the human proteome and are implicated in numerous diseases, according to the preprint.

The authors hypothesised that a model trained on evolutionary patterns would treat disordered and folded regions differently. Their analysis supports that expectation: ESM-2 shows reduced attention on disordered regions. In this context, attention refers to how strongly the model weights relationships among positions in a sequence while processing it.

Signals remain visible inside disordered regions

Reduced attention did not mean that the model contained no useful information about IDPs. The researchers report heightened attention at disease-relevant residues, including when those residues occur in regions with high levels of disorder. This suggests that the model’s internal representation distinguishes some biologically important positions even when the surrounding protein sequence lacks a stable structure.

The team also reports recovering two physical characteristics of IDPs from ESM-2’s internal signals. The first is the radius of gyration, a measure of how widely a molecule’s mass is distributed around its centre. The second is an individual dynamic contact map, which describes changing contacts between parts of a protein as its shape shifts. The authors obtained these properties from the model’s logits—its pre-output scores—and embeddings.

Together, the results indicate that a sequence-trained model can encode information related not only to residue patterns, but also to aspects of the dynamic physical behaviour associated with disordered proteins. That makes mechanistic interpretability useful here: instead of treating ESM-2 only as a tool that produces predictions, researchers can inspect its internal signals to ask what biological properties those signals represent.

What the finding shows

The preprint provides evidence that ESM-2 contains recoverable information about IDP structure and function despite a bias toward more structured residues. For protein biology, this matters because disordered regions are important in cellular regulation and disease but are harder to describe using a single static structure.

The result is an early model-interpretability finding rather than a clinical or therapeutic result. The work concerns information encoded in ESM-2’s representations; it does not report a patient study or a treatment intervention. The record is a bioRxiv preprint, and the supplied abstract does not provide quantitative benchmark values, sample sizes or experimental validation. How consistently the findings generalise to other protein language models or to broader experimental prediction tasks is not stated in the supplied record.

Sources