Summary

A UCSF-led bioRxiv preprint introduces CELL-FM, a bidirectional generative model that connects protein sequences and cellular context with fluorescence microscopy images. The authors use it for virtual localization, motif analysis and mutagenesis of intrinsically disordered peptides.

A research team at the University of California, San Francisco has proposed a generative modelling framework that connects protein sequence information with fluorescence microscopy images. The system, called CELL-FM, is designed to recreate experimental readouts from a protein sequence and cellular context, then use those synthetic readouts for downstream biological analysis.

The work appears as a bioRxiv preprint posted on 18 September 2026. In an illustrative application, the authors use CELL-FM for in-silico protein-localisation prediction, functional motif analysis and large-scale virtual mutagenesis.

From experimental readouts to virtual experiments

Large screening projects can measure how genetic or other perturbations affect cells, producing libraries that combine inputs such as sequences with readouts such as microscopy images. Turning those measurements into biological mechanisms often requires a separate predictive task for each phenotype or label.

The approach described by the UCSF researchers changes that workflow. Instead of training a model only to predict a particular label, they train a generative model to recreate the richer experimental readout while conditioning it on the relevant experimental context. Established analysis models can then be applied to the synthetic data.

This separates representation learning—the process of learning useful structure from the data—from the creation of a task-specific annotation. It also allows one experimental modality to support several later analyses.

How CELL-FM connects sequence and images

CELL-FM is described as a bidirectional sequence–image generative framework. In one direction, it maps a protein sequence together with cellular context to a fluorescence microscopy image. In the other direction, it uses image information to connect observed cellular organisation back to sequence-level analysis.

Fluorescence microscopy can show where molecules are located and how they are organised inside cells. Those spatial patterns contain more information than a single hand-crafted label such as “present” or “absent” in a particular compartment. By modelling images directly, the framework is intended to retain details about cellular organisation that may be discarded when images are reduced to predefined categories.

Virtual mutagenesis of disordered peptides

One application involved intrinsically disordered peptides. Unlike many compact proteins, intrinsically disordered sequences do not maintain one fixed three-dimensional structure under ordinary cellular conditions. Their sequence can still influence how they interact and organise inside cells.

The authors used virtual mutagenesis to alter amino-acid features computationally and generate predicted microscopy readouts for the resulting sequences. They report that this analysis revealed amino-acid features controlling condensate formation in the intrinsically disordered peptides studied.

A condensate is a concentrated, organised cellular assembly formed when molecules gather through biochemical interactions, often without a surrounding membrane. Predicting how sequence changes affect condensate formation can help connect molecular composition with visible cellular behaviour.

Why the framework matters

The central idea is to treat a model-generated image as a reusable virtual experiment rather than as the final prediction. The same synthetic readout can potentially be examined for localisation, morphology or other phenotypes with different downstream tools. For large-scale sequence exploration, this provides a route to test many hypothetical mutations computationally before selecting experiments for laboratory measurement.

The preprint presents CELL-FM as an illustrative framework rather than a clinical system. It is a computational method built around biological imaging data, and the supplied abstract does not specify quantitative accuracy across different cell types, proteins or experimental conditions. The findings are also reported in a bioRxiv preprint and therefore are presented before peer review.

Sources