ExplainerAI and biological discovery

What is a virtual cell, and what can it predict?

A virtual-cell model predicts a measurement after researchers change a cell. Its usefulness depends on which experiment it can predict, which examples it has already seen and whether the prediction survives a new biological context.

Five illustrated cells connected to a panel containing simplified computational representations.
Measurements from cells supply the data for a computational model.

A virtual cell is a computational model of a specified cellular behaviour. In the research discussed here, it predicts how gene expression changes after a genetic, chemical or signalling intervention. Researchers can compare those predictions with measurements from physical cells.

A model trained to predict RNA measurements gives an answer about RNA. Protein activity, metabolism and tissue function each need their own data and experimental test.

Imagine a researcher deciding which genes to suppress in an ageing cell culture. A model could rank proposed interventions by their predicted effect on an expression pattern. That would help choose experiments. Whether the cells repair damage, divide appropriately or perform their specialised job remains a question for those experiments. This is an illustrative use, not a reported rejuvenation result.

SchematicA virtual-cell prediction meets a held-out experiment
  1. Measure cellular responsesInterventions, context and assay define the examples.
  2. Predict a withheld responseState what the model has not seen during training.
  3. Compare with physical cellsScore the measured assay; test function separately.

Predicting RNA measurements does not supply an unmeasured tissue outcome. Sources for this account.

The experiment behind the training data

Perturb-seq combines genetic interventions with RNA sequencing of individual cells. In the 2016 work by Dixit and colleagues, expressed barcodes connected a cell's RNA measurements with its assigned CRISPR perturbation. Pooling many interventions allowed researchers to examine responses across large numbers of cells.

Each training example links an intervention and cellular context to measured responses. A prediction model needs to learn how those parts relate. The 2022 genome-scale study by Replogle and colleagues extended this experimental approach across expressed genes and supplied data used in subsequent model evaluations.

An expression profile records which RNA molecules the assay detected and their measured abundance. Different cells receiving the same intervention can produce different profiles. The distribution of responses therefore deserves attention alongside the average. A small population responding strongly can disappear in an average dominated by cells that hardly respond.

Suppressing expression with CRISPR interference, activating expression and removing a gene create different interventions. Measurement time affects what the assay captures. A model's answer should name the intervention and the time represented by its training data.

What State attempts

Arc's State paper appeared online in Cell on 31 August 2026, following its earlier preprint. The authors report improved prediction against their chosen baselines across genetic, chemical and signalling datasets.

Arc's tool description separates an embedding component, which represents individual cells, from a transition component operating on sets of cells. The design addresses variation within a population and between contexts. These are defined evaluations of a particular model and dataset collection. They do not establish reliable prediction of every response in an arbitrary human tissue.

A developer evaluation is useful evidence, especially when the methods and comparison code are available. An external test asks a further question: will the result persist when another group chooses the experiments, keeps their results hidden and applies a declared scoring method?

Four ways to make a prediction test harder

A test should explain what the model has never seen. Randomly withholding some cells from an experiment leaves the model much closer to the answer than withholding the entire intervention. Changing the cellular context adds another demand.

An editorial map of prediction tests; these are different claims, not interchangeable measures of accuracy.

Scroll the table sideways to read all columns.

An editorial map of prediction tests; these are different claims, not interchangeable measures of accuracy.
Held out from trainingWhat the test asksWhat a good result leaves open
Some cells from a measured conditionCan the model reproduce more cells from a familiar experiment?A new intervention or cell type.
A genetic interventionCan it predict a response to an intervention omitted from training?Transfer to another cellular context.
A combination of interventionsCan it predict their combined response from the available examples?Combinations with different components or stronger interactions.
A cellular contextCan it predict responses in cells without matching perturbation examples?Other donors, environments and functional outcomes.

A public test in unfamiliar cells

Arc's 2026 Virtual Cell Challenge asks entrants to predict CRISPR-interference responses across six cell lines. Participants receive unperturbed profiles and target-gene identifiers, while the matching experimental responses stay hidden. Three lines support validation; three are reserved for final testing. The final scoring combines six metrics.

As checked on 4 October, the final test release is scheduled for 22 October and submissions for 5 November. Final results are not yet available. The test asks whether models can predict responses to familiar interventions in cellular contexts lacking matching response examples.

For longevity research, the useful next step is an experiment selected because of a prediction, followed by a functional result in the stated model. A paper can show that the model narrowed a search even when most candidates fail. Report the failures and the number actually tested; those figures tell readers what was gained.

Sources

  1. Paper · 15 Dec 2016Perturb-Seq: Dissecting Molecular Circuits with Scalable Single-Cell RNA Profiling of Pooled Genetic Screens

    Primary pooled genetic-screen and single-cell RNA measurement study; methods and publication record checked.

    Checked 4 Oct 2026
  2. Paper · 9 Jun 2022Mapping information-rich genotype-phenotype landscapes with genome-scale Perturb-seq

    Primary genome-scale Perturb-seq study; online article date precedes the 7 July issue date. Publication record checked; no numerical result is reconstructed here.

    Checked 4 Oct 2026
  3. Paper · 31 Aug 2026Predicting cellular responses to perturbation across diverse contexts with State

    Peer-reviewed Cell article, available online as a corrected proof. Publisher summary and journal metadata checked; no claim of a complete supplementary-methods audit.

    Checked 4 Oct 2026
  4. Institution · 26 Jun 2025Arc Virtual Cell Model: State

    Institution-maintained description of State's embedding and transition components. Date refers to initial model release, not a dated update of this maintained page.

    Checked 4 Oct 2026
  5. Institution · 20 Aug 2026The 2026 Virtual Cell Challenge: predicting perturbation responses in cell contexts a model has never seen

    Official description of the held-out contexts, measurements, evaluation and planned October/November dates. Final results are still future events.

    Checked 4 Oct 2026
Editorial responsibility

Dr T Smith, organic chemist and science educator. Report a correction.

Search the publication

Search programme histories and research articles.