InvestigationAI and biological discovery

How should we test an AI model of a cell?

A 2025 benchmark found that tested deep-learning models failed to beat simple predictions. An October 2026 study shows that scoring choices can conceal useful performance. Their comparison turns on the task, the controls and what a metric rewards.

A cell population beside three different computational representations.
Different models can predict responses in the same cell population. The test determines which predictions are useful.

An AI model of gene expression needs to beat a useful comparison on a declared task. The comparison can be surprisingly simple: give every intervention the same average response, or combine two measured effects by addition. A large model earns its place by improving the prediction researchers need.

The dispute concerns the scoring instrument too. If a score cannot distinguish a prediction containing an experimental signal from an average with no intervention-specific information, failure to improve that score has an ambiguous interpretation. The prediction and the scoring method both need examination.

What the 2025 benchmark established

Ahlmann-Eltze, Huber and Anders compared five foundation models and two further deep-learning approaches in Nature Methods on 4 August 2025. Their tests concerned single and double genetic perturbations. Some foundation models were adapted to this task through a decoder, which converts the model's representation into predicted expression.

For combinations in one experiment, all measured single interventions were available for training, alongside half the measured pairs. The other half formed the test. The models failed to beat an additive prediction. Tests of unseen single perturbations likewise found no consistent improvement over simple mean or linear predictions. The authors examined additional gene subsets and scores, so the criticism rested on several comparisons.

The study used four datasets from cancer cell lines. That scope supports a criticism of the tested methods under those conditions. It cannot determine the performance of every later model or of a different biological population.

What a whole-profile score can reward

Consider an invented example with 1,000 measured genes. An intervention changes ten strongly; the remaining 990 barely move. A prediction that reproduces the quiet majority can obtain a small average error while missing every strongly changed gene. A second prediction might capture those ten responses yet make small errors across many unchanged genes. Which prediction scores better depends on how the errors are counted.

There is nothing mysterious about that arithmetic. Mean squared error squares each prediction error, adds the squared errors and divides by the number of measurements. It answers an average numerical question. Choosing the measurements, their scale and their weights changes the question. An experimenter seeking a response in a small pathway needs to know whether the score can detect improvement in that pathway.

A weighted score gives selected errors more influence. A retrieval score can ask whether a predicted response identifies the correct intervention among alternatives. These approaches need their own checks. A model should not win by exaggerating a few changes, producing impossible profiles or ignoring the rest of the assay.

The October 2026 calibration study

Miller and colleagues analysed 14 datasets and 18 metrics in Nature Biotechnology on 1 October 2026. They tested whether a scoring method separates an uninformative mean from a positive control containing measured perturbation signal. Their positive-control construction combines information from split experimental cells with the mean; it is a calibration reference, not a model predicting an unknown experiment.

Under metrics with better calibration, tested deep-learning models could outperform uninformative baselines. The study addressed unseen interventions and unseen combinations. It left unseen cellular contexts untested.

The authors disclose employment at Shift Bioscience and an AI leadership role at Xaira Therapeutics. Their methods deserve assessment on their merits, with those interests visible.

How to read the two results together

Both papers support keeping simple baselines. The later paper adds a demand to check whether the scoring method is sensitive to the biological signal sought. It does not provide a controlled ranking of every model published since 2025, and a win against an uninformative mean should not be restated as a win against every linear approach.

Name what was withheld, the information available during training and the experimental population. Keep absolute profile error beside tests of changed genes or intervention identity. Show uncertainty across perturbations, because a mean score can hide a model that succeeds on a few responses and fails elsewhere.

An independently measured follow-up would answer a further question: did the prediction lead researchers to an intervention that produced the intended function? A score on RNA measurements and a functional benefit are separate results.

Questions to ask before accepting a perturbation-model comparison.

Scroll the table sideways to read all columns.

Questions to ask before accepting a perturbation-model comparison.
CheckWhat to look for
Training boundaryWhich interventions, combinations and contexts were absent?
Negative comparisonA declared mean, no-change, additive or linear prediction appropriate to the task.
Metric sensitivityA positive control that the score should recognise, alongside safeguards against implausible outputs.
Biological outcomeThe measured assay and any separate functional experiment.
ReproducibilityData, code, exact model version and a test set kept outside development.

Transfer to unfamiliar cells remains a test

Arc's ongoing 2026 challenge asks models to predict responses in cell lines without matching perturbation examples. That addresses the context question excluded from the October calibration study. Its final test release is scheduled for 22 October. The final results are not yet available.

Sources

  1. Paper · 4 Aug 2025Deep-learning-based gene perturbation effect prediction does not yet outperform simple linear baselines

    Full primary account checked, including task splits, alternative scoring, four-dataset cancer-cell scope and stated limitations. Authors declare no competing interests.

    Checked 4 Oct 2026
  2. Paper · 1 Oct 2026Deep learning perturbation models can outperform baselines on calibrated metrics

    Primary full text checked. Calibration uses idealised positive controls; unseen contexts remain unaddressed. Disclosed interests include Shift Bioscience employees and Xaira Therapeutics' Chief AI Scientist.

    Checked 4 Oct 2026
  3. Institution · 20 Aug 2026The 2026 Virtual Cell Challenge: predicting perturbation responses in cell contexts a model has never seen

    Official ongoing challenge design. The final test and announcement are future events as checked on 4 October.

    Checked 4 Oct 2026
Editorial responsibility

Dr T Smith, organic chemist and science educator. Report a correction.

Search the publication

Search programme histories and research articles.