
An AI model of gene expression needs to beat a useful comparison on a declared task. The comparison can be surprisingly simple: give every intervention the same average response, or combine two measured effects by addition. A large model earns its place by improving the prediction researchers need.
The dispute concerns the scoring instrument too. If a score cannot distinguish a prediction containing an experimental signal from an average with no intervention-specific information, failure to improve that score has an ambiguous interpretation. The prediction and the scoring method both need examination.
What the 2025 benchmark established
Ahlmann-Eltze, Huber and Anders compared five foundation models and two further deep-learning approaches in Nature Methods on 4 August 2025. Their tests concerned single and double genetic perturbations. Some foundation models were adapted to this task through a decoder, which converts the model's representation into predicted expression.
For combinations in one experiment, all measured single interventions were available for training, alongside half the measured pairs. The other half formed the test. The models failed to beat an additive prediction. Tests of unseen single perturbations likewise found no consistent improvement over simple mean or linear predictions. The authors examined additional gene subsets and scores, so the criticism rested on several comparisons.
The study used four datasets from cancer cell lines. That scope supports a criticism of the tested methods under those conditions. It cannot determine the performance of every later model or of a different biological population.
What a whole-profile score can reward
Consider an invented example with 1,000 measured genes. An intervention changes ten strongly; the remaining 990 barely move. A prediction that reproduces the quiet majority can obtain a small average error while missing every strongly changed gene. A second prediction might capture those ten responses yet make small errors across many unchanged genes. Which prediction scores better depends on how the errors are counted.
There is nothing mysterious about that arithmetic. Mean squared error squares each prediction error, adds the squared errors and divides by the number of measurements. It answers an average numerical question. Choosing the measurements, their scale and their weights changes the question. An experimenter seeking a response in a small pathway needs to know whether the score can detect improvement in that pathway.
A weighted score gives selected errors more influence. A retrieval score can ask whether a predicted response identifies the correct intervention among alternatives. These approaches need their own checks. A model should not win by exaggerating a few changes, producing impossible profiles or ignoring the rest of the assay.
The October 2026 calibration study
Miller and colleagues analysed 14 datasets and 18 metrics in Nature Biotechnology on 1 October 2026. They tested whether a scoring method separates an uninformative mean from a positive control containing measured perturbation signal. Their positive-control construction combines information from split experimental cells with the mean; it is a calibration reference, not a model predicting an unknown experiment.
Under metrics with better calibration, tested deep-learning models could outperform uninformative baselines. The study addressed unseen interventions and unseen combinations. It left unseen cellular contexts untested.
The authors disclose employment at Shift Bioscience and an AI leadership role at Xaira Therapeutics. Their methods deserve assessment on their merits, with those interests visible.
How to read the two results together
Both papers support keeping simple baselines. The later paper adds a demand to check whether the scoring method is sensitive to the biological signal sought. It does not provide a controlled ranking of every model published since 2025, and a win against an uninformative mean should not be restated as a win against every linear approach.
Name what was withheld, the information available during training and the experimental population. Keep absolute profile error beside tests of changed genes or intervention identity. Show uncertainty across perturbations, because a mean score can hide a model that succeeds on a few responses and fails elsewhere.
An independently measured follow-up would answer a further question: did the prediction lead researchers to an intervention that produced the intended function? A score on RNA measurements and a functional benefit are separate results.
Scroll the table sideways to read all columns.
| Check | What to look for |
|---|---|
| Training boundary | Which interventions, combinations and contexts were absent? |
| Negative comparison | A declared mean, no-change, additive or linear prediction appropriate to the task. |
| Metric sensitivity | A positive control that the score should recognise, alongside safeguards against implausible outputs. |
| Biological outcome | The measured assay and any separate functional experiment. |
| Reproducibility | Data, code, exact model version and a test set kept outside development. |
Transfer to unfamiliar cells remains a test
Arc's ongoing 2026 challenge asks models to predict responses in cell lines without matching perturbation examples. That addresses the context question excluded from the October calibration study. Its final test release is scheduled for 22 October. The final results are not yet available.
Sources
- Paper · 4 Aug 2025Deep-learning-based gene perturbation effect prediction does not yet outperform simple linear baselines
Full primary account checked, including task splits, alternative scoring, four-dataset cancer-cell scope and stated limitations. Authors declare no competing interests.
Checked 4 Oct 2026 - Paper · 1 Oct 2026Deep learning perturbation models can outperform baselines on calibrated metrics
Primary full text checked. Calibration uses idealised positive controls; unseen contexts remain unaddressed. Disclosed interests include Shift Bioscience employees and Xaira Therapeutics' Chief AI Scientist.
Checked 4 Oct 2026 - Institution · 20 Aug 2026The 2026 Virtual Cell Challenge: predicting perturbation responses in cell contexts a model has never seen
Official ongoing challenge design. The final test and announcement are future events as checked on 4 October.
Checked 4 Oct 2026
Dr T Smith, organic chemist and science educator. Report a correction.