Miscalibrated benchmarks may have understated deep learning perturbation models
A peer-reviewed paper finds common perturbation benchmarks miscalibrated; deep models beat baselines once metrics are fixed.
Nature Biotechnology published a peer-reviewed paper on October 1 finding that common benchmarking metrics are miscalibrated, and that deep-learning genetic perturbation models can beat uninformative baselines once metrics are calibrated.
Earlier benchmarks found a simple mean baseline matched or beat deep models on MAE, MSE and related metrics, casting doubt on in silico screening. This paper introduces positive and negative controls across 14 datasets and 18 metrics, showing the common metrics are often insensitive to real perturbation signal.
The authors also propose a calibration measure called dynamic range fraction (DRF) and released code. The result concerns relative performance under calibrated metrics, not whether the models yet reach usable predictive accuracy.
Sources:https://www.nature.com/articles/s41587-026-03307-w?code=4f3fcff5-c696-4b1b-8d0e-e14db5f5dc0b&error=cookies_not_supportedhttps://github.com/shiftbioscience/Perturbation-Models-Outperform-Baselineshttps://www.nature.com/articles/s41587-026-03307-w?code=b11dba9d-e21c-44d4-99ed-23e2f753ca58&error=cookies_not_supported