Miscalibrated benchmarks may have understated deep learning perturbation models
Common perturbation benchmarks are miscalibrated; deep models beat baselines once metrics are fixed.
BedeutungLokalBeweisE3 überprüfbarAufbereitungSchnell
Readers can now know that common genetic perturbation benchmarking metrics are miscalibrated: with positive and negative controls introduced across 14 datasets and 18 metrics, the metrics prove often insensitive to real perturbation signal, and deep-learning models can beat uninformative baselines once metrics are calibrated.
Earlier benchmarks found a simple mean baseline matched or beat deep models on MAE, MSE and related metrics, casting doubt on in silico screening.
A peer-reviewed paper published in Nature Biotechnology on October 1 proposes a calibration measure called dynamic range fraction (DRF) and released code. The result concerns relative performance under calibrated metrics, not whether the models yet reach usable predictive accuracy.