Benchmark results
Leaderboard
Rank models across the included datasets using each dataset's default split. Choose an overall, prediction-task, or dataset view.
- Complete overall rankings
- --
- Models evaluated
- --
- Datasets
- --
- Model-target summaries
- --
Loading coverage summary.
All submitted datasets
Overall ranking
Mean standardized metric across source datasets. Higher is better.
No models match that search.
How ranking works
From split seeds to mean z-score
- 01
Standardize each target
Within each seed, dataset, and target, model performance is converted to a z-score. Pearson correlations receive a Fisher transform first.
- 02
Combine related results
Within each seed, z-scores are averaged across targets, then across related sub-datasets such as cell lines or species.
- 03
Average split seeds
Each model is evaluated on ten train-test splits. The aggregated z-scores are then averaged across those seeds.
- 04
Compute the overall score
Complete models are ranked by their mean z-score across the 11 source datasets represented by the 19 registered dataset IDs.
Standardization makes AUPRC and Pearson correlation comparable only as relative performance within the same target. The overall score is the manuscript's mean standardized metric, not an absolute measure of biological performance. Regression uses cross-validated Ridge results.