Results explorer

Every method, every tissue, all thirteen metrics. All values are oriented so that higher is better.

Select All tissues to see each metric averaged across tissues alongside the mean aggregate rank, or pick a single tissue to see its raw values and that tissue’s aggregate ranks. Click any column heading to sort by it.

Loading the benchmark tables…

What the columns mean

Aggregate rank (one tissue) is the manuscript’s per-tissue figure, lower is better: every method is ranked on each of the twelve scored metrics among all 20 methods, the ranks are averaged within the four families, and the family ranks are weighted — cluster-label concordance 0.325, local cell-type neighbourhood structure 0.325, batch mixing 0.25, continuous cell-type separation 0.10 — with an undefined batch family dropped and the weights renormalised.

Mean aggregate rank (all tissues) is that figure averaged over the 26 tissues, the headline number of the results page. Ranks are always computed among all 20 methods; the Method type filter hides rows and does not re-rank.

Biological conservationARI, NMI, HOM, COM, FMI (cluster agreement, from a Leiden resolution scan), ASW (geometric separation), Acc@kNN, cLISI, GC (local neighbourhood structure).

Batch mixingkBET, BRAS, CiLISI. Defined only on the 19 tissues with more than one technical batch; blank () elsewhere.

iLISI is shown for transparency but is excluded from the overall score. It measures nearly the same thing as CiLISI, and including both would double-weight LISI-style mixing against kBET and BRAS.

Full definitions are on the evaluation metrics page.

Caveats

Warning

Integration methods consume batch labels; zero-shot methods never see them. They are ranked in the same table because the manuscript ranks them together, but a gap between the two kinds measures what the batch labels buy, not model quality alone. Use the Method type filter to look at one kind at a time.

Rows marked baseline are the reference points rather than foundation models: PCA among the zero-shot methods; scVI trained de novo, Harmony and Seurat among the integration methods. The other scVI row is the CELLxGENE Census model pretrained on tens of millions of cells and run zero-shot — a pretrained model, not a baseline. A foundation model that does not clear the relevant baseline has not earned its GPU hours.

Rows marked community, when present, are methods whose numbers were supplied by their contributor under the published protocol and have not yet been reproduced by the maintainers — see Contributing a method. None of the rows on this page carry that badge.

Single-tissue values are raw metric values and are not comparable across tissues — a tissue with more cell types produces systematically lower cluster-agreement scores whatever the method. That is why the aggregated view ranks within tissues before averaging.

See also