PathTrove data atlas

What public pathology data exists, and what it can really be used for

355
data resources
30
anatomical sites
13,490
report entries
16,936
source quotations

Scale and growth

83% of resources were released in 2020 or later

Resources per release year. 2026 only covers January to June.

2025 alone added 73 resources, more than everything released before 2019.

18
13
28
21
33
35
41
66
73
21
≤2017201820192020202120222023202420252026 H1

Most resources are datasets; one in four comes from a challenge or benchmark

Challenges and benchmarks often reuse images from existing datasets.

Resource type
2616727
DatasetChallengeBenchmark

Coverage

12 sites have 30 or more resources each; 7 sites have fewer than 10

A resource can be tagged with several sites.

Breast
102
Lung
81
Colorectum
77
Brain
56
Kidney
56
Prostate
48
Liver
38
Lymph node
36
Skin
35
Bladder
35
Stomach
33
Cervix
31
Bone
24
Esophagus
22
Blood
21
Ovary
21
Pancreas
21
Adrenal gland
18
Heart
16
Uterus
14
Head and neck
12
Soft tissue
11
Bile duct
11
Eye
9
Spleen
8
Pleura
7
Testis
6
Thyroid
5
Thymus
4
GI tract
3
≥ 30 resources10–29 resources< 10 resources

58 resources involve 23 rare diseases (Orphanet codes); glioblastoma has the most

The 12 rare diseases with the most resources; a resource can involve several.

Glioblastoma
20
Cholangiocarcinoma
11
Thymoma
10
Diffuse large B-cell lymphoma
8
Uveal melanoma
7
Adrenocortical carcinoma
7
Multiple myeloma
7
Cervical squamous cell carcinoma
5
Ependymoma
4
Clear cell ovarian carcinoma
3
Endometrioid ovarian carcinoma
3
Mucinous adenocarcinoma
3

Cases come mainly from Europe, East Asia and North America

A resource can include cases from several regions. 117 resources do not say whether they are single- or multi-center.

Europe
≈90
East Asia
40
North America
35
South Asia
20
Five more regions (Southeast Asia, Latin America, Middle East, Africa, Oceania) have fewer than 10 resources each.
Single or multiple centers
114124
Multi-centerSingle-center

Image objects and modalities

From 3D images to patches, counts differ by four orders of magnitude

Counts are image instances, not patients.

1K
10K
100K
1M
10M
100M
3D images
1,475
15 resources
Field-of-view images
13,235
3 resources
TMA cores
25,353
7 resources
Whole-slide images
343,865
136 resources
ROI images
11.14M
84 resources
Patches
28.71M
81 resources
HorizontalDot area

Beyond H&E, six other kinds of stain, imaging and molecular data

A resource can contain several modalities.

H&E
268
Special stains
55
IHC / mIHC
52
IF / mIF
29
Molecular data
20
Spatial transcriptomics
13
CODEX
4

Among resources with molecular data, SurGen has the most patients and is fully open; sequencing data mostly requires an application

Bars show patients per resource; tags show how open the original data is.

Example of controlled access: PTRC-HGSOC has 158 patients, with 120 WGS and 106 RNA-seq samples deposited in dbGaP.

SurGen (colorectal, MSI/KRAS status)
843Fully open
BOEHMK (ovary, DNA/mutations)
444Partially open
SegLungTCGA (lung, DNA/mutations)
411Fully open
MSKMINDProjectM (lung, DNA/mutations)
247Partially open
PD-L1 (lung, DNA/mutations)
217Partially open
TransNEO (breast, DNA and RNA)
168Partially open
PTRC-HGSOC (ovary, WGS and RNA-seq)
158Partially open
FUSCC bladder cancer (WES)
75Partially open
AURORA (metastatic breast cancer, DNA and RNA)
55Partially open

Two aggregated sets are far larger than the newly acquired spatial transcriptomics data: 6 new cohorts total 207 samples

Samples per resource (slides for STimage-1K4M). The aggregated sets collect published data and can overlap with the newly acquired resources, so the numbers cannot be added.

13spatial transcriptomics resources
207samples in the 6 newly acquired resources
HEST-1k (multi-organ)
1,276aggregated set · samples
STimage-1K4M (multi-organ)
1,149aggregated set · slides
ST-Net (breast)
68newly acquired
GSE210616 (breast)
43newly acquired
her2st (breast)
36newly acquired
STHELAR (multi-organ)
31partly reuses existing data
CellHIST-Bench (multi-organ)
25newly acquired
GSE111672 (pancreas)
23newly acquired
CAMEO-Lung
23processed from existing data
CAMEO-Thymus
19processed from existing data
Human DLPFC Visium (brain)
12newly acquired
CAMEO-Breast
7processed from existing data
iSCALE (stomach, brain)
5partly reuses existing data

Among fluorescence and CODEX multiplex imaging resources, two exceed 1,000 patients

Bars show patients per resource.

SegPath (multi-organ)
1,583patients · 158,687 patches
cell-niches-data (lung)
1,168patients · 4,429 TMA cores
CRC FFPE CODEX (colorectal)
35patients · 251 ROIs
CODEX-HCC (liver)
15patients · 37 markers · 646 images

Pairing and tasks

138 resources link images across stains or modalities

Resources by type of correspondence.

Pixel-aligned images support registration and virtual staining; different images of the same case link morphology to omics or other images.

Pixel-aligned
43
Different images of the same case
30
Partly aligned regions
28
Algorithm-generated pairs
28
Several markers on one section
9

Classification and segmentation dominate; clinical and molecular prediction tasks are rare

Resources per officially declared task; a resource can declare several. Pink: clinical and molecular prediction tasks.

Classification
197
Segmentation
120
Detection
55
Survival
25
Regression
23
Generation
18
VQA
16
Retrieval
15
Registration
14
Counting
10
Captioning
8
Staining
6
Reasoning
4
Molecular prediction
3
Treatment response
3

Clinical and acquisition information

Along the patient timeline, data exists from diagnosis to follow-up outcomes

Each node is a separate resource count.

Diagnosispathology
63
Biomarkers / receptors
29
Stage
variable
23
Grade
variable
Treatmentbaseline, perioperative, therapy
21
Pre-treatment / baseline
4
Intra- / perioperative
36
Treatment
Follow-upafter treatment
53
Post-treatment / follow-up
21
Overall survival
outcome
21
Recurrence / relapse
outcome

187 resources have no structured clinical variables; 168 do, with 694 fields in total

Horizontal: clinical fields per resource; bar height: number of resources.

Each field's meaning and values are written in the report, so the table can be searched by variable.

187
43
45
50
18
12
012–34–67–1011+
Structured clinical fields per resource
19DLBCL-Morph (lymphoma): IHC, FISH, IPI, survival
12HMU-CRC-Hist550K (colorectal cancer): TNM, MSI, OS

239 resources record the scanner maker; 23 contain slides from several scanners

Resource counts; a resource can record several scanner makers.

Multi-scanner designs range from controlled rescans (Multi-Scanner SCC, skin squamous cell carcinoma: 44 samples × 5 scanners) to routine archives.

Leica / Aperio
79
3DHISTECH
38
Hamamatsu
37
Olympus
21
Zeiss
17
Nikon
11
Philips
5
Others: Akoya 4 · Huron 3 · KFBio 3 · Ventana 2
Specimen sampling and processing
FFPE
100
Surgical resection
68
Biopsy
66
Frozen section
15

Quality, origin and access

Quality control is mostly manual review; 92 resources do not describe any

78 resources name what they checked, most often annotation quality, then staining, focus, registration and label consistency.

Reported quality control
101766992
ManualManual + automatedOnly some steps documentedAutomated only 9Not described

Most labels are new, but 138 resources reuse existing labels

Image origin and label origin are recorded separately: new labels on reused images are common.

Where the labels come from
20710132
NewPartly reuse existing labelsProcessed from existing labelsReorganized existing labels 5
Where the images come from
205506228
Newly acquiredPartly reuse existing imagesProcessed from existing imagesReorganized existing images
140 resources reuse existing images and may share cases with the datasets they came from

185 fully open, 149 partially open, 14 not open; most downloads still need extra steps

How open the original data is
53.2%42.8%4.0%
Fully openPartially openNot open 4.0%

按本图已分类记录计算占比。

Download test from mainland China (September 2026)
44.6%45.1%8.2%2.2%
Direct download worksNeeds a proxyCannot downloadOther 2.2%
Of the 30 that cannot be downloaded, 19 first require registration or approval

按本图已分类记录计算占比。

Data is hosted on many platforms, not just a few big ones

Counted by each guide's primary hosting platform. "Other" covers Dataverse, Grand Challenge, NCI GDC, NCBI GEO, Mendeley Data, EBI BioStudies and institutional servers.

Zenodo
79
Hugging Face
50
Figshare
36
TCIA
28
GitHub
26
Google Drive
25
AWS S3
13
Kaggle
11
Other
87