从临床研究需求, Toward Ready-to-Use Clinical Data Services 到开箱即用的病理数据 for Computational Pathology
PathTrove 提供什么 What PathTrove provides
用实际需求查找数据 Find data from your actual needs
开展深度可行性研究 Run in-depth feasibility studies
355
30
105
找到相关资源,也看清研究条件。 Find relevant resources, and see their research conditions.
检索与可行性研究都建立在这套数据库上。
Search and feasibility studies are built on this database.
30 个解剖部位都有数据,常见癌种之外也有专科部位 All 30 anatomical sites have data, including specialist sites beyond the common cancers
沿着病人的诊疗时间线,从诊断到随访结局都能找到数据 Along the patient timeline, data exists from diagnosis to follow-up outcomes
常规病理图像之外,还有哪些数据? Beyond routine pathology images, what other data is there?
数据能支持什么,逐项查清楚 What the data can support, checked item by item
从读懂数据,到准备实验。 From understanding the data to preparing experiments.
逐项核对数据条件,
让资源可以比较、结论可以追溯。 Every data condition checked,
so resources can be compared and conclusions traced.
让资源可以比较、结论可以追溯。
so resources can be compared and conclusions traced.
每条附来源原文 Every entry quotes its source
来源不一致时两边都保留 When sources disagree, both are kept
区分研究描述与实际文件 Study descriptions are separated from the actual files
研究任务 Task 由 H&E 图像块预测 MSIH / nonMSIH。 Predict MSIH / nonMSIH from H&E tiles. 图像条件 Image conditions 颜色归一化肿瘤图像块;512 × 512 像素,0.5 μm/像素。 Color-normalized tumor tiles; 512 × 512 pixels at 0.5 μm/pixel. 官方划分 Official split TRAIN:281 名患者、19,557 个图像块。
TEST:142 名患者、32,361 个图像块。
训练集为平衡类别做了下采样,测试集没有,因此测试集图像块更多。TRAIN: 281 patients, 19,557 tiles.
TEST: 142 patients, 32,361 tiles.
The training set was downsampled to balance the classes and the test set was not, so the test set has more tiles.两种口径 Two counts 论文记载 TCGA 子队列 426 例;公开压缩包清单识别出 423 个唯一患者条码。报告同时记录两者。 The paper records 426 patients in the TCGA subcohort; the public archive listings contain 423 unique patient barcodes. The report records both. 临床信息 Clinical information 论文附表只有年龄、性别、分期等队列汇总,没有逐患者临床表。 The paper's supplementary tables give only cohort summaries such as age, sex and stage; there is no per-patient clinical table. 证据依据 Evidence Zenodo 3832231、官方 ZIP 文件清单、Echle 等(Gastroenterology,2020)及补充表。 Zenodo 3832231, the official ZIP file listings, Echle et al. (Gastroenterology, 2020) and its supplementary tables.
从上游平台出发,
看清数据如何组织和获取。 Start from the upstream platform,
see how its data is organized and accessed.
看清数据如何组织和获取。
see how its data is organized and accessed.
数据源概览与项目组成 Source overview and project composition 图像、临床与组学的关联路径 How images, clinical data and omics are linked 官方访问入口与获取条件 Official entry points and access conditions
先理解组织关系,再定位研究材料 Understand how things are organized, then locate research material
覆盖内容 Coverage 综合报告梳理 33 个 TCGA 项目,涉及临床、样本、病理图像和多组学数据。 The combined report covers 33 TCGA projects, spanning clinical data, samples, pathology images and multi-omics. 关联方式 How things link 从项目与病例进入样本和文件,核对图像与目标分子或临床信息是否可以对应。 Go from project and case to sample and file, and check whether images can be matched to the target molecular or clinical information. 获取条件 Access 区分公开文件与受控文件,记录 GDC 及相关下游平台入口。 Open and controlled files are kept apart, and the entry points of GDC and related downstream platforms are recorded. 转成任务前还需做什么 Before it becomes a task 项目级数据没有统一的视觉任务标签和固定实验划分,需要根据研究问题进一步整理。 Project-level data has no unified visual task labels or fixed splits; these have to be organized for each research question.
以数据为中心,
重新读懂一篇研究。 Re-reading a study,
centered on its data.
重新读懂一篇研究。
centered on its data.
每项任务的数据、标签、划分与结果 Data, labels, splits and results for each task 数据集的来源、角色与公开范围 Where each dataset comes from, its role and how open it is 结果的统计单位与复用时的注意事项 The unit behind each result and caveats for reuse
空间转录组与图学习,如何用于脑肿瘤诊断? How are spatial transcriptomics and graph learning used for brain tumor diagnosis?
Spatially resolved transcriptomics and graph-based deep learning improve accuracy of routine CNS tumor diagnostics
拆到具体任务 Broken into tasks 8 项分析:前 4 项(质控、虚拟免疫组化、拷贝数、细胞组成)不训练模型;后 4 项(区域分割、甲基化亚类、MGMT、CDKN2A/B)训练图神经网络,按患者划分数据。 8 analyses: the first 4 (quality control, virtual immunohistochemistry, copy number, cell composition) train no model; the last 4 (region segmentation, methylation subclass, MGMT, CDKN2A/B) train graph neural networks with patient-level splits. 还原数据分工 Data roles NePSTA 的 107 个样本是主数据;Ren 等的 11 个样本只作质控对照;GBMap 是细胞类型字典,不是测试集。 NePSTA's 107 samples are the main data; the 11 samples from Ren et al. are only a quality-control reference; GBMap is a cell-type dictionary, not a test set. 核对划分与结果 Splits and results MGMT 预测:一个队列 30 人训练,另一队列 23 人评估。23 人上的患者级准确率不可能是 99%,论文报告的 99% 应为子图级。 MGMT prediction: trained on 30 patients from one cohort, evaluated on 23 from another. With 23 patients, patient-level accuracy cannot be 99%; the reported 99% should be at the subgraph level. 核查复用条件 Reuse conditions 数据可用性声明只涉及验证队列,不能据此推定训练队列也已公开。 The data availability statement covers only the validation cohort; it cannot be taken to mean the training cohort is public too.
大型数据集,
拆成可以执行的获取步骤。 Large datasets,
broken into steps you can run.
拆成可以执行的获取步骤。
broken into steps you can run.
# 1 表格:三个小文件,直接下载
B=https://www.cancerimagingarchive.net/wp-content/uploads
M=https://pathdb.cancerimagingarchive.net/system/files/collectionmetadata/202405
curl -fSLO "$B/new_CA125-data_20230207.xlsx"
curl -fSLO "$B/Final-patient_list.xlsx"
curl -fSL -o manifest.csv \
"$M/Ovarian%20Bevacizumab%20Response_05-28-2024.csv"
# 2 切片:Aspera 命令行整包接收(需 Ruby ≥ 3.1)
gem install aspera-cli && ascli conf ascp install
PKG='https://faspex.cancerimagingarchive.net/aspera/faspex/public/package?context=eyJyZXNvdXJjZSI6InBhY2thZ2VzIiwidHlwZSI6ImV4dGVybmFsX2Rvd25sb2FkX3BhY2thZ2UiLCJpZCI6Ijc0NCIsInBhc3Njb2RlIjoiMWZiZTNmMzgwNmY1MTEyNjBlODkxYjk0MWVjYzdkZTY2MGQwZGNkYSIsInBhY2thZ2VfaWQiOiI3NDQiLCJlbWFpbCI6ImhlbHBAY2FuY2VyaW1hZ2luZ2FyY2hpdmUubmV0In0='
ascli faspex5 packages receive --url="$PKG" \
--to-folder=./ovarian_bev/slides \
--http-options=@json:'{"ssl_options":["IGNORE_UNEXPECTED_EOF"]}' \
--ts=@json:'{"target_rate_kbps":5000,"min_rate_kbps":100}' # 1 Tables: three small files, direct download
B=https://www.cancerimagingarchive.net/wp-content/uploads
M=https://pathdb.cancerimagingarchive.net/system/files/collectionmetadata/202405
curl -fSLO "$B/new_CA125-data_20230207.xlsx"
curl -fSLO "$B/Final-patient_list.xlsx"
curl -fSL -o manifest.csv \
"$M/Ovarian%20Bevacizumab%20Response_05-28-2024.csv"
# 2 Slides: receive the whole package with the Aspera CLI (needs Ruby ≥ 3.1)
gem install aspera-cli && ascli conf ascp install
PKG='https://faspex.cancerimagingarchive.net/aspera/faspex/public/package?context=eyJyZXNvdXJjZSI6InBhY2thZ2VzIiwidHlwZSI6ImV4dGVybmFsX2Rvd25sb2FkX3BhY2thZ2UiLCJpZCI6Ijc0NCIsInBhc3Njb2RlIjoiMWZiZTNmMzgwNmY1MTEyNjBlODkxYjk0MWVjYzdkZTY2MGQwZGNkYSIsInBhY2thZ2VfaWQiOiI3NDQiLCJlbWFpbCI6ImhlbHBAY2FuY2VyaW1hZ2luZ2FyY2hpdmUubmV0In0='
ascli faspex5 packages receive --url="$PKG" \
--to-folder=./ovarian_bev/slides \
--http-options=@json:'{"ssl_options":["IGNORE_UNEXPECTED_EOF"]}' \
--ts=@json:'{"target_rate_kbps":5000,"min_rate_kbps":100}' 切片没有普通下载链接,只经 Aspera 分发;命令行代替网页插件,可在服务器上直接运行。 Slides have no ordinary download link and are distributed only through Aspera; the command line replaces the browser plug-in and runs directly on a server. 只能整包接收:指定子目录会返回服务器 500 错误。 Only the whole package can be received: asking for a subfolder returns a server 500 error. 不加 --ts速率参数,传输会停在 0 字节。Without the --tsrate parameter, the transfer stays at 0 bytes.中断后重跑同一条命令即续传;包内若有官方 MD5 清单( *.sums),用md5sum -c逐文件核对。After an interruption, rerun the same command to resume; if the package includes an official MD5 list ( *.sums), check each file withmd5sum -c.
样本、标签与划分,
在同一份文件里对应。 Samples, labels and splits,
matched in one file.
在同一份文件里对应。
matched in one file.
类别名称与官方文件一致 Class names match the official files 官方给了划分就沿用,没有就如实留空 Official splits are kept when provided; otherwise the split is left empty 与官方页面不一致的计数、已移除的文件,在说明中注明 Counts that differ from the official page, and removed files, are noted in the description
TCIA · new_CA125-data_20230207.xlsx
工作表 Ovary.effective-162 → 160 行 工作表 Ovary.invalid-126 → 126 行 列:Patient ID · Treatment effect · Image No. Sheet Ovary.effective-162 → 160 rows Sheet Ovary.invalid-126 → 126 rows Columns: Patient ID · Treatment effect · Image No.
{
"sample_id": "1612595C.svs",
"image_name": "1612595C.svs",
"split": null,
"label": {
"bevacizumab_treatment_effectiveness": {
"class_name": "effective",
"class_id": 0
}
}
}sample_idsplitlabel从研究需求,到可以开始实验的数据。 From a research need to data you can start experimenting with.
再把其中一项的数据准备到可以直接使用。
then prepare the data for one of them until it is ready to use.
用病理图像预测卵巢癌患者对药物的反应,公开数据能支持哪些评估? Predicting how ovarian cancer patients respond to drugs from pathology images: which evaluations can public data support?
铂类化疗 Platinum chemotherapy
贝伐珠单抗 Bevacizumab
PARP 抑制剂 PARP inhibitors
论文报告:把“药物敏感性”拆成 8 项可以评估的任务。 Paper reports: break “drug sensitivity” into 8 tasks that can be evaluated.
贝伐珠单抗疗效需要什么 What bevacizumab response needs
同一终点的先例 Precedent on the same endpoint
69.79%数据库:找到候选,查清每个资源能承担什么。 Database: find candidates and check what each resource can do.
器官含卵巢的对象 Objects whose organ includes ovary
有药物疗效标签 With drug-response labels
PTRC-HGSOC铂类 · 158 人 PTRC-HGSOCplatinum · 158 patients Ovarian-Bev贝伐 · 78 人 Ovarian-Bevbevacizumab · 78 patients ATEC23贝伐 · 180 个组织芯片点 ATEC23bevacizumab · 180 TMA cores
Ovarian-Bev
外部检索:确认缺口是真缺口,不是漏检。 External search: confirm the gaps are real, not missed.
数据在别处,不在同一批患者 The data exists elsewhere, for other patients
有数据,但不公开 The data exists but is not public
找不到独立外部验证 No independent external validation
下载指南与划分 JSON:准备到可以开始实验。 Download guide and split JSON: ready to start experiments.
{
"sample_id": "1612595C.svs",
"patient_id": "2909711",
"split": "train",
"label": {
"bevacizumab_treatment_effectiveness": {
"class_name": "effective", "class_id": 0
}
}
}