数据源报告
TCGA 数据源报告
The Cancer Genome Atlas · 33 个项目 · 官方入口 GDC Data Portal
一、数据源概览
The Cancer Genome Atlas(TCGA)是由美国国家癌症研究所(NCI)和 National Human Genome Research Institute(NHGRI)共同支持的大型癌症分子图谱计划。TCGA 以癌种或组织部位划分项目,汇集临床、biospecimen、病理图像、DNA/RNA 测序、突变、拷贝数、甲基化、转录组、蛋白组和结构变异等数据。
当前 TCGA 数据主要通过 NCI Genomic Data Commons(GDC)项目页面和 API 分发,也可通过 cBioPortal PanCancer Atlas 等下游分析平台访问处理后的分子 profile。GDC 项目级数据包含 33 个 TCGA 癌种/疾病项目报告,覆盖白血病、脑胶质瘤、乳腺癌、肺癌、肾癌、肝癌、消化系统肿瘤、妇科肿瘤、肉瘤和黑色素瘤等范围。
TCGA 的主要定位是癌症多组学和临床病理研究数据源。部分项目包含 SVS 格式数字病理切片,可支持计算病理研究;项目级数据通常没有统一的视觉任务标签、固定 train/validation/test 划分或官方 CPath benchmark 协议。
二、来源与开放状态
| 项目 |
内容 |
| 数据源名称 |
The Cancer Genome Atlas(TCGA) |
| 官方数据入口 |
GDC Data Portal |
| 官方项目 API |
https://api.gdc.cancer.gov/projects/{project_id} |
| 主要下游平台 |
cBioPortal、UCSC Xena、各项目论文补充数据 |
| 数据组织 |
TCGA project → case → sample → portion/analyte → file |
| 主要数据类型 |
Clinical、biospecimen、pathology image、sequencing、mutation、CNV、methylation、expression、RPPA、structural variation |
| 访问状态 |
Partially Open |
| 开放文件 |
部分病理图像、临床/样本 metadata、部分处理后组学文件和公开分析对象 |
| 受控文件 |
大量 WXS/WGS、BAM、原始测序和部分临床数据需要受控访问 |
| 统一许可证 |
未确认覆盖全部 TCGA 项目的统一 SPDX 数据许可证 |
| 统一任务划分 |
未发现覆盖全部癌种的官方固定数据划分 |
GDC 公开项目页面、病例分面、样本 metadata、文件清单和部分开放文件可直接访问。controlled access 文件通常需要 dbGaP 相关授权、GDC 账号或数据传输工具配置。cBioPortal 的 analysis-ready profile 适合快速探索和下游建模,原始或项目级文件应回到 GDC 核对来源和访问状态。
三、项目组成与规模
本报告覆盖 33 个 TCGA 项目。下表列出当前 GDC 项目报告中的代表性病例规模、主要组织范围和病理图像信息;病例数和 Slide Image 数量来自不同 GDC 聚合口径,使用时应记录查询时间。
| 项目 |
GDC ID |
主要组织/疾病 |
病例规模(约) |
病理图像概况 |
| Acute Myeloid Leukemia |
TCGA-LAML |
急性髓系白血病 |
200 |
0 个 Slide Image |
| Adrenocortical Carcinoma |
TCGA-ACC |
肾上腺皮质癌 |
92 |
323 个 Slide Image |
| Bladder Urothelial Carcinoma |
TCGA-BLCA |
膀胱尿路上皮癌 |
412 |
926 个 Slide Image |
| Brain Lower Grade Glioma |
TCGA-LGG |
低级别胶质瘤 |
516 |
1,572 个 Slide Image |
| Breast Invasive Carcinoma |
TCGA-BRCA |
乳腺浸润性癌 |
1,098 |
3,112 个 Slide Image |
| Cervical Squamous Cell Carcinoma and Endocervical Adenocarcinoma |
TCGA-CESC |
宫颈鳞癌与腺癌 |
307 |
604 个 Slide Image |
| Cholangiocarcinoma |
TCGA-CHOL |
胆管癌 |
51 |
110 个 Slide Image |
| Colon Adenocarcinoma |
TCGA-COAD |
结肠腺癌 |
461 |
1,442 个 Slide Image |
| Esophageal Carcinoma |
TCGA-ESCA |
食管癌 |
185 |
396 个 Slide Image |
| Glioblastoma Multiforme |
TCGA-GBM |
胶质母细胞瘤 |
617 |
2,053 个 Slide Image |
| Head and Neck Squamous Cell Carcinoma |
TCGA-HNSC |
头颈部鳞癌 |
528 |
1,263 个 Slide Image |
| Kidney Chromophobe |
TCGA-KICH |
肾嫌色细胞癌 |
113 |
326 个 Slide Image |
| Kidney Renal Clear Cell Carcinoma |
TCGA-KIRC |
肾透明细胞癌 |
537 |
2,173 个 Slide Image |
| Kidney Renal Papillary Cell Carcinoma |
TCGA-KIRP |
肾乳头状细胞癌 |
291 |
773 个 Slide Image |
| Liver Hepatocellular Carcinoma |
TCGA-LIHC |
肝细胞癌 |
377 |
870 个 Slide Image |
| Lung Adenocarcinoma |
TCGA-LUAD |
肺腺癌 |
585 |
1,608 个 Slide Image |
| Lung Squamous Cell Carcinoma |
TCGA-LUSC |
肺鳞癌 |
504 |
1,612 个 Slide Image |
| Lymphoid Neoplasm Diffuse Large B-cell Lymphoma |
TCGA-DLBC |
弥漫大 B 细胞淋巴瘤 |
58 |
103 个 Slide Image |
| Mesothelioma |
TCGA-MESO |
恶性间皮瘤 |
87 |
175 个 Slide Image |
| Ovarian Serous Cystadenocarcinoma |
TCGA-OV |
卵巢浆液性肿瘤 |
608 |
1,481 个 Slide Image |
| Pancreatic Adenocarcinoma |
TCGA-PAAD |
胰腺腺癌 |
185 |
466 个 Slide Image |
| Pheochromocytoma and Paraganglioma |
TCGA-PCPG |
嗜铬细胞瘤/副神经节瘤 |
179 |
385 个 Slide Image |
| Prostate Adenocarcinoma |
TCGA-PRAD |
前列腺腺癌 |
500 |
1,172 个 Slide Image |
| Rectum Adenocarcinoma |
TCGA-READ |
直肠腺癌 |
172 |
530 个 Slide Image |
| Sarcoma |
TCGA-SARC |
成人软组织肉瘤 |
261 |
890 个 Slide Image |
| Skin Cutaneous Melanoma |
TCGA-SKCM |
皮肤黑色素瘤 |
470 |
950 个 Slide Image |
| Stomach Adenocarcinoma |
TCGA-STAD |
胃腺癌 |
443 |
1,197 个 Slide Image |
| Testicular Germ Cell Tumors |
TCGA-TGCT |
睾丸生殖细胞肿瘤 |
263 |
663 个 Slide Image |
| Thymoma |
TCGA-THYM |
胸腺上皮肿瘤 |
124 |
318 个 Slide Image |
| Thyroid Carcinoma |
TCGA-THCA |
甲状腺癌 |
507 |
1,158 个 Slide Image |
| Uterine Carcinosarcoma |
TCGA-UCS |
子宫癌肉瘤 |
57 |
154 个 Slide Image |
| Uterine Corpus Endometrial Carcinoma |
TCGA-UCEC |
子宫体子宫内膜癌 |
560 |
1,371 个 Slide Image |
| Uveal Melanoma |
TCGA-UVM |
葡萄膜黑色素瘤 |
80 |
150 个 Slide Image |
| 合计(33 个项目) |
— |
— |
11,428 |
30,326 个 Slide Image |
表中「合计」行为 33 个 TCGA 项目子报告同类数量(病例规模、Slide Image 数)的直接加总:病例合计 11,428、Slide Image 合计 30,326。该合计为 GDC 当前 release 口径的项目级求和,不等于去重后的唯一患者/切片数(同一项目内病例与切片为多对多,跨项目不重叠),也与论文冻结队列数量不同。
不同项目的 GDC 文件总量从数千到数万不等。GDC 项目文件数量、病例数量和论文冻结队列数量经常存在差异:项目统计反映当前 release,论文数字反映特定研究设计和数据冻结版本,二者需要分别引用。
四、癌种与临床病理范围
4.1 主要癌种分组
- 血液和淋巴系统:TCGA-LAML、TCGA-DLBC。
- 中枢神经系统:TCGA-GBM、TCGA-LGG。
- 乳腺:TCGA-BRCA。
- 肺和胸膜:TCGA-LUAD、TCGA-LUSC、TCGA-MESO。
- 肾脏和泌尿系统:TCGA-KICH、TCGA-KIRC、TCGA-KIRP、TCGA-BLCA。
- 消化系统:TCGA-CHOL、TCGA-COAD、TCGA-READ、TCGA-ESCA、TCGA-STAD、TCGA-LIHC、TCGA-PAAD。
- 妇科和女性生殖系统:TCGA-CESC、TCGA-OV、TCGA-UCS、TCGA-UCEC。
- 头颈部:TCGA-HNSC。
- 男性生殖系统:TCGA-PRAD、TCGA-TGCT。
- 皮肤和眼部:TCGA-SKCM、TCGA-UVM。
- 内分泌、神经内分泌和间叶组织:TCGA-ACC、TCGA-PCPG、TCGA-THCA、TCGA-THYM、TCGA-SARC。
4.2 分子和病理亚型
TCGA 项目论文和 GDC metadata 支持多种疾病特异性分类。例如:乳腺癌常见分子亚型、胃癌 EBV/MSI/GS/CIN 分型、子宫内膜癌 POLE/MSI/copy-number 分型、肾乳头状细胞癌 Type 1/Type 2、皮肤黑色素瘤 BRAF/RAS/NF1/Triple-WT 分型、软组织肉瘤 DDLPS/LMS/UPS/MFS/MPNST/SS 亚型,以及结直肠和食管癌的病理与分子类别。
这些分类来自不同论文、数据冻结版本和分析流程。跨项目合并时应保留原始 diagnosis、disease type、histological type、molecular subtype 和 publication cohort 字段,建立单独的标准化映射表。
五、数据模态与图像结构
5.1 数字病理图像
多数 TCGA 项目包含 GDC Slide Image 文件,常见格式为 .svs,包括 Tissue Slide 和 Diagnostic Slide。病理切片主要来自 H&E 染色的组织学复核;部分项目论文还说明了病理专家复核、肿瘤细胞比例、坏死比例、样本质量和组织处理流程。
存在明显例外:TCGA-LAML 当前 GDC 文件清单确认无 Slide Image(0 个)。对于其他项目,是否存在图像、图像数量、可访问状态和文件格式应以当前 GDC manifest 为准。GDC Slide Image 的存在也不代表已有像素级 mask、细胞级标注或统一的病理任务标签。
5.2 基因组、转录组和表观组
常见数据对象包括:
- WGS、WXS/WES 和 RNA-Seq。
- miRNA-Seq、基因表达矩阵和转录本数据。
- 体细胞突变、拷贝数变异、结构变异和融合。
- DNA methylation、SNP array 和其他表观遗传对象。
- RPPA 蛋白表达,少数项目还提供 ATAC-Seq 或其他分子层数据。
- clinical、biospecimen、pathology report、follow-up 和 sample metadata。
原始测序数据、BAM/FASTQ 等文件常有 controlled access;处理后的矩阵、临床分面、部分突变和图像文件可能公开。研究者应根据 data category、data type、access、experimental strategy 和 file format 逐项筛选。
5.3 空间组学
33 个项目的报告均未确认 TCGA 33 个项目形成统一的 Visium、Xenium、MERFISH、GeoMx 或其他空间转录组交付体系,也没有确认统一的 ST spot/bin 分辨率。TCGA 的病理 WSI 可作为组织形态数据使用,但 WSI 像素坐标与空间转录组 spot/cell 坐标具有不同的数据语义。
六、临床、样本和分类 metadata
GDC 的常见 metadata 层级包括:
case:参与者或病例级身份及临床关联。
sample:肿瘤、正常、复发、转移和其他样本类型。
portion、analyte:样本处理和分子实验层级。
file:文件名、格式、数据类别、实验策略、访问级别和实体 ID。
diagnoses:primary diagnosis、disease type、tumor stage、tumor grade 和组织部位。
demographic:age、gender、race、ethnicity、vital status 等。
项目报告还记录了多种病例级和论文级分类字段,例如 AJCC stage、WHO grade、histological type、molecular subtype、tumor/normal 状态、肿瘤细胞比例、坏死比例、RIN 和样本质量标记。字段覆盖率、编码大小写、缺失值和时间点因项目而异。
推荐将 GDC 原始 metadata、论文定义的 cohort 变量和 cBioPortal analysis-ready profile 分开保存。cBioPortal 适合快速读取 mutation、CNA、mRNA、methylation、RPPA、clinical 和 case list;它的样本范围可能小于当前 GDC 项目全量病例。
七、质量控制与数据处理建议
TCGA 项目论文通常描述样本筛选、病理复核、肿瘤细胞比例、坏死比例、RNA 质量、测序深度、平台批次和多组学数据冻结条件。不同项目的 QC 规则并不统一,GDC 文件清单也无法替代图像和矩阵级质量检查。
建议执行以下流程:
- 固定 GDC 项目 ID、API 查询时间、data release 和 manifest。
- 按 case、sample、portion、analyte、file 建立关联索引。
- 过滤并记录文件 access level、data category、data type、experimental strategy 和格式。
- 对 SVS 图像检查金字塔层级、像素尺寸、组织覆盖、空白、失焦、折叠、扫描伪影和重复切片。
- 对测序和组学数据检查参考版本、批次、缺失值、覆盖度、表达矩阵版本、样本配对和分子 profile 来源。
- 对临床字段检查年龄、性别、分期、分级、诊断、肿瘤/正常标记和缺失编码。
- 以 case 或 sample 为单位切分数据,避免同一病例的 WSI、patch、组学文件或重复样本跨分区泄漏。
- 记录论文冻结队列与 GDC 当前项目队列的差异,不用论文病例数替代当前下载文件的真实规模。
八、人口统计与代表性
TCGA 覆盖多个癌种和研究中心,但病例构成由各癌种项目的入组、组织获取和实验设计决定。年龄、性别、种族、族裔、临床阶段和样本类型的覆盖度因项目变化明显。部分项目存在特定性别或年龄结构,例如前列腺、睾丸、卵巢、乳腺和子宫项目具有明确的疾病相关人群特征;LAML、DLBC 等血液肿瘤项目的样本来源也不同于实体瘤。
公平性分析应至少按癌种、项目、中心、年龄、性别、种族、族裔、样本类型和临床阶段分层。公开 metadata 缺失或受控时,应报告可用字段覆盖率,避免把 TCGA 研究队列直接解释为普通人群代表性样本。
九、适用研究方向
TCGA 适合支持:
- 癌种内和跨癌种的基因组、转录组、甲基化与蛋白组整合。
- H&E WSI 与临床、分子亚型、突变、CNV、表达和预后关联。
- 病理图像表征学习、组织区域检测和病例级多模态建模。
- 肿瘤/正常比较、分子亚型识别和临床分层。
- 生物标志物、通路、驱动基因和治疗靶点探索。
- 不同癌种的组织形态、分子状态和临床结局比较。
- cBioPortal analysis-ready profile 与 GDC 原始文件的交叉验证。
TCGA 项目适合研究型分析和资源复用。现有项目报告没有确认统一的计算病理任务、固定标签、官方 benchmark split 或 leaderboard,研究者需要自行制定病例级划分、标签规则和外部验证方案。
十、使用边界与关键限制
- 访问分层:开放图像和 metadata 与 controlled 测序、BAM/FASTQ、部分临床数据的访问条件不同。
- 版本差异:GDC 当前项目统计与历史 TCGA 论文冻结队列数量可能不同。
- 项目异质性:33 个项目的癌种、样本处理、测序平台、病理审阅、批次和临床变量存在差异。
- 病例级依赖:同一病例可能有多张切片、多种样本和多个分子文件,随机 patch 划分会造成泄漏。
- 病理标注有限:Slide Image 文件通常缺少统一的像素级 mask、细胞级标注和病理任务标签。
- 图像覆盖差异:TCGA-LAML 当前报告确认无 Slide Image(0 个),其余项目的图像数量也应以当前 manifest 复核。
- 标签语义不统一:diagnosis、histology、molecular subtype、stage 和 grade 的定义由项目和论文决定。
- 人口代表性有限:癌种特异入组、公开字段缺失和项目规模差异会影响公平性和外推性。
- 跨平台偏差:GDC、cBioPortal、论文补充数据和其他分析平台的样本范围与处理版本可能不同。
- 存储和计算成本:WSI、原始测序和多组学文件规模较大,下载前应规划存储、传输和计算资源。
十一、推荐使用流程
建议先确定癌种项目和研究模态,再通过 GDC 项目页或 API 查询病例、样本、文件和访问状态。需要快速探索分子 profile 时可使用 cBioPortal,但正式分析应保留 GDC project ID、manifest、文件版本和样本筛选记录。下载后先建立 case–sample–file 关系表,分别完成临床、图像、测序和组学 QC,再按病例级切分并开展建模。结果中应注明 GDC data release、查询日期、开放/受控文件范围、论文冻结队列与当前项目队列差异。
十二、引用建议
TCGA 数据分析应同时引用 TCGA/GDC 官方入口、具体项目论文、实际使用的 GDC data release 和文件清单。涉及下游分析时,可补充引用 cBioPortal PanCancer Atlas 或 UCSC Xena 的具体 study/profile 版本。病理图像研究应注明项目 ID、Slide Image 文件类型、图像筛选规则和病例级切分策略。
Data-source report
TCGA data-source report
The Cancer Genome Atlas · 33 projects · official entry: GDC Data Portal
1. Source overview
The Cancer Genome Atlas (TCGA) is a large cancer molecular-profiling program jointly supported by the US National Cancer Institute (NCI) and the National Human Genome Research Institute (NHGRI). TCGA organizes its projects by cancer type or tissue site and brings together clinical, biospecimen, pathology image, DNA/RNA sequencing, mutation, copy number, methylation, transcriptome, proteome and structural variation data.
TCGA data is now distributed mainly through the project pages and API of the NCI Genomic Data Commons (GDC); processed molecular profiles can also be reached through downstream analysis platforms such as the cBioPortal PanCancer Atlas. GDC project-level data covers 33 TCGA cancer or disease project reports, spanning leukemia, glioma, breast, lung, kidney and liver cancer, digestive-system tumors, gynecologic tumors, sarcoma and melanoma.
TCGA is primarily a data source for cancer multi-omics and clinico-pathological research. Some projects include digital pathology slides in SVS format that can support computational pathology research; project-level data usually has no unified visual task labels, fixed train/validation/test split or official CPath benchmark protocol.
2. Sources and access
| Item |
Details |
| Source name |
The Cancer Genome Atlas (TCGA) |
| Official data entry |
GDC Data Portal |
| Official project API |
https://api.gdc.cancer.gov/projects/{project_id} |
| Main downstream platforms |
cBioPortal, UCSC Xena, supplementary data of the project papers |
| Data organization |
TCGA project → case → sample → portion/analyte → file |
| Main data types |
Clinical, biospecimen, pathology image, sequencing, mutation, CNV, methylation, expression, RPPA, structural variation |
| Access status |
Partially Open |
| Open files |
Some pathology images, clinical/sample metadata, some processed omics files and public analysis objects |
| Controlled files |
Much WXS/WGS, BAM, raw sequencing and some clinical data require controlled access |
| Unified license |
No unified SPDX data license covering all TCGA projects was confirmed |
| Unified task split |
No official fixed data split covering all cancer types was found |
Public GDC project pages, case facets, sample metadata, file listings and some open files can be accessed directly. Controlled-access files usually require dbGaP authorization, a GDC account or a configured data transfer tool. cBioPortal's analysis-ready profiles suit quick exploration and downstream modeling; raw or project-level files should be checked back against GDC for source and access status.
3. Project composition and scale
This report covers 33 TCGA projects. The table lists representative case counts, main tissue scope and pathology image information from the current GDC project reports; case counts and Slide Image counts come from different GDC aggregations, so record the query time when using them.
| Project |
GDC ID |
Main tissue / disease |
Cases (approx.) |
Pathology images |
| Acute Myeloid Leukemia |
TCGA-LAML |
Acute myeloid leukemia |
200 |
0 Slide Images |
| Adrenocortical Carcinoma |
TCGA-ACC |
Adrenocortical carcinoma |
92 |
323 Slide Images |
| Bladder Urothelial Carcinoma |
TCGA-BLCA |
Bladder urothelial carcinoma |
412 |
926 Slide Images |
| Brain Lower Grade Glioma |
TCGA-LGG |
Lower-grade glioma |
516 |
1,572 Slide Images |
| Breast Invasive Carcinoma |
TCGA-BRCA |
Breast invasive carcinoma |
1,098 |
3,112 Slide Images |
| Cervical Squamous Cell Carcinoma and Endocervical Adenocarcinoma |
TCGA-CESC |
Cervical squamous cell carcinoma and adenocarcinoma |
307 |
604 Slide Images |
| Cholangiocarcinoma |
TCGA-CHOL |
Cholangiocarcinoma |
51 |
110 Slide Images |
| Colon Adenocarcinoma |
TCGA-COAD |
Colon adenocarcinoma |
461 |
1,442 Slide Images |
| Esophageal Carcinoma |
TCGA-ESCA |
Esophageal carcinoma |
185 |
396 Slide Images |
| Glioblastoma Multiforme |
TCGA-GBM |
Glioblastoma |
617 |
2,053 Slide Images |
| Head and Neck Squamous Cell Carcinoma |
TCGA-HNSC |
Head and neck squamous cell carcinoma |
528 |
1,263 Slide Images |
| Kidney Chromophobe |
TCGA-KICH |
Kidney chromophobe |
113 |
326 Slide Images |
| Kidney Renal Clear Cell Carcinoma |
TCGA-KIRC |
Kidney renal clear cell carcinoma |
537 |
2,173 Slide Images |
| Kidney Renal Papillary Cell Carcinoma |
TCGA-KIRP |
Kidney renal papillary cell carcinoma |
291 |
773 Slide Images |
| Liver Hepatocellular Carcinoma |
TCGA-LIHC |
Liver hepatocellular carcinoma |
377 |
870 Slide Images |
| Lung Adenocarcinoma |
TCGA-LUAD |
Lung adenocarcinoma |
585 |
1,608 Slide Images |
| Lung Squamous Cell Carcinoma |
TCGA-LUSC |
Lung squamous cell carcinoma |
504 |
1,612 Slide Images |
| Lymphoid Neoplasm Diffuse Large B-cell Lymphoma |
TCGA-DLBC |
Diffuse large B-cell lymphoma |
58 |
103 Slide Images |
| Mesothelioma |
TCGA-MESO |
Mesothelioma |
87 |
175 Slide Images |
| Ovarian Serous Cystadenocarcinoma |
TCGA-OV |
Ovarian serous cystadenocarcinoma |
608 |
1,481 Slide Images |
| Pancreatic Adenocarcinoma |
TCGA-PAAD |
Pancreatic adenocarcinoma |
185 |
466 Slide Images |
| Pheochromocytoma and Paraganglioma |
TCGA-PCPG |
Pheochromocytoma and paraganglioma |
179 |
385 Slide Images |
| Prostate Adenocarcinoma |
TCGA-PRAD |
Prostate adenocarcinoma |
500 |
1,172 Slide Images |
| Rectum Adenocarcinoma |
TCGA-READ |
Rectum adenocarcinoma |
172 |
530 Slide Images |
| Sarcoma |
TCGA-SARC |
Adult soft tissue sarcoma |
261 |
890 Slide Images |
| Skin Cutaneous Melanoma |
TCGA-SKCM |
Skin cutaneous melanoma |
470 |
950 Slide Images |
| Stomach Adenocarcinoma |
TCGA-STAD |
Stomach adenocarcinoma |
443 |
1,197 Slide Images |
| Testicular Germ Cell Tumors |
TCGA-TGCT |
Testicular germ cell tumors |
263 |
663 Slide Images |
| Thymoma |
TCGA-THYM |
Thymoma |
124 |
318 Slide Images |
| Thyroid Carcinoma |
TCGA-THCA |
Thyroid carcinoma |
507 |
1,158 Slide Images |
| Uterine Carcinosarcoma |
TCGA-UCS |
Uterine carcinosarcoma |
57 |
154 Slide Images |
| Uterine Corpus Endometrial Carcinoma |
TCGA-UCEC |
Uterine corpus endometrial carcinoma |
560 |
1,371 Slide Images |
| Uveal Melanoma |
TCGA-UVM |
Uveal melanoma |
80 |
150 Slide Images |
| Total (33 projects) |
— |
— |
11,428 |
30,326 Slide Images |
The “Total” row is a direct sum of the same quantities (case counts, Slide Image counts) across the 33 TCGA project sub-reports: 11,428 cases and 30,326 Slide Images. It is a project-level sum under the current GDC release, not a count of unique patients or slides (cases and slides are many-to-many within a project and do not overlap across projects), and it differs from the frozen cohorts in the papers.
Total GDC file counts range from thousands to tens of thousands per project. GDC file counts, case counts and the frozen cohort sizes in the papers often differ: project statistics reflect the current release, while paper numbers reflect a specific study design and data freeze, so the two should be cited separately.
4. Cancer types and clinico-pathological scope
4.1 Main cancer groups
- Blood and lymphatic system: TCGA-LAML, TCGA-DLBC.
- Central nervous system: TCGA-GBM, TCGA-LGG.
- Breast: TCGA-BRCA.
- Lung and pleura: TCGA-LUAD, TCGA-LUSC, TCGA-MESO.
- Kidney and urinary system: TCGA-KICH, TCGA-KIRC, TCGA-KIRP, TCGA-BLCA.
- Digestive system: TCGA-CHOL, TCGA-COAD, TCGA-READ, TCGA-ESCA, TCGA-STAD, TCGA-LIHC, TCGA-PAAD.
- Gynecologic and female reproductive system: TCGA-CESC, TCGA-OV, TCGA-UCS, TCGA-UCEC.
- Head and neck: TCGA-HNSC.
- Male reproductive system: TCGA-PRAD, TCGA-TGCT.
- Skin and eye: TCGA-SKCM, TCGA-UVM.
- Endocrine, neuroendocrine and mesenchymal tissue: TCGA-ACC, TCGA-PCPG, TCGA-THCA, TCGA-THYM, TCGA-SARC.
4.2 Molecular and pathological subtypes
TCGA project papers and GDC metadata support many disease-specific classifications, for example common molecular subtypes of breast cancer, the EBV/MSI/GS/CIN classes of gastric cancer, POLE/MSI/copy-number classes of endometrial cancer, Type 1/Type 2 papillary renal cell carcinoma, BRAF/RAS/NF1/Triple-WT classes of cutaneous melanoma, DDLPS/LMS/UPS/MFS/MPNST/SS subtypes of soft tissue sarcoma, and the pathological and molecular categories of colorectal and esophageal cancer.
These classifications come from different papers, data freezes and analysis pipelines. When merging across projects, keep the original diagnosis, disease type, histological type, molecular subtype and publication cohort fields and build a separate standardized mapping table.
5. Data modalities and image structure
5.1 Digital pathology images
Most TCGA projects include GDC Slide Image files, usually in .svs format, as Tissue Slide and Diagnostic Slide. The slides come mainly from H&E-stained histology review; some project papers also describe pathologist review, tumor cell fraction, necrosis fraction, sample quality and tissue processing.
There is a clear exception: the current GDC file listing for TCGA-LAML confirms no Slide Images (0). For the other projects, whether images exist, how many, their access status and file format should be taken from the current GDC manifest. The presence of GDC Slide Images does not mean pixel-level masks, cell-level annotations or unified pathology task labels exist.
5.2 Genome, transcriptome and epigenome
Common data objects include:
- WGS, WXS/WES and RNA-Seq.
- miRNA-Seq, gene expression matrices and transcript data.
- Somatic mutations, copy number variation, structural variants and fusions.
- DNA methylation, SNP arrays and other epigenetic objects.
- RPPA protein expression; a few projects also provide ATAC-Seq or other molecular layers.
- Clinical, biospecimen, pathology report, follow-up and sample metadata.
Raw sequencing data and files such as BAM/FASTQ are often under controlled access; processed matrices, clinical facets and some mutation and image files may be public. Researchers should filter item by item by data category, data type, access, experimental strategy and file format.
5.3 Spatial omics
None of the reports on the 33 projects confirms a unified Visium, Xenium, MERFISH, GeoMx or other spatial transcriptomics delivery across TCGA, or a unified ST spot/bin resolution. TCGA pathology WSIs can be used as tissue morphology data, but WSI pixel coordinates and spatial transcriptomics spot/cell coordinates mean different things.
6. Clinical, sample and classification metadata
Common GDC metadata levels include:
case: participant- or case-level identity and clinical links.
sample: tumor, normal, recurrent, metastatic and other sample types.
portion, analyte: sample processing and molecular assay levels.
file: file name, format, data category, experimental strategy, access level and entity ID.
diagnoses: primary diagnosis, disease type, tumor stage, tumor grade and tissue site.
demographic: age, gender, race, ethnicity, vital status and similar.
The project reports also record many case-level and paper-level classification fields, such as AJCC stage, WHO grade, histological type, molecular subtype, tumor/normal status, tumor cell fraction, necrosis fraction, RIN and sample quality flags. Field coverage, coding and capitalization, missing values and time points vary by project.
Keep the original GDC metadata, the cohort variables defined in papers and the cBioPortal analysis-ready profiles separate. cBioPortal is convenient for quickly reading mutation, CNA, mRNA, methylation, RPPA, clinical data and case lists; its sample scope may be smaller than all cases of the current GDC project.
7. Quality control and processing advice
TCGA project papers usually describe sample selection, pathologist review, tumor cell fraction, necrosis fraction, RNA quality, sequencing depth, platform batches and multi-omics data-freeze conditions. QC rules differ between projects, and a GDC file listing cannot replace image- and matrix-level quality checks.
Suggested workflow:
- Fix the GDC project ID, API query time, data release and manifest.
- Build a linked index by case, sample, portion, analyte and file.
- Filter and record each file's access level, data category, data type, experimental strategy and format.
- For SVS images, check pyramid levels, pixel size, tissue coverage, blank areas, out-of-focus regions, folds, scanning artifacts and duplicate slides.
- For sequencing and omics data, check reference versions, batches, missing values, coverage, expression-matrix versions, sample pairing and the source of molecular profiles.
- For clinical fields, check age, sex, stage, grade, diagnosis, tumor/normal flags and missing-value codes.
- Split data by case or sample so that WSIs, patches, omics files or duplicate samples of the same case do not leak across partitions.
- Record the differences between the frozen cohorts in the papers and the current GDC project cohorts; do not use paper case counts in place of the real size of the files downloaded now.
8. Demographics and representativeness
TCGA covers many cancer types and research centers, but case composition is determined by each project's enrollment, tissue acquisition and experimental design. Coverage of age, sex, race, ethnicity, clinical stage and sample type varies markedly between projects. Some projects have specific sex or age structures; for example the prostate, testis, ovary, breast and uterus projects have clear disease-related population characteristics, and the samples of hematologic projects such as LAML and DLBC also differ from solid tumors.
Fairness analyses should at least stratify by cancer type, project, center, age, sex, race, ethnicity, sample type and clinical stage. Where public metadata is missing or controlled, report the coverage of the available fields, and do not interpret TCGA study cohorts as representative samples of the general population.
9. Suitable research directions
TCGA is suited to:
- Integrating genome, transcriptome, methylation and proteome within and across cancer types.
- Linking H&E WSIs with clinical data, molecular subtypes, mutations, CNV, expression and prognosis.
- Pathology image representation learning, tissue region detection and case-level multimodal modeling.
- Tumor/normal comparison, molecular subtype identification and clinical stratification.
- Exploring biomarkers, pathways, driver genes and therapeutic targets.
- Comparing tissue morphology, molecular status and clinical outcomes across cancer types.
- Cross-validating cBioPortal analysis-ready profiles against the original GDC files.
TCGA projects suit research analyses and resource reuse. The existing project reports confirm no unified computational pathology task, fixed labels, official benchmark split or leaderboard; researchers need to define their own case-level splits, label rules and external validation plans.
10. Boundaries and key limitations
- Tiered access: open images and metadata have different access conditions from controlled sequencing, BAM/FASTQ and some clinical data.
- Version differences: current GDC project statistics may differ from the frozen cohorts in historical TCGA papers.
- Project heterogeneity: the 33 projects differ in cancer type, sample processing, sequencing platform, pathology review, batches and clinical variables.
- Case-level dependence: one case can have several slides, sample types and molecular files, so random patch splits cause leakage.
- Limited pathology annotation: Slide Image files usually lack unified pixel-level masks, cell-level annotations and pathology task labels.
- Image coverage differs: the current report confirms TCGA-LAML has no Slide Images (0), and image counts for the other projects should also be rechecked against the current manifest.
- Label semantics are not unified: definitions of diagnosis, histology, molecular subtype, stage and grade are set by each project and paper.
- Limited population representativeness: cancer-specific enrollment, missing public fields and differing project sizes affect fairness and generalizability.
- Cross-platform bias: GDC, cBioPortal, paper supplementary data and other analysis platforms may differ in sample scope and processing version.
- Storage and compute cost: WSIs, raw sequencing and multi-omics files are large; plan storage, transfer and compute before downloading.
11. Recommended workflow
First decide the cancer project and research modality, then query cases, samples, files and access status through the GDC project page or API. cBioPortal is useful for quick exploration of molecular profiles, but formal analyses should keep the GDC project ID, manifest, file versions and sample-selection records. After downloading, first build a case–sample–file relation table, run clinical, image, sequencing and omics QC separately, and then split by case and start modeling. Results should state the GDC data release, query date, the scope of open/controlled files, and the differences between the paper's frozen cohort and the current project cohort.
12. Citation advice
TCGA analyses should cite the official TCGA/GDC entry, the specific project papers, and the GDC data release and file listing actually used. For downstream analyses, also cite the specific study/profile version of the cBioPortal PanCancer Atlas or UCSC Xena. Pathology image studies should state the project IDs, Slide Image file types, image selection rules and case-level split strategy.