数据源报告

TCGA 数据源报告

The Cancer Genome Atlas · 33 个项目 · 官方入口 GDC Data Portal

一、数据源概览

The Cancer Genome Atlas(TCGA)是由美国国家癌症研究所(NCI)和 National Human Genome Research Institute(NHGRI)共同支持的大型癌症分子图谱计划。TCGA 以癌种或组织部位划分项目,汇集临床、biospecimen、病理图像、DNA/RNA 测序、突变、拷贝数、甲基化、转录组、蛋白组和结构变异等数据。

当前 TCGA 数据主要通过 NCI Genomic Data Commons(GDC)项目页面和 API 分发,也可通过 cBioPortal PanCancer Atlas 等下游分析平台访问处理后的分子 profile。GDC 项目级数据包含 33 个 TCGA 癌种/疾病项目报告,覆盖白血病、脑胶质瘤、乳腺癌、肺癌、肾癌、肝癌、消化系统肿瘤、妇科肿瘤、肉瘤和黑色素瘤等范围。

TCGA 的主要定位是癌症多组学和临床病理研究数据源。部分项目包含 SVS 格式数字病理切片,可支持计算病理研究;项目级数据通常没有统一的视觉任务标签、固定 train/validation/test 划分或官方 CPath benchmark 协议。

二、来源与开放状态

项目 内容
数据源名称 The Cancer Genome Atlas(TCGA)
官方数据入口 GDC Data Portal
官方项目 API https://api.gdc.cancer.gov/projects/{project_id}
主要下游平台 cBioPortal、UCSC Xena、各项目论文补充数据
数据组织 TCGA project → case → sample → portion/analyte → file
主要数据类型 Clinical、biospecimen、pathology image、sequencing、mutation、CNV、methylation、expression、RPPA、structural variation
访问状态 Partially Open
开放文件 部分病理图像、临床/样本 metadata、部分处理后组学文件和公开分析对象
受控文件 大量 WXS/WGS、BAM、原始测序和部分临床数据需要受控访问
统一许可证 未确认覆盖全部 TCGA 项目的统一 SPDX 数据许可证
统一任务划分 未发现覆盖全部癌种的官方固定数据划分

GDC 公开项目页面、病例分面、样本 metadata、文件清单和部分开放文件可直接访问。controlled access 文件通常需要 dbGaP 相关授权、GDC 账号或数据传输工具配置。cBioPortal 的 analysis-ready profile 适合快速探索和下游建模,原始或项目级文件应回到 GDC 核对来源和访问状态。

三、项目组成与规模

本报告覆盖 33 个 TCGA 项目。下表列出当前 GDC 项目报告中的代表性病例规模、主要组织范围和病理图像信息;病例数和 Slide Image 数量来自不同 GDC 聚合口径,使用时应记录查询时间。

项目 GDC ID 主要组织/疾病 病例规模(约) 病理图像概况
Acute Myeloid Leukemia TCGA-LAML 急性髓系白血病 200 0 个 Slide Image
Adrenocortical Carcinoma TCGA-ACC 肾上腺皮质癌 92 323 个 Slide Image
Bladder Urothelial Carcinoma TCGA-BLCA 膀胱尿路上皮癌 412 926 个 Slide Image
Brain Lower Grade Glioma TCGA-LGG 低级别胶质瘤 516 1,572 个 Slide Image
Breast Invasive Carcinoma TCGA-BRCA 乳腺浸润性癌 1,098 3,112 个 Slide Image
Cervical Squamous Cell Carcinoma and Endocervical Adenocarcinoma TCGA-CESC 宫颈鳞癌与腺癌 307 604 个 Slide Image
Cholangiocarcinoma TCGA-CHOL 胆管癌 51 110 个 Slide Image
Colon Adenocarcinoma TCGA-COAD 结肠腺癌 461 1,442 个 Slide Image
Esophageal Carcinoma TCGA-ESCA 食管癌 185 396 个 Slide Image
Glioblastoma Multiforme TCGA-GBM 胶质母细胞瘤 617 2,053 个 Slide Image
Head and Neck Squamous Cell Carcinoma TCGA-HNSC 头颈部鳞癌 528 1,263 个 Slide Image
Kidney Chromophobe TCGA-KICH 肾嫌色细胞癌 113 326 个 Slide Image
Kidney Renal Clear Cell Carcinoma TCGA-KIRC 肾透明细胞癌 537 2,173 个 Slide Image
Kidney Renal Papillary Cell Carcinoma TCGA-KIRP 肾乳头状细胞癌 291 773 个 Slide Image
Liver Hepatocellular Carcinoma TCGA-LIHC 肝细胞癌 377 870 个 Slide Image
Lung Adenocarcinoma TCGA-LUAD 肺腺癌 585 1,608 个 Slide Image
Lung Squamous Cell Carcinoma TCGA-LUSC 肺鳞癌 504 1,612 个 Slide Image
Lymphoid Neoplasm Diffuse Large B-cell Lymphoma TCGA-DLBC 弥漫大 B 细胞淋巴瘤 58 103 个 Slide Image
Mesothelioma TCGA-MESO 恶性间皮瘤 87 175 个 Slide Image
Ovarian Serous Cystadenocarcinoma TCGA-OV 卵巢浆液性肿瘤 608 1,481 个 Slide Image
Pancreatic Adenocarcinoma TCGA-PAAD 胰腺腺癌 185 466 个 Slide Image
Pheochromocytoma and Paraganglioma TCGA-PCPG 嗜铬细胞瘤/副神经节瘤 179 385 个 Slide Image
Prostate Adenocarcinoma TCGA-PRAD 前列腺腺癌 500 1,172 个 Slide Image
Rectum Adenocarcinoma TCGA-READ 直肠腺癌 172 530 个 Slide Image
Sarcoma TCGA-SARC 成人软组织肉瘤 261 890 个 Slide Image
Skin Cutaneous Melanoma TCGA-SKCM 皮肤黑色素瘤 470 950 个 Slide Image
Stomach Adenocarcinoma TCGA-STAD 胃腺癌 443 1,197 个 Slide Image
Testicular Germ Cell Tumors TCGA-TGCT 睾丸生殖细胞肿瘤 263 663 个 Slide Image
Thymoma TCGA-THYM 胸腺上皮肿瘤 124 318 个 Slide Image
Thyroid Carcinoma TCGA-THCA 甲状腺癌 507 1,158 个 Slide Image
Uterine Carcinosarcoma TCGA-UCS 子宫癌肉瘤 57 154 个 Slide Image
Uterine Corpus Endometrial Carcinoma TCGA-UCEC 子宫体子宫内膜癌 560 1,371 个 Slide Image
Uveal Melanoma TCGA-UVM 葡萄膜黑色素瘤 80 150 个 Slide Image
合计(33 个项目) — — 11,428 30,326 个 Slide Image

表中「合计」行为 33 个 TCGA 项目子报告同类数量(病例规模、Slide Image 数)的直接加总:病例合计 11,428、Slide Image 合计 30,326。该合计为 GDC 当前 release 口径的项目级求和,不等于去重后的唯一患者/切片数(同一项目内病例与切片为多对多,跨项目不重叠),也与论文冻结队列数量不同。

不同项目的 GDC 文件总量从数千到数万不等。GDC 项目文件数量、病例数量和论文冻结队列数量经常存在差异:项目统计反映当前 release,论文数字反映特定研究设计和数据冻结版本,二者需要分别引用。

四、癌种与临床病理范围

4.1 主要癌种分组

4.2 分子和病理亚型

TCGA 项目论文和 GDC metadata 支持多种疾病特异性分类。例如:乳腺癌常见分子亚型、胃癌 EBV/MSI/GS/CIN 分型、子宫内膜癌 POLE/MSI/copy-number 分型、肾乳头状细胞癌 Type 1/Type 2、皮肤黑色素瘤 BRAF/RAS/NF1/Triple-WT 分型、软组织肉瘤 DDLPS/LMS/UPS/MFS/MPNST/SS 亚型,以及结直肠和食管癌的病理与分子类别。

这些分类来自不同论文、数据冻结版本和分析流程。跨项目合并时应保留原始 diagnosis、disease type、histological type、molecular subtype 和 publication cohort 字段,建立单独的标准化映射表。

五、数据模态与图像结构

5.1 数字病理图像

多数 TCGA 项目包含 GDC Slide Image 文件,常见格式为 .svs,包括 Tissue Slide 和 Diagnostic Slide。病理切片主要来自 H&E 染色的组织学复核;部分项目论文还说明了病理专家复核、肿瘤细胞比例、坏死比例、样本质量和组织处理流程。

存在明显例外:TCGA-LAML 当前 GDC 文件清单确认无 Slide Image(0 个)。对于其他项目,是否存在图像、图像数量、可访问状态和文件格式应以当前 GDC manifest 为准。GDC Slide Image 的存在也不代表已有像素级 mask、细胞级标注或统一的病理任务标签。

5.2 基因组、转录组和表观组

常见数据对象包括:

原始测序数据、BAM/FASTQ 等文件常有 controlled access;处理后的矩阵、临床分面、部分突变和图像文件可能公开。研究者应根据 data category、data type、access、experimental strategy 和 file format 逐项筛选。

5.3 空间组学

33 个项目的报告均未确认 TCGA 33 个项目形成统一的 Visium、Xenium、MERFISH、GeoMx 或其他空间转录组交付体系,也没有确认统一的 ST spot/bin 分辨率。TCGA 的病理 WSI 可作为组织形态数据使用,但 WSI 像素坐标与空间转录组 spot/cell 坐标具有不同的数据语义。

六、临床、样本和分类 metadata

GDC 的常见 metadata 层级包括:

项目报告还记录了多种病例级和论文级分类字段,例如 AJCC stage、WHO grade、histological type、molecular subtype、tumor/normal 状态、肿瘤细胞比例、坏死比例、RIN 和样本质量标记。字段覆盖率、编码大小写、缺失值和时间点因项目而异。

推荐将 GDC 原始 metadata、论文定义的 cohort 变量和 cBioPortal analysis-ready profile 分开保存。cBioPortal 适合快速读取 mutation、CNA、mRNA、methylation、RPPA、clinical 和 case list;它的样本范围可能小于当前 GDC 项目全量病例。

七、质量控制与数据处理建议

TCGA 项目论文通常描述样本筛选、病理复核、肿瘤细胞比例、坏死比例、RNA 质量、测序深度、平台批次和多组学数据冻结条件。不同项目的 QC 规则并不统一,GDC 文件清单也无法替代图像和矩阵级质量检查。

建议执行以下流程:

  1. 固定 GDC 项目 ID、API 查询时间、data release 和 manifest。
  2. 按 case、sample、portion、analyte、file 建立关联索引。
  3. 过滤并记录文件 access level、data category、data type、experimental strategy 和格式。
  4. 对 SVS 图像检查金字塔层级、像素尺寸、组织覆盖、空白、失焦、折叠、扫描伪影和重复切片。
  5. 对测序和组学数据检查参考版本、批次、缺失值、覆盖度、表达矩阵版本、样本配对和分子 profile 来源。
  6. 对临床字段检查年龄、性别、分期、分级、诊断、肿瘤/正常标记和缺失编码。
  7. 以 case 或 sample 为单位切分数据,避免同一病例的 WSI、patch、组学文件或重复样本跨分区泄漏。
  8. 记录论文冻结队列与 GDC 当前项目队列的差异,不用论文病例数替代当前下载文件的真实规模。

八、人口统计与代表性

TCGA 覆盖多个癌种和研究中心,但病例构成由各癌种项目的入组、组织获取和实验设计决定。年龄、性别、种族、族裔、临床阶段和样本类型的覆盖度因项目变化明显。部分项目存在特定性别或年龄结构,例如前列腺、睾丸、卵巢、乳腺和子宫项目具有明确的疾病相关人群特征;LAML、DLBC 等血液肿瘤项目的样本来源也不同于实体瘤。

公平性分析应至少按癌种、项目、中心、年龄、性别、种族、族裔、样本类型和临床阶段分层。公开 metadata 缺失或受控时,应报告可用字段覆盖率,避免把 TCGA 研究队列直接解释为普通人群代表性样本。

九、适用研究方向

TCGA 适合支持:

TCGA 项目适合研究型分析和资源复用。现有项目报告没有确认统一的计算病理任务、固定标签、官方 benchmark split 或 leaderboard,研究者需要自行制定病例级划分、标签规则和外部验证方案。

十、使用边界与关键限制

  1. 访问分层:开放图像和 metadata 与 controlled 测序、BAM/FASTQ、部分临床数据的访问条件不同。
  2. 版本差异:GDC 当前项目统计与历史 TCGA 论文冻结队列数量可能不同。
  3. 项目异质性:33 个项目的癌种、样本处理、测序平台、病理审阅、批次和临床变量存在差异。
  4. 病例级依赖:同一病例可能有多张切片、多种样本和多个分子文件,随机 patch 划分会造成泄漏。
  5. 病理标注有限:Slide Image 文件通常缺少统一的像素级 mask、细胞级标注和病理任务标签。
  6. 图像覆盖差异:TCGA-LAML 当前报告确认无 Slide Image(0 个),其余项目的图像数量也应以当前 manifest 复核。
  7. 标签语义不统一:diagnosis、histology、molecular subtype、stage 和 grade 的定义由项目和论文决定。
  8. 人口代表性有限:癌种特异入组、公开字段缺失和项目规模差异会影响公平性和外推性。
  9. 跨平台偏差:GDC、cBioPortal、论文补充数据和其他分析平台的样本范围与处理版本可能不同。
  10. 存储和计算成本:WSI、原始测序和多组学文件规模较大,下载前应规划存储、传输和计算资源。

十一、推荐使用流程

建议先确定癌种项目和研究模态,再通过 GDC 项目页或 API 查询病例、样本、文件和访问状态。需要快速探索分子 profile 时可使用 cBioPortal,但正式分析应保留 GDC project ID、manifest、文件版本和样本筛选记录。下载后先建立 case–sample–file 关系表,分别完成临床、图像、测序和组学 QC,再按病例级切分并开展建模。结果中应注明 GDC data release、查询日期、开放/受控文件范围、论文冻结队列与当前项目队列差异。

十二、引用建议

TCGA 数据分析应同时引用 TCGA/GDC 官方入口、具体项目论文、实际使用的 GDC data release 和文件清单。涉及下游分析时,可补充引用 cBioPortal PanCancer Atlas 或 UCSC Xena 的具体 study/profile 版本。病理图像研究应注明项目 ID、Slide Image 文件类型、图像筛选规则和病例级切分策略。

Data-source report

TCGA data-source report

The Cancer Genome Atlas · 33 projects · official entry: GDC Data Portal

1. Source overview

The Cancer Genome Atlas (TCGA) is a large cancer molecular-profiling program jointly supported by the US National Cancer Institute (NCI) and the National Human Genome Research Institute (NHGRI). TCGA organizes its projects by cancer type or tissue site and brings together clinical, biospecimen, pathology image, DNA/RNA sequencing, mutation, copy number, methylation, transcriptome, proteome and structural variation data.

TCGA data is now distributed mainly through the project pages and API of the NCI Genomic Data Commons (GDC); processed molecular profiles can also be reached through downstream analysis platforms such as the cBioPortal PanCancer Atlas. GDC project-level data covers 33 TCGA cancer or disease project reports, spanning leukemia, glioma, breast, lung, kidney and liver cancer, digestive-system tumors, gynecologic tumors, sarcoma and melanoma.

TCGA is primarily a data source for cancer multi-omics and clinico-pathological research. Some projects include digital pathology slides in SVS format that can support computational pathology research; project-level data usually has no unified visual task labels, fixed train/validation/test split or official CPath benchmark protocol.

2. Sources and access

Item Details
Source name The Cancer Genome Atlas (TCGA)
Official data entry GDC Data Portal
Official project API https://api.gdc.cancer.gov/projects/{project_id}
Main downstream platforms cBioPortal, UCSC Xena, supplementary data of the project papers
Data organization TCGA project → case → sample → portion/analyte → file
Main data types Clinical, biospecimen, pathology image, sequencing, mutation, CNV, methylation, expression, RPPA, structural variation
Access status Partially Open
Open files Some pathology images, clinical/sample metadata, some processed omics files and public analysis objects
Controlled files Much WXS/WGS, BAM, raw sequencing and some clinical data require controlled access
Unified license No unified SPDX data license covering all TCGA projects was confirmed
Unified task split No official fixed data split covering all cancer types was found

Public GDC project pages, case facets, sample metadata, file listings and some open files can be accessed directly. Controlled-access files usually require dbGaP authorization, a GDC account or a configured data transfer tool. cBioPortal's analysis-ready profiles suit quick exploration and downstream modeling; raw or project-level files should be checked back against GDC for source and access status.

3. Project composition and scale

This report covers 33 TCGA projects. The table lists representative case counts, main tissue scope and pathology image information from the current GDC project reports; case counts and Slide Image counts come from different GDC aggregations, so record the query time when using them.

Project GDC ID Main tissue / disease Cases (approx.) Pathology images
Acute Myeloid Leukemia TCGA-LAML Acute myeloid leukemia 200 0 Slide Images
Adrenocortical Carcinoma TCGA-ACC Adrenocortical carcinoma 92 323 Slide Images
Bladder Urothelial Carcinoma TCGA-BLCA Bladder urothelial carcinoma 412 926 Slide Images
Brain Lower Grade Glioma TCGA-LGG Lower-grade glioma 516 1,572 Slide Images
Breast Invasive Carcinoma TCGA-BRCA Breast invasive carcinoma 1,098 3,112 Slide Images
Cervical Squamous Cell Carcinoma and Endocervical Adenocarcinoma TCGA-CESC Cervical squamous cell carcinoma and adenocarcinoma 307 604 Slide Images
Cholangiocarcinoma TCGA-CHOL Cholangiocarcinoma 51 110 Slide Images
Colon Adenocarcinoma TCGA-COAD Colon adenocarcinoma 461 1,442 Slide Images
Esophageal Carcinoma TCGA-ESCA Esophageal carcinoma 185 396 Slide Images
Glioblastoma Multiforme TCGA-GBM Glioblastoma 617 2,053 Slide Images
Head and Neck Squamous Cell Carcinoma TCGA-HNSC Head and neck squamous cell carcinoma 528 1,263 Slide Images
Kidney Chromophobe TCGA-KICH Kidney chromophobe 113 326 Slide Images
Kidney Renal Clear Cell Carcinoma TCGA-KIRC Kidney renal clear cell carcinoma 537 2,173 Slide Images
Kidney Renal Papillary Cell Carcinoma TCGA-KIRP Kidney renal papillary cell carcinoma 291 773 Slide Images
Liver Hepatocellular Carcinoma TCGA-LIHC Liver hepatocellular carcinoma 377 870 Slide Images
Lung Adenocarcinoma TCGA-LUAD Lung adenocarcinoma 585 1,608 Slide Images
Lung Squamous Cell Carcinoma TCGA-LUSC Lung squamous cell carcinoma 504 1,612 Slide Images
Lymphoid Neoplasm Diffuse Large B-cell Lymphoma TCGA-DLBC Diffuse large B-cell lymphoma 58 103 Slide Images
Mesothelioma TCGA-MESO Mesothelioma 87 175 Slide Images
Ovarian Serous Cystadenocarcinoma TCGA-OV Ovarian serous cystadenocarcinoma 608 1,481 Slide Images
Pancreatic Adenocarcinoma TCGA-PAAD Pancreatic adenocarcinoma 185 466 Slide Images
Pheochromocytoma and Paraganglioma TCGA-PCPG Pheochromocytoma and paraganglioma 179 385 Slide Images
Prostate Adenocarcinoma TCGA-PRAD Prostate adenocarcinoma 500 1,172 Slide Images
Rectum Adenocarcinoma TCGA-READ Rectum adenocarcinoma 172 530 Slide Images
Sarcoma TCGA-SARC Adult soft tissue sarcoma 261 890 Slide Images
Skin Cutaneous Melanoma TCGA-SKCM Skin cutaneous melanoma 470 950 Slide Images
Stomach Adenocarcinoma TCGA-STAD Stomach adenocarcinoma 443 1,197 Slide Images
Testicular Germ Cell Tumors TCGA-TGCT Testicular germ cell tumors 263 663 Slide Images
Thymoma TCGA-THYM Thymoma 124 318 Slide Images
Thyroid Carcinoma TCGA-THCA Thyroid carcinoma 507 1,158 Slide Images
Uterine Carcinosarcoma TCGA-UCS Uterine carcinosarcoma 57 154 Slide Images
Uterine Corpus Endometrial Carcinoma TCGA-UCEC Uterine corpus endometrial carcinoma 560 1,371 Slide Images
Uveal Melanoma TCGA-UVM Uveal melanoma 80 150 Slide Images
Total (33 projects) — — 11,428 30,326 Slide Images

The “Total” row is a direct sum of the same quantities (case counts, Slide Image counts) across the 33 TCGA project sub-reports: 11,428 cases and 30,326 Slide Images. It is a project-level sum under the current GDC release, not a count of unique patients or slides (cases and slides are many-to-many within a project and do not overlap across projects), and it differs from the frozen cohorts in the papers.

Total GDC file counts range from thousands to tens of thousands per project. GDC file counts, case counts and the frozen cohort sizes in the papers often differ: project statistics reflect the current release, while paper numbers reflect a specific study design and data freeze, so the two should be cited separately.

4. Cancer types and clinico-pathological scope

4.1 Main cancer groups

4.2 Molecular and pathological subtypes

TCGA project papers and GDC metadata support many disease-specific classifications, for example common molecular subtypes of breast cancer, the EBV/MSI/GS/CIN classes of gastric cancer, POLE/MSI/copy-number classes of endometrial cancer, Type 1/Type 2 papillary renal cell carcinoma, BRAF/RAS/NF1/Triple-WT classes of cutaneous melanoma, DDLPS/LMS/UPS/MFS/MPNST/SS subtypes of soft tissue sarcoma, and the pathological and molecular categories of colorectal and esophageal cancer.

These classifications come from different papers, data freezes and analysis pipelines. When merging across projects, keep the original diagnosis, disease type, histological type, molecular subtype and publication cohort fields and build a separate standardized mapping table.

5. Data modalities and image structure

5.1 Digital pathology images

Most TCGA projects include GDC Slide Image files, usually in .svs format, as Tissue Slide and Diagnostic Slide. The slides come mainly from H&E-stained histology review; some project papers also describe pathologist review, tumor cell fraction, necrosis fraction, sample quality and tissue processing.

There is a clear exception: the current GDC file listing for TCGA-LAML confirms no Slide Images (0). For the other projects, whether images exist, how many, their access status and file format should be taken from the current GDC manifest. The presence of GDC Slide Images does not mean pixel-level masks, cell-level annotations or unified pathology task labels exist.

5.2 Genome, transcriptome and epigenome

Common data objects include:

Raw sequencing data and files such as BAM/FASTQ are often under controlled access; processed matrices, clinical facets and some mutation and image files may be public. Researchers should filter item by item by data category, data type, access, experimental strategy and file format.

5.3 Spatial omics

None of the reports on the 33 projects confirms a unified Visium, Xenium, MERFISH, GeoMx or other spatial transcriptomics delivery across TCGA, or a unified ST spot/bin resolution. TCGA pathology WSIs can be used as tissue morphology data, but WSI pixel coordinates and spatial transcriptomics spot/cell coordinates mean different things.

6. Clinical, sample and classification metadata

Common GDC metadata levels include:

The project reports also record many case-level and paper-level classification fields, such as AJCC stage, WHO grade, histological type, molecular subtype, tumor/normal status, tumor cell fraction, necrosis fraction, RIN and sample quality flags. Field coverage, coding and capitalization, missing values and time points vary by project.

Keep the original GDC metadata, the cohort variables defined in papers and the cBioPortal analysis-ready profiles separate. cBioPortal is convenient for quickly reading mutation, CNA, mRNA, methylation, RPPA, clinical data and case lists; its sample scope may be smaller than all cases of the current GDC project.

7. Quality control and processing advice

TCGA project papers usually describe sample selection, pathologist review, tumor cell fraction, necrosis fraction, RNA quality, sequencing depth, platform batches and multi-omics data-freeze conditions. QC rules differ between projects, and a GDC file listing cannot replace image- and matrix-level quality checks.

Suggested workflow:

  1. Fix the GDC project ID, API query time, data release and manifest.
  2. Build a linked index by case, sample, portion, analyte and file.
  3. Filter and record each file's access level, data category, data type, experimental strategy and format.
  4. For SVS images, check pyramid levels, pixel size, tissue coverage, blank areas, out-of-focus regions, folds, scanning artifacts and duplicate slides.
  5. For sequencing and omics data, check reference versions, batches, missing values, coverage, expression-matrix versions, sample pairing and the source of molecular profiles.
  6. For clinical fields, check age, sex, stage, grade, diagnosis, tumor/normal flags and missing-value codes.
  7. Split data by case or sample so that WSIs, patches, omics files or duplicate samples of the same case do not leak across partitions.
  8. Record the differences between the frozen cohorts in the papers and the current GDC project cohorts; do not use paper case counts in place of the real size of the files downloaded now.

8. Demographics and representativeness

TCGA covers many cancer types and research centers, but case composition is determined by each project's enrollment, tissue acquisition and experimental design. Coverage of age, sex, race, ethnicity, clinical stage and sample type varies markedly between projects. Some projects have specific sex or age structures; for example the prostate, testis, ovary, breast and uterus projects have clear disease-related population characteristics, and the samples of hematologic projects such as LAML and DLBC also differ from solid tumors.

Fairness analyses should at least stratify by cancer type, project, center, age, sex, race, ethnicity, sample type and clinical stage. Where public metadata is missing or controlled, report the coverage of the available fields, and do not interpret TCGA study cohorts as representative samples of the general population.

9. Suitable research directions

TCGA is suited to:

TCGA projects suit research analyses and resource reuse. The existing project reports confirm no unified computational pathology task, fixed labels, official benchmark split or leaderboard; researchers need to define their own case-level splits, label rules and external validation plans.

10. Boundaries and key limitations

  1. Tiered access: open images and metadata have different access conditions from controlled sequencing, BAM/FASTQ and some clinical data.
  2. Version differences: current GDC project statistics may differ from the frozen cohorts in historical TCGA papers.
  3. Project heterogeneity: the 33 projects differ in cancer type, sample processing, sequencing platform, pathology review, batches and clinical variables.
  4. Case-level dependence: one case can have several slides, sample types and molecular files, so random patch splits cause leakage.
  5. Limited pathology annotation: Slide Image files usually lack unified pixel-level masks, cell-level annotations and pathology task labels.
  6. Image coverage differs: the current report confirms TCGA-LAML has no Slide Images (0), and image counts for the other projects should also be rechecked against the current manifest.
  7. Label semantics are not unified: definitions of diagnosis, histology, molecular subtype, stage and grade are set by each project and paper.
  8. Limited population representativeness: cancer-specific enrollment, missing public fields and differing project sizes affect fairness and generalizability.
  9. Cross-platform bias: GDC, cBioPortal, paper supplementary data and other analysis platforms may differ in sample scope and processing version.
  10. Storage and compute cost: WSIs, raw sequencing and multi-omics files are large; plan storage, transfer and compute before downloading.

11. Recommended workflow

First decide the cancer project and research modality, then query cases, samples, files and access status through the GDC project page or API. cBioPortal is useful for quick exploration of molecular profiles, but formal analyses should keep the GDC project ID, manifest, file versions and sample-selection records. After downloading, first build a case–sample–file relation table, run clinical, image, sequencing and omics QC separately, and then split by case and start modeling. Results should state the GDC data release, query date, the scope of open/controlled files, and the differences between the paper's frozen cohort and the current project cohort.

12. Citation advice

TCGA analyses should cite the official TCGA/GDC entry, the specific project papers, and the GDC data release and file listing actually used. For downstream analyses, also cite the specific study/profile version of the cBioPortal PanCancer Atlas or UCSC Xena. Pathology image studies should state the project IDs, Slide Image file types, image selection rules and case-level split strategy.