NePSTA 队列
4 个德国中心 107 个样本,含空间转录组、H&E 与 EPIC 甲基化等诊断结果;8 项任务的主数据。
论文声明验证队列已在 Zenodo 与 Dryad 公开;训练队列未见公开存放。
数据入口 ↗论文报告 · 以数据为中心
NePSTA · Nature Cancer,2025 · 8 项任务的数据、划分与结果
这项研究把 Visium 空间转录组与图神经网络结合,用于中枢神经系统肿瘤的常规诊断。共 8 项分析。复用前要分清:前 4 项检查数据质量并与常规检测对照,不训练模型;后 4 项训练模型,按患者划分数据。
4 个德国中心 107 个样本,含空间转录组、H&E 与 EPIC 甲基化等诊断结果;8 项任务的主数据。
论文声明验证队列已在 Zenodo 与 Dryad 公开;训练队列未见公开存放。
数据入口 ↗11 个样本,只在质控任务中作外部对照。
处理后数据在 GEO(GSE194329)与 Figshare 公开;原始数据需向 GSA-Human 申请,约 3–6 周。
数据入口 ↗整合 16 个数据集、110 名患者的胶质母细胞瘤单细胞数据,用作细胞类型字典。
只在任务 4、8 中定义细胞类型,不是测试集。GEO 公开(GSE211376)。
数据入口 ↗| 任务 | 数据与规模 | 标签或参照 | 划分 | 主要结果 |
|---|---|---|---|---|
| 空间转录组质控 比较不同处理方式、组织类型下的数据质量 | NePSTA 107 个样本 + Ren 等 11 个外部样本 | 无标签,按处理方式、组织类型分组比较 | 不划分 | 各组每 spot UMI 数无显著差异;spot 数主要取决于活检还是整块切除 |
| 虚拟免疫组化 由空间表达推断 Ki67、GFAP、NeuN 的蛋白丰度 | NePSTA 中 12 名有相邻切片 IHC 的参与者 | 相邻切片的 IHC 信号强度 | 不划分,与参照对比 | 相关系数:Ki67 0.47,GFAP 0.32,NeuN 0.57 |
| 拷贝数变异推断 由空间表达推断染色体增益和缺失 | NePSTA 中有甲基化芯片拷贝数结果的参与者 | 甲基化芯片的拷贝数判定 | 不划分,与参照对比 | 一致率 81.2%;三分类 AUC 80.37% |
| 细胞组成与微环境 估计每个 spot 的细胞类型,按甲基化亚类比较微环境 | NePSTA 全部样本 + GBMap(16 个数据集、110 名患者) | GBMap 细胞类型特征;甲基化亚类用于分组 | 不划分 | 间充质亚型免疫细胞丰富;RTK 亚型缺少 T 细胞和巨噬细胞浸润 |
| 组织区域分割 把切片分为主瘤、坏死、浸润区、白质和皮层 | NePSTA:队列 1 的 41 人、队列 2 的 27 人 | 神经病理专家的 H&E 区域标注 | 按患者:41 人训练(内部十折),27 人验证 | 验证集准确率 87.43% |
| 甲基化亚类分类 由空间数据预测肿瘤区域的 DNA 甲基化亚类 | NePSTA 中有 EPIC 甲基化结果的参与者;9.7 万个肿瘤子图 | EPIC 甲基化芯片的亚类 | 按患者五折交叉验证 | 患者级准确率 0.893;3 跳邻域子图准确率 0.999,无邻域 0.409 |
| MGMT 启动子甲基化 由空间数据预测 MGMT 启动子甲基化状态 | NePSTA 中有 MGMT 结果的 53 人 | EPIC 甲基化芯片的 MGMT 状态 | 一个队列 30 人训练,另一队列 23 人评估 | 准确率 99%;F1、精确率、召回率均为 1.0 |
| CDKN2A/B 缺失 预测 CDKN2A/B 纯合或杂合缺失,并比较缺失区域的细胞组成 | NePSTA 中有 CDKN2A/B 结果的 53 人;GBMap 用于邻域分析 | 甲基化芯片的 CDKN2A/B 状态 | 一个队列 30 人训练,另一队列 23 人评估 | 子图准确率 85.4%;缺失区域富集单核细胞与血管内皮细胞 |
Source location: Paper | PDF p. 11
“For data analysis and quality control, we used the Cell Ranger pipeline from 10X Genomics.”
“Standard import procedures include normalizing gene expression, which is achieved using the Seurat version 4.0 package.”
Source location: Paper | PDF p. 2
“Our study involved samples from four German medical centers (Heidelberg, n = 50; Freiburg, n = 41; Mannheim, n = 15; Memmingen, n = 1), along with external controls published recently9 (n = 11) and 12 healthy cortex samples, totaling 130 samples.”
“The first cohort (n = 66) comprised samples from three centers, encompassing a range of CNS pathologies from highly malignant glioblastomas (GBs) to epilepsy-associated (glio)neuronal tumors.”
“These samples underwent comprehensive spatial transcriptomics profiling, alongside state-of-the-art diagnostic workups, including morphology inspection, EPIC methylation arrays, classifier predictions, IHC and NGS.”
Source location: Paper | PDF p. 2
“In cohort 1, the majorit of samples (n = 42, 63.6%) were processed from paraffin-embedded tis sue, with only 24 samples (36.6%) processed from freshly frozen tissue Conversely, cohort 2 included approximately half of the samples (n = 18 43.9%) processed from paraffin-embedded tissue.”
“First, we aimed to explore and quantify potential confounders of the technology, as clinical applications demand robust and consistent quality readouts.”
Source location: Paper | PDF p. 2; Paper | PDF p. 24
“Our study involved samples from four German medical centers (Heidelberg, n = 50; Freiburg, n = 41; Mannheim, n = 15; Memmingen, n = 1), along with external controls published recently9 (n = 11) and 12 healthy cortex samples, totaling 130 samples.”
“The first cohort (n = 66) comprised samples from three centers, encompassing a range of CNS pathologies from highly malignant glioblastomas (GBs) to epilepsy-associated (glio)neuronal tumors.”
“The second cohort (n = 41) was a single-center cohort from Freiburg, similarly profiled (Fig. 1b).”
“In this study we report a total of 130 samples”
Source location: Paper | PDF p. 2; Paper | PDF p. 12
“We start by converting the IHC image to grayscale by extracting the blue-channel (img[,,3]) and transform into data.frame format using the reshape2::melt() function. The pixel intensity values were extracted and rescaled to a range of 0–1.”
“When juxtaposed with consecutive participant sections, our virtual stainings exhibited robust alignment with IHC-derived results (Fig. 2b).”
Source location: Paper | PDF p. 2; Paper | PDF p. 5
“To this end, we devised an innovative computational module named ‘inferred IHC’. This module harnesses super-resolution spatial transcriptomics through the Bayesian inference11 to forecast protein abundance, rendering it a viable diagnostic surrogate for traditional IHC (Fig. 2a,b).”
“Paired IHC and spatia transcriptomics (for analysis of a–d) was available for 12 participants.”
Source location: Paper | PDF p. 2; Paper | PDF p. 5
“Paired IHC and spatia transcriptomics (for analysis of a–d) was available for 12 participants.”
“When juxtaposed with consecutive participant sections, our virtual stainings exhibited robust alignment with IHC-derived results (Fig. 2b).”
Source location: Paper | PDF p. 5
“Paired IHC and spatia transcriptomics (for analysis of a–d) was available for 12 participants.”
Source location: Paper | PDF p. 11
“Chromosomal bins were created using the SPATAwrapper::Create.ref.bins() function, with a bin size of 1 Mbp used for this study. Data were then rescaled and interpolated over a 10-kbp window, with normalization achieved using a loess regression model”
Source location: Paper | PDF p. 4; Paper | PDF p. 4
“Direct comparison of methylation-based CNV against our inferred CNV detection revealed a consensus of CNV alterations in 81.2% cases (consensus no alterations, 70.3%; consensus gains, 5.5%; consensus loss, 5.4%), only 0.05% divergent gains (gains detected by Visium and loss detected in the methylation-based CNV analysis) and 0.02% divergent losses (Fig. 3d).”
“We performed multiclass receiver operating characteristic (ROC) analysis to validate the accuracy of detecting gains and losses or diploid chromosome sets, demonstrating an overall area under the curve (AUC) of 80.37% (Fig. 3f).”
Source location: Paper | PDF p. 4
“To more precisely examine the potential variability in accuracy of CNV detection across different chromosomal regions”
Source location: Paper | PDF p. 4
“Direct comparison of methylation-based CNV against our inferred CNV detection revealed a consensus of CNV alterations in 81.2% cases (consensus no alterations, 70.3%; consensus gains, 5.5%; consensus loss, 5.4%), only 0.05% divergent gains (gains detected by Visium and loss detected in the methylation-based CNV analysis) and 0.02% divergent losses (Fig. 3d).”
Source location: Paper | PDF p. 12
“We trained the model on a graphics processing unit for 500 iterations to ensure computational efficiency. After training, we used the export\_posterior function to extract the posterior distribution of cell type proportions”
Source location: Paper | PDF p. 4
“Leveraging the robust and well-validated Cell2lo cation algorithm, combined with an extensive GB reference dataset (GBMap), we predicted the abundance of myeloid, T cell and stromal subpopulations across all methylation classes (Fig. 4a).”
Source location: Paper | PDF p. 4
“The meth ylation Mes subtype exhibited immune-rich microenvironments with an immunosuppressive myeloid profile, while others (RTKI and RTKII) displayed a notable absence of T cell and tumor-associated macrophage infiltration (Extended Data Fig. 3).”
Source location: Paper | PDF p. 4
“Leveraging the robust and well-validated Cell2lo cation algorithm, combined with an extensive GB reference dataset (GBMap), we predicted the abundance of myeloid, T cell and stromal subpopulations across all methylation classes (Fig. 4a).”
Source location: Paper | PDF p. 13
“Expression profiles (top 5,000 genes), CNAs, histological annotations (one-hot encoded) and encoded H&E image vectors were included”
“Edge features were defined by spatial proximity with up to six neighbors per node”
Source location: Paper | PDF p. 6; Paper | PDF p. 11
“A pivotal step before subgroup prediction involved automated spatial transcriptomics data segmentation, aiming to classify subgraphs histologically.”
“Manual segmentation of the histological regions was performed by a neuropathologist according to the morphological features of the H&E scan.”
Source location: Paper | PDF p. 6
“We used a tenfold cross-validation to train on n = 41”
“samples with sufficient histological segmentation)”
Source location: Paper | PDF p. 6
“We used a tenfold cross-validation to train on n = 41”
“Model validation on n = 27 participants”
Source location: Paper | PDF p. 13; Paper | PDF p. 13
“Expression profiles (top 5,000 genes), CNAs, histological annotations (one-hot encoded) and encoded H&E image vectors were included”
“Edge features were defined by spatial proximity with up to six neighbors per node”
“A three-hop neighborhood for a query spot included all spots within three edges, capturing the spatial context and connectivity within the defined distance.”
Source location: Paper | PDF p. 7; Paper | PDF p. 13
“From 107 EPIC-characterized participants, datasets were divided into training and validation subsets. For multibiopsy samples, datasets were split by biopsy cores; for single-specimen samples, manual segmentation was performed using SPATA2 to create regions.”
“We divided the samples into two distinct spatial transcriptomic datasets. For samples with multiple biopsies, we treated the raw data from each biopsy as separate datasets. In cases where the entire tissue was on a single slide, we segmented the samples as illustrated in Fig. 5a.”
Source location: Paper | PDF p. 7; Paper | PDF p. 13
“From 107 EPIC-characterized participants, datasets were divided into training and validation subsets. For multibiopsy samples, datasets were split by biopsy cores; for single-specimen samples, manual segmentation was performed using SPATA2 to create regions.”
“To avoid bias from similar normalization and preprocessing in both the training and the validation cohorts, we processed each dataset individually after splitting the data at the count level.”
Source location: Paper | PDF p. 13 — plus Web | publisher supplementary PDF, Supplementary Table 1
“From 107 EPIC-characterized participants, datasets were divided into training and validation subsets. For multibiopsy samples, datasets were split by biopsy cores; for single-specimen samples, manual segmentation was performed using SPATA2 to create regions.”
“Training data comprised 97,000 subgraphs from tumor datasets and 12,000 subgraphs from healthy controls, created using a three-hop neighborhood approach.”
“Supplementary Table 1: Evaluation of different models to explore the impact of extended neighborhoods for the predictive power of NEPSTA GIN: Graph isomorphic network, GAN: Graph attention network. The model evaluation was performed on 84 patients.”
Source location: Paper | PDF p. 13
“Expression profiles (top 5,000 genes), CNAs, histological annotations (one-hot encoded) and encoded H&E image vectors were included”
“Edge features were defined by spatial proximity with up to six neighbors per node”
Source location: Paper | PDF p. 7; Paper | PDF p. 21
“Extended Data Fig. 6 | Heterogeneity of CDKN2A/B alterations. (a) Schematic of the GNN workflow Since the MGMT and CDKN2A/B status is available in both cohorts (n = 53 patients) we train (n = 30 patients) and evaluate (n = 23 patients) on separate cohorts.”
“To harness this insight, we augmented our GNN framework, incorporating multiple MLPs specifically designed to predict MGMT promoter methylation and the loss of CDKN2A/B”
Source location: Paper | PDF p. 21
“Extended Data Fig. 6 | Heterogeneity of CDKN2A/B alterations. (a) Schematic of the GNN workflow Since the MGMT and CDKN2A/B status is available in both cohorts (n = 53 patients) we train (n = 30 patients) and evaluate (n = 23 patients) on separate cohorts.”
Source location: Paper | PDF p. 21
“Extended Data Fig. 6 | Heterogeneity of CDKN2A/B alterations. (a) Schematic of the GNN workflow Since the MGMT and CDKN2A/B status is available in both cohorts (n = 53 patients) we train (n = 30 patients) and evaluate (n = 23 patients) on separate cohorts.”
Source location: Paper | PDF p. 13
“Expression profiles (top 5,000 genes), CNAs, histological annotations (one-hot encoded) and encoded H&E image vectors were included”
“Edge features were defined by spatial proximity with up to six neighbors per node”
Source location: Paper | PDF p. 8; Paper | PDF p. 21
“and further differentiation between homozygous and heterozygous deletion were found to be impossible in cases when the loss was not associated with chr9p loss (Fig. 6a,b).”
“Extended Data Fig. 6 | Heterogeneity of CDKN2A/B alterations. (a) Schematic of the GNN workflow Since the MGMT and CDKN2A/B status is available in both cohorts (n = 53 patients) we train (n = 30 patients) and evaluate (n = 23 patients) on separate cohorts.”
Source location: Paper | PDF p. 10; Paper | PDF p. 21
“Contrary to our expectations, our findings revealed that the majority of samples with heterozygous deletions (two of three) displayed a mixed genotype, encompassing both partial and complete loss of CDKN2A/B across different biopsy specimens.”
“Extended Data Fig. 6 | Heterogeneity of CDKN2A/B alterations. (a) Schematic of the GNN workflow Since the MGMT and CDKN2A/B status is available in both cohorts (n = 53 patients) we train (n = 30 patients) and evaluate (n = 23 patients) on separate cohorts.”
Source location: Paper | PDF p. 10
“Contrary to our expectations, our findings revealed that the majority of samples with heterozygous deletions (two of three) displayed a mixed genotype, encompassing both partial and complete loss of CDKN2A/B across different biopsy specimens.”
Ritter 等,Nature Cancer,2025。
Spatially resolved transcriptomics and graph-based deep learning improve accuracy of routine CNS tumor diagnostics
Paper report · data-centered
NePSTA · Nature Cancer, 2025 · data, splits and results of 8 tasks
This study combines Visium spatial transcriptomics with graph neural networks for routine diagnosis of central nervous system tumors, in 8 analyses. Before reuse, keep them apart: the first 4 check data quality and compare against routine assays without training a model; the last 4 train models with patient-level splits.
107 samples from 4 German centers, with spatial transcriptomics, H&E and diagnostic results such as EPIC methylation; the main data for all 8 tasks.
The paper states that the validation cohort is public on Zenodo and Dryad; no public deposit of the training cohort was found.
Data entry ↗11 samples, used only as an external reference in the quality-control task.
Processed data is public on GEO (GSE194329) and Figshare; raw data requires an application to GSA-Human, about 3–6 weeks.
Data entry ↗Integrates glioblastoma single-cell data from 16 datasets and 110 patients, used as a cell-type dictionary.
Used only to define cell types in tasks 4 and 8; it is not a test set. Public on GEO (GSE211376).
Data entry ↗| Task | Data and scale | Labels or reference | Split | Main results |
|---|---|---|---|---|
| Spatial transcriptomics QC Compare data quality across processing methods and tissue types | 107 NePSTA samples + 11 external samples from Ren et al. | No labels; groups compared by processing method and tissue type | No split | No significant difference in UMIs per spot between groups; the number of spots depends mainly on biopsy versus en bloc resection |
| Virtual immunohistochemistry Infer Ki67, GFAP and NeuN protein abundance from spatial expression | 12 NePSTA participants with IHC on adjacent sections | IHC signal intensity on adjacent sections | No split; compared with the reference | Correlation: Ki67 0.47, GFAP 0.32, NeuN 0.57 |
| Copy number inference Infer chromosomal gains and losses from spatial expression | NePSTA participants with copy number results from methylation arrays | Copy number calls from methylation arrays | No split; compared with the reference | Agreement 81.2%; three-class AUC 80.37% |
| Cell composition and microenvironment Estimate cell types per spot and compare the microenvironment across methylation subclasses | All NePSTA samples + GBMap (16 datasets, 110 patients) | GBMap cell-type signatures; methylation subclasses for grouping | No split | The mesenchymal subtype is rich in immune cells; the RTK subtype lacks T-cell and macrophage infiltration |
| Tissue region segmentation Divide a section into main tumor, necrosis, infiltration zone, white matter and cortex | NePSTA: 41 patients in cohort 1, 27 in cohort 2 | Neuropathologist H&E region annotations | By patient: 41 for training (internal 10-fold), 27 for validation | Validation accuracy 87.43% |
| Methylation subclass classification Predict the DNA methylation subclass of tumor regions from spatial data | NePSTA participants with EPIC methylation results; 97,000 tumor subgraphs | Subclass from EPIC methylation arrays | Patient-level 5-fold cross-validation | Patient-level accuracy 0.893; 3-hop subgraph accuracy 0.999, no neighborhood 0.409 |
| MGMT promoter methylation Predict MGMT promoter methylation status from spatial data | 53 NePSTA patients with MGMT results | MGMT status from EPIC methylation arrays | 30 patients from one cohort for training, 23 from the other for evaluation | Accuracy 99%; F1, precision and recall all 1.0 |
| CDKN2A/B deletion Predict homozygous or heterozygous CDKN2A/B deletion and compare the cell composition of deleted regions | 53 NePSTA patients with CDKN2A/B results; GBMap for the neighborhood analysis | CDKN2A/B status from methylation arrays | 30 patients from one cohort for training, 23 from the other for evaluation | Subgraph accuracy 85.4%; deleted regions are enriched in monocytes and vascular endothelial cells |
Source location: Paper | PDF p. 11
“For data analysis and quality control, we used the Cell Ranger pipeline from 10X Genomics.”
“Standard import procedures include normalizing gene expression, which is achieved using the Seurat version 4.0 package.”
Source location: Paper | PDF p. 2
“Our study involved samples from four German medical centers (Heidelberg, n = 50; Freiburg, n = 41; Mannheim, n = 15; Memmingen, n = 1), along with external controls published recently9 (n = 11) and 12 healthy cortex samples, totaling 130 samples.”
“The first cohort (n = 66) comprised samples from three centers, encompassing a range of CNS pathologies from highly malignant glioblastomas (GBs) to epilepsy-associated (glio)neuronal tumors.”
“These samples underwent comprehensive spatial transcriptomics profiling, alongside state-of-the-art diagnostic workups, including morphology inspection, EPIC methylation arrays, classifier predictions, IHC and NGS.”
Source location: Paper | PDF p. 2
“In cohort 1, the majorit of samples (n = 42, 63.6%) were processed from paraffin-embedded tis sue, with only 24 samples (36.6%) processed from freshly frozen tissue Conversely, cohort 2 included approximately half of the samples (n = 18 43.9%) processed from paraffin-embedded tissue.”
“First, we aimed to explore and quantify potential confounders of the technology, as clinical applications demand robust and consistent quality readouts.”
Source location: Paper | PDF p. 2; Paper | PDF p. 24
“Our study involved samples from four German medical centers (Heidelberg, n = 50; Freiburg, n = 41; Mannheim, n = 15; Memmingen, n = 1), along with external controls published recently9 (n = 11) and 12 healthy cortex samples, totaling 130 samples.”
“The first cohort (n = 66) comprised samples from three centers, encompassing a range of CNS pathologies from highly malignant glioblastomas (GBs) to epilepsy-associated (glio)neuronal tumors.”
“The second cohort (n = 41) was a single-center cohort from Freiburg, similarly profiled (Fig. 1b).”
“In this study we report a total of 130 samples”
Source location: Paper | PDF p. 2; Paper | PDF p. 12
“We start by converting the IHC image to grayscale by extracting the blue-channel (img[,,3]) and transform into data.frame format using the reshape2::melt() function. The pixel intensity values were extracted and rescaled to a range of 0–1.”
“When juxtaposed with consecutive participant sections, our virtual stainings exhibited robust alignment with IHC-derived results (Fig. 2b).”
Source location: Paper | PDF p. 2; Paper | PDF p. 5
“To this end, we devised an innovative computational module named ‘inferred IHC’. This module harnesses super-resolution spatial transcriptomics through the Bayesian inference11 to forecast protein abundance, rendering it a viable diagnostic surrogate for traditional IHC (Fig. 2a,b).”
“Paired IHC and spatia transcriptomics (for analysis of a–d) was available for 12 participants.”
Source location: Paper | PDF p. 2; Paper | PDF p. 5
“Paired IHC and spatia transcriptomics (for analysis of a–d) was available for 12 participants.”
“When juxtaposed with consecutive participant sections, our virtual stainings exhibited robust alignment with IHC-derived results (Fig. 2b).”
Source location: Paper | PDF p. 5
“Paired IHC and spatia transcriptomics (for analysis of a–d) was available for 12 participants.”
Source location: Paper | PDF p. 11
“Chromosomal bins were created using the SPATAwrapper::Create.ref.bins() function, with a bin size of 1 Mbp used for this study. Data were then rescaled and interpolated over a 10-kbp window, with normalization achieved using a loess regression model”
Source location: Paper | PDF p. 4; Paper | PDF p. 4
“Direct comparison of methylation-based CNV against our inferred CNV detection revealed a consensus of CNV alterations in 81.2% cases (consensus no alterations, 70.3%; consensus gains, 5.5%; consensus loss, 5.4%), only 0.05% divergent gains (gains detected by Visium and loss detected in the methylation-based CNV analysis) and 0.02% divergent losses (Fig. 3d).”
“We performed multiclass receiver operating characteristic (ROC) analysis to validate the accuracy of detecting gains and losses or diploid chromosome sets, demonstrating an overall area under the curve (AUC) of 80.37% (Fig. 3f).”
Source location: Paper | PDF p. 4
“To more precisely examine the potential variability in accuracy of CNV detection across different chromosomal regions”
Source location: Paper | PDF p. 4
“Direct comparison of methylation-based CNV against our inferred CNV detection revealed a consensus of CNV alterations in 81.2% cases (consensus no alterations, 70.3%; consensus gains, 5.5%; consensus loss, 5.4%), only 0.05% divergent gains (gains detected by Visium and loss detected in the methylation-based CNV analysis) and 0.02% divergent losses (Fig. 3d).”
Source location: Paper | PDF p. 12
“We trained the model on a graphics processing unit for 500 iterations to ensure computational efficiency. After training, we used the export\_posterior function to extract the posterior distribution of cell type proportions”
Source location: Paper | PDF p. 4
“Leveraging the robust and well-validated Cell2lo cation algorithm, combined with an extensive GB reference dataset (GBMap), we predicted the abundance of myeloid, T cell and stromal subpopulations across all methylation classes (Fig. 4a).”
Source location: Paper | PDF p. 4
“The meth ylation Mes subtype exhibited immune-rich microenvironments with an immunosuppressive myeloid profile, while others (RTKI and RTKII) displayed a notable absence of T cell and tumor-associated macrophage infiltration (Extended Data Fig. 3).”
Source location: Paper | PDF p. 4
“Leveraging the robust and well-validated Cell2lo cation algorithm, combined with an extensive GB reference dataset (GBMap), we predicted the abundance of myeloid, T cell and stromal subpopulations across all methylation classes (Fig. 4a).”
Source location: Paper | PDF p. 13
“Expression profiles (top 5,000 genes), CNAs, histological annotations (one-hot encoded) and encoded H&E image vectors were included”
“Edge features were defined by spatial proximity with up to six neighbors per node”
Source location: Paper | PDF p. 6; Paper | PDF p. 11
“A pivotal step before subgroup prediction involved automated spatial transcriptomics data segmentation, aiming to classify subgraphs histologically.”
“Manual segmentation of the histological regions was performed by a neuropathologist according to the morphological features of the H&E scan.”
Source location: Paper | PDF p. 6
“We used a tenfold cross-validation to train on n = 41”
“samples with sufficient histological segmentation)”
Source location: Paper | PDF p. 6
“We used a tenfold cross-validation to train on n = 41”
“Model validation on n = 27 participants”
Source location: Paper | PDF p. 13; Paper | PDF p. 13
“Expression profiles (top 5,000 genes), CNAs, histological annotations (one-hot encoded) and encoded H&E image vectors were included”
“Edge features were defined by spatial proximity with up to six neighbors per node”
“A three-hop neighborhood for a query spot included all spots within three edges, capturing the spatial context and connectivity within the defined distance.”
Source location: Paper | PDF p. 7; Paper | PDF p. 13
“From 107 EPIC-characterized participants, datasets were divided into training and validation subsets. For multibiopsy samples, datasets were split by biopsy cores; for single-specimen samples, manual segmentation was performed using SPATA2 to create regions.”
“We divided the samples into two distinct spatial transcriptomic datasets. For samples with multiple biopsies, we treated the raw data from each biopsy as separate datasets. In cases where the entire tissue was on a single slide, we segmented the samples as illustrated in Fig. 5a.”
Source location: Paper | PDF p. 7; Paper | PDF p. 13
“From 107 EPIC-characterized participants, datasets were divided into training and validation subsets. For multibiopsy samples, datasets were split by biopsy cores; for single-specimen samples, manual segmentation was performed using SPATA2 to create regions.”
“To avoid bias from similar normalization and preprocessing in both the training and the validation cohorts, we processed each dataset individually after splitting the data at the count level.”
Source location: Paper | PDF p. 13 — plus Web | publisher supplementary PDF, Supplementary Table 1
“From 107 EPIC-characterized participants, datasets were divided into training and validation subsets. For multibiopsy samples, datasets were split by biopsy cores; for single-specimen samples, manual segmentation was performed using SPATA2 to create regions.”
“Training data comprised 97,000 subgraphs from tumor datasets and 12,000 subgraphs from healthy controls, created using a three-hop neighborhood approach.”
“Supplementary Table 1: Evaluation of different models to explore the impact of extended neighborhoods for the predictive power of NEPSTA GIN: Graph isomorphic network, GAN: Graph attention network. The model evaluation was performed on 84 patients.”
Source location: Paper | PDF p. 13
“Expression profiles (top 5,000 genes), CNAs, histological annotations (one-hot encoded) and encoded H&E image vectors were included”
“Edge features were defined by spatial proximity with up to six neighbors per node”
Source location: Paper | PDF p. 7; Paper | PDF p. 21
“Extended Data Fig. 6 | Heterogeneity of CDKN2A/B alterations. (a) Schematic of the GNN workflow Since the MGMT and CDKN2A/B status is available in both cohorts (n = 53 patients) we train (n = 30 patients) and evaluate (n = 23 patients) on separate cohorts.”
“To harness this insight, we augmented our GNN framework, incorporating multiple MLPs specifically designed to predict MGMT promoter methylation and the loss of CDKN2A/B”
Source location: Paper | PDF p. 21
“Extended Data Fig. 6 | Heterogeneity of CDKN2A/B alterations. (a) Schematic of the GNN workflow Since the MGMT and CDKN2A/B status is available in both cohorts (n = 53 patients) we train (n = 30 patients) and evaluate (n = 23 patients) on separate cohorts.”
Source location: Paper | PDF p. 21
“Extended Data Fig. 6 | Heterogeneity of CDKN2A/B alterations. (a) Schematic of the GNN workflow Since the MGMT and CDKN2A/B status is available in both cohorts (n = 53 patients) we train (n = 30 patients) and evaluate (n = 23 patients) on separate cohorts.”
Source location: Paper | PDF p. 13
“Expression profiles (top 5,000 genes), CNAs, histological annotations (one-hot encoded) and encoded H&E image vectors were included”
“Edge features were defined by spatial proximity with up to six neighbors per node”
Source location: Paper | PDF p. 8; Paper | PDF p. 21
“and further differentiation between homozygous and heterozygous deletion were found to be impossible in cases when the loss was not associated with chr9p loss (Fig. 6a,b).”
“Extended Data Fig. 6 | Heterogeneity of CDKN2A/B alterations. (a) Schematic of the GNN workflow Since the MGMT and CDKN2A/B status is available in both cohorts (n = 53 patients) we train (n = 30 patients) and evaluate (n = 23 patients) on separate cohorts.”
Source location: Paper | PDF p. 10; Paper | PDF p. 21
“Contrary to our expectations, our findings revealed that the majority of samples with heterozygous deletions (two of three) displayed a mixed genotype, encompassing both partial and complete loss of CDKN2A/B across different biopsy specimens.”
“Extended Data Fig. 6 | Heterogeneity of CDKN2A/B alterations. (a) Schematic of the GNN workflow Since the MGMT and CDKN2A/B status is available in both cohorts (n = 53 patients) we train (n = 30 patients) and evaluate (n = 23 patients) on separate cohorts.”
Source location: Paper | PDF p. 10
“Contrary to our expectations, our findings revealed that the majority of samples with heterozygous deletions (two of three) displayed a mixed genotype, encompassing both partial and complete loss of CDKN2A/B across different biopsy specimens.”
Ritter et al., Nature Cancer, 2025.
Spatially resolved transcriptomics and graph-based deep learning improve accuracy of routine CNS tumor diagnostics