Patient Stratification
5 min
the patient stratification heatmap shows the gene expression pattern of the graph genes across the samples contributing to the graph construction it helps you quickly see which genes are high or low in expression, and how variable patient cohorts or genes relate to each other the patient stratification card contains further details and inclusion options on the datasets contributing to the heatmap note that in some networks, sample selection is required to ensure that no more than 1,500 samples are included in the heatmap, in order to reduce computational time and maintain a clear visualization the heatmap intensity goes from blue (low expression levels) to red (high expression levels) the genes are shown over the rows with row labels showing the gene details, and the samples are shown over columns with column labels showing the annotated sample details the heatmap can thus be used to detect gene profiles and potential biomarkers for all variable patient cohorts with curated metadata details when a second network is added to the graph creation, as described under two networks comparison docid\ tkqki5g opbwejgrivtj3 , first select the nework based on which data you want the heatmap to be generated from, before clicking the "generate heatmap" button note that the genes within the heatmap will include all genes within the graph if you get this note "n genes from the graph are not included in the heatmap datasets ", it indicates the number of genes not identified in the data contributing to heatmap production we are working on updating the heatmaps within mavatar discovery, with the intention to include more data and to have a higher level metadata integration the majority of heatmaps now are either based on one dataset representative of the tissue and disease, or a copy from the disease specific networks respective tissue – general network if your network of interest contains a heatmap based on only one dataset, or contains more diseases than your network of interest, be aware that an improved version is underway read more about how the heatmaps are produced under expression heatmap construction docid\ kv178h6llc16gg2xzfbgh heatmap management and functions manage the heatmap by selecting which gene and sample details to show under show/hide tracks hover over the gene and sample labels to see the details click on a gene to highlight in within the graph this will also show it within the gene information docid\ xtq51aiwkyz4fk97ay eg , conditions expression chart docid\ bo2l7qhkrocn1ezlqhaiv , and cell type explorer docid\ xfu1ggawgnttln4ypfw0v cards right click on a gene docid 2cxkfmzdutkkjqeigyscd for further actions sample category enrichment analysis the statistical analysis can be used to test whether any of the sample categories are significantly enriched or depleted within any of the hierarchical clusters in the heatmap to run it, define the number of clusters (any value between two and 50) and select the metadata set to perform the statistics on a global test is first performed to evaluate whether any of the categories show variable expression across any of the clusters a pairwise comparison is then performed to check, for each category and cluster, whether the genes are enriched or depleted the global test used depends on the smallest sample size involved, as well as the number of clusters and categories being compared the standard procedure is the chi squared test (χ²) applied to the full contingency table however, if the sample size is smaller than five in any of the categories, a fisher's exact test is used instead for 2×2 tables, or a monte carlo simulation of the χ² statistic when there are more categories to compare the output includes information on which statistical test was used in each case for each category within each cluster, a pairwise comparison is then performed by computing the odds ratio (or) to quantify the association between that category and the cluster, relative to all other categories and clusters combined if any of the groups contain zero counts, a haldane anscombe correction is applied before the or is calculated the category is defined as enriched or depleted within the cluster based on whether the or is greater than or less than 1, respectively a fisher's exact test is then performed to evaluate whether this enrichment or depletion is statistically significant, with p values adjusted for multiple comparisons using the benjamini hochberg (fdr) correction the results are presented in a downloadable table