Using Decision Trees to predict 5 year mortality

Using GUIDE to study survival of patients for 5 years after the diagnosis date and determine factors affecting survival.

RGUIDESurvival AnalysisDecision TreesCLLMBL

Overview

This project explores patients and genomic dataset of 1148 patients of Chronic Lymphocytic Leukemia (CLL) and Monoclonal B-cell Lymphocytosis (MBL). CLL (1100 patients) is a form of leukemia related cancer and MBL (54 patients) is a precursor condition to CLL.

This report used GUIDE to study survival of patients for 5 years after the diagnosis date. This project also used GUIDE output to plot survival plots and determine factors affecting survival from the data used.

The objective of this project was to study factors contributing to patient mortality.

Data

The main dataset folder contained 14 data files formatted in .txt file format. Each datafile was studied along with corresponding meta data text files to select columns that would contribute to determination of factors contributing towards the Survival Analysis of patients.

From 14 data files, 6 data files were used to create 2 data files:

  1. A data file containing clinical and laboratory information (df_1)
  2. A data file containing clinical, laboratory and genetic information (df_2)

Two separate datafiles were made for purposes of Study Design (univariate/multivariate analysis).

Gene information from the original data file contained in rows representing singular genes were transposed to later become a dataset where each column represented a particular gene (~20000 genes) with dummy variable (0/1) signalling presence/absence of gene in a patient.

Each data point in a data file was modified to remove all instances of space, commas and any special characters that might cause GUIDE to misread the datafile.

Merging of the datafile was performed using the left merge function in R and attention to detail was maintained to not miss any rows of patient data while merging two data files.

New variables surv_5yrs and death_5yr were constructed to represent survival length and status for patients for a 5 year time frame.

Methods

Study Design

Models of univariate and multivariate nature were performed to observe singular and multivariate effects of factors towards survival prediction and analysis of patient data.

Univariate survival analysis of predictors of poor prognosis of CLL as well as multivariate survival analysis using all columns was performed.

An important score analysis was performed on df_1 and df_2 separately to observe clinical-laboratory variables’ importance score and then clinical-laboratory-genetic importance score on survival status.

GUIDE was also used to build and analyze classification tree to predict survival status of patients:

A single tree was used for all analysis with output being classification unless otherwise stated.

Regression was performed on censored responses with dependent variables being surv_5yrs.

The Kaplan Meier Curve was drawn to study the survival curves from the fit.txt file generated from GUIDE.

Quantile regression was also performed using GUIDE to study factors present in patients who are at 90th percentile or above at the risk of death.

Traditional Methods

Traditional methods like Random Forest, C-trees, Ranger and RPart were used to study prediction using simulation shown in class.

Logistic Regression was attempted early in the semester but later abandoned due to a dataset containing categorical variables of factor level 1 and imputation requiring clinical input.

Results

GUIDE model identified X17p_status as the most significant predictor of 5-year mortality (Rank 1 in importance score: 2.586), followed closely by X8p_status (Rank 2: 1.972) and IGHV mutation status (Rank 3: 1.790) in importance score for classification of survival of patients.

Age at sampling was consistently identified across analyses as a meaningful contributor to survival prediction.

It is important to distinguish why variable selection via importance scoring differs from node selection in a decision tree. While chi-squared tests and optimal splits are used to determine node selection, importance scores are calculated on the basis of chi-square tests with Bonferroni corrections and strongest adjusted association with optimal split with minimal squared error.

It was observed that rare genetic variants such as ATM and TP53 (< 25 instances each in gene pool of ~20000), which are known to contribute to poor prognosis and diagnosis of patients' conditions, did not appear prominently in either the importance scores or tree nodes.

Additionally, survival curve analysis of variables like del(17p) status, IGHV mutation status and RAI stage was conducted with inconsistent results when univariate analysis was done, i.e. surv_5yrs as dependent variable and Variable X as “c”.

Whilst at the same time, traditional survival curve attempted in 01_univariate_model.pdf gave an expected result.

Quantile Regression

Quantile regression was performed on the 5-year survival time (continuous variable) to study the 10th, 50th, and 90th percentiles of patients and observe factors affecting patient survival.

Censored response analysis was also performed on chromosomes’ arm structure variables using GUIDE to determine specific factors affecting survival. This analysis was conducted in a manner analogous to univariate survival analysis.

It was seen that TREATMENT_STATUS_AT_SAMPLING is the primary predictor for 5-year survival across 10th, 50th and 90th quantiles with patients who were untreated at the time of sampling generally showing better outcomes.

The 10th percentile model reveals the X17p_status genetic marker is the most critical factor for patients in the highest-risk category.

Across all quantiles, additional factors such as IGHV_MUTATION_STATUS and patient age refine survival estimates, though a significant portion of the population reaches the maximum observed survival value of 60 units.

This is likely because CLL and MBL are slow progressing diseases that sometimes require years of monitoring before treatment is started.

Survival Curve

X17p_status was identified as the most significant predictor at the root node, followed by X8p_status.

Patients with this genetic loss are assigned to Node 2 with a log-relative hazard coefficient of 1.174, while those without the loss (classified as "Unchanged" or other) are assigned to Node 3 with a lower coefficient of -0.133.

This indicates a lower relative risk of death within five years.

Despite these differing risk profiles, the median survival time for both groups in this specific model was calculated at the maximum observed value of 60 units.

Importance Score

While the important score remained mainly of variables mentioned in Quantile Regression and Cna_survival_curve_guide - it should be noted that rare variants which are known in the literature to be significant contributors to the disease are not present even in top 10 most important predictors.

This is not a critique of GUIDE but of the modeling technique used with an expectation to find genetic rare variants.

Traditional R Method

Both IGHV mutation status and X17p status demonstrate high statistical significance (p < 0.0001), with unmutated IGHV patients facing a 2.24-fold higher hazard of death and those with 17p loss experiencing a dramatic 3.7-fold increase in risk.

Conversely, the RAI stage at diagnosis fails to show a statistically significant association with 5-year survival (p ≈ 0.16–0.63), characterized by overlapping Kaplan Meier curves and a weak hazard ratio that suggests it is not a robust independent predictor in this dataset.

Simulation

After 50 simulations:

MethodError
ctree0.2029
cforest0.1957
rpart0.2231
ranger0.1901
gcon0.2212
gforest0.2259

Conclusion

This project allowed me to study patient mortality using decision trees, survival analysis and genomic information.

GUIDE identified X17p_status, X8p_status and IGHV mutation status as important predictors of survival. Age at sampling was also consistently identified across analyses as a meaningful contributor to survival prediction.

An important observation from this project was that rare genetic variants such as ATM and TP53 did not appear prominently in the importance scores or tree nodes despite being known to contribute to poor prognosis.

The project also demonstrated how different methods can produce different results when studying survival and mortality in clinical and genomic datasets.