Stochastic Analytical Predictive Models for Life Sciences and Crop Production Process
Graduation Year
2024
Document Type
Dissertation
Degree
Ph.D.
Degree Name
Doctor of Philosophy (Ph.D.)
Degree Granting Department
Mathematics and Statistics
Major Professor
Chris P. Tsokos, Ph.D.
Committee Member
Kandethody M. Ramachandran, Ph.D.
Committee Member
Lu Lu, Ph.D.
Committee Member
Yicheng Tu, Ph.D.
Keywords
3-Parameter Fatigue Life Distribution, Arterial Stenosis, Artificial Neural Network, Desirability Function, Power Law Process, 3-P-Log-logistic distribution
Abstract
Analytical predictive modeling uses algorithm-based mathematical, probabilistic and statistical methods to anticipate future events by identifying patterns in past data. It is a technique for predicting outcomes and a key application of statistical analysis in real-world scenarios. Real data-driven predictive models in various fields, such as life sciences, economics,or production industry, help individuals and institutions make data-driven informed decisions, which is essential for business success, providing companies with a competitive edge.
One of every five adult deaths is caused by heart disease and one person dies every 33 from cardiovascular disease in the United States according to the Center for Disease Control and Prevention (CDC) Our study is concerned about one of the major causes of heart attack and stroke, the Percentage Stenosis of the Arteries of the Heart, (PSAH). The percentagestenosis is a measure of the degree of narrowing or closeness of a blood vessel, usually an artery, expressed as a percentage of the vessel’s original diameter. The most common disease caused by stenosis is coronary artery disease (CAD), which is the leading cause of death worldwide, affecting millions of people. In this dissertation, our main goal is to address the key problem in cardiovascular disease, the Percentage Stenosis of the Arteries of the Heart, PSAH. we proposed an analytical predictive model that identifies the risk factors, individually and interactively, that drive the PSAH of patients likely to experience cardiovascular events such as heart attacks or strokes and predicts the PSAH with at least a 96% accuracy. we obtained the ranks of the individual and interaction risk factors according to the percentage contribution PSAH. we performed a step-by-step method of optimizing (minimum) the PSAH using the risk factors that drive PSAH. Through desirability function approach we obtained optimal values of the risk factors of PSAH for three grading stages of stenosis severity: Extreme (Severe), Moderate, and Mild. We further developed a parametric analysis by defining the probability behavior that probabilistically characterizes the behavior of PSAH. This useful information provides guidance for treatment decisions concerning patients with high risk of PSAH.
The time it takes for an individual or patient diagnosed with a particular type of cancer to survive (survival time) has become a significant area of research. Some cancers have long survival periods of up to 120 months, while others have much shorter survival times, as brief as 6 months. Our study focuses on one of the cancers with the lowest average survival rates, Malignant Gliomas (Primary Brain Tumor). According to the National Cancer Institute (NCI), long-term survival (5-year survival rate) is achieved by less than 5% of patients, with a median survival ranging from 7 to 24 months.
The Kaplan-Meier and Cox Proportional Hazard (Cox-PH) models have traditionally been used for survival analysis of cancer data. These methods are derived from nonparametric and semi-parametric approaches, respectively, which are less robust compared to parametric approaches and supervised machine learning algorithms. In the era of Big Data, the importance and use of supervised machine learning models, like artificial neural networks (ANN), have greatly increased in modern statistics. In this dissertation, we propose ANN based survival analysis models to predict the survival times of malignant glioma patients. The model considers the risk factors contributing to the survival times and subsequently identified the top 10 significant risk factors in predicting patient survival times. This makes our ANN-based survival analysis models efficient alternatives to conventional survival analysis models, offering improved predictive power. Additionally, we conducted parametric survival analysis, identifying a well-defined 3-Parameter Generalized Pareto Probability Distribution to obtain the survival function of patients and compare its estimates with the Kaplan-Meier estimator.
Another part of the research study in this dissertation focuses on the soybean production process, which involves identifying, optimizing, and monitoring the production returns of soybeans in the United States. Soybeans are the second most valuable crop in the United States, contributing about $48.6 billion in cash crop receipts to U.S. farmers in 2021. The United States was the leading producer and exporter of soybeans for decades, from 1990 until 2018, until Brazil took over in 2019. One of the key agendas of the United States Department of Agriculture (USDA) is to regain its position as the leading producer of soybeans. In this study, we develop a real data-driven analytical model for the production returns of soybeans in the United States. The model identifies and ranked six significant individual contributable factors and five significant contributable interaction terms that accurately predict the production returns of soybeans in the United States with at least 97% accuracy.
We further performed an optimization analysis to maximize the Production Returns of Soybeans in the United States. we determined the optimal combination of the attributable factors or risk factors that maximize the desired outcome, that is, the production returns of soybeans. We utilized 2-D contour plots and 3-D surface plots to explore the optimal combination of two input or risk factors that maximize the production returns while holding other factors constant and provide guidance on how to achieve a target value of the response while controlling the other risk factors. monitor, assess, and evaluate the production returns of soybeans in the United States. We employed the Power Law Process (PLP) to develop two (2) very important methods, the Stochastic Production Intensity Function (SPIF) and the Stochastic Production Monitoring Indicator (SPMI). The SPIF captures the probabilistic behavior and the rate of change of the production returns as a function of time while the SPMI serves as an indicator that monitors the rate of change and signals the behavior of the production returns at a specific time. The interpretation of the SPMI signals is explained in enabling the manager to determine whether production returns are increasing, decreasing, or remaining unchanged. This information is extremely important to production managers to implement appropriate strategic decisions to ensure that the annual returns are profitable.
Scholar Commons Citation
Tetteh-Bator, Erasmus, "Stochastic Analytical Predictive Models for Life Sciences and Crop Production Process" (2024). USF Tampa Graduate Theses and Dissertations.
https://digitalcommons.usf.edu/etd/11156
